DAYJOB benchmark finds AI agents struggle with long-horizon professional work
A new benchmark tested AI agents on 130 long-horizon tasks in healthcare and finance, and even its strongest tested model passed fewer than one in four attempts. The researchers say a recurring failure was accepting a wrong premise in sourc
Open discussion →