At a recent Modal conference, AI leaders made the case for software agents that work towards outcomes, while developers and company executives described the gap between AI doing much of a task and taking it off people’s hands. The question is no longer just what models can do, but whether that capability is becoming useful, dependable automation.
Watch Desk analysis
What happened
The Information reports that Cognition chief executive Scott Wu argued that more data centres, memory and sandboxes could help companies move towards “virtual employees” working on outcomes rather than tasks. He pointed to progress in mathematical reasoning as a source of confidence in that direction.
But the conference also heard a more practical account of what remains unfinished. Anthropic’s Claude Code product chief Cat Wu said teams can still find themselves doing the same tasks after an AI assistant has completed much of the work. Diogo Almeida, who now leads TypeSafe AI, framed the problem bluntly: “Where the misfile is all the automation?”
The article also describes open-source developer Dax Raad’s concern that engineers are not paying enough attention to code written by AI. Separately, it notes that economists have yet to see signs of AI-driven productivity in major statistics. That is a broad economic picture, not proof that no individual team is benefiting.
The Information’s account of the conference
Why it matters
These are two different measures of progress. Models may become more capable, while the work around them still needs people to check results, manage hand-offs and deal with systems that block bots. For software teams, an assistant completing 80 per cent of a job is useful, but it is not the same as removing the job from the queue.
That distinction also matters for claims about productivity. A striking model result or a polished agent demo can show what is possible; it does not, by itself, show that organisations are getting more done. The conference accounts offer a useful view of the friction, though they are not a systematic measure of industry-wide adoption.
Our read
The strongest test of AI automation is not whether a model can finish a task in a controlled demo. It is whether people can hand over a real process, understand what the system did and spend less time supervising or repairing the result. The industry has plenty of ambition. The more interesting question is whether the hand-offs are getting boringly reliable.
What to watch
- Whether companies can show measured productivity gains beyond individual demonstrations.
- How software teams balance agent-written code with meaningful human review.
- Whether AI agents can complete useful workflows when websites and services resist automated access.
Discussion spark: Should companies judge AI progress by what a model can do in a demo, or by whether workers can hand over a real process with less supervision?
Sources and evidence
- AI's Simmering Question: "Where the F— Is All the Automation?" (8 October 2026, 00:20 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.