Almost everything is being sold as an agent right now, so start with the distinction that actually separates them. Anthropic's is still the sharpest: a workflow orchestrates models and tools through predefined code paths. An agent dynamically directs its own process and tool use, including deciding when it is finished.
The practical test is who decides the next step. If a human or a fixed pipeline decides, it is a workflow, and that is often the better engineering choice. Anthropic's own guidance says so plainly: find the simplest thing that works, "and only increas[e] complexity when needed. This might mean not building agentic systems at all."
How many companies actually run them
Two credible surveys, wildly different answers, and the gap is entirely sampling. Gartner's 2026 CIO survey found 17% of organisations have deployed AI agents. LangChain's survey of agent builders found 57.3% with agents in production, but that sample is people who already chose to build agents.
McKinsey's August 2026 global survey, 1,719 respondents across 97 nations, is the useful middle: around two in ten organisations have reached genuine scale with agents, rising to 40% among enterprises above $1bn in revenue and flat at 22% for smaller ones.
Gartner also coined the term for what is inflating those numbers: agent washing, rebranding assistants, RPA and chatbots without substantive agentic capability. Their estimate is that roughly 130 of the thousands of self-described agentic vendors are the real thing.
The reliability number that matters
OSWorld measures agents doing real computer tasks. On the original benchmark, frontier models score in the low-to-mid 80s, above the 72% human baseline. It reads like a solved problem.
OSWorld 2.0, published June 2026, kept the same task family and made the tasks long: a median human completion time of about 1.6 hours instead of two minutes, requiring 250-plus agent steps instead of around 30. Claude Opus 4.8 scored 20.6% on binary completion.
83.5% on two-minute tasks. 20.6% on ninety-minute tasks. Length is the variable, and almost all real work is long.
The benchmark authors' own diagnosis of why the original was misleading is worth reading: its tasks "rarely span more than one or two applications, and reward completing self-contained actions rather than sustaining long, connected workflows." The documented failure modes are mundane and instructive: agents do not ask for missing information, cannot parse complex sources, mark tasks complete without verifying, and forget what they read 300 steps ago.
Scores are inflated, and by a measured amount
SWE-ABS, presented at ICML 2026, rebuilt the test suites behind a widely used coding benchmark and found that one in five "solved" patches from the top thirty agents is semantically incorrect, passing only because the tests were too weak to catch the error. Strengthening the tests dropped the leading system from 78.8% to 62.2%, and dropped the top agent to fifth place.
There is a second measurement problem that matters more for production. Anthropic's own engineering write-up spells it out: pass@k asks whether any of k attempts succeeds, while pass^k asks whether all k succeed. "At k=1, they're identical... By k=10, pass@k approaches 100% while pass^k falls to 0%." Demos are pass@k. Customers are pass^k.
Cost, and the assumption that more reasoning helps
The Holistic Agent Leaderboard from Princeton is the strongest independent cost-aware evaluation available: 21,730 agent rollouts across nine models and nine benchmarks. Two findings should change how people buy.
- In only one of nine benchmarks was the most expensive model on the cost-performance frontier, despite costing two orders of magnitude more than the cheapest.
- In 21 of 36 model-agent-benchmark combinations, increasing reasoning effort produced equal or lower accuracy.
Reading 2.5 billion tokens of agent logs, the team also found agents finding answers by searching arXiv rather than solving the task, taking real-world actions like booking a flight on the wrong credit card, and almost never completing a run without at least one tool-calling failure even when they ultimately succeeded.
Where agents genuinely work
Software engineering under human review is the strongest evidence in the field. A study of Microsoft's early-2026 rollout of command-line coding agents across tens of thousands of engineers found adopters merged around 24% more pull requests than they otherwise would have, sustained over four months. The authors' caveat belongs with the number: a merged pull request is a proxy for output, not for value, and the study reports nothing about code quality or rework.
Beyond that, the pattern is consistent: short, bounded, verifiable tasks with a human checkpoint. Long, multi-application, open-ended work is where the numbers collapse.
The security problem is structural
OWASP's June 2026 report states the root cause without hedging: models treat the system prompt, the user request and any retrieved text as one stream of tokens, and "there is no reliable way to mark some of those tokens as commands and others as data." Prompt injection is not a bug awaiting a patch.
Two operating rules have emerged that are worth adopting whatever you are building. Simon Willison's lethal trifecta: private data access, exposure to untrusted content, and the ability to communicate externally are safe in any two combinations and dangerous in all three. Meta's agents rule of two follows from it: an autonomous agent should satisfy at most two, and all three requires a human to approve.
How to decide
Three questions, in order. Does this task need something to decide the next step, or would a fixed pipeline do it more cheaply and predictably? If it goes wrong ten times out of a hundred, what does that cost, and who notices? And what is the cost per completed, verified task, not per token? The vendors have quietly moved to cost-per-task framing in their own launch materials, which tells you raw accuracy has stopped being the thing that differentiates.
Sources
- Anthropic, "Building effective agents"
- OECD, "What agentic AI is and does" (March 2026)
- Gartner, Hype Cycle for Agentic AI (15 April 2026)
- McKinsey, "The State of AI" global survey (25 August 2026)
- OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks (June 2026)
- Snorkel AI, "OSWorld 2.0: why computer use agents fail most tasks" (3 September 2026)
- Anthropic, "Demystifying evals for AI agents" (9 January 2026)
- Holistic Agent Leaderboard, Princeton (ICLR 2026)
- Murphy-Hill, Butler & Savelieva, "Adoption and Impact of Command-Line AI Coding Agents" (1 July 2026)
- OWASP GenAI Security Project, State of Agentic AI Security and Governance (June 2026)