What is an AI agent, actually? A definition and a buyer’s test

A training invitation arrived in my inbox on 18 August. Build Your First GTM AI Agent with Claude Cowork. Good trainer, real substance, worth the hour.

Read the description and the agent finds LinkedIn commenters, enriches them through Apollo, ranks them against your ICP, and drafts the outreach. Read the last line of the same description and the thing you have built is called a reusable skill.

That gap between the label on the tin and the thing inside it has a name now. Gartner calls it agent washing.

What is an AI agent?

An AI agent is a system where the model decides its own next step.

Anthropic published the cleanest version of this distinction in December 2024 and it has held up better than anything written since. They group everything under agentic systems, then split it in two.

Workflows are systems where models and tools are orchestrated through predefined code paths. Agents are systems where models dynamically direct their own processes and tool usage, keeping control over how they accomplish tasks.

In practice the agent thinks, calls a tool, reads what comes back, and decides again, and it keeps going until it has enough. Anthropic’s plainer version: agents are typically just models using tools based on environmental feedback in a loop.

The reason this matters is what a model cannot otherwise do. A model cannot decide halfway through an answer that it needs more information and go and get it. Every word it produces comes from what is already in front of it. The loop is what lets it act on the world, look at the result, and adjust.

Perplexity Deep Research works this way. So do ChatGPT agent mode, Claude Code, and Cursor.

What is the difference between an agent and a workflow?

Who decides the sequence.

In a workflow, a person wrote the steps in advance. Scrape the post, enrich the list, score against the ICP, draft the email. The model does the writing and the judgment inside each step. It does not choose which step comes next, and it cannot invent a step nobody planned for.

In an agent, the model chooses. It might search twice, discard both results, try a different tool, and stop early. Nobody wrote that path down.

Almost everything currently sold as an agent is a workflow. That is not automatically a scam, and workflows are frequently the better buy. They cost less, they fail in ways you can predict, and you can read them and know what they will do on Tuesday. Anthropic’s own advice to developers is to find the simplest solution possible and only add complexity when it demonstrably improves the outcome, which sometimes means not building an agentic system at all.

Mike Bayly of The AI Corner makes the practical version of this argument for New Zealand businesses. Teach people to build a skill, not an agent. A skill in his definition is a reusable recipe: the instructions, the company context, the code for calculations that must be exact, and where to go for information. Agents, he writes, bring added complexities that not everyone learning to use AI needs to think about.

What is agent washing?

Gartner’s term, from June 2025, for vendors rebranding existing AI assistants, robotic process automation and chatbots as agents without substantial agentic capability underneath.

Their estimate is the number worth remembering. Of the thousands of agentic AI vendors in the market, Gartner reckons only around 130 are real. They also predict that more than 40% of agentic AI projects will be cancelled by the end of 2027, on escalating costs, unclear business value, or inadequate risk controls.

Do real agents work? Yes, and here are the numbers

Kieran Flanagan leads Agentic GTM and Systems at HubSpot, a role he points out did not exist six months before he wrote about it in May 2026. HubSpot’s published figures are the best public evidence in the category, and they are self-reported, so read them as a company describing its own results.

  • Their Inbound Agent handles 82% of all inbound chats with no human involved.
  • The Prospecting Agent books over 10,000 meetings per quarter.
  • The Demand Agent added 345,000 accounts to their addressable market.
  • Their answer engine agent grew qualified leads from AI-generated answers by 1,850% between Q1 2025 and Q1 2026, and those leads convert at up to three times the rate of traditional search.

So the technology is real. The interesting part is what that team did before building anything.

Why do most agent projects fail?

Two reasons, and neither is the model.

The first is data. HubSpot invested heavily in enriching contacts, removing defunct company records and building golden record standards across the CRM before the Demand Agent produced anything. Flanagan’s line on it is worth repeating exactly: if your CRM is messy, your agents will be confidently wrong at scale. If you are not sure whether your own foundations are ready for this, that is what the BUILD Scorecard is for.

The second is reliability, and this is the finding most people have not caught up with.

A Princeton team including Sayash Kapoor and Arvind Narayanan published a paper in February 2026, accepted at ICML, called Towards a Science of AI Agent Reliability. They measured 15 models across two benchmarks using twelve metrics grouped into four dimensions: consistency, robustness, predictability, and safety. Their finding, in their own words, is that recent capability gains have only yielded small improvements in reliability.

The distinction they draw is the one benchmarks hide. When a model scores 70% on a task, does that mean it works every single time on 70% of cases, or that it fails unpredictably three times in ten on any case? A single success metric cannot tell you which. For a marketing deliverable that difference is annoying. For an agent emailing your prospect list, it is the entire risk profile.

What should you ask before you buy one?

One question, in two halves. What does this thing decide on its own, and what happens when it decides wrong?

The second half is where the money is, and August produced hard evidence for it. Anthropic tested Claude Code across 1,053 people before making auto mode the default on 14 August.

What the testing found

A screening classifier caught 89% of dangerous commands. Human reviewers caught 13.6%. People approve 97% of permission prompts, and 62% had turned the prompts off altogether.

Anthropic’s own explanation for that gap is reflexive clicking rather than scrutiny. The approval dialog was never doing the work everyone assumed it was doing.

There is a live example of what that costs. On 10 August, ABC News reported what it called Australia’s first known autonomous AI cyberattack. A man named Andrew asked OpenClaw, an agent platform running on Claude, to manage his gym bookings. Sitting fourth on a waitlist, he also asked whether he could move up. The agent found the booking system’s API had no authorisation checks at all and cancelled the reservation of the person in first place. Andrew had not asked for that. In his words, the API has zero authorisation checks on cancelling other people’s reservations, and he tested it on waitlist position one, and it went through.

Nobody wrote a step that said remove a stranger from a queue. The agent chose it, because it was pursuing the goal it had been given and the goal said nothing about other people.

Will New Zealand customers accept it?

Selectively, and the pattern is clear.

Financial Services Council research released at their 2026 conference found 61% of Kiwis comfortable with AI detecting fraud and scams, 58% with it explaining financial products, and 57% with it answering questions. Comfort with AI making financial decisions for them drops to 35%, with 44% actively uncomfortable. Seventy percent are concerned about privacy and how their personal data gets used.

Kiwi customers will accept an agent that helps them. They are much slower to accept one that decides for them. Which puts the design question in front of the technology question every time.

So before you buy the agent, map the decision. Write down what the thing chooses without asking, and what it costs you if it chooses badly. That list is the first page of an AI governance policy, and if you want a structured version, the AI Governance Check walks you through it. If the list comes back empty, you are buying a workflow, and you should pay workflow prices.

Which decisions in your business are you genuinely willing to hand over? If you want to talk one through, let’s talk.


Sources

Share the Post:

YOU MIGHT ALSO LIKE

Secret Link