← BlogAI Consultancy

Agent Washing Is Real: How to Tell a Genuine AI Agent Vendor From a Rebranded Chatbot

By Aditya JhaOctober 2, 202610 min read

Agent Washing Is Real: How to Tell a Genuine AI Agent Vendor From a Rebranded Chatbot

A Series B fintech's CFO in London sits through a 45-minute demo from an "AI agent" vendor. The agent handles customer refund requests end to end, the sales engineer promises. It looks seamless: a chat window, a refund gets approved, a confirmation email goes out. Three weeks into the pilot, her ops lead discovers the "agent" is a fixed decision tree wired to a chatbot front end, every edge case silently falls back to a human inbox, and the vendor's own engineers can't explain what happens when two steps fail in the same session. She isn't an edge case. Gartner estimates only about 130 of the thousands of vendors now selling agentic AI are delivering anything that meets the definition, and most of the rest are engaged in what Gartner itself calls "agent washing": the rebranding of existing chatbots, RPA bots and assistants without substantial agentic capabilities.

What actually separates an AI agent from a chatbot with extra steps?

It comes down to who decides the next step. Anthropic's own engineering definition is the cleanest line in the industry: "workflows are systems where LLMs and tools are orchestrated through predefined code paths," while "agents are systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks." Everything else people argue about, autonomy, memory, judgment, follows from that one distinction.

A vendor demo where every branch was scripted in advance by a human is a workflow wearing a chat interface, no matter how fluent the conversation sounds. A real agent decomposes a goal it has never seen phrased that exact way before, picks its own next action, and changes course when a tool call fails, the mechanism we unpacked in how AI agents work: function calling and architecture. If you're trying to decide whether a use case even needs this, n8n vs. AI agents: when to use each is the companion read.

Why would a vendor rebrand a chatbot as an agent in the first place?

Budget and buzz move faster than engineering. Gartner's Anushree Verma put it plainly: most agentic AI propositions "lack significant value or return on investment (ROI), as current models don't have the maturity and agency to autonomously achieve complex business goals or follow nuanced instructions over time," and "many use cases positioned as agentic today don't require agentic implementations." Rebranding an existing chatbot is cheaper than rebuilding one, and most buyers never ask the technical question that would catch it.

This isn't a minor labeling dispute. The same Gartner analysis found that over 40% of agentic AI projects will be canceled by the end of 2027 due to escalating costs, unclear business value, or inadequate risk controls, almost entirely projects sold on a capability the underlying system never actually had. A January 2025 Gartner poll of 3,412 webinar attendees shows how much of the market is still guessing at that gap rather than confirming it.

Source: Gartner, January 2025 poll of 3,412 webinar attendees, cited in Gartner's June 25, 2025 press release on agentic AI project cancellations.
Source: Gartner, January 2025 poll of 3,412 webinar attendees, cited in Gartner's June 25, 2025 press release on agentic AI project cancellations.

What's the fastest way to test a vendor's agent claim in the room?

Don't ask what it does, ask what happens when it's wrong. A practical six-axis scorecard for spotting agent washing gives you the exact dimensions to probe live in a demo, each scored from 0 (washed) to 5 (genuinely agentic):

  • **Autonomy** — does it only respond when a human prompts it, or does it initiate steps and escalate at defined checkpoints on its own?
  • **Planning** — is the flow fixed and rewritten by a developer every time the process changes, or does it decompose a novel goal into steps and replan when one fails?
  • **Tool use** — does a human execute every system action elsewhere, or does it independently select the right API, set the parameters, and handle the error itself?
  • **Memory** — does every session reset to zero, or does it retain durable state across sessions and make different decisions because of it?
  • **Feedback loop** — can it evaluate its own output against the goal and retry, or does a human have to catch every mistake downstream?
  • **Human-in-the-loop boundaries** — can the vendor state, as a number, what percentage of production runs need human intervention, or is that rate simply undocumented?

What should you actually ask the vendor before signing?

Four questions expose most agent washing in a single meeting, and a vendor building a genuine agent will answer them with specifics, usually logs, while a vendor who rebranded a chatbot gets vague or defensive:

  • "Which steps in this demo were pre-scripted by your team, and which were planned by the model at run time?"
  • "Show me a session where a tool call failed. What did the system do next, and who decided that, the model or a hardcoded fallback?"
  • "What's your current human-intervention rate in production, and will you put that number in the contract?"
  • "If I change the underlying goal mid-task, does it replan, or does the session just fail?"

How AIBOOTSTRAPPER helps

We built VitalPulse (see case studies) to pass exactly this kind of scrutiny, not just a demo. The triage assistant decides for itself how to route a patient conversation, ask more questions, flag urgency, or hand off, reads live vitals from a connected wearable as a genuine tool call rather than a static form, and hands off to a human doctor only at a defined checkpoint with a complete structured summary. That planning-plus-tool-use-plus-documented-handoff pattern is what cut consultation prep time by 68% while running bilingually in Arabic and English, 24/7, and it's the bar we hold our own builds to.

If you're mid-pilot with a vendor and can't get a straight answer to the four questions above, book a call and we'll help you run the audit, or see our AI consultancy services.

Want this done for you?

Book a free strategy call and we'll show you how to build and market your business with AI.

FAQ

Questions, answered

Everything you might want to know before we hop on a call.

Gartner's term for vendors rebranding existing chatbots, RPA bots, and virtual assistants as "AI agents" without adding the underlying planning, tool-use, and autonomy that actually defines an agent.

Gartner estimates only about 130 of the thousands of vendors now marketing agentic AI products are genuine, as of its June 2025 analysis on agentic AI project cancellations.

Per Anthropic's definition, a chatbot or workflow follows a path predetermined by a developer at build time; an agent lets the model itself decide the next step and which tools to call, based on the task, at run time.

Ask which steps are pre-scripted versus model-planned, request a session where a tool call failed, demand their current human-intervention rate in writing, and test whether the system can replan when the goal changes mid-task.

Keep reading

Let's talk

Ready to build and sell with AI?

Book a free 30 minute strategy call. We'll map the highest ROI AI move for your business, no pitch, just value.