A COO opens the Q3 board deck to the slide on the AI agent the company approved budget for in Q1. The slide says "performing well." Someone asks for a number, resolution rate, cost per ticket, hours saved, anything, and the room goes quiet, because nobody actually instrumented the thing before turning it on. That gap between "it feels like it's helping" and "here is the number that proves it" isn't a one-off embarrassment. Liferay's 2026 Agentic AI Adoption and Governance Report, based on a survey of 500 U.S. professionals involved in AI decision-making, found that 54% of organizations already have AI agents in production or active piloting, but only 25% consistently measure that agent's impact using defined KPIs. Most companies running agents today genuinely cannot tell you, in a number, whether the thing is working.
How many companies actually have AI agents live right now, and how many can prove it?
More than half. 54% of the organizations Liferay surveyed report AI agents in production or active piloting, led by technology companies at 72%, healthcare at 66%, and transportation and logistics at 64%. Data analysis and reporting is the single most common use case, cited by 35% of companies, followed by customer support at 32% and IT service desk work at 21%.
But adoption and proof are two different numbers, and the second one is far smaller. Only 25% of those organizations consistently measure AI agent impact with defined KPIs, and just 24% have a formal, company-wide AI usage policy governing how agents get deployed and reviewed in the first place. That's three-quarters of deployed agents running without a KPI anyone agreed on in advance, which is a very different problem than "the agent isn't good enough," the one most post-mortems reach for by default.

Why doesn't "it feels like it's helping" count as an answer?
Because it can't survive the next budget review. 30% of organizations in the same survey cite security or privacy concerns as a leading barrier to realizing agent value, and "unclear ROI" shows up as a barrier too, just further down the list than cost or training gaps, which tells you the measurement problem isn't rare enough to be an edge case. It's common enough to be its own line item in enterprise adoption research.
An agent that genuinely isn't working and an agent that's working fine but was never instrumented look identical from the outside: a quiet Slack channel, no escalations, nobody complaining. Without a number, you can't tell those two situations apart, which means you also can't tell a founder or a board whether to scale the thing, fix it, or kill it, the exact decision the broader production-vs-pilot gap shows most companies are making on instinct rather than data.
What should you actually track? A four-layer KPI framework
- **Adoption (leading indicator):** active usage against the people or workflows it was built for, not a deployment checkbox. An agent nobody routes real tickets to isn't failing, it's unmeasured on the metric that matters first.
- **Experience:** first contact resolution, customer effort score, and CSAT or NPS movement. Liferay's own recommended targets are concrete enough to set as real goals: first contact resolution up 15%, customer effort score down 10%, CSAT or NPS up 5 points.
- **Operational and business outcome (lagging indicator):** time-to-resolution, error or rework rate, and anything that ties directly to revenue or cost, lead conversion, procurement cycle time, complaint rate. These are the numbers a board actually reads, and they should be set as explicit targets before launch, not reverse-engineered from whatever the agent happens to produce.
- **Risk:** the rate of human interventions and corrections the agent needed, compliance violations, and data privacy incidents. A fast agent that quietly needs constant human correction isn't actually saving the time its adoption number implies.
How do you actually instrument this, mechanically, instead of just picking metric names?
Three things have to exist before an agent goes live, not after someone asks for a number. First, a baseline: what did resolution time, error rate, or cost per task look like before the agent, measured the same way you'll measure it after, so the comparison means something. Second, an event schema: every agent action that matters gets logged as a structured event tied to an outcome, a ticket resolved, a lead converted, a human override triggered, not just "agent responded," which proves activity but nothing about value delivered. Third, attribution: where possible, a control group or staggered rollout, so a metric improving isn't silently explained by something unrelated that changed at the same time.
None of this requires exotic tooling. It requires deciding, before launch, which of the four layers above actually matters for this specific agent, and wiring the logging for those numbers into the build instead of trying to reconstruct them from logs after a board asks.
How AIBOOTSTRAPPER helps
This is the same discipline behind VitalPulse, the remote patient monitoring and AI consultation assistant we built for a digital health startup in Dubai: the result wasn't a vague "doctors are happier," it was a measured 68% faster consultation prep time, because the baseline and the metric were defined before the AI assistant went live, not reconstructed afterward from anecdote.
If you're about to launch an agent and aren't sure what to measure, or you already have one live and can't actually answer whether it's working, book a call, or see how we scope AI consultancy and product engagements with the KPI defined on day one, not discovered in a board meeting six months later.
Want this done for you?
Book a free strategy call and we'll show you how to build and market your business with AI.
