All posts

Comparisons & pricing

89% automated resolution, measured the strict way — at about half the per-resolution price

  • AI metrics
  • Pricing
  • Comparisons
By the PilotPM team3 min read

Here is a number our sales conversations keep coming back to, so I'm putting it in writing with its full definition attached:

Across our production workspaces over the last 30 days, 89% of the conversations our AI answered were resolved — no human takeover, no further customer message.

And here is what makes that number different from the ones you'll see on other vendors' sites: we measure it stricter than they do, and we'll show you the parts they leave out.

The same words don't mean the same thing

"Automated resolution" sounds like a standard. It isn't. As of July 2026, from the vendors' own documentation:

  • Intercom Fin counts a resolution when the customer confirms — or simply goes quiet for 24 hours after Fin's answer. A teammate stepping in afterwards does not void it. Fin's published average is around 65–67% on that counting.
  • Zendesk requires 72 hours of silence, an LLM check that the request was actually satisfied, and no escalation to a human. Customer stories land roughly between 39% and 60%.
  • Ada disqualifies any human involvement and grades every conversation for relevance, accuracy, and safety. They cite an industry average around 52%.

Our rule sits at the strict end of that spectrum: if a human takes over after the AI answered, we count it as a failure. If the customer writes back — even weeks later — the resolution is voided. No 24-hour assumptions, no quiet denominator changes.

Measured that way, our production number is 89%. Under Fin's looser counting, it would read higher. We publish the strict one because it's the one a buyer can audit.

The number most vendors won't put next to it

Every resolution rate in this industry — including ours — is measured over conversations the AI answered, not over everything your customers sent. So the honest version of any resolution claim needs a second number: involvement — how much of your inbound the AI actually handles.

Ours is deliberately small today: under 5% of total inbound across production workspaces. That's not a limitation of the AI — it's the safety model. Every category starts in draft mode, where a human reviews everything. A category only graduates to auto-send when its acceptance record earns it, behind per-category confidence thresholds, a 50% canary test, and a statistical tripwire that automatically rolls back anything that makes replies worse.

So the pitch isn't "89% of everything." The pitch is: when this system is allowed to answer, the conversation ends — 89% of the time, measured strictly. Turning the dial up is then an evidence question, not a leap of faith: you expand the AI's territory category by category, watching the resolution rate hold as you go. That's how you get to the big vendors' coverage without ever having handed your inbox to an unproven bot.

What it costs — and how the billing actually works

The big vendors bill per resolution — roughly $0.99 at Intercom, and reportedly $1.50–2.00 at Zendesk. Remember who defines "resolution" in that deal: the party charging for it. Twenty-four hours of customer silence, billed.

We don't bill per resolution. Plans include a monthly credit allocation, and every AI action has a visible credit price. Use more than your allocation and you pay for the extra credits at the same visible rates — or move up a tier. Your bill scales with usage you can audit line by line, never with a metric the vendor controls.

At our current production automation levels, the effective cost works out to about $0.45 per resolved conversation — roughly half the category's per-resolution rates — and unlike a per-resolution meter, it gets cheaper per outcome as the AI improves, not more expensive.

Why the strict measurement is the product

The reason we can publish an auditable number is that the whole system is built around measurement:

  • Every AI draft your team edits becomes training signal — an improvement engine mines those edits every 6 hours and proposes fixes: reply rules, knowledge-base corrections, sometimes code changes.
  • Nothing ships on vibes. Every proposed change replays against a golden set of your real past conversations, launches to a 50% canary, and auto-rolls-back on statistical evidence of degradation.
  • Every number on the dashboard carries its own definition, time window, and sample size — including the unflattering ones.

If a vendor's number can't survive those questions, it isn't a number. It's a slide.

The five questions to take into any vendor evaluation

  1. What counts as "resolved" — confirmation, or silence? How long is the window?
  2. If a human teammate replies after the AI, does it still count?
  3. What's the denominator — AI-engaged conversations or all inbound? Has it changed recently?
  4. What happens when a "resolved" customer comes back?
  5. Can I see the involvement rate next to the resolution rate?

We'll answer all five on a live dashboard, on your data, in the first week of a pilot. Bring the same list to everyone else you're evaluating.

More from the PilotPM blog — comparisons, pricing breakdowns, and field notes from building an AI-native customer support stack.