Article

The AI resolution rate gap:
vendors say 76%, our data says 46%.

By Marija Jovanović

Updated

8 minute read

Every vendor of AI support leads with one number. For Fin it’s 76%. For Help Scout, it’s 73%. For Lyro, it’s 67%. For Ada, it’s 84%. For Freshdesk, it’s up to 86%. For Forethought, it’s up to 98%. For us, it’s 45.9%, and you can see it on our public page, with the denominator next to it. So either Outlearn resolves thirty points fewer conversations than everyone else, or the numbers aren’t measuring the same thing. We checked, and mostly, they don’t.

This is a post about how a resolution is counted. It uses two things that nobody has put together before: a dated audit of what the dozen vendors claim, and a recount of one real production corpus (29,003 conversations) under the definitions those vendors use.

What we audited

On 5 September 2026 we read the public marketing and help pages of twelve AI support vendors and wrote down three things for each: the headline resolution number, what kind of number it is, and whether the vendor anywhere defines what a resolution is. Vendors are named as plain text; the sources are in our research notes and every figure below is quoted as published.

Published AI resolution claims, twelve vendors, read on 5 September 2026. Figures quoted as the vendor publishes them.
VendorHeadline numberKind of numberDefines “resolution”?When the customer goes silent
Intercom Fin76% average across 12,000+ customersAverageYesCounts as an “assumed resolution”; reversed if the customer comes back
Zendesk AI agentsUp to 80%“Up to”YesCounts after 72 hours of inactivity if an LLM judges the reply relevant
Freshdesk FreddyUp to 86%“Up to”NoNot stated
Ada84%HeadlineNoNot stated
ForethoughtUp to 98%“Up to”NoNot stated
Help Scout AI Answers73% averageAverageYesCounts if the customer takes no further action
Tidio Lyro67% averageAverageDefines a conversation, not a resolutionNot stated
Decagon70% and 80% in customer case studiesCase studiesNoNot stated
Kustomer70% in one customer quoteCase studyNoNot stated
Gorgias60% “resolved instantly”HeadlineNoNot stated
SierraNone publishedNoneNoneNone
Zoho Desk ZiaNone publishedNoneNoneNone
Outlearn45.9% of 29,003 conversationsNumerator and denominatorYesDoes not count; published separately as 48.4% “neither”
“Not stated” means the page carrying the number does not say what happens to a customer who stops replying. “None” means the vendor publishes no resolution rate at all.

Four things fall out of that table.

Ten of the twelve publish a number. Three define it. Intercom, Zendesk and Help Scout explain what a resolution is. The other seven publish a percentage of something they don’t describe. Tidio defines its billing unit, a conversation with at least one AI reply, but not the resolution rate it advertises.

Nobody publishes a denominator. Not one of the ten tells you how many conversations the percentage was calculated over. Intercom comes closest with a numerator, two million weekly resolutions, and a customer count. A rate with no denominator cannot be checked, and it cannot be compared to another vendor’s rate either.

Three of the ten lead with “up to”. Zendesk, Freshdesk and Forethought. “Up to 98%” is a statement about the best deployment the vendor has seen, not about what a typical customer gets, and it is doing the work of an average in every comparison table where it appears.

The headline figures run from 60% to 98%, median 74.5%. Our 45.9% is not in that range. The next section is about why.

Published headline AI resolution rates, ten vendors, against Outlearn’s strict number and its ceiling under the industry’s silence-counts rule. Read on 5 September 2026.

0%25%50%75%100%Forethought98%Freshdesk Freddy86%Ada84%Zendesk AI agents80%Intercom Fin76%Help Scout AI Answers73%Decagon70%Kustomer70%Tidio Lyro67%Gorgias60%Outlearn46%94%
  • Average, headline or case-study figure
  • “Up to” claim (a maximum, not an average)
  • Outlearn, strict count
  • Outlearn if silence counted as resolved (dashed)
Rounded to whole numbers. The figures are the vendors’ own, as published; the exact numbers and sources are in the table above.

The word that moves the number: silence

Every published definition has to decide what to do when the AI answers and the customer says nothing more. This is not an edge case. In our corpus, it is the largest single outcome.

Here is what the three vendors who define the term do with it.

  • Intercom counts a resolution when, after Fin’s last answer, the customer “either confirms the answer was satisfactory (confirmed resolution), or exits the conversation without requesting further assistance (assumed resolution)”. Silence is a resolution. If the customer comes back later, the resolution is reversed.
  • Zendesk counts an automated resolution “after 72 hours of inactivity if AI evaluation has confirmed that the AI agent’s response was relevant”. Positive feedback counts, and so does sharing a help-centre article, without the customer clicking it if the article came in a generative reply. Silence plus a relevance check is a resolution.
  • Help Scout counts a resolution “only if the customer receives an AI response and doesn’t use escalation, search the knowledge base, ask more questions, or indicate that they need more help”. Silence, with no other action, is a resolution.

The three rules are not identical. Zendesk’s relevance check is a real attempt to tell “helped” from “gave up”, and Help Scout’s rule, which excludes anyone who went on to search the knowledge base, is slightly stricter than Intercom’s. But all three land in the same place: a person who leaves without complaining is counted as resolved.

Our benchmark does the opposite. It counts a conversation as resolved automatically only if the agent closed it without human involvement. A conversation that was abandoned by the visitor, or one where the visitor read the answer and left the conversation, is neither considered resolved nor handed off, and we publish that line rather than folding it into either side.

One corpus, three answers

However, as we are reporting all three outcomes, we can recount the same conversations under the rules of the industry. The corpus is every conversation handled in Outlearn workspaces from October 2025 to August 2026.

Every conversation handled in Outlearn workspaces, October 2025 to August 2026, by outcome. The same three lines are published live on the AI support benchmark.
OutcomeConversationsShare
Resolved automatically (agent closed it, no human)13,29945.9%
Handed to a human1,6625.7%
Neither (visitor left, or stopped after a quick answer)14,04248.4%
Total29,003100%

The same 29,003 conversations as one bar. The bracket is what the corpus reads if every silent exit is counted as a resolution.

46%48%94% if silence counted as resolved
  • Resolved automatically 46%
  • Handed off 6%
  • Neither 48%
Rounded to whole numbers; the counts are in the table above. Resolved and handed off do not add up to the total. The remainder ended without either outcome, and we publish it rather than fold it into either side.

Now apply the definitions.

The same 29,003 conversations, recounted under four definitions of a resolution.
DefinitionWhat countsThe same corpus reads
Strict (ours)Agent closed the conversation, no human45.9%
Confirmed onlyCustomer explicitly said it helpedBelow 45.9%; we do not publish this split
Silence counts (Intercom, Help Scout style)Closed, or customer left without asking moreUp to 94.3% (27,341 of 29,003)
Silence plus relevance check (Zendesk style)As above, filtered by an LLM’s judgement of the last replyBetween the two; depends on the model doing the judging

The 94.3% is a ceiling, not a claim. It assumes that every one of the 14,042 silent conversations was a satisfied customer (which is certainly not true), but it is a number that a vendor using Intercom’s rule would be entitled to print about our conversations, and it is 18 points above the highest average anyone in the table publishes.

What this does not prove

It doesn’t prove Outlearn resolves as many conversations as Fin or AI Answers. Our number could be lower than theirs because our agent is worse, and the public data cannot tell the difference. Also, two other things weaken a direct comparison: 31% of our corpus is from a single month, March 2026, and our conversations come by our website widget, email, Slack, Teams and Google Chat, where chat-first vendors measure chat.

What it does prove, however, is more limited but more useful. The gap between any two published resolution rates cannot be read as a product difference until both vendors have answered the same five questions.

The five questions to ask any vendor

  1. What is the denominator? How many conversations, over what period, on which channels.
  2. What happens to a customer who goes silent? Resolved, unresolved, or excluded.
  3. Is there a verification step, and what is it? A customer confirmation, an LLM judgement, a survey, or nothing.
  4. Is a resolution ever reversed? If the customer returns with the same problem a day later, does the count go down.
  5. Is the resolution rate also the billing unit? Help Scout charges $0.75 per resolution and Intercom prices Fin per resolution too. When the definition of a resolution is the invoice, the silence rule is a price.

If a vendor cannot answer the first two, the number on their homepage is decoration.

How we count, and what we would change

We publish the numerator, the denominator, the time window, the handoff rate and the “neither” share, and we update it every week. That is the benchmark we ask others to meet, so we hold ourselves to it in public at the AI support benchmark.

We would like to add the confirmed-only split, the strictest reading, to the table above. For now, we stand by 45.9% as the number, and 94.3% is the number we could print and won’t.

For the metrics definitions used in this post, the glossary has resolution rate, deflection rate, containment rate and first contact resolution with their formula and the way they get inflated.

FAQ

What is a good AI resolution rate?

Not yet knowable, because everyone counts differently.

Vendor-reported averages run from 73% to 76%, and all of them count a silent customer as resolved. Under a strict rule that counts only conversations the AI closed, one production dataset of 29,003 conversations shows 45.9%. Ask for the definition before you ask for the number.

What is an “assumed resolution”?

A resolution that is counted because the customer left without asking for more help, rather than because the customer explicitly confirmed it worked.

Intercom uses the term. Zendesk and Help Scout apply a similar rule but with additional conditions. It is the single largest reason for differences between vendors’ reported rates.

Why does Outlearn publish a lower resolution rate than other vendors?

Because we count only conversations the AI agent closed with no human involved, and we publish the 48.4% of conversations that ended without a resolution or a handoff as a separate line instead of counting them as resolved. If we recount under the rule most vendors use, the same conversations would read up to 94.3%.

See what the agent resolves on your own docs.

Connect a knowledge base, ask it the questions your customers ask, and read the answers it gives. Free plan, no card.

Keep reading

AI Resolution Rate: Claims vs Production Data | Outlearn