---
title: "Outcome-Based AI Pricing: Who Counts the Outcomes?"
description: "Intercom bills an AI outcome when the customer goes quiet after an answer. Zendesk refuses to. Same month, a 2.9x invoice spread. What to demand before you sign"
image: "https://storage.googleapis.com/promptmetrics-uploads/website/posts/1786540603312-190903699.jpg"
author: "Izzy"
category: "AI Cost & ROI"
publishedAt: "2026-05-14T00:00:00.000Z"
updatedAt: "2026-08-19T09:30:30.133Z"
canonical: "https://www.promptmetrics.dev/blog/why-ai-cant-be-priced-like-saas-and-what-comes-next"
---

# Outcome-Based AI Pricing: Who Counts the Outcomes?

Your AI vendor is about to offer you outcome-based pricing, and it will sound like they're taking on your risk. They're not. They're handing you an accounting job you don't have the records to do.

I spent a year building governed agents into other people's CRMs before PromptMetrics existed. The pattern I keep hitting isn't that outcome pricing is a scam. It's that the word in the middle of the contract, the one the whole invoice hangs on, means different things to different vendors, and nobody makes the buyer read it.

Two vendors in this market bill the exact same event in opposite directions. I'll show you both, from their own documentation. Then I'll show you what happens to one month of tickets when you bill it three ways.

> **The short version**
> 
> *   Intercom bills a Fin outcome when a customer goes quiet after an answer
>     ([Fin docs](https://fin.ai/help/en/articles/13975800-fin-pricing-outcomes), 2026).
>     Zendesk calls that same silence a contained resolution and refuses to charge
>     ([Zendesk](https://support.zendesk.com/hc/en-us/articles/9570369117338-About-automated-resolution-tiers), 2026).
>     
> *   Salesforce charges $2 per conversation and never defines "conversation."
>     
> *   One 400-ticket month, three plausible definitions: 318, 246 or 108 billable.
>     A 2.9x spread on identical work.
>     

## Why won't your next AI contract have seats in it?

Because your vendor's margins broke. In a survey fielded in mid-2025, seven in ten software companies building AI features said delivery costs are eroding their margins. Another 52% said they were planning new pricing models to mitigate cloud spend (Revenera, _Monetization Monitor: Software Monetization Models & Strategies, 2026 Outlook_, October 2025, n=501 product leaders). They aren't moving off seats to be generous.

The direction of travel is real. Gartner expects at least 40% of enterprise SaaS spend to shift toward usage, agent, or outcome-based pricing by 2030 (Gartner, cited in Deloitte, _TMT Predictions 2026: SaaS and AI agents_, November 2025). And the logic is genuinely good. Seat licences priced AI backwards. If an agent lets four people do the work of ten, a per-seat contract punishes you for the savings. Paying for results instead of headcount is the honest fix.

So I'm not here to tell you outcome pricing is wrong. I'm telling you it moves something onto your desk, and almost nobody has noticed what.

## What does outcome pricing actually move onto you?

Outcome pricing doesn't transfer risk to the vendor. It transfers accounting to you. Under a seat licence you count people, which you already do, in a system you already trust: payroll. Under outcome pricing you count events inside a workflow, and the events are generated by software the vendor controls.

Think about what verifying an invoice now requires. Someone has to define the billable event, in writing, before the first one fires. Then count them independently. Then find the ones that shouldn't have counted at all, which is the hard part, because a false positive looks exactly like a success until you check. And when you dispute a line, you need a record good enough to win.

That's four capabilities. Most teams signing these contracts have none of them. This is the same failure mode I've written about before: the problem isn't the tool you picked, [it's the shape of the work around it](/blog/it-s-not-the-tools-it-s-the-shape).
The vendor, meanwhile, has all four, because the vendor built the instrumentation. Which sets the default outcome of every dispute. The vendor's number wins, not because anyone is being dishonest, but because theirs is the
only number that exists.

## Two vendors, one word, opposite invoices

Two established vendors define a billable AI resolution in ways that bill the same customer behaviour in opposite directions, and both publish it openly.

Intercom charges $0.99 per outcome for Fin. Its documentation defines a resolution as an outcome where "the customer confirms Fin resolved the issue or does not request more help after Fin answers" ([Fin pricing docs](https://fin.ai/help/en/articles/13975800-fin-pricing-outcomes),
retrieved August 2026). Read the second half again. Once Fin has answered, a customer who never comes back is billed as a success.

Credit where it's due: Intercom publishes an exclusion list, and it's a real one. Escalations, procedure failures, frustration detection, explicit requests for a human, and conversations abandoned after Fin asks a clarifying question are all unbilled. The gap is narrower than the headline definition suggests. It is still a gap, and it sits exactly where silence lives.

### Zendesk names the same event and refuses to bill it

Zendesk goes the other way. Its automated resolution tiers charge only for requests "successfully resolved by an AI agent, without any escalation to a human agent," verified through a 72-hour window and an LLM evaluation of whether the answer actually landed. The silence case gets its own name, a contained resolution, and Zendesk does not charge for it.

Salesforce publishes $2 per conversation for Agentforce and, as of August 2026, no definition of a conversation on either its pricing page or its help documentation. Not a message, not a session, not a case thread. The page carries a footnote saying it "is subject to change."

### The vendor that changed its own definition

One more detail, and it's the one I'd put in front of a CFO.

In March 2026 Intercom renamed its billing unit. "Resolution" became "outcome," and the stated reason was that the old metric "only recognized value when the conversation ended in a full AI resolution." Read that as a buyer. A vendor changed the definition of the thing you were paying for, mid-contract, because the definition had stopped suiting them.

That's not a scandal. It's evidence the unit isn't stable enough to price against without your own records.

Three vendors. Three units. You cannot build a per-unit cost model across them.

## Can you prove the resolution happened?

Not from the vendor's dashboard alone. Here's why, in one month of tickets.

Take four hundred of them. One coding agent front-lining the queue, humans behind it, one line in the contract reading €1.00 per resolution. That rate is a round number chosen for clean arithmetic, not a conversion of anyone's list price.

Now bill that month three times.

D1 Touched D2 Deflected D3 Durable 318 · €318 246 · €246 108 · €108

Illustrative dataset, 400 tickets, €1.00 per resolution. PromptMetrics, 2026.

| 
Definition

 | 

What counts as a resolution

 | 

Billable

 | 

Share of 400

 | 

Invoice at €1.00

 |
| --- | --- | --- | --- | --- |
| 

**D1 Touched**

 | 

The agent replied at least once and the ticket later closed.

 | 

318

 | 

79.5%

 | 

€318.00

 |
| 

**D2 Deflected**

 | 

D1, and no human replied after the agent.

 | 

246

 | 

61.5%

 | 

€246.00

 |
| 

**D3 Durable**

 | 

D2, and not auto-closed, no reopen within 7 days, no new ticket from the same contact within 7 days.

 | 

108

 | 

27.0%

 | 

€108.00

 |

Same month. Same tickets. Same agent. Three invoices, and the widest is 2.9 times the narrowest.

None of the three is a trick. I could write any of them into a contract and defend it with a straight face:

*   **D1 Touched.** The agent replied at least once and the ticket later closed.
    
*   **D2 Deflected.** D1, and no human replied after the agent.
    
*   **D3 Durable.** D2, and it wasn't auto-closed, wasn't reopened within seven
    days, and the same contact didn't open a new ticket within seven days.
    

D1 is close to what most vendor dashboards report, because "the agent touched it and it closed" is the cheapest thing to measure. D2 is what most buyers think they're buying. D3 is the only one where the customer's problem actually went away.

### What each gate removes

Here's what the gates strip out of the 318 tickets the agent touched and closed.

Touched + closed Human replied after = Deflected Auto-closed on timeout Reopened within 7 days New ticket, same contact = Durable 318 −72 246 −70 −41 −27 108

Illustrative dataset. PromptMetrics, 2026.

Look at 72, 70, and 68 (41 reopened plus 27 who opened a fresh ticket). Three roughly equal causes. That's the part that should worry you, because you can't patch this with one clause. Once the human-reply gate is in place, closing the auto-close loophole recovers another 70 of the 210 tickets separating D1 from D3.

**Yes, the euro numbers here are small.** €318 against €108 on a 400-ticket month is a rounding error, and I picked a small queue on purpose so the numbers stay holdable. The ratio is what scales. At 10,000 tickets a month the same three definitions bill €7,950, €6,150 and €2,700. That's €5,250 a month of pure definitional drift, €63,000 a year, on one workflow, decided by a word nobody argued about at signing.

Run your own volume through it. The multiplier doesn't change.

### Which definition is right?

D3, the durable one. Pay for durable resolutions, or don't pay per resolution.

A ticket that reopens on Thursday wasn't resolved on Tuesday. A ticket that timed out was never resolved by anyone. And if your own team wrote the reply that closed it, that's payroll, not vendor value. Every gate in D3 strips out something you wouldn't have paid a human for either. That's the whole test.

Note that Zendesk already sells something close to D3, with a 72-hour window instead of my seven days. So the strict definition isn't hypothetical. One vendor ships it.

## The baseline is the whole negotiation

Take the baseline before anything gets built, because it's the only number in the argument that belongs to you. You cannot demonstrate value against a figure you never took, and you cannot dispute a vendor's rate without a denominator of your own.

Time the workflow by observation, not by survey. How long does the dreaded task take today, measured with a timer, over one or two weeks, on data you own. Two reasons this matters more under outcome pricing than it ever did under seats. It's your only independent denominator when the vendor reports a rate, so without it you are quoting their arithmetic back at them. And it's the only version of the number that survives a procurement challenge, because you took it yourself and you can show your method, which is a thing no dashboard export can do for you.

This is the discipline every credible practitioner agrees with and most pilots skip. Not a clever technique. Just the one nobody does.

## What has to be in the audit record?

Six fields, all held on your side rather than in the vendor's dashboard. A durable resolution (D3) isn't harder to write into a contract than a touched one (D1). It's harder to _prove_, and that's the real reason vendors don't lead with it.

To compute a durable resolution you need all six on every gated action:

1.  What the agent proposed to do, and the payload it sent.
    
2.  Who replied last, and whether they were a person or the agent.
    
3.  Why it closed. Resolved, or aged out.
    
4.  An idempotency key. This is the one that sounds technical and isn't: it's just a unique stamp on each action, so that if the agent retries after a timeout you can prove the same reply went out once rather than twice. Without it, a flaky connection is indistinguishable from two billable events.
    
5.  Reopen events, with timestamps.
    
6.  Contact-level linkage, so a follow-up ticket points back at the original.
    

The first four are the approval record we design before a skill goes live. We [put the gate inside the tool](/blog/the-gate-belongs-inside-the-tool-giving-an-ai-agent-write-access-to-a-production-crm) rather than bolting an approval button onto an opaque workflow, which is what makes the record replayable. Fields five and six are already sitting in your CRM and almost nobody queries them.

If you can't produce those six, you aren't auditing an invoice. You're approving one.

## Does outcome pricing change your AI Act position?

Yes, and the deadline already passed. Article 50 transparency duties have been live since 2 August 2026, per Cooley's August briefing. The same date gave the EU AI Office power to enforce against general-purpose AI providers, which Wilson Sonsini covers. Break those duties and the fine runs to €15 million or 3% of
global turnover, whichever is higher. That's Article 99(4)(g). It names Article 50 by name.

### What the marking deadline means for you

There's a timing detail most people miss. Was your AI system live before 2 August 2026? Then you have until 2 December to add machine-readable marking. Wilson Sonsini traces that window to the AI Omnibus, the 8 July 2026 amendment. Cooley confirms the date but doesn't name the instrument. Deploy anything new from August onward and you get no window at all, which means the team who shipped in June has an easier quarter than you do. If you're standing up an outcome-priced agent this quarter, your deadline is harder than the team who shipped in June.

Here's the connection to everything above. Enforcement is expected to open with technical compliance dialogues rather than fines (Wilson Sonsini, _EU AI Act Enforcement Phase Begins_, August 2026). A dialogue means producing documentation: what the system does, which actions it takes, who reviewed them. That's the same evidence you'd need to audit a per-resolution invoice.

Which means the record isn't overhead you're buying for compliance. It's one artifact doing two jobs, and your vendor's billing dashboard does neither. Anyone [evaluating EU AI implementation partners](/blog/ai-implementation-agencies-in-the-eu-how-to-evaluate-them-2026) should ask who owns that record before asking about price. Our own [data processing and residency documentation](/security#dpa) is where we put ours,
because a DPO asks for it before anything reads production data.

## Which six questions belong in the contract?

Put these in the RFP, in writing, before the pricing conversation. Copy them.

Before discussing price, require the vendor to write the billable unit into the contract with every exclusion listed, state whether customer silence bills, state whether a human reply voids the charge, name the reopen window, commit to definition-change notice, and export the underlying event log. Six asks. All six belong in writing.

| 
#

 | 

Ask the vendor

 | 

What a bad answer looks like

 |
| --- | --- | --- |
| 

1

 | 

Write your definition of a billable unit into the contract, including every excluded case.

 | 

A link to a help-centre page the vendor can edit without telling you.

 |
| 

2

 | 

If the customer goes silent, is that billable? State it explicitly.

 | 

"It depends on the conversation flow."

 |
| 

3

 | 

If a human replies after the agent, is it still billable? What if the human resolved it?

 | 

Anything other than no.

 |
| 

4

 | 

What reopen window voids a billed resolution, and who watches it?

 | 

No window, or a window the vendor measures alone.

 |
| 

5

 | 

What notice do we get if you change the definition, and does a change reopen pricing?

 | 

Silence. Intercom redefined its unit in March 2026.

 |
| 

6

 | 

Export the underlying event log to us, on our schedule, in a format we can query.

 | 

A dashboard, a PDF, or read-only access.

 |

Question six is the one that decides the other five. Every definition above is unenforceable without the log.

## Where this argument still leaks

The weakest part of this piece is the dataset, and I'd rather say so than have you find it. I chose the gate proportions myself. A vendor could reasonably argue I tuned them to make the gap look bad. The mechanism holds at any proportions, but the specific 2.9x is mine, not a measurement of your queue or anyone else's.

Two more honest problems.

Strict definitions have a cost. Push every vendor to durable-only and some will refuse outcome pricing and go back to seats. If you preferred seats you wouldn't have read this far, but reduced choice is a real consequence and it lands on you, not on them. And seven days is arbitrary. I picked it because it's a week. Zendesk picked 72 hours. Fourteen days would strip more. Whatever number you choose, choose it before you sign, not after the first invoice you want to dispute.

One more thing. The replayable approval record I described is designed and specified. We haven't shipped it as a standalone product, and we have no closed-won data behind any of this yet.

## What to do before the next vendor call

Do one thing this week: open your current AI vendor's pricing documentation and find the sentence that defines the billable unit. If you can't find one, that's your answer.

Then, in order:

1.  Time your target workflow for two weeks. Observation, not survey. You own the
    
    data.
    
2.  Pull the six fields above out of your CRM for last month and compute your own D
    

1.  Most of them are already there.
    
2.  Send the six questions before you discuss price, not after.
    

## Frequently Asked Questions

### Isn't outcome pricing still better than per-seat for AI?

Usually yes. Seat licences punish you for the headcount savings the agent creates, which is why 52% of software producers said they were planning new pricing models to mitigate cloud spend (Revenera, 2026 Outlook). The argument here isn't to refuse outcome pricing.
It's to refuse an outcome you can't count, and to write the definition down before you sign rather than after the first disputed invoice.

### What if the vendor won't give us the event log?

Treat it as a pricing signal, not a technical limitation. A vendor confident in its definition has no reason to withhold the data that proves it. If the log is off the table, negotiate a flat fee or a capped commitment instead, so your exposure doesn't depend on a number only one party can see.

### We already signed. What now?

Compute your own D3 from last month's CRM data and compare it against what you were billed. Most of the six fields are already sitting there. If the gap is material you now have a renegotiation conversation grounded in numbers you took yourself, which is a different conversation from one grounded in theirs. Annual review is when definitions get rewritten.

### Does this apply outside customer support?

The mechanism does. Anywhere a vendor bills per completed thing, the definition of "completed" carries the price: qualified leads, matched invoices, reviewed documents, recovered claims. Support is just where the definitions are published, so it's where you can prove the problem instead of asserting it.

## The word is the price

Outcome pricing is the right idea. The contracts aren't ready.

The industry shipped the pricing model and left out the definition that makes it work. Two vendors bill the same silence in opposite directions. A third won't say what its unit is. There's no standard here yet.

The fix is small and boring. Vendors write the unit into the contract, list what they don't charge for, and hand over the log. Buyers time the workflow before they buy it, and keep their own record of what the agent did.

Nobody has to build anything new. This is a paperwork problem. It just doesn't look like one when it arrives as a price.

What does your outcome-priced invoice actually look like? Which definition is buried in it, and was it what you expected?
