Skip to main content

Your AI Costs Per Outcome. Whose Outcome?

Zendesk bills $1.50 per automated resolution and confirms it after 72 hours of silence. Three vendors bill AI by outcome, and each defines it differently.

Your AI Costs Per Outcome. Whose Outcome?

Three companies now bill your AI by the outcome. A fourth bills per conversation and calls it the same thing. Not one of them means what the others mean.

You probably didn't choose a model provider. No OpenAI account, no API key. Your AI arrived pre-installed in the tools you already ran: the CRM, the help desk, the sequencer. So when people write about single-provider risk, they're describing somebody else's problem. A platform team. An SLA. A pager going off at 3 am.

Your version is quieter, and it costs more. Here's what can actually break an AI workflow you depend on, roughly in order of how likely each one is to be the thing that gets you.

Key Takeaways

  • Zendesk moved first, publicly reported in September 2024, from $1.50 per automated resolution, counted after 72 hours of ticket inactivity.

  • Fin bills $0.99 when a customer walks away and stays away for 24 hours. Walking away isn't the same as being helped.

  • HubSpot joined in April 2026 at $0.50 per resolved conversation, which is cheaper per unit than what it replaced.

  • You can't fail over from a definition. The fix isn't a second model provider. It's a workflow written down somewhere you can read it.

What does single-vendor AI reliance actually cost an operator?

This could limit your ability to know what happened. For operators, single-vendor AI risk isn't downtime; it's inspectability. The AI arrives bundled inside the CRM or the help desk, the vendor defines the billable unit, and the workflow itself is never written down anywhere you can read. Three vendors now bill by outcome, each using its own definition of one.

That difference decides the fix. An outage is loud, rare, and somebody else's engineering problem. A workflow nobody can read is silent and permanent, and it's yours.

Written out, ranked by how likely each is to be the thing that gets you:

  1. The person who built it. One head, one login, no artifact. High risk.

  2. The workflow, never written down. It lives inside a vendor's builder. High risk.

  3. The SaaS vendor's AI feature. Repriced or redefined without your consent. High risk.

  4. The automation layer. Zapier, Make, n8n. Fails silently by default. Medium risk.

  5. The model version. Deprecated, or quietly swapped underneath you. Medium risk.

  6. The seat, API key, or OAuth grant. Leaves when the builder leaves. Medium risk.

  7. The model vendor itself. OpenAI or Anthropic goes down. Low risk.

Five and six are the ones people nod at and never check. Both are real. Neither is what ends you. The one everybody writes about is seventh.

Who decides when your AI did its job?

The vendor does, and each one decides differently. Zendesk moved first, introducing per-resolution pricing that Computer Weekly reported in September 2024: additional AI agent resolutions from $1.50 depending on plan, or $2.00 pay-as-you-go, and no charge at all when the agent escalates to a human. Zendesk's own documentation sets the trigger: an automated resolution is counted after 72 hours of inactivity, once its AI has judged the response relevant. Three competitors have since written their own definitions.

Vendor

Price

Billable unit

Confirmation window

Who defines it

Zendesk

From $1.50, plan-dependent / $2.00 PAYG

Automated resolution

72 hours inactive

Zendesk

Fin (formerly Intercom)

$0.99, on a $49/mo base including 50

Resolution or procedure handoff

24 hours of silence

Fin

HubSpot Breeze

$0.50

Resolved conversation

72 hours, no human handoff

HubSpot

Salesforce Agentforce

$2.00, or ~$0.10 per action

Conversation, or credit-metered action

None defined

Salesforce

Fin's version deserves a close read. An outcome bills when Fin resolves an issue or completes a procedure that hands off to a human, and a resolution counts either when the customer confirms it helped or when they exit without asking for more. Fin's docs are admirably blunt about the second case: "A customer leaving after receiving Fin's answer without requesting further help is counted as an Assumed Resolution and billed at $0.99," with the clock closing after 24 hours of disengagement (Fin pricing documentation, June 2026). But a customer leaving isn't a customer being helped. People abandon chats because they gave up, got distracted, or found the answer somewhere else. Gleap, a competing support vendor, puts real-world Fin resolution rates at 42 to 50% (Gleap, Intercom Fin AI Pricing Explained, 2026). Fin's own claim, in Salesforce's acquisition announcement, is roughly 76% of support volume resolved end-to-end.

Treat both numbers as interested parties talking. That's the point: you're billed against the vendor's figure, and there's no neutral referee.

Now the honest part. HubSpot's April 2026 change moved Breeze Customer Agent from $1.00 per conversation to $0.50 per resolved conversation (HubSpot, April 2026). Per unit, that's a cut. HubSpot also publishes that the agent resolves 65% of conversations and reduces resolution time by 39% across more than 8,000 customers who have switched it on. Paying for results beats paying per seat for software nobody opens.

So the pricing model isn't the problem. The problem is that each vendor writes its own definition, and the definitions don't agree with each other. Zendesk and HubSpot both landed on 72 hours. Fin closes the question after 24. Nobody's customers voted on any of those numbers.

The edges are where it gets strange. HubSpot's own documentation says a conversation also counts as resolved when a lead is marked qualified, partially qualified, or not qualified, so a lead your agent rules out is still a billable resolution. Fin runs the opposite way on second thoughts: if a customer returns to a resolved conversation asking for more help, even in a later billing period, Fin deducts that resolution and doesn't charge for it. Same event, opposite treatment, and nothing on either invoice tells you which rule you're under.

If you're paying $3,000 a month for automations you can't take with you, you don't have an AI system. You have a landlord.

And the billing change isn't even your worst dependency.

The most fragile part of your AI stack is a person

The highest-probability failure in an operator's AI stack isn't a model outage. It's an undocumented workflow held in one colleague's head, running under one login, with no readable artifact anywhere. That's layers one and two, and between them they break more workflows than every provider incident combined.

You've seen this. Someone capable built the thing that now runs quietly in the background. It works. Nobody else has opened it since. Ask why step four exists and the honest answer is "I'd have to look." No document means no handover, nothing to review, nothing to hand an auditor.

That's uncomfortable if you're the person who got told to make AI work here. Nobody gave you a playbook. You built something that worked, and the reward was becoming the single point of failure for it. Worth noting: that's a fixable problem, and unlike the pricing, it's fixable by you alone. We've written before about why the shape of the work changed faster than the org chart did.

You just accepted that a vendor can redefine your bill without asking. The bigger exposure is internal.

What happens when your AI vendor gets acquired?

On 15 June 2026, Salesforce signed a definitive agreement to acquire Fin, formerly Intercom, for approximately $3.6 billion, with closing expected in the fourth quarter of Salesforce's fiscal 2027 and Fin's technology folding into Agentforce (Salesforce, June 2026).

Think about who that lands on. Teams that picked Intercom precisely because it wasn't Salesforce now have Salesforce, on a timeline they don't control.

I won't speculate about what happens to pricing. The deal hasn't closed. The narrower point is hard to argue with: the vendor you evaluated can become a different vendor, and no amount of diligence at signing prevents that.

Your automation can fail without telling you.

Zapier doesn't proactively tell you when a Zap breaks. Detecting failures is something you configure yourself (LowCode Agency, 2026). n8n behaves the same way: no built-in notification, error handling opt-in, and a workflow that halts at the first failing node, leaving partly processed data and no recovery path (Speedrun Ventures, Complete Guide to n8n Workflow Monitoring and Error Handling, January 2026).

This is the operator's outage, and it's worse than a pager. A pager wakes someone up. A silent stop just accumulates until somebody notices three weeks later that the enrichment stopped, with no log of what got missed in between.

It shows up in the aggregate too. Gartner forecast in 2024 that 30% of generative AI projects would be abandoned after proof of concept; by its 2026 update,e the actual figure was at least 50% (Gartner, GenAI project failure). Mostly not because models failed. Because nobody could see what was happening. That's the same argument we made about why eval datasets matter more than model choice.

You already have multi-provider AI. That's the problem.

Your vendors route across models behind the scenes. Fin runs Apex, a support model it built in-house and claims outperforms frontier models from OpenAI and Anthropic on resolution rate (Salesforce, June 2026). You didn't pick that either. So you already have provider redundancy, with no visibility into it and no say over it.

Which is why the standard advice doesn't fit you. The usual prescription is an AI gateway: middleware between your code and the vendors, circuit breakers, automatic failover from one provider to another when one degrades. Real engineering, and it works, for products with an SLA and a team to run it.

In January, on this blog, I told you to build one. For most people reading this, that was the wrong advice, and I'd rather say so here than quietly stop mentioning it.

A gateway makes the model swappable. It does nothing to make the workflow swappable, because your workflow isn't in code. It's in a vendor's visual builder, a Slack thread, and somebody's head. Swap the model underneath something unreadable, and all you've changed is which black box you can't inspect. The trap has a name now: orchestration lock-in, where workflow formats and agent architectures can't be exported or rebuilt anywhere else (Zylos Research, AI Agent Ecosystem Fragmentation, April 2026).

What I don't have a clean answer for yet: how you review something an agent did when the agent also writes the record of what it did. That problem shows up in billing and in approvals, and I suspect they're the same problem. I'll write about it when I understand it better.

What does a portable AI workflow actually look like?

A governed skill is a markdown file. That's the whole trick, and it sounds unimpressive right up until you need it. A portable workflow is one written as a readable file rather than assembled inside a vendor's interface, which means repricing becomes a configuration change instead of a rebuild from no leverage.

You can read it. You can diff it in Git and see what changed last Tuesday. You can hand it to a new hire, to a different model, or to an auditor who wants to know how a decision got made.

Compare the artifacts. Lead routing built as a Zapier zap exists as boxes on a canvas inside an account. Repricing means renegotiating from no leverage. The builder leaving means the logic leaves too. The same routing as a markdown skill is a file in your repository: repricing becomes a question of which model you call, and the builder leaving becomes a pull request. If you want the longer version of that argument, we compared the CLI, MCP, and custom-skill approaches for HubSpot in detail.

We publish ours, including the HubSpot MCP server and the rest of the open-source skills. The whole argument is that you should be able to read the thing running your business, so making it from behind a login would be strange.

Where should the human approval gate go?

Before the irreversible step, and staffed by a person who can see what's actually being approved. In January 2026, attackers compromised executive devices at Step Finance, a Solana portfolio manager. What turned that breach into a fatal one was that its AI trading agents held permission to move large transfers with no human approval: 261,000 SOL left, worth $27 to $30 million; $4.7 million came back, the token fell 97%, and the company shut down (Beam, 5 Real AI Agent Security Breaches in 2026, May 2026; corroborated by NeuralTrust). Beam's own headline puts the loss at $40 million, so treat the exact figure as contested and the outcome as not.

Read that carefully, because the lesson is precise. The agents weren't jailbroken or tricked. They did exactly what they were built to do, once someone else held the controls. Excessive permission plus no gate is the failure, and the breach only decided the timing.

There's a sharper way to put the reviewer's problem, from a practitioner discussion on r/AI_Agents in June 2026: the reviewer isn't approving an operation; they're approving the story the agent told about the operation. Treat that as a practitioner's framing rather than research, because that's what it is. It also connects straight back to the invoice. If "resolved" is a story an agent tells about its own work, the review problem and the billing problem are the same.

Practically, decide in advance what happens when the cheap path isn't good enough. For anything consequential, the system should stop and ask a person, not guess with a weaker model. How we implement that gate is documented, because a gate nobody can inspect isn't a gate.

What the EU AI Act actually requires in August 2026

Article 50 transparency obligations, Article 4 AI literacy duties, and the rules for general-purpose AI. That's it. High-risk obligations for stand-alone Annex III systems were deferred to 2 December 2027 under the Digital Omnibus, given final approval by the Council of the EU on 29 June 2026 (Gibson Dunn, EU AI Act Omnibus Agreement, 2026).

One wrinkle worth knowing: the Article 50(2) watermarking requirement shifts to 2 December 2026 for systems already on the market, and none of the changes take legal effect until the Omnibus is published in the Official Journal, expected before 2 August (Covington, Inside Privacy, 2026). Plenty of posts published this year still call 2 August the high-risk deadline. It isn't.

Article 4 is the one that ties back to layer two. It obliges you to ensure AI literacy among the people deploying these systems, which is hard to evidence when the workflow only exists in somebody's head. Ask what you'd actually hand an assessor, and you're asking this post's question again. Check your own scope rather than trusting a blog, including this one. Our read on EU data and the Act is public.

How do I check my own exposure?

Five checks establish operator AI exposure: whether anyone besides the builder can explain the workflow, whether the logic can be exported in readable form, whether you know how your vendor defines its billable unit, how long silent failures go unnoticed, and who approves before an automation writes to your CRM. One per layer. Answer them out loud, which is harder.

  1. If the person who built your most important automation left tomorrow, who could explain step four?

  2. Can you get that workflow out of the vendor's interface as something a human being can read?

  3. Do you know how your vendor defines the unit it bills your AI in, and what resets that clock?

  4. Last time something failed silently, how long before anyone noticed?

  5. Who approves before an automation writes to your CRM?

Two or more "I don't know" answers mean the problem isn't which model you're on.

Frequently Asked Questions

What counts as a "resolved" conversation in AI pricing?

It depends entirely on the vendor. Zendesk counts an automated resolution after 72 hours of ticket inactivity. HubSpot requires the agent to act with no human handoff inside 72 hours, and also counts a lead marked not qualified. Fin closes the question after 24 hours of silence. One word, three definitions, each written by the party sending the invoice.

Should I use OpenAI or Claude for my workflows?

For most operators, this is the wrong question. You usually aren't choosing the model. Your SaaS vendor is, and it changes that choice without telling you. The better question: can you read, export, and rebuild the workflow if that vendor's pricing, terms, or ownership change?

Do I need an engineer to make a workflow portable?

Not for the artifact. It's a markdown file describing what happens, in order, in plain language. You'll want engineering help for API authentication and for building the approval gate properly. Writing down what the workflow actually does is the part that doesn't require code.

Is a markdown file really enough to run a business process?

On its own, no. It's the readable description, not the runtime. You still need something to execute it, credentials that don't belong to one person, and a human approval step before anything irreversible. What it does is stop the logic being trapped where you can't inspect or move it.

Newsletter

Get the next field note

One email per week. No content calendar — just what we’re building, what broke, and what we changed our minds about.

Community

Build the fluency once. Keep it.

This is the thinking we teach live in the Real-Work Cohort, and continue, between cohorts, in Operator Stack.