Field notes.
Real opinions about what AI is getting right and wrong for operators in 2026, written by us, in our own voice, not a content calendar. Plus the occasional teardown of something Izzy built and what it taught us.
Articles

Prompt Caching vs. Fine-Tuning: Stop Wasting AI Budget
Is fine-tuning inflating your LLM bill? Discover why Prompt Caching is the superior architecture for context injection and how to save 90% on input tokens.

Your AI Costs Per Outcome. Whose Outcome?
Zendesk bills $1.50 per automated resolution and confirms it after 72 hours of silence. Three vendors bill AI by outcome, and each defines it differently.

Why Only 5% of AI Projects Reach Production (And the "Evaluation Gap" Behind It)
Industry data shows only 5% of AI projects reach full production. Discover the 5 hidden evaluation gaps from RAG black boxes to compliance risks that stall the rest.

LLM Observability Costs 2026: Pricing, Categories & The APM Tax
Is your APM bill hiding a €50k/month "Observability Tax"? We break down the 4 tool categories, 2026 pricing models, and how to choose the right hybrid stack.

Your AI Bill Is a Workflow Problem, Not a Hardware Problem
Torn between dedicated and serverless GPU? Our CTO guide offers a data-driven breakdown, TCO calculations, and a strategy for optimizing your AI infrastructure.

The 5 Most Common Problems with Agentic AI in Production - And How to Solve Them
Gartner predicts 40% of AI agents will fail. Discover the 5 top production pitfalls from hidden cost spirals to compliance risks and the architectural fixes you need.

The "Redundancy Tax": How Prompt Caching & The Rule of 3 Fix AI Margins
Stop paying full price to re-process static data. Discover how Prompt Caching reduces LLM costs by 90%—but only if you follow the "Rule of 3" break-even math.

The Architecture of Autonomy: Why Human-in-the-Loop Is Permanent Infrastructure
HITL isn't temporary it's essential for Level 3 Autonomy. Learn architectural patterns like Interruption Gateways and Risk-Tiered Routing to secure Agentic AI.

Defensible AI: The CTO’s Guide to Reliable "LLM-as-a-Judge" Evaluations
Stop relying on "vibe checks." This CTO guide covers how to build reliable LLM-as-a-Judge evaluations, enforce strict rubrics, and block AI regressions in CI/CD.
Field notes, in your inbox
One email per week. No content calendar — just what we’re building, what broke, and what we changed our minds about.
Operator Stack
The free community where we continue these conversations between posts.