Someone handed you an AI mandate. Maybe it came from a board deck, maybe from a CEO who said "we need to be doing AI" and looked at you. Either way, you're the one now expected to turn benchmark hype into something that runs in production, on a deadline, without a technical co-pilot.
Here's the number that should be on your desk before you commit to anything: the best AI agent in the world completes 15.8% of real freelance projects to a standard a paying client would accept. That's up from 2.5% when the benchmark launched less than a year ago (Center for AI Safety, July 2026). Six times better than last year. Still failing 84 out of every 100 real jobs.
The bottom line: The Remote Labor Index tested frontier AI agents (including Claude Opus 4.8 and Fable 5) on 240 real freelance projects. The best score, as of July 2026, is 15.8% success, up from 2.5% at launch (CAIS). The gap between that number and 90%+ academic benchmark scores isn't a rounding error. It maps almost exactly onto four skills we build into every engagement: Delegation, Description, Discernment, and Diligence. Skip any one of them and the mandate you were handed turns into the zombie integration nobody wants to own.
Delegation: what should actually get handed to AI?
The Remote Labor Index, built by Scale Labs and the Center for AI Safety, tests AI agents against 240 real freelance projects worth over $140,000 in professional time, evaluated by human judges against a professional's actual deliverable (CAIS, July 2026). The best score as of July 2026 is 15.8% (Table 5), up from 2.5% at the October 2025 launch, a real trajectory that still leaves 84 of every 100 real jobs failing to meet a paying client's bar. Frontier models score 90%+ on academic benchmarks like MMLU and HumanEval in the same period. Not "can it answer a question." Can it do the job?
That distinction is Delegation, the first of the four thinking skills we build into every engagement: knowing which work to hand to AI and which to keep in human hands. An LLM has plenty of intelligence. It's agency, the ability to sustain judgment across dozens of steps without supervision, that RLI shows is still thin on the ground.
The RLI's 240 projects span 3D and CAD design, architecture, video and audio production, data analysis, and web development, drawn from real freelance briefs. Real client, real brief, real deliverable, evaluated by a human who's done that work themselves, not an automated grader.

What this means for your mandate: the trajectory is real, and it's fast. What it doesn't mean is that you can point at a benchmark screenshot and tell your team AI is ready to run a workflow unsupervised.
The 74-point gap between a 90%+ benchmark score and a 15.8% real-job completion rate isn't a temporary bug that better prompting fixes. It's the reason every skill we build starts with mapping your actual workflows first, so we automate the step that has high AI-leverage, not the step that looks impressive in a demo.
Description: How do you keep an AI agent on-brief?
In the RLI's original evaluation round, researchers clustered evaluators' written justifications for failure and found 35.7% of submitted projects were incomplete or malformed (missing components, truncated files, absent source assets), a further 14.8% contained internal inconsistencies across files, and 45.6% had general quality issues (Mazeika et al., "Remote Labor Index," arXiv:2510.26787, October 2025). The common thread across all three failure modes: the agent lost track of what the client actually asked for somewhere between the brief and the finished file.
Say you've correctly delegated a task to an AI agent. The next failure mode shows up mid-project, not at the start. As an agent works through a multi-step deliverable, the context window fills with intermediate steps, tool outputs, and error logs. The original brief, the one describing what the client actually asked for, gets diluted with every step, and the agent starts reacting to the noise instead of the goal.
This is Description: writing and rewriting, instructions clear and persistent enough that an AI system keeps producing what you actually asked for, not what the last error message nudged it toward. A coding agent calls a library it never imported. An architecture agent's floor plan contradicts the 3D render it generated two steps earlier. Neither failure is a reasoning problem. Both are a description problem: the brief stopped being the thing the agent was optimizing for.
Larger context windows don't fix this on their own. A million-token window still isn't a persistent project model; it's a longer scratchpad that fills with noise the agent starts reacting to instead of the original brief. We handle this by separating "global state" (the brief, the constraints, the success criteria) from "local state" (the current step, the error log), and re-injecting the global state at every inference step, not just the first one. That's not a nice-to-have. It's the difference between a skill that survives contact with a real client's edge cases and one that quietly drifts off-brief by hour three.
Discernment: can an AI judge grade AI work reliably?
When CAIS tested whether an automated LLM judge could replace human evaluators on the Remote Labor Index, the judge overestimated GPT-5.5's real-world quality by roughly 2.9x (17.9% vs. a human-verified 6.25%), and Opus 4.8's by roughly 2.3x (18.8% vs. 8.33%), while still ranking models correctly relative to each other (CAIS, July 2026). Automated grading tracks relative progress, not absolute reliability, and that gap widened specifically on the newest, most capable models, exactly where the stakes of trusting a bad grade are highest.
Model | Human evaluation (ground truth) | Automated LLM judge | Overestimate |
|---|---|---|---|
GPT-5.5 | 6.25% | 17.9% | ~2.9x |
Opus 4.8 | 8.33% | 18.8% | ~2.3x |
Earlier models (calibration set) | 3.3% | ~3% | ~1x |
Source: Center for AI Safety, "A Significant Increase in Digital Labor Automation," July 2026
As newer models started scoring above the human evaluators' original calibration set, CAIS tested whether an automated judge could replace human review to keep evaluation costs down. It couldn't, and CAIS's own explanation is worth sitting with: evaluating an AI deliverable is itself a demanding, agentic task, one that requires opening real files in real applications and judging them the way a client would. An AI judge inherits the same weaknesses as the AI worker it's grading.
This is Discernment: knowing when to trust an output, when to question it, and when to override it, and knowing that you can't fully delegate that judgment to another AI either. It's why every workflow we build puts a human at the review gate, not an automated grader standing in for one. Track the full cost of a workflow (agent time, plus review time, plus fix time), and you'll often find that below some reliability threshold, the "time saved" story doesn't hold up once someone senior has to actually check the work.
Diligence: does the EU AI Act apply if you're not an enterprise?
Yes, and the size of your company has nothing to do with it. Under the EU AI Act, non-compliance with high-risk AI system obligations (Articles 9 through 15, Article 26 for deployers) carries penalties up to €15 million or 3% of global turnover, effective August 2, 2026 (EU AI Act, Art. 99(4); Legiscope, 2026). That's a distinct, lower tier from the €35 million or 7% penalty reserved for prohibited practices under Article 5, a distinction worth getting right before it reaches your legal team, since rounding it up doesn't make the deadline more real;l, it just makes the wrong argument land.
Under Annex III, AI systems used in employment, worker management, and access to self-employment (assigning tickets, triaging requests, evaluating performance) are explicitly high-risk, regardless of whether the company running them has fifty people or five thousand. If that system fails at anything close to the rate the RLI documents, and the failure costs someone an unfair workload or a missed deadline, the liability sits with whoever deployed it, mandate carrier included.
This is Diligence: responsible AI interaction, meaning you know what data leaves the building, who approved what, and what your audit trail looks like when someone asks. Risk management (Art. 9), technical documentation, event logging, and human oversight (Art. 14) aren't optional add-ons here. We build the audit hooks and approval gates into an engagement from day one, not as a compliance patch bolted on after a system's already live, because retrofitting documentation onto a system that's already deployed is measurably harder than building it in from the start.
Why this only works as fluency, not as four separate fixes
Notice the pattern: Delegation without Description gets you a well-scoped task and a diluted brief. Discernment without Diligence gets you a human reviewer with no audit trail to show why they approved what they approved. These four skills don't stack neatly; they compound, or they don't.
That compounding is also where most AI rollouts quietly die. Junior team members traditionally build judgment by doing the small, low-stakes tasks AI agents now handle first. If AI absorbs "draft this," "clean this," and "format this" without a plan for how juniors develop discernment some other way, you save time today and create a judgment shortage in three to five years. Turning that risk into a training asset means shifting juniors from doing the small task to reviewing the AI's attempt at it, which is exactly the kind of "meaningful human oversight" Article 14 requires, and exactly what we teach explicitly in the AI Fluency Cohort.
We don't have a published case study with these exact numbers yet. We're building our first ones with our first customers right now, which is the honest answer to "can I see proof this works," not a polished logo wall. What we can tell you is which failure mode the RLI data maps to which one of these four skills, because we've built our entire practice around developing all four together, not selling a tool that handles one and hoping the rest sorts itself out.
If you've been handed the mandate and don't know where to start: figuring out what to build first is a free conversation, not a sales call. If you know exactly which workflow is bleeding time and want it built with a human review gate from day one, that's a First Skill Sprint, one to two weeks. If you want your team to own the capability, not just receive the output, that's a Flagship Pilot or the AI Fluency Cohort.
FAQ
What is the Remote Labor Index? The Remote Labor Index (RLI) is a benchmark built by Scale Labs and the Center for AI Safety that tests AI agents on 240 real, paid freelance projects rather than multiple-choice questions. Human evaluators judge each deliverable against a professional's actual work. The best model scored 15.8% as of July 2026, up from 2.5% at launch (CAIS).
Does the EU AI Act actually apply to a small or mid-sized company, or just large enterprises? Yes. The EU AI Act applies based on what the AI system does, not how large the company deploying it is. Annex III classifies AI used in employment, worker management, or self-employment access as high-risk regardless of company size, with obligations enforceable from August 2, 2026 (Legiscope, 2026).
What's the actual penalty for a non-compliant high-risk AI system? Up to €15 million or 3% of global annual turnover, whichever is higher, under Article 99(4). That's a different, lower tier than the €35 million / 7% penalty reserved for prohibited practices under Article 5, a distinction worth getting right before it reaches your legal team.
Why do AI agents score 90%+ on benchmarks but fail most real jobs? Academic benchmarks like MMLU test knowledge retrieval: can the model answer a question? The RLI tests sustained execution: can the model deliver a multi-step, real-world project a paying client would accept? Those are different capabilities, and the RLI shows the second one, what we call agency, still lags the first one badly.
I've just been handed an AI mandate. Where do I actually start? Start by mapping your real workflows, not by picking a tool. Identify which task has genuinely high AI-leverage, build a human review gate into it from the first version, and measure a baseline before you touch anything, so you can prove what changed. That mapping conversation is free.



