For most of the current AI boom, the debate about whether models can do real jobs has run on anecdote and vibes. The Remote Labor Index was built to replace that with a number. Created by Scale AI and the Center for AI Safety, it measures how often an AI agent can finish an actual paid freelance project to the standard a human professional would deliver. When the benchmark first appeared in late 2025, the answer was humbling: the best agent cleared just 2.5 percent of the work. By the summer of 2026, the top score had jumped to 16.1 percent. That leap, and the careful way the researchers arrived at it, is the most grounded read we have on how close AI agents are to earning a living.
- The Remote Labor Index scores AI agents on 240 real freelance projects worth more than $140,000, spanning 23 sectors of remote knowledge work.
- At launch in October 2025, the top automation rate was 2.5 percent, achieved by the agent Manus.
- An updated run published on July 1, 2026 lifted the frontier to 16.1 percent for Fable 5, with Opus 4.8 at 8.3 percent and GPT-5.5 at 6.3 percent.
- The score is deliberately strict: a project counts as automated only when trained evaluators judge the AI deliverable at least as good as a human professional's finished work.
What the Remote Labor Index measures
Most AI benchmarks test fragments of ability. They ask a model to answer trivia, solve a math problem, or fix an isolated bug in a code file. Those tests are useful, but they say little about whether a system can take a messy client request and hand back something a paying customer would accept. The Remote Labor Index was designed to close that gap. It evaluates end-to-end performance on economically valuable remote work, the kind of assignment a freelancer would actually pick up on a marketplace and get paid to complete.
The distinction matters because knowledge work is rarely a single clean task. A logo commission involves reading a brief, interpreting taste, producing files in the right formats, and iterating toward something usable. The benchmark keeps that full shape intact. Rather than slicing a job into a tidy question with one right answer, it hands the agent a complete project and asks for a finished deliverable, then compares that output against what a human professional produced for the same brief. The paper describing the method, published on arXiv on October 30, 2025 under the lead of Mantas Mazeika and a team of 47 authors, frames this as evaluating agents in practical settings rather than on toy problems.
How the benchmark is built
The Remote Labor Index draws on 240 professional freelance projects sourced from real marketplaces, with a combined value north of $140,000. According to the Scale Labs project page, these projects span 23 sectors of remote work, from graphic design and video editing to data analysis, architecture, and software. Each project is a self-contained unit of paid work: a median assignment takes a human professional roughly 11.5 hours and carries a median price around $200. That is a meaningful chunk of labor, not a five-minute microtask. Sourcing projects from live marketplaces also means the tasks carry the same ambiguity, missing context, and imperfect instructions that real clients hand over, which is exactly the friction that trips up systems trained on cleaner data.
Every project ships with a gold-standard deliverable, the actual work a human freelancer was paid to produce. When an AI agent attempts a project, its output goes head to head against that human reference. Trained human evaluators judge the two side by side and decide which they would rather receive as the client. To keep the scoring honest, three independent evaluators review each pairing and a majority vote settles the outcome. The researchers report an inter-annotator agreement of 94.4 percent, meaning the graders overwhelmingly agree with one another, which is the sign of a rubric that measures something real rather than grader mood.
The headline metric is the automation rate: the share of projects where the AI deliverable is judged at least as good as the human professional's gold-standard work. A 16 percent automation rate does not mean a model is 16 percent good at everything. It means that on roughly one in six of these complete, paid projects, the agent's finished output matched or beat a professional's.
The first results sat near the floor
The original October 2025 run delivered a reality check to anyone expecting agents to be quietly replacing freelancers already. Across the frontier systems tested, performance clustered just above zero. The top automation rate was 2.5 percent, posted by the agent framework Manus. Grok 4 and Claude Sonnet 4.5 followed at 2.1 percent each, GPT-5 landed at 1.7 percent, OpenAI's ChatGPT Agent reached 1.3 percent, and Gemini 2.5 Pro trailed at 0.8 percent.
Those figures were far below the confident claims circulating at the time about agents automating whole job categories. The gap is instructive. Systems that ace narrow coding benchmarks and score well on professional exams still struggled to ship a single finished freelance project that a client would prefer over a human's. Passing a bar exam question is not the same as running a legal matter, and generating a plausible code snippet is not the same as delivering a working web application to spec. The Remote Labor Index exposed that distance in a way tidier benchmarks had hidden.
The July 2026 jump on the Remote Labor Index
Eight months later the picture changed sharply. In an update published by the Center for AI Safety on July 1, 2026, the top automation rate reached 16.1 percent, achieved by Anthropic's Fable 5. Opus 4.8 came in at 8.3 percent and OpenAI's GPT-5.5 at 6.3 percent. The prior leader, Opus 4.6, had scored 4.17 percent, so the frontier more than quadrupled in under eight months. The Decoder summarized the shift bluntly in its coverage, noting that agents had gone from completing 2.5 percent of freelance jobs at professional quality to 16 percent, and framed it as movement outpacing many industry predictions.
The update also exposed how much compute now goes into a single serious attempt. Models were run with extended reasoning settings and given up to 24 hours of wall-clock time per project, with a default spending budget of $50 per task. Fable 5 was allotted $150 per project and completed 218 of the 240 projects before access restrictions cut its run short. Agents even had access to an NVIDIA A100 GPU when a task called for it. These are not chatbot replies dashed off in seconds. They are long, resource-heavy runs that mirror how a professional might grind through a difficult commission.
The frontier has more than quadrupled in under eight months.Center for AI Safety
How the Remote Labor Index compares to other benchmarks
The Remote Labor Index is not the only attempt to measure AI against real work, and understanding its neighbors helps explain why its numbers look so different from cheerier headlines. Most agent benchmarks in wide use focus heavily on software engineering, where success is easy to check with automated tests. That convenience skews the field. As one analysis covered by The Decoder put it, agent benchmarks obsess over coding while ignoring the vast majority of the US labor market. A model can look near-superhuman on a coding leaderboard and still flounder at the graphic design, video editing and audio work that fills real freelance marketplaces.
OpenAI's GDPval benchmark took a different slice, testing model knowledge across a broad range of occupations to gauge general professional competence. That approach captures breadth, but it leans on knowledge questions rather than finished deliverables, so a strong GDPval score does not guarantee a client-ready result. Other efforts, sometimes grouped under names like APEX-Agents, push toward sustained performance in a narrow set of high-value professions. Each design makes a tradeoff between realism and measurability.
The Remote Labor Index plants its flag firmly on the realism side. By insisting on complete projects, human gold-standard comparisons, and trained graders, it accepts a harder and slower evaluation in exchange for a score that maps to something a business owner cares about: would I pay for this output instead of a freelancer's? That is also why it is worth revisiting an earlier, sobering result. A related study from the Center for AI Safety and Scale AI found that no tested agent could complete more than a low single-digit percentage of simulated freelance work, earning only a tiny fraction of the money on offer. The July 2026 update did not overturn that finding so much as extend its curve upward.
Why the headline number is smaller than the hype
A 16 percent automation rate is a large jump, but it is worth sitting with what the other 84 percent represents. On more than four out of five real freelance projects, the best agent in the world still could not match a human professional's finished work. The Remote Labor Index is valuable precisely because it refuses to grade on a curve. Its bar is client-ready output, not partial credit for a decent attempt.
The researchers were also candid about a trap that inflates many other evaluations. They found that judging an RLI deliverable is itself a demanding, agentic task, and that automated large language model judges are unreliable for measuring absolute quality on these projects. In their testing, LLM graders overestimated agent performance by a factor of 2.3 to 2.9 times compared with careful human review. Those automated judges were still useful for ranking systems against each other, but not for pinning down a true automation rate. That finding is a quiet warning about the flood of self-reported benchmark scores where a model grades its own homework.
There is a selection effect to keep in mind as well. Fable 5's 16.1 percent came from a run that stopped at 218 of 240 projects because of access restrictions. A score computed over a slightly different project set is not perfectly comparable to one computed over the full 240. The direction of travel is not in doubt, but the precise decimal deserves the same skepticism the researchers themselves model.
Which work is falling to AI, and which is holding
Because the benchmark spans many sectors, it offers a rough map of where agents are gaining ground. The projects that agents handle best tend to be the ones with clear specifications and verifiable outputs, such as structured data work and self-contained software tasks where correctness can be checked. Creative and production work with unambiguous requirements has also moved within reach as models improved at following detailed briefs.
The work that keeps resisting automation is the work that depends on context an agent cannot easily acquire. Highly nuanced, client-specific assignments that require deep domain knowledge or a particular brand voice remain hard, because the winning answer lives in a human relationship and a set of unstated preferences rather than in the brief. This is the same pattern seen across knowledge work more broadly: the more a task can be pinned down in writing, the more automatable it becomes, and the more it relies on tacit judgment, the longer it holds.
It is also worth watching how the gap between top systems widened. In the first run, the leaders sat within a couple of percentage points of each other, all bunched near the floor. In the July 2026 update the spread opened up, with Fable 5 at 16.1 percent nearly doubling the next-best score and leaving a clear ladder down through Opus 4.8 and GPT-5.5. When a benchmark starts to separate the field like that, it usually means the test has enough headroom to reward genuine capability differences rather than noise. That separation is a sign the Remote Labor Index still has room to run before it saturates.
That framing helps explain why the jump from 2.5 to 16.1 percent should be read as a trajectory rather than a verdict. The CAIS update noted an expectation that a majority of freelance tasks could become completable by AI at professional quality within a couple of years if the current pace holds. Predictions like that have a poor track record in AI, in both directions, so the honest posture is to treat it as a hypothesis the next benchmark run will test.
A better way to argue about AI and jobs
The lasting contribution of the Remote Labor Index is not any single score. It is a method that turns a shouting match into a measurement. Instead of asking whether AI can do knowledge work in the abstract, it asks a specific, falsifiable question: on this fixed set of real paid projects, judged by trained humans against professional work, how often does the agent win? That question can be re-run every few months, and the resulting curve is far more informative than a press release claiming a model can replace an entire profession.
Read that way, the story of the past year is genuine and measured progress rather than sudden replacement. Agents went from clearing one freelance project in forty to clearing roughly one in six, using long runs and real tools, on a benchmark engineered to be hard to game. That is a serious result. It is also a long way from the world where an agent quietly takes over a freelancer's client list. The value of the Remote Labor Index is that it lets both of those truths sit in the same table, backed by the same evidence, so the argument can finally move forward on data instead of adjectives.
For anyone trying to plan around this technology, the practical lesson is to distrust round numbers that arrive without a method attached. A vendor claiming its agent can do most of a given job is making a Remote Labor Index style claim without the Remote Labor Index style rigor. The useful questions to ask are the ones the benchmark bakes in. What was the complete task, not the cherry-picked subtask? Who judged the output, and were they comparing it against professional work or against a low bar? How much time and compute did the agent consume to get there, since a result that takes 24 hours and a dedicated GPU has very different economics from one that takes a minute? Those questions turn a marketing figure back into something you can reason about.
The next run of the benchmark will be the real test of the trend. If the top automation rate climbs from 16 percent toward the thirties on the same 240 projects, the case that agents are absorbing a growing slice of remote work gets much stronger. If it stalls, the 2026 jump will look more like a burst tied to a single strong model than a smooth march toward full automation. Either outcome is worth knowing, and the Remote Labor Index is one of the few tools built to tell the difference honestly. That is a rare and useful thing in a field where most of the loudest claims still come with no scoreboard at all.