The expert line: what OpenAI's GDPval benchmark measured

The Measured View ยท 30 July 2026

Most AI benchmarks are exams. They ask a model to solve problems that are hard to answer and easy to score: competition mathematics, graduate physics, puzzles with one right answer. They tell you something real about capability and almost nothing about whether a model can do the work you are paid for.

In September 2025, OpenAI published GDPval, an evaluation built around a different question: how do models perform on the actual work products professionals are paid to produce? Legal briefs. Engineering designs. Nursing care plans. Financial analyses. Real deliverables, with the reference files and context that come with them, drawn from real occupations.

The construction was serious. GDPval covers 44 occupations across the nine sectors that contribute most to US GDP, from software developers and lawyers to registered nurses and mechanical engineers. The full set holds 1,320 tasks, each written by a professional with 14 years of experience on average and each passed through around five rounds of expert review. Grading was blind: experienced professionals compared model output against the human expert's own deliverable without being told which was which. The blinding was not perfect, and graders may sometimes have inferred which was the model from stylistic tells.

The best model at release, Claude Opus 4.1, produced deliverables rated as good as or better than the expert's in just under half of head-to-head comparisons, 47.6 percent. Performance had more than doubled between GPT-4o in spring 2024 and GPT-5 in summer 2025, on a roughly linear trend. One detail is worth noting: on OpenAI's own benchmark, the strongest performer at release was a competitor's model.

Share of blind comparisons rated as good as or better than the industry expert's

Best frontier model at release (Claude Opus 4.1)
47.6%
Parity with human experts
50%
Data: GDPval, Patwardhan et al., OpenAI (2025), arXiv:2510.04374

Models also completed these tasks roughly a hundred times faster and more cheaply than the experts, though that figure counts inference alone and excludes the human oversight real work requires. And 47.6 percent is the September 2025 release number. Newer models have scored higher since.

What the benchmark cannot see matters as much as what it measures. GDPval is one-shot: a well-specified prompt, the reference files, one deliverable. It skips the drafts, the feedback loops and the client conversations that surround professional work. OpenAI says so openly. A lawyer often has to work out that a brief is the right move before writing one, and GDPval never asks that question. It measures execution on well-specified tasks, and much of professional value sits upstream of execution, in deciding what the task is. No benchmark measures that yet.

So the finding is task-level, not job-level. A job is a bundle: specified deliverables, ambiguous problems, relationships, judgment calls, accountability. GDPval says the specified-deliverable share of that bundle is approaching expert quality at a fraction of the cost. It says little about the rest.

Two moves follow. Inventory your deliverables: list what you produce in a typical month and mark the items that are well specified, with clear inputs. That column is where models compete first, and where you should already be delegating. Then move your value upstream, because specification, judgment and accountability all gain value as execution gets cheap. Become the person who defines the task and signs off on the result.

The expert line is no longer far away. On real deliverables from 44 occupations, blind expert graders picked or tied the model in just under half of comparisons, and the trend points up. Execution is becoming cheap. Knowing what to execute, and standing behind it, is not.

The paper: GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Originally published on LinkedIn.

Back to The Measured View