OpenAI says GPT-5 stacks up to humans in a wide range of jobs. The company used a new benchmark called GDPval to measure model performance on tasks that matter to the economy. The test looks at work tasks across dozens of occupations and asks professionals to compare AI outputs to human outputs. The early results show that GPT 5 and some competitor models are closing the gap on many common tasks. The findings are important for anyone who follows the future of work and the real-world value of AI.
What GDPval shows
GDPval evaluates model performance across occupations that matter most to gross domestic product. The first version of the test focuses on written deliverables and analysis tasks. Experts in each field compare reports or briefs created by humans and by the models. The test covers 44 occupations in nine broad industries. In these head-to-head comparisons, GPT 5 scored at a rate that put it close to the work quality of seasoned professionals on many tasks. Another leading model from a rival lab also performed strongly on the same tests.

OpenAI reports that a high compute variant of GPT-5 won or tied against human professionals on a sizable share of tasks. That is not a claim that the model can do every job. It is a specific result about narrow kinds of work. GDPval looks at sampled tasks that professionals actually perform in the workplace. It does not yet measure long-running workflows, interpersonal judgment calls, or complex teamwork. OpenAI says it plans more extensive tests in future versions.
What this means
The most practical takeaway is that AI models are now useful assistants for many routine knowledge work tasks. People in those roles can use a model to draft reports, create initial research briefs, and speed up parts of their workflow. When a model can produce a high-quality first draft, professionals can focus on higher-value activities such as strategy, checking for errors, and making final decisions.

At the same time, this progress raises hard questions. Models can make mistakes that look plausible. They can omit context that a human expert would notice. They can produce outputs that are biased or incomplete. That means companies need guardrails and new review practices if they adopt these tools widely. Relying on models without strong verification can create real risks for organizations and for consumers.
How firms may react, some businesses will accelerate adoption. Teams that write reports, prepare analyses, or draft communications may embed models into their workflow today. Other organizations will move more slowly. They will require human-in-the-loop checks and clearer audit trails. Public sector groups and regulated industries may demand additional transparency before they rely on AI output in decisions that affect people.