Ai just beat human economists · The Machines Just Wrote an Economics Paper. It Ranked Higher Than the Humans.
A Federal Reserve economist ran the experiment. The humans finished last.
Serafin Grundl, an economist at the Federal Reserve Board of Governors, did something most of the profession has been quietly dreading. He took a serious empirical research task, the kind that takes a PhD economist weeks of careful work, and handed it to three agentic AI systems. He then gave the same task to 146 human research teams and had AI judges rank every submission against each other.
The ranking was identical across every reviewer model: Codex with GPT-5.4 came first, Codex with GPT-5.3-Codex second, Claude Code with Opus 4.6 third, and the human researchers finished last. The order held across every judge and across every task variant, which rules out the easy explanation that one particular model was flattering itself or punishing the humans.
This is not a benchmark and it is not a toy task. It is a serious head-to-head comparison of agentic AI systems and human economists doing real empirical work, and it quietly closes the chapter on the idea that AI is merely a helper for coding and drafting.
What the study actually measured
The research question was genuine applied microeconomics, namely estimating the causal effect of DACA eligibility on the probability of working full-time among Hispanic-Mexican, Mexican-born individuals in the United States, using American Community Survey data from 2006 to 2016. Grundl ran the comparison across three difficulty levels that progressively constrained the researcher, starting with full freedom and moving to a prescribed research design and then to a pre-cleaned dataset.
The human side of the comparison was serious. 146 research teams were recruited, 87 percent held PhDs, 67 percent were faculty, and each team was paid 2,000 dollars to complete all three tasks. This was not undergraduates doing a homework set, it was practicing applied economists producing the kind of work they would submit to a journal.
Grundl then ran 100 independent instances of each AI system on each task, for a total of 900 AI replications. Every run produced actual code, actual data cleaning choices, actual regressions, and an actual written report in the style of an academic submission. Three reviewer models, Gemini 3.1 Pro, Opus 4.6, and GPT-5.4, then read 300 bundles of four submissions each and wrote comparison reports ranking them on identification, robustness, and substantive quality.
The detail that should end the debate
The more interesting finding is not that AI won the ranking. It is how it won, because the victory looks different depending on which part of the distribution you inspect.
If you just compare median estimates, the humans and the AI systems land in roughly the same place, which suggests on first glance that the two groups produce similar work on average. The distributions, however, tell a very different story. Human estimates have substantially wider tails, with larger standard deviations and more extreme outliers, while the AI estimates concentrate more tightly in the middle even though individual runs can occasionally flip sign. More importantly, the written submissions themselves, when evaluated on substance rather than on a point estimate, were consistently ranked higher for the AI systems than for the humans.
The cleanest evidence against reviewer bias sits inside the paper itself. Claude Code with Opus 4.6 was one of the reviewers, and it consistently ranked the Codex submissions above its own, which is the opposite of what you would expect if the AI judges were simply flattering their own kind. If there is a bias in this tournament, it is a bias against the reviewing model’s own work, and the humans still finished behind.
What just changed
For the past three years, the dominant narrative about AI in knowledge work has been a story of assistance, a faster intern who helps you write code, draft emails, and summarize documents. Grundl’s paper ends that narrative, because empirical research is the hardest, most credentialed, and most carefully reviewed work the knowledge economy produces, and agentic AI has now been shown to do it at or above human quality in a fair fight.
The paper’s closing observation is worth reading in its original form, where Grundl writes that “agentic AI systems will allow us to scale empirical research in economics.” The operative word is scale, not assist, not accelerate at the margin, but change the order of magnitude of analytical output that a single researcher or a single institution can produce.
What this is not
It would be easy to read Grundl’s ranking as a verdict against human researchers, and that reading would miss the actual point entirely. The 146 human research teams in this study are not mediocre, they are among the strongest applied economists in the profession, with 87 percent holding PhDs, 67 percent serving as tenure-track or tenured faculty, and roughly 40 percent working directly in labor or immigration economics. These are the people who review the papers, teach the methods, and produce the empirical evidence that shapes policy, and they are not being left behind because they are bad at their jobs.
The ranking is better understood as a preview of what happens when that same level of human expertise builds on top of agentic systems rather than working alongside them. A PhD economist who runs a research question through 100 independent AI replications, inspects the forks in the analysis path, selects the robust specifications, and writes the interpretation is not being replaced by AI, she is doing research that no single researcher on the planet could have done alone three years ago, and she is doing it with a depth of robustness checking that used to be reserved for referee-response rounds.
The real implication of Grundl’s paper is not that human researchers are obsolete but that the gap between augmented and unaugmented researchers is about to become the dominant variable in knowledge work, and the humans who pull ahead in the next decade will not be the ones who defend their turf against AI, they will be the ones who figured out how to build on it first and with the most judgment. Every research organization, every consultancy, every analytics team now needs to decide which side of that gap it wants to be on, because the technology no longer tolerates fence-sitting.
What this means for the next twelve months
Research becomes a volume problem rather than a talent problem, because a single researcher with access to agentic systems can now run hundreds of independent analyses in the time it previously took one PhD team to run one, which means robustness checks and multi-specification sensitivity analysis stop being a chore reserved for referee responses and become the default mode of working.
Peer review is likely to be rebuilt around the same tools, because the Grundl paper also demonstrates that AI reviewers rank submissions consistently across different models, which is a better inter-rater reliability than most human review panels achieve, and the journals that ignore this will be outrun by workflows that embrace it.
The bottleneck moves decisively to judgment, because when analysis itself is cheap, the scarce resource becomes the question you choose to ask, the framing you bring to it, and your ability to recognize which answers would actually change a decision. That part of the work is still human, and in the new workflow it is the only part that is reliably human.
Smaller organizations finally catch up to larger ones on analytical throughput, because a three-person team equipped with agentic tools now has the research capacity of a twenty-person department from three years ago, and the competitive gap in knowledge-intensive industries is collapsing from the bottom rather than widening from the top.
The European angle no one wants to say out loud
Europe has a demographic problem that will not solve itself, with a structural shortage of millions of skilled workers projected across the next fifteen years, and the productivity tooling we build right now is not a nice-to-have but the bridge across a labor gap that no amount of hiring or immigration policy will close in time.
Agentic research of the kind Grundl documents is one of those bridges, because a system that can take a research question, translate it into a design, write the code, run the analysis, inspect its own results, debug its own mistakes, and produce a written report that peer-ranks above human work is not really a tool in the traditional sense, it is analytical capacity that can be provisioned the way we used to provision cloud computing. Organizations that treat this as capacity rather than as a novelty will pull ahead quickly, and the ones still debating whether to pilot a chatbot will find themselves in two years explaining to their boards how they missed the moment their competitors doubled analytical throughput without hiring a single new person.
Back to the paper
Grundl is careful in the way Federal Reserve economists tend to be careful, and he is explicit that AI systems make mistakes, that different runs of the same model can produce estimates with opposing signs, and that none of this is magic. He also notes, correctly, that the same problems apply to human research teams, where inter-analyst variation in empirical economics has been documented for years and is the whole reason the Huntington-Klein et al. many-analysts study exists in the first place.
The insight that matters more than the caveats is what becomes possible when you can run an analysis 100 times for the cost of one human team-month, because in that regime you stop needing any single run to be perfect and you start aggregating, inspecting dispersion, and surfacing the forks in the analysis path that produce different conclusions, which gives you better empirical work rather than worse because variance becomes observable instead of hidden.
That is the new standard for serious empirical work, not AI replacing humans, but a workflow in which human judgment selects the question and reads the results while hundreds of parallel analyses fill in the space between. The humans finished last in Grundl’s ranking, but the humans also designed the experiment, chose the task that made the whole comparison visible, and wrote the paper that forced the rest of us to notice, which remains the part of the job that is still worth doing.
Gerhard Kürner is CEO of 506.ai. He writes about enterprise AI, the European productivity question, and what actually changes when agentic systems cross quality thresholds.
Reference: Grundl, S. (2026). A Comparison of Agentic AI Systems and Human Economists. Federal Reserve Board of Governors. claude-code-economist.com



