Tech●●●●●Difficulty 4 of 5

How do you test whether an AI agent can really do a job?

In 2023 the best model fixed 1.96% of real software bugs. By 2026 the top score on a checked subset was above 80%, yet studies found leaked answers and weak tests behind such scores.

▶ Start the story

You give it real jobs that people have already done, hide the answer key, and check the result with tests. One example is SWE-bench, which has 2,294 tasks taken from software issues that were actually filed, and actually fixed, in 12 Python repositories. The agent is shown the code as it stood before the fix, plus the text of the issue, and has to write a patch. It never sees the tests that will judge the result.

When it was introduced in October 2023, researchers at Princeton University and the University of Chicago found that the strongest model they tried, Claude 2, fixed only 1.96% of the tasks. By February 2026 the leading published result on a human-checked subset called SWE-bench Verified was 80.9%.

1.96%

Share of SWE-bench tasks fixed by the strongest model tested in October 2023 (Claude 2)

Then the benchmark itself was questioned. OpenAI said it would stop quoting the Verified figure: the models being scored had read the repositories the tasks come from, and every frontier model OpenAI tried in a sample could reproduce wording from the problem statements or from the developers' own fix. A separate study, by Reem Aleithan and colleagues, of the system then top of the leaderboard in 2024 found that in 32.67% of its supposedly solved tasks the fix was written out in the issue or its comments, and another 31.08% passed only because the tests were too weak. Take both out and its score dropped from 12.47% to 3.97%.

Wikipedia's article on language model benchmarks calls this Goodhart's law: if models are designed or selected to score highly on a benchmark, it may cease to be a good indicator of model quality. These figures are as of early 2026.

Quiz me

0/3

  1. 1.In SWE-bench, what does the system under test have to produce?
  2. 2.What is the problem called when benchmark answers already appear in a model's training data?
  3. 3.What did the study of 2024 leaderboard patches find?

Recap

A benchmark score can only be trusted as far as the test is unseen and sound.

💡 A trick to remember it · A score is a thermometer: it only tells the truth if the test is not rigged or already seen.

Surprising fact · The first top score in October 2023 was 1.96%; the leading published Verified score reached 80.9% by February 2026, and OpenAI, which had helped build Verified, said it would stop quoting it.

Sources (5)

No source, no claim. Every fact in this lesson (15 claims) cites at least one of these.

  1. [1]SWE-bench · Wikipedia
  2. [2]SWE-bench: Can Language Models Resolve Real-World GitHub Issues? · arXiv (Jimenez et al.)
  3. [3]Language model benchmark · Wikipedia
  4. [4]Building effective agents · Anthropic Engineering
  5. [5]AI agent · Wikipedia
More lessons in 💻 Tech (3) See all tech lessons →

One more light on your map.

Get one lesson like this every day, about the things you love. Free, in two or five minutes.

Get the share card for this lesson ↗