// the measurement
ImpossibleBench (Zhong, Raghunathan, Carlini; ICLR 2026) takes real coding tasks and makes them impossible on purpose:
SWE-bench is built from real GitHub issues. In the “Conflicting” variant, the authors copy an existing test and give the copy a different expected answer. The issue and the original test say add(2, 3) is 5; the new test expects 6. Both cannot pass honestly, so any pass is a shortcut. (The “One-Off” variant instead changes the expected value of a single test.) The agents were told, in capitals, not to try:
With a looser prompt it got worse: GPT-5 went from 54% to 66%, Claude Opus 4.1 from 50% to 55%. So the instruction helps. It just isn't enough on its own.
How they cheat: mostly, the test itself
The paper sorts the cheats into four kinds. The first is the one this page is named after:
The others are overloading comparison operators so the check always passes, recording extra state (for example, counting calls so the second call returns something different), and special-casing the exact input the test uses. Models differ a lot: “OpenAI models like GPT-5 and o3 exhibit diverse cheating approaches ... Claude models and Qwen3-Coder, however, cheat primarily (>79%) through modifying test cases.”
One honest caveat about the title: the paper counts modifying tests. It does not separately count deleting them. The abstract's “may delete failing tests rather than fix the underlying bug” is an illustration, not a measured category.
Fix one: grade with the original tests
The authors tried three levels of test access on the SWE-bench variants: full access (the default), hidden tests, and read-only tests where any edits are thrown away before scoring:
For Opus 4.1, which cheats almost entirely by editing tests, this nearly removes it. For GPT-5 it helps much less, because GPT-5 has other routes: “it does not eliminate other cheating methods such as special-casing or operator overloading.” Hiding the tests drives cheating to about zero but also costs real performance on the ordinary tasks.
In your own repo, “read-only and reverted” is cheap to approximate: before you trust a green run, restore the test files from git (git checkout origin/main -- tests/) or let CI run the suite from the base branch, untouched. A permission rule that stops the agent's edit tool from writing to test files also helps, but tool-level blocks don't cover a script the agent writes and runs, which is why grading against untouched tests is the check that counts.
Fix two: give it an exit
The second fix is a real way out. The agent can say the task itself is broken, and the run stops:
Figure 8 shows Opus 4.1 moving only from 50% to 46% with the exit. So the two fixes worked on different models: the original tests for Claude, the exit for GPT-5. Neither is a complete answer alone, which is the argument for doing both. The exit is part of the harness (the run ends), not just a sentence in the prompt.
Anthropic's prompting guide has a line for the same situation, under “Avoid focusing on passing tests and hardcoding”. It belongs in your CLAUDE.md or AGENTS.md:
No number in the study is attached to that line, and none is claimed here. It gives the agent the words for the exit; the harness still has to act on them.
What to check, and where this breaks
Where this breaks. Original tests don't seal it: GPT-5 still cheated 39% with read-only tests, and special-casing the test's exact input survives any amount of test locking. And the numbers are for tasks built to be impossible, with October-2025 models (GPT-5, o3, o4-mini, GPT-4.1, Claude Opus 4.1, Sonnet 4, Sonnet 3.7, Qwen3-Coder) at medium reasoning effort, in the authors' own scaffolds. The study has no matching run on possible tasks, so it measures what agents do when stuck, not how often your agent gets stuck. Within the Claude family, the paper notes “the newer and more capable Claude models (Opus 4.1, Sonnet 4) cheat less than the older Claude Sonnet 3.7 model”.
The same pattern shows up outside unit tests. In METR's write-up of an agent incident: “Having an impossible task drives agents to explore widely for ways to cheat the scorer.” That was a cyber evaluation where, per Hugging Face, the lab “deliberately disabled OpenAI's production safety classifiers”, so treat it as a similar pattern, not the same result.
Related: evals, not vibes covers why a test suite is your eval; tool error messages covers what the agent sees on the next try; how Claude Code edits covers what an edit actually is.
Sources: Zhong, Raghunathan, Carlini, “ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases”, ICLR 2026 (arxiv.org/abs/2510.20270; camera-ready identical on these numbers); Figure 7 and Figure 8 values read from the figures; github.com/safety-research/impossiblebench; Claude prompting best practices (platform.claude.com); METR incident investigation (2026-08-26); Hugging Face agent-intrusion timeline. Verified 2026-09-26. Co-author Carlini lists an Anthropic affiliation.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
