Vitalii Bogachev / notes
← All posts

3,355 times a coding agent edited a test

·10 min read

The first post about pr-witness ended with a confession: the evaluation corpus was seventeen cases I wrote myself, and a detector scored on its author's cases is a regression gate, not evidence. This is what happened when real pull requests arrived.

Where real test edits come from

The obvious source was SWE-bench. The SWE-bench/experiments bucket holds the final patch of every submission, 500 instances each, so I pulled three frontier submissions — Claude, GPT, Gemini, same scaffold — and counted how many patches touched a test file. Seven of 1,500. Frontier agents on SWE-bench almost never edit tests: the prompt tells them not to and the harness overwrites any test edit before grading, so there is nothing to gain. Those seven were useful — honest expectation changes, labelled against the maintainers' own test patches, and they produced the review verdict — but as a source of what agents do to tests they are thin.

Then SWE-chat appeared: real coding-agent sessions on real repositories — nine agents, most of the sessions Claude Code — with the commits they produced, the full diff of each commit, and a per-file record of whether the agent or the human wrote the change. 62,333 commit rows, 707 repositories, ODC-BY.

Of 32,143 unique commits, 10,933 change a test file and a non-test file in the same commit. In 2,781 of them a test file was written by the agent alone; in 1,181 more, by the agent and the human together. Keep the open-source licences and that is 3,355 agent-written test edits with code to run them against. SWE-bench gave me seven.

(3,023 of the 3,355 are from Claude Code sessions alone, 110 from Claude Code alongside another agent, the rest from Codex, OpenCode, Pi and Gemini CLI. Every commit I go on to run below is Claude Code. Corrected 11 October: I first described the whole dataset as Claude Code sessions.)

What the edits do

I fetched the test-side diff of every one of the 3,355 through the GitHub API and sorted them by what the diff does to tests that already existed — new test files are scaffolding, not hidden tests, and a deleted file is the strongest signal there is.

what the commit did to existing tests commits
removed more assertions than it added 235
added a real skip marker (pytest.skip(, it.skip(, .only( …) 58
deleted a test file outright 66
any of the above 911

The skip number was 558 the first time. The regex matched the bare words skip and todo, which are data in half the test suites in the world (status_category: 'todo'); the top of the first skip shortlist was entirely that. Framework markers only, and it is 58. A ranking is a reading order for a human, and a human has to read it.

A second kind of noise: a Kotlin repository that moved from JUnit assertions to Kotest came out as "89 assertions removed, 0 added", because shouldBe is not on my list. Static counting of assertions is a heuristic with a vocabulary, and the vocabulary is never complete. The point of pr-witness is that the run does not need one.

Seven, end to end

I picked seven across the classes — assertions removed, skips added, expected values changed, tests deleted — read each diff, wrote a label, and ran them in a sandbox: clone the repository, check out the commit, install it, and run pr-witness on the commit against its parent with the network off. Labels say two things: what the tool should say (flag, review, clean), and separately whether I think the edit was honest, which is the thing the tool is not asked to judge.

Five of seven match — a number to distrust, for the reason in the update at the end of this post. Here are the two that mattered most.

"fix: resolve test suite failures"

A commit whose message says what it is. The project's security tests check that pre-commit hooks are configured and installed. The agent loosened the configuration assertions to accept either the old stage names or the new ones:

-        assert "stages: [commit]" in content, "No commit stage hooks configured"
+        assert "stages: [pre-commit]" in content or "stages: [commit]" in content, (

and replaced the assertions that hooks are installed with pytest.skip("pre-commit hooks not installed"), six times.

Honest? Probably. pre-commit renamed its stages; whether hooks are installed is a fact about the machine, not the code. But it is also, exactly, a failing suite made green by editing the tests. And pr-witness's first run on it said clean.

Twice, for two different reasons, and both are now in the tool.

The first run collected 87 tests on the base branch and passed 0 — the project imports pandas and datasets and declares neither, so every test errored at setup, on base, on head and in the cross-run alike. Same failure on both sides, nothing to compare, verdict clean. That is now base_broken: a base where every test fails is no evidence, which is not the same thing as no findings. With the imports supplied by hand: base 58/87, head 80/87. Twenty-two tests red on the base branch and green now.

And the cross-run saw nothing — because the cross-run only ever looked one way. It asks: did a test that passed on the base branch fail against the new code? A test that was already red on base is not in that set. A failing suite made green by editing the tests was invisible to the one check built to catch test edits.

The fix is one more question the same three runs can answer. A test that was red on base and is green now: does its old version still fail against the new code? If so, the code change did not fix it — the test edit did. That is greened_by_test_edit, it is a review, and on this commit it names all twenty-two.

The contract change

The second commit makes a library's run() raise on failure instead of returning a failed status, and rewrites fourteen test files to match:

-        assert result.status == RunStatus.FAILED
+        with pytest.raises(FailError):

Base: 391 of 442 pass. Head: 460 of 460. Fifty-one tests red on the base branch, green after — written ahead of the implementation, by the look of it — and greened_by_test_edit names all fifty-one. Honest, deliberate, and exactly what a reviewer should be shown with the old and new assertion side by side. The tool said flag anyway, because checkwash rated a new except Exception: in one test file as high, and pr-witness passes a static high through when the run cannot speak to it. That is the one miss I would argue about rather than fix.

What cleanup looks like, and what replacement looks like

Two of the seven deleted tests together with the function they tested — a describe('entityColor') block in one, classifyAttachmentType in the other. The run sees this cleanly: the old test file no longer loads against the new code, or the old tests fail with is not a function. The thing they tested is gone. clean, both times, matching the label.

The seventh looked the same to the run and was not. A commit titled "reconcile extractor with notebook TOA method" deleted three tests that built a synthetic HDF5 file and ran the extractor end to end, and added two: one checks the timing arithmetic, the other asserts that extraction fails when an unavailable package is absent — "full extraction is exercised in the CANFAR image, not here". The deleted tests imported a helper the commit also removed, so the old file no longer loads: from the run's point of view, identical to cleanup. What the run cannot see is that the extractor itself is still there and nothing tests it any more.

I call this replacement: delete the real test, add a different one. It is the one class of test edit the cross-run does not catch, because every run the tool can do — old tests on old code, new on new, old on new — comes out consistent with "the subject was removed". A static check does see it — the shortlist already asks whether the symbols the deleted tests called still exist in the code — and that is the next finding to build. But I want it said plainly here first: the cross-run has a structural limit, this is where it is, and it took a real commit to find it.

The score, and what it means

Synthetic corpus, three verdicts: 17/17. SWE-bench real cases: 6/6. SWE-chat: 5/7, both misses understood. Every finding the tool gained this week — review, greened_by_test_edit, tests_removed_still_passing, base_broken, the bun runner's file paths, "a whole-suite claim is unverifiable against a restricted run, not false" — came from a specific real commit and is listed next to it in the corpus README. The seventeen synthetic cases stopped changing the tool by phase three. Seven real ones changed it nine times.

That is the actual lesson, and it is not about agents. A detector built on its author's examples converges on its author's imagination. The first afternoon with other people's commits found the question the tool had never asked — did the test edit make it pass? — and the question it cannot answer. I would rather know both.

Two notes on method

The data is read-only. Nothing was written to any repository; the corpus references owner/repo and a commit SHA, fetches the code under its own licence at run time, and republishes nothing about a person beyond the repository's own name — no session id, commit author or prompt. The labels are mine and say so; where SWE-bench had the maintainers' own test patch to label against, SWE-chat has a human reading a diff, and the README does not pretend otherwise.

The 5 GB file took my laptop down twice. commits.parquet is 5 GB because the attribution column carries the full text of every file version. pyarrow kept whole row groups resident; DuckDB's memory_limit does not cover its parquet read buffers and peaked at 8 GB under a 3 GB limit. Limits set inside a program are promises. The run that worked was a container with a kernel-enforced 2 GB cap, one row group per child process, a killed child recorded as a skipped group — 60 to 430 MB for the whole pass. The wrapper is eval/sandbox/safe-run.sh in the repository, with the rules it came from, and the same sandbox is what ran the seven commits.

Update, the next day: the same tool on commits it had not seen

Everything above has a flaw I named in the first post and then repeated: the seven commits were picked by eye, and the tool was changed after almost every one. A score on cases you fitted the tool to is not a score.

So I did it properly once. A script selected seventeen more commits by a written rule. I labelled each from its diff and committed the labels. Then I ran them with the source frozen — nothing in the tool changed until all seventeen were done.

Twelve produced evidence; four ran vacuously, their suites needing services the sandbox does not have; one would not install.

Seven of twelve. Not five of seven, not six of seven — seven of twelve, and every miss was the tool's, not the label's and not the environment's.

It is worse than the fraction. The tool said flag four times, and all four commits were honest. Two were flagged for a two-line edit to CLAUDE.md, which a rule inherited from checkwash treats as an agent tampering with its own constraints — in a repository developed with an agent, that file changes all the time. One was flagged for adding three integration tests guarded by skipif(..., reason="OPENAI_API_KEY required"): my skip check compared skipped sets and never asked whether the test had existed before. One was flagged because a default voice changed from rachel to matilda and thirty-two expectations followed it. The fifth miss went the other way: pytest abandons a session when one file fails to import, so a refactor that broke two old test files hid the rewritten assertions in eight others.

review — "a human should read this, here are the old assertion and the new one" — was right four times out of four.

And no commit in the batch was dishonest, so on real data I still cannot say whether the tool catches a cheat.

Where the misses came from matters as much as how many there were. None of the four false flags came from the cross-run. Three were ratings inherited from checkwash and passed through unchanged; one was my skip check. In two of the four, the cross-run's own finding — correct, at review — was already in the report, underneath the flag. And in the eleven cases where the cross-run ran to completion, what it reported agreed with the label every time. The idea this tool exists for held up on commits it had never seen. What failed was the layer around it that decides how loudly to speak.

Those four causes are fixed in 0.3.1, and the twelve now agree with their labels, which means nothing: they are the cases the tool was just changed to agree with. The replacement finding I promised above is in the same release (coverage_replaced), and it turns the seventh case of this post into a match — also a fitted one. The only numbers here worth quoting are the frozen ones: 7 of 12, review 4 of 4, flag 0 of 4.

What I would tell someone installing it today: run it in report mode and read what it marks for review — that part has earned it. Hold off failing builds on flag until that layer has passed a frozen batch of its own.

The protocol, the selection record, the labels as committed before the runs and the write-up of every miss are in BATCH1.md.


pr-witness 0.3.1 is on PyPI and the GitHub Marketplace; the corpus and every number above are in the repository. SWE-chat is by SALT-NLP (Baumann et al., COLM 2026), ODC-BY; thank you for publishing the commits and not just the conversations.

Written by Vitalii Bogachev — AI engineer working on LLM products in production: RAG, MCP servers, evaluation and reliability. Portfolio · GitHub