**
What I Benchmarked
**
I keep seeing posts about using AI to review pull requests before they get merged. That made me want to test one specific thing: can an LLM catch a security bug when nobody tells it to look for one?
So I built a benchmark called "Plausible PR." It has 10 pull request diffs. Each one is written to look like a harmless cleanup: collapsing an if-statement, swapping a comparison, "simplifying" a query. Every single diff actually removes something that was protecting the app: a permission check, a rate limit, a constant-time comparison, an input boundary check.
The prompt is always the same: "Is this PR safe to merge?" I never tell the model to look for security issues. I just want to know if it notices on its own, the way a careful human reviewer would.
**
Models Tested
**
I ran the benchmark against four models:
- Claude Sonnet 5 (Anthropic)
- GPT-5.5 (OpenAI)
- Gemini 3.7 Flash (Google)
- DeepSeek-R1 (DeepSeek)
I picked one model from each major lab, plus an open-weight model, so the comparison covers the range people actually choose between when they wire an LLM into a review workflow.
Findings
Here is the full board, pass or fail per task per model:

Claude Sonnet 5 came out on top with 9/10. GPT-5.5 trailed at 6/10.
The result that surprised me most is task 2. It is a one-line change: a token check switches from comparing bytes to comparing a plain string with secrets.compare_digest. That function throws a TypeError on non-ASCII text instead of returning False. So a malformed login attempt no longer gets a clean 401, it crashes the request instead. All four models, including the one that scored 9/10 on everything else, called this a safe simplification and approved it.
The pattern I noticed: models are good at catching bugs that have a keyword to react to. "SQL" and string interpolation together set off an obvious alarm, so every model caught the SQL injection (task 4) and the path traversal (task 9). But bugs that only show up when you trace what happens to a slightly unusual input, like a malformed header or a non-ASCII string, get waved through even by the strongest model. Nothing about the diff "looks dangerous." It just quietly changes what happens on an edge case nobody wrote a test for.
If I ran this again, I would add a second round: give the model the same crash bug, but this time ask "does this handle all input types correctly," and see if a more targeted prompt catches what the open-ended "is this safe to merge" prompt misses. That would tell me whether this is a knowledge gap or a hint-dependence problem.
My Benchmark
Full benchmark on Kaggle:
https://www.kaggle.com/benchmarks/kudzaimurimi/plausible-pr-security-bugs-hidden-in-clean-refact
Each task includes the full diff, the model's response, and the judge criteria used to score it, so you can see exactly why something passed or failed.