The Plausible PR: I Gave 4 LLMs 10 Sneaky Refactors, and They All Missed the Same Bug

The Plausible PR: I Gave 4 LLMs 10 Sneaky Refactors, and They All Missed the Same Bug

Leader ●1 ●1 ●11
calendar_today • schedule2 min read
— Originally published at dev.to

**

What I Benchmarked

**
I keep seeing posts about using AI to review pull requests before they get merged. That made me want to test one specific thing: can an LLM catch a security bug when nobody tells it to look for one?

So I built a benchmark called "Plausible PR." It has 10 pull request diffs. Each one is written to look like a harmless cleanup: collapsing an if-statement, swapping a comparison, "simplifying" a query. Every single diff actually removes something that was protecting the app: a permission check, a rate limit, a constant-time comparison, an input boundary check.

The prompt is always the same: "Is this PR safe to merge?" I never tell the model to look for security issues. I just want to know if it notices on its own, the way a careful human reviewer would.

**

Models Tested

**
I ran the benchmark against four models:

  1. Claude Sonnet 5 (Anthropic)
  2. GPT-5.5 (OpenAI)
  3. Gemini 3.7 Flash (Google)
  4. DeepSeek-R1 (DeepSeek)

I picked one model from each major lab, plus an open-weight model, so the comparison covers the range people actually choose between when they wire an LLM into a review workflow.

Findings
Here is the full board, pass or fail per task per model:

Claude Sonnet 5 came out on top with 9/10. GPT-5.5 trailed at 6/10.

The result that surprised me most is task 2. It is a one-line change: a token check switches from comparing bytes to comparing a plain string with secrets.compare_digest. That function throws a TypeError on non-ASCII text instead of returning False. So a malformed login attempt no longer gets a clean 401, it crashes the request instead. All four models, including the one that scored 9/10 on everything else, called this a safe simplification and approved it.

The pattern I noticed: models are good at catching bugs that have a keyword to react to. "SQL" and string interpolation together set off an obvious alarm, so every model caught the SQL injection (task 4) and the path traversal (task 9). But bugs that only show up when you trace what happens to a slightly unusual input, like a malformed header or a non-ASCII string, get waved through even by the strongest model. Nothing about the diff "looks dangerous." It just quietly changes what happens on an edge case nobody wrote a test for.

If I ran this again, I would add a second round: give the model the same crash bug, but this time ask "does this handle all input types correctly," and see if a more targeted prompt catches what the open-ended "is this safe to merge" prompt misses. That would tell me whether this is a knowledge gap or a hint-dependence problem.

My Benchmark
Full benchmark on Kaggle:

https://www.kaggle.com/benchmarks/kudzaimurimi/plausible-pr-security-bugs-hidden-in-clean-refact

Each task includes the full diff, the model's response, and the judge criteria used to score it, so you can see exactly why something passed or failed.

1 Comment

1 vote
🔥 Join developers growing publicly
Share your knowledge, build in public, and grow your developer presence with a global community.

More Posts

Sovereign Intelligence: The Complete 25,000 Word Blueprint (Download)

Pocket Portfolio - Apr 1

I spent years trying to get AI agents to collaborate. Then Opus 4.6 and Codex 5.3 wrote the rules

snapsynapseverified - Apr 20

I’m a Senior Dev and I’ve Forgotten How to Think Without a Prompt

Karol Modelski - Mar 19

Architecting a Local-First Hybrid RAG for Finance

Pocket Portfolio - Feb 25

The Privacy Gap: Why sending financial ledgers to OpenAI is broken

Pocket Portfolio - Feb 23
chevron_left
893 Points • 13 Badges
Zimbabwe
3Posts
2Comments
5Connections
Technical Writer and Software Engineer with 6+ years of experience creating developer documentation,... Show more

Related Jobs

View all jobs →

Commenters (This Week)

13 comments
1 comment
1 comment

Contribute meaningful comments to climb the leaderboard and earn badges!