I Taught AI to Review My Code So I Never Miss Leg Day Again

I Taught AI to Review My Code So I Never Miss Leg Day Again

●40 ●61 ●144
calendar_today ago • schedule7 min read

It's 10:48 p.m. on a Sunday. I'm folded over my desk like a question mark, reviewing a 1,400-line pull request titled "small cleanup."

My neck has developed opinions. My quads have filed a bug report titled "Not used since the last release." Status: won't fix.

Somewhere in this diff, around line 812, there's a bug. I know it's there the way you know there's a spider in the room. I just can't see it, because I've been reading for ninety minutes and my eyes have quietly switched from reading to scrolling with intent.

That was the night I made a rule: the code works for me, or it doesn't work here.

TL;DR: Your evenings aren't disappearing because you're slow. They're disappearing because AI made code cheap and review expensive. The fix is a four-layer pipeline where machines reject what machines can reject, an AI does the first pass, and you only spend human attention on the stuff that deserves it. Config snippets and a "steal this" checklist are at the bottom.

First, an Honest Word About the Title

Yes, I close pull requests in five seconds. Yes, I have done it from a plank. GitHub Mobile, shaking arms, steady conscience.

But I don't do it because a robot typed "LGTM." I do it because the previous four hours and fifty-nine minutes of review already happened, automatically, before the pull request ever reached my phone.

The five-second merge is the output of a system, not a leap of faith. A leap of faith is how you end up writing a postmortem.

Why Your Evenings Are Disappearing (It's the Volume, Not You)

Here's the part nobody put in the AI-coding marketing deck. Faros AI analyzed telemetry from more than 10,000 developers in 2025, and teams with high AI adoption completed about 21% more tasks and merged about 98% more pull requests. Great. But pull requests also got roughly 154% larger, and median review time went up about 91%, while organization-level delivery metrics stayed essentially flat.

DORA's 2025 research landed in the same neighbourhood. Roughly 90% of respondents now use AI at work, AI adoption is finally associated with higher throughput, and it's still associated with worse delivery stability. DORA's framing, which I can't improve on: AI is an amplifier. It amplifies a healthy delivery system and it amplifies a chaotic one, with equal enthusiasm.

Translation: the bottleneck moved. It used to be writing code. Now it's you, at 11 p.m., reading code that nobody typed. Congratulations, you're the world's most expensive linter.

So let's stop being one.

The Four Layers (a.k.a. How I Get to Leg Day)

Layer 0: Don't Review What a Machine Can Reject

The single biggest waste of human review time is arguing about things a tool can settle in four seconds: formatting, unused imports, type errors, a failing test, a leaked secret, a dependency with a known CVE. If a human has ever typed "nit: missing newline," your pipeline is broken, not your colleague.

Make all of it a required check that blocks the merge button. Here's a minimal, boring example for a Go project:

name: ci
on: [pull_request]

permissions:
  contents: read

jobs:
  checks:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4   # pin to a full commit SHA in real life
      - uses: actions/setup-go@v5   # same here
        with:
          go-version: stable
      - run: test -z "$(gofmt -l .)"
      - run: go vet ./...
      - run: go test -race ./...

About that pinning comment: if you read my supply-chain post, you know why. Tags are mutable, commit SHAs are not. Your CI pipeline is the most privileged thing in your repo, so treat its dependencies accordingly. Add a secret scanner and a dependency audit to the same job and you've deleted roughly half of what used to eat your Sunday night.

Layer 1: The AI First Pass (Your Tireless, Slightly Overconfident Junior)

This is where an AI reviewer earns its keep: summarizing what a 1,400-line diff actually does, catching the boring-but-real stuff (ignored error returns, missing nil checks, tests that don't test), and giving you a map before you enter the cave.

Rules I follow so it stays useful instead of noisy:

  • Tell it what you care about, in writing. Several AI reviewers read repo-level instruction files such as AGENTS.md or .github/copilot-instructions.md and treat them as review criteria. Check your tool's docs for the exact filename. Mine looks like this:
# Review rules
- Flag any new heap allocation in files under /hotpath.
- Flag any syscall wrapper that ignores its error return.
- Flag changes under /crypto that don't touch a test file.
- Do not comment on formatting. CI already handles that.
  • Expect false positives. Industry estimates for typical tools tend to land somewhere around 5 to 15%, and a tool that cries wolf gets muted within a week. Muted tools protect nobody.
  • Know the meter is running. Since GitHub's billing change this June, each Copilot review consumes both AI credits and Actions minutes. "Free robot reviewer" now comes with a taxi meter.
  • Run the Leg Day Test before you trust anything. Take a representative PR, seed it with four bugs you know about, and open a second, correct PR to see whether the tool invents a problem that isn't there. Count real catches and false alarms separately. Re-run it every few months, because vendors swap the underlying models and the reviewer you tested in spring is not necessarily the one running in autumn.

The AI is a junior who never sleeps. It is not the person who signs off.

Layer 2: Humans Review Intent, Not Syntax

With layers 0 and 1 done, your job shrinks to the part machines are genuinely bad at. Four questions, in order:

  1. Should this exist? Is this the right change for the actual problem?
  2. What breaks at 3 a.m.? Failure modes, rollbacks, observability.
  3. What's the blast radius? One endpoint, or the whole auth flow?
  4. What would an attacker do with this? (If you're a regular here, you already ask this one by reflex.)

Keep the diff small, because attention has a shelf life. The classic SmartBear/Cisco study of roughly 2,500 reviews found defect detection falls off sharply beyond a few hundred lines and after about an hour of continuous reviewing. A 1,400-line "small cleanup" isn't reviewable. It's a hostage situation. Send it back and ask for three PRs.

This is the layer that makes plank-merging safe instead of reckless:

# .github/CODEOWNERS
*            @you
/auth/       @you @second-pair-of-eyes
/crypto/     @you @second-pair-of-eyes
/.github/    @you @second-pair-of-eyes
  • Branch protection with required checks. Nothing merges red. No exceptions for "it's just a typo," because the typo is always in the config file.
  • Mandatory second human on risky paths: auth, crypto, payments, infrastructure, the CI config itself.
  • Auto-merge only for low-risk changes with green checks: docs, patch-level dependency bumps, generated files.
  • No special trust for AI-authored code. If anything it gets more scrutiny, since nobody on the team can say "I wrote it, I remember why."

Now the five-second approval isn't a gamble. It's a rubber stamp on a document that's already been notarized three times.

Discipline Is a Pipeline

Here's the philosophical bit, and then I'll let you go.

Willpower is a terrible architecture. It's a flaky test: passes when you're rested, fails at 11 p.m. on a Sunday, and nobody can reproduce it. Psychologists have studied this for decades. Precommitment and "if X, then Y" implementation plans beat raw motivation because they move the decision out of the moment where you're weakest.

A review pipeline is precommitment for your career. You decided once, in daylight, what "good enough to merge" means. After that, the system enforces it while you're tired, rushed, or hanging off a pull-up bar.

The gym works the same way. You don't "feel like" brushing your teeth every morning; it's a routine. Leg day is a CI job for your body: scheduled, boring, and fails visibly when you skip it. Last post I said sleep is a build step. This one's the sequel: the gym is your deploy window. Both are non-negotiable, and both die first when the code starts running you.

Where This Goes Wrong (Because Something Always Does)

  • Rubber-stamping. If the AI reviewer says "looks good" and you stop reading, you haven't automated review, you've automated not reviewing. Layer 2 exists so a human still owns the decision.
  • Giant PRs sneaking back in. AI makes it easy to generate 2,000 lines before lunch. Enforce a size limit, or review capacity collapses again and you're right back at the desk at 11 p.m.
  • Amplifier effect. A weak test suite plus a fast code generator plus an AI reviewer is just a faster way to ship bugs with a confident summary attached.

Steal This Setup

  • [ ] Formatter, linter, type-checker, tests, secret scan, dependency audit as required checks
  • [ ] CI actions pinned to commit SHAs, with minimal permissions
  • [ ] AI reviewer configured with written rules for your codebase
  • [ ] The Leg Day Test run on your own repo before you trust the tool
  • [ ] PR size limit (a few hundred lines, then split)
  • [ ] CODEOWNERS with a second human on auth, crypto, payments, infra and CI config
  • [ ] Auto-merge only for low-risk changes with green checks
  • [ ] A calendar block for the gym that gets the same respect as a production incident

The Point

Code should work for you, not the other way around. Automate the boring review work, spend your attention where judgment actually matters, and use the time you win back for the things that keep you functional: sleep, sunlight, and heavy objects that don't compile.

Your pull request queue will always be full. Your quads, on the other hand, are on a deadline.


Your turn: What's the most mind-numbing chore in your workflow that a pipeline could be doing instead? Drop it in the comments and let's crowdsource the automation.

And if you're reading this hunched over a diff right now: close the tab, stand up, do ten squats. The PR will still be there. Honestly, so will the bug on line 812. But your legs will be thrilled.

4 Comments

2 votes
1
2 votes
1
🔥 Join developers growing publicly
Share your knowledge, build in public, and grow your developer presence with a global community.

More Posts

I’m a Senior Dev and I’ve Forgotten How to Think Without a Prompt

Karol Modelski - Mar 19

Your Tech Stack Isn’t Your Ceiling. Your Story Is

Karol Modelski - Apr 9

Why Prompt Engineering Is Just an Expensive Way to Be Incompetent

Karol Modelski - May 21

Why “Building in Public” Is Hollowing Out Your Developer Career

Karol Modelski - Jun 18

TypeScript Complexity Has Finally Reached the Point of Total Absurdity

Karol Modelski - Apr 23
chevron_left
4.3k Points • 245 Badges
Canada • t.co/4fpTf3dL1D
38Posts
76Comments
12Connections
Writing ForgeZero: Fixing the mess of modern build systems.
Performance overhead is my personal ene... Show more

Related Jobs

View all jobs →

Commenters (This Week)

9 comments
1 comment

Contribute meaningful comments to climb the leaderboard and earn badges!