I Spent a Month Being Deliberately Lazy With AI. Here Are the Receipts.

I Spent a Month Being Deliberately Lazy With AI. Here Are the Receipts.

●34 ●52 ●128
calendar_today ago • schedule9 min read

For one month, every working day, I ran an experiment on myself. Not at my job, obviously. On my own projects, in the languages I can actually defend in a code review.

The rules were simple and, I'll admit, a little pathetic. Write a prompt. Sit back. Drink coffee. Click Keep Changes. Repeat until a working product appears or my will to live leaves the building, whichever comes first.

The goal wasn't to prove AI is bad. The goal was to find out what happens when a senior engineer voluntarily turns off the part of his brain that usually says "wait, that's wrong."

Spoiler: the coffee got cold. A lot.

I ran the experiment across three projects. One of them a fair number of you already know. Here's how each one went.

Project One: ForgeMed-AI (Score: 3/10)

ForgeMed-AI is my medical AI project. I never presented it publicly. I only built it, quietly, and I'm telling you about it now for the first time.

The stack was a proper polyglot circus:

  • Go for the backend, handling the interaction between the model and the website.
  • Python for training the model.
  • C for a driver that regulates the temperature of the actual device the neural network trains on.
  • JavaScript for the frontend.

Yes, a hardware thermal driver. If you've never written code where a mistake means a physical object gets warmer than it should, let me paint you the actual picture: a bug in a C thermal driver is the moment your device gets hot enough to fry an egg on the casing while the room starts smelling faintly of scorched PCB and your own dying nerves. Nothing focuses a code review quite like the low, distant fear that you're one off-by-one away from explaining to your landlord why there's a burn mark on the desk. It also means "looks plausible" is not a standard of correctness you're allowed to use.

The verdict: the model itself was decent, and the code it produced was not. It wrote a remarkable amount of confident nonsense. I spent the month saying "redo this" and "fix that," on a loop, like a man politely asking a vending machine to reconsider. A single file could eat several hours. The C was the worst of it, so I gradually started writing that part myself.

And then, at the end of the month, I did something I recommend to no one who values their weekend: I deleted all of the AI-written source and rebuilt the project by hand.

Was that faster than fixing it? Ask me again when the hangover from that decision wears off. But I understand every line now, which is the entire difference between owning a system and merely being its custodian.

(If you'd like a full post about ForgeMed-AI, say so in the comments. If enough of you ask, I'll write it.)

Project Two: ForgePanel (Score: 1/10)

I got interested in the whole modern proxy scene: VLESS with REALITY, Trojan, Shadowsocks. The idea was to build a panel for spinning up your own VPN on top of the Xray core. Vite on the frontend, Go on the backend, because I am a man of simple tastes.

The AI couldn't do the frontend. Not "did it badly." Could not do it. At one point it spent what felt like three actual hours trying to centre a single div, gave up on CSS as a concept entirely, and just started hallucinating Tailwind classes that don't exist — confidently sprinkling justify-center-super and its cousins across roughly five hundred lines, as if enough made-up utility classes would eventually summon the correct one out of sheer volume. It did not.

One out of ten, and that point is a courtesy.

Project Three: ForgeZero (Score: 6/10)

This is the one that matters, because this is where I couldn't afford to be lazy.

For a month I used the model on ForgeZero, and the "sit back and click Keep Changes" plan died on contact with reality. Every line of generated code went through my verification process, which I will affectionately describe as surgical and which the model would describe, if it had feelings, as hostile.

I ran tests on different hardware. I spun up virtual machines to exercise different environments. Nothing was trusted until it had survived contact with a machine that didn't care about its feelings.

What I found, in order of how much it hurt:

Documentation: bad. This one genuinely stings. It's a language model. Writing prose about code is supposed to be the home game. Instead, I had to feed it my own facts and my own code, and it still produced documentation I had to rework. The task I gave it was simple: describe what's here, from the material provided. It found a way to make that hard.

Go: rewritten roughly fifty times. I'm not being poetic. Some pieces of Go went through so many revisions that I started to lose track of which version had been the original crime.

Plan 9 assembly: a special circle of difficulty. For anyone who hasn't met it: Plan 9 assembly is the assembler dialect used inside the Go toolchain. It traces back to the assembler of the Plan 9 operating system from Bell Labs. It is quirky by design. Operand order is the reverse of what most people expect, registers like FP, SP and SB are pseudo-registers with their own rules, and function declarations carry frame-size and argument-size annotations that must agree with reality or things go wrong in ways that are deeply unpleasant to debug. Go's go vet has an assembly checker precisely because it is so easy to get these subtly wrong. Now add different architectures, ARM and x86-64, each with its own calling conventions and instruction sets, and you have a task where "it looks about right" is the fastest way to a mystery crash.

I've said it before: the model is weak at C, and worse at assembly, particularly across platforms. My suspicion, and this is my read rather than a measured fact, is that there simply isn't much high-quality Plan 9 assembly in the world to learn from. A model is only as fluent as what it has seen.

Tests: quantity over insight. I asked it to cover every module, and I'll concede I may have been a harsh taskmaster about it. But the pattern was baffling. It would write tests for a module, and then write more tests for the same module that exercised the same behaviour under a different name. Same input class, same code path, different function name. If the scenarios differ, wonderful, that's coverage. If they don't, that's not a test suite, it's a word count.

Coverage percentage is a number you can inflate. Confidence is not.

It's afraid of hardcore code. This was the most interesting failure, so let me give it room.

My philosophy for ForgeZero is zero allocations. When I see allocations in a hot path, I lose my composure. So I forced the model to work at the lowest level available: manual byte iteration, no standard library where it could be avoided, and my own implementations instead. That's the philosophy that eventually led me to replace BLAKE3 in the non-cryptographic path with my own hash, BB64 (Boom-Boom 64), alongside Boom-Boom Hash (BB), the cryptographic one built on BLAKE3.

The model kept flinching. Left alone, it reaches for the idiomatic, comfortable, well-trodden option, because that's what the overwhelming majority of code it learned from does. Here's a deliberately tiny illustration of the gap. This is not from ForgeZero, just the shape of the argument:

// What the model reaches for first: readable, idiomatic, allocates.
func countLines(b []byte) int {
return len(strings.Split(string(b), "\n"))
}

// What the zero-allocation philosophy demands: one pass, no garbage.
func countLines(b []byte) int {
n := 1
for _, c := range b {
if c == '\n' {
n++
}
}
return n
}

Both are correct. Only one of them makes a benchmark say 0 B/op. Idiomatic code is the statistical centre of everything the model has read, and "zero allocations at any cost" lives on the far tail of that distribution. You have to drag it out there, kicking, by writing long prompts that spell out the how as well as the what.

It is pathologically polite about its own crimes. This deserves its own line because it happened constantly and it never stopped being funny. It would hand me two hundred lines of allocation-riddled, half-working garbage with the emotional confidence of someone who just cured a disease: "Certainly! Here is the perfect solution for your use case." Not "here's an attempt," not "this might need tuning." Perfect. Delivered with a smile you can somehow hear through plain text. I started treating that specific sentence as a warning label, the software equivalent of a smoke detector going off two rooms away.

Which brings me to the real cost of all this.

The prompts were long. Not "write me a function" long. Multi-paragraph, task-decomposed, here's-the-constraint, here's-why, here's-what-I-already-tried long. At which point a fair question presents itself: if I have to specify the implementation this precisely, how much of the work did the model actually do?

The Part Where Science Shows Up Uninvited

If you're thinking "maybe I'm just bad at this," fair enough, but consider the most-discussed data point in this whole debate.

In July 2025, the research nonprofit METR ran a randomized controlled trial: 16 experienced open-source developers, 246 real tasks on mature codebases they already knew well. Some tasks allowed AI tools, some didn't. The result was that the developers took 19% longer with AI.

Here's the part that should really bother you. Before the study, they predicted AI would make them 24% faster. After the study, having been measurably slower, they still estimated they'd been about 20% faster.

Their self-assessment was off by roughly forty percentage points, and they didn't notice.

Now, honesty requires the caveats, because I dislike people who quote one study like scripture. It was a small sample. The tools have moved quickly since then. METR itself has since said it's changing its follow-up design, and later updates noted selection effects, since some invited developers declined to participate if they couldn't use AI at all. Other research on large organizations finds real gains in other settings, particularly on well-scoped work.

So the honest reading is not "AI makes you slower." It's narrower and more useful: on complex code you already understand deeply, the tool's help can cost more in verification than it saves in typing, and your own feeling of speed is a terrible instrument for measuring it.

That matches my month almost uncomfortably well. Where I understood the code deeply and the standards were high, I spent my time reviewing, correcting and re-explaining. It felt productive, because something was always happening on screen. Whether it was faster is a different question, and I only trust the answer I can measure.

What I'd Actually Do Differently

After a month of self-inflicted evidence, here's my working rulebook.

  • Use it hardest where mistakes are cheapest. Boilerplate, scaffolding, throwaway scripts, a first draft you fully intend to rewrite. The cost of a wrong answer there is a shrug.
  • Use it least where mistakes are invisible. Hardware drivers, assembly, concurrency, anything where wrong looks exactly like right until it isn't. In those places, the verification cost is the whole cost.
  • Treat every generated line as a pull request from a stranger. Confident, fast, occasionally brilliant, and with no stake in what happens at 3 a.m. Review it accordingly.
  • Write the tests yourself, or at least design them yourself. A model that writes the code and the tests that judge the code is grading its own homework.
  • Never let the tool set the standard. If your standard is zero allocations, say so, enforce it with a benchmark, and let the machine fail the benchmark as many times as it takes. 0 allocs/op is not a matter of opinion.
  • Measure, don't feel. If you can't show a before-and-after number, you don't know whether it helped. You just know it was busy.

The Verdict

Scores again, for the record: ForgeMed-AI 3/10, ForgePanel 1/10, ForgeZero 6/10.

Notice which project scored highest. It wasn't the one where I relaxed. It was the one where I refused to.

That's the whole lesson, and it's a little annoying because it's so unglamorous. The tool got more useful the less I trusted it. The moment I stopped treating it as a colleague and started treating it as a fast, tireless, overconfident intern with an encyclopedic memory and no sense of consequence, the results improved. Not to magic. To six out of ten, which, for the record, is a perfectly respectable grade for an intern.

The dream of sitting back with coffee while the machine ships your product is real for a certain class of task. For the code that matters most, the coffee goes cold, the diff gets read line by line, and the engineer stays employed.

Which, if I'm being honest, I find deeply reassuring.

Want the full ForgeMed-AI story, thermal driver and all? Tell me in the comments. And if your own AI experiment went better than mine, I genuinely want to hear how, and more importantly, how you measured it.

🔥 Join developers growing publicly
Share your knowledge, build in public, and grow your developer presence with a global community.

More Posts

Frameworks Are Institutional Memory

Ken W. Algerverified - Sep 17

I’m a Senior Dev and I’ve Forgotten How to Think Without a Prompt

Karol Modelski - Mar 19

Sovereign Intelligence: The Complete 25,000 Word Blueprint (Download)

Pocket Portfolio - Apr 1

I spent years trying to get AI agents to collaborate. Then Opus 4.6 and Codex 5.3 wrote the rules

snapsynapseverified - Apr 20

The Sovereign Vault — A Comprehensive Guide to Protocol-Driven AI

Ken W. Algerverified - Jun 4
chevron_left
3.9k Points • 214 Badges
Canada • t.co/4fpTf3dL1D
34Posts
65Comments
12Connections
Writing ForgeZero: Fixing the mess of modern build systems.
Performance overhead is my personal ene... Show more

Related Jobs

View all jobs →

Commenters (This Week)

2 comments
1 comment
1 comment

Contribute meaningful comments to climb the leaderboard and earn badges!