Every organization that lets an AI agent ship code has a sign-off step: a name on the release, a reviewer marking the pull request approved, a ticket moved to done. A human is accountable for that call, always. What most teams cannot answer is whethe...
Reliability in an AI agent is a harness property, not a model property.
The cleanest proof arrived at the bottom of the model-size ladder: a 688 MB model controlling a smart home, showcased by the model's own maker. The part worth studying is the 25...
The obvious way to improve a coding agent is to make it more capable: a stronger model, a wider context window, more tools, more room to act on its own. That is not where my problems come from. My agents seldom fail because they reason badly. They fa...
The AI agent used to be the star of every demo.
Now it's on the shutdown list. Not because the model got worse.
The most valuable asset in your AI program is in none of the quotes you ever signed.
A demo is a showroom. Good light, everything polis...
I changed one model string in ten cron jobs last night. 4.7 to 4.8. Then I went to bed.
The benchmark threads can wait until morning. My agents can't. They fire whether I'm awake or not: a briefing at 6:30, follow-up drafts at 11:45, a sync at 4 AM ...
Forrestchang's andrej-karpathy-skillshttps://github.com/forrestchang/andrej-karpathy-skills CLAUDE.md is four rules aimed at the moment Claude is writing code. They work. What they don't cover is the moment Claude is running. Once a Claude-driven pip...
Everyone's sharing their skill libraries right now. "Here are my 20 custom slash commands." "Check out my prompt template collection." "This skill saves me 2 hours a day."
I use skills too. I have about a dozen. They handle cover letters, content pi...