Interesting idea basically treating agent rules like product contracts instead of README notes. That distinction actually makes a lot of sense.
Beyond CLAUDE.md and AGENTS.md: when your coding agent needs a behavior spec
9 Comments
@[BlackSpecter] Thanks — yeah that's the gap exactly. README/instruction files describe how the agent should work; behavior specs describe what the product should do. Both useful, different layers. The interesting failure mode is when teams use the first as a substitute for the second and don't notice until something silently drifts in production.
Please log in to add a comment.
The layer distinction lands, and the trust-signal point especially.
One thing I keep running into: a house style rule of ours lived in CLAUDE.md for weeks, got violated anyway, and a commit hook stopped it the same day. Prose the model reads and a check that can fail are different instruments.
So on your third ceiling "No verification path": a .pbc.md is still Markdown the agent reads as context. What closes the verification gap in practice, a checker that parses the contract and fails CI, or is it still a human reading the diff?
@[jaafarabazid] You're right, and I won't dodge it: prose the model reads is not a check that can fail. Your hook did in a day what the file didn't do in weeks. That comparison is fair and the file lost it.
But I'd push back on the framing of the either/or, because the spec doesn't claim the job you're asking about. Its own three-line split:
PRD explains why.
PBC specifies what.
Code/tests/runtime prove how.
Proving is the HOW layer's job. Your hook is a HOW-layer instrument — a good one. What it had no way to point at was a stated WHAT.
So, concretely, two different things get checked:
The contract itself is typed and lintable, so a checker exists today — pbc validate . --fail-on-warnings fails CI on a dangling state transition, a behavior that lost its edge cases, a section that rotted while the code moved on. That's contract integrity.
Conformance stays with your tests, hooks and runtime. The format's contribution is that behaviors carry stable ids, so an external evidence system references the id instead of rewriting the contract, and the result comes back attached as provenance — kind: test|runtime|code, confidence: verified|inferred|assumed. One rule matters more than the rest: failing evidence never rewrites the contract. It downgrades the confidence and waits for a human to decide whether the promise changed or the code did.
Which is the real difference from prose, and it isn't Markdown vs Markdown: you can't verify against prose because prose has no addressable units. A behavior with a stable id does. Your hook could report against one; it can't report against a bullet in CLAUDE.md.
The part I'd stay honest about: deciding what the product promises is not automatable. Deriving that from the implementation is circular — the implementation is what you're checking.
Fair pushback @[Vinh Nguyen], and the addressable-units point is the one I'll keep. My hook only matches text, so it can't say which promise it's enforcing.
That showed today: it blocked a commit twice because a check in the same command searched for the em dash, while the commit message itself was clean. A check scoped to a stated behaviour ("commit messages contain no em dashes") would have passed it. A character match can't tell a violation from a mention.
How do you keep an id stable when one behaviour later splits into two?
@[jaafarabazid] I don’t think the ID itself has to survive the split.
If B1 turns out to contain two independent behaviors, I’d keep B1 in history and give the two new behaviors their own IDs, with explicit lineage back to it.
Something like:
B1 -> B2 + B3
The important part is that the split is recorded, not hidden by renaming or reusing IDs. Old evidence can still point to B1 at the revision where it existed, while new tests/evidence target B2 and B3.
So “stable ID” probably means stable identity while that behavior remains the same thing. Once its meaning genuinely changes or branches, preserving the lineage matters more than preserving the literal ID.
That settles it, thanks @[Vinh Nguyen]. Recording the split rather than reusing the ID is the part I was missing.
It matches what I have here: one promise, "no em dashes", is already two behaviours. One check runs on file writes, another on commit commands, same script, different scope, and the false positive came from the second one enforcing the first one's wording. So B1 -> B2 + B3, with the old evidence still pointing at B1.
The next thing I'd want is that lineage being machine-readable, so a failing check can name the behaviour it enforces rather than the characters it matched.
Please log in to add a comment.
Please log in to comment on this post.
More Posts
- © 2026 Coder Legion
- Feedback / Bug
- Privacy
- About Us
- Contacts
- You Tube
- Tiktok
- Premium Subscription
- Terms of Service
- Early Builders
More From Vinh Nguyen
Related Jobs
- Earn From Your Mobile Phone Usage - No Surveys RequiredNielsen Pulse · Full time · Warsaw, NC
- Earn From Your Mobile Phone Usage - No Surveys RequiredNielsen Pulse · Full time · Vevay, IN
- Earn From Your Mobile Phone Usage - No Surveys RequiredNielsen Pulse · Full time · Italian Republic
Commenters (This Week)
Contribute meaningful comments to climb the leaderboard and earn badges!