The closing line is the one that got me — communication was never the hard part. But I'd push the identity example a bit further: say you actually resolve it, and the HR ID, the GitHub handle and the full name all turn out to be the same person. Your agents now agree on who. They still don't agree on what's true about them. A shared knowledge graph isn't shared truth, it's a shared pile of claims, and every line in it looks equally settled by default — including the one some agent wrote at 2am off a stale cache. Which makes me think the thing that should be moving between agents isn't the fact at all, it's the receipt for it: who issued this, over exactly which bytes, when, and does it expire. Make that the unit of exchange and the graph stops being a heap of anonymous assertions. Though I'd be honest about what that buys you — a receipt like that proves origin and integrity, not honesty. Whoever holds the signing key can write whatever they want and seal it perfectly validly.
Still worth building. It's just not the same thing as "verified," and I suspect plenty of systems are going to ship that badge without ever noticing the difference.
AI Agents Have a Communication Protocol Now. They Still Don't Remember Anything.
7 Comments
@[Ali Ulu] That's a sharper distinction than the one I made, and it's the right one. "They agree on who" and "they agree on what's true about who" aren't the same claim, and I ran them together.
The receipt framing is a good move. A knowledge graph built from provenance — who issued this, over what exact bytes, when, does it expire — at least tells you where a claim came from and whether it's stale. That's more than most systems in production are tracking today. But you're right to flag the ceiling on it fast: a signature proves origin and integrity, not accuracy. A wrong or malicious agent with a valid key can produce a record that's perfectly verifiable and perfectly false. That distinction matters, because "signed" and "verified" are getting used interchangeably in a lot of vendor language right now, and they're not the same word.
Where I'd push it one step further: if the receipt is the unit of exchange, you still need a policy for what happens when two valid, signed, contradictory receipts show up about the same fact. In a system running hundreds of agents, that's not an edge case. That's Tuesday. Provenance tells you who to ask. It doesn't tell you who's right.
This is worth its own follow-up. Appreciate you taking the closing line somewhere I didn't.
@[Tom Smith] You've named the thing I'd built but never actually articulated, and that's a genuinely useful thing to have pointed out.
The contradiction case is handled in the system I work on, but until I read your comment I couldn't have told you why it's handled the way it is. Two conflicting claims don't produce a winner. The conflict gets typed — agent-vs-agent is a named case, not an exception path — both sides' evidence is carried forward, and the recommendation is flag. Never reject, never an automatic pick. It goes to a human triage surface.
I'd assumed that was a limitation I'd get around to fixing. Your framing makes clear it's the correct ceiling, not a gap: provenance tells you who to ask, not who's right. So the honest output of a trust layer isn't a verdict, it's an escalation — and a system that quietly resolves the conflict for you is doing the one thing it has no standing to do.
Worth noticing the field is called recommendation and not verdict. Past me apparently knew something present me hadn't put into words yet. That's the part you added.
The follow-up is worth writing. If you do, I'd read it closely.
Please log in to add a comment.
The thread above lands on flag to human triage, and I want to sit with that for a second, because it is the right call and it is also the part that breaks quietly.
A flag queue is an alert queue. Alert queues have a well documented failure mode: conflict volume grows with the number of agents, triage capacity stays flat because it is made of people, and the queue gets drained by whoever is on call that week rather than by whoever understands the claim. Six months in, somebody writes a rule that auto-resolves the noisy category, and the escape hatch you designed so carefully is now closed for the most common case. I have watched that happen with paging alerts more than once, and nothing about agents changes the arithmetic.
So the design question I would ask early is not how conflicts get resolved. It is what the system does during the window before anyone looks. If agent A and agent B disagree about a fact at 2am and triage happens at 9am, something has to serve reads for those seven hours. Blocking, last known good, and serving the claim with a contested marker attached are three very different products, and the choice gets made by default if nobody makes it deliberately.
The other lever is expiry, which Ali already named as part of the receipt. A claim with a short and honest TTL turns a whole class of conflicts into non-events, because the stale side ages out before anyone has to adjudicate it. That will not touch genuine disagreement, but a lot of what looks like disagreement in these systems is just one agent holding something longer than it had any right to.
Good thread. The move from what is true to who said it, when, and for how long is the useful one.
@[Mike Dabydeen] The paging analogy is dead on, and it's the piece I didn't push far enough. A flag queue only works if triage capacity scales with alert volume, and it never does. "Somebody writes a rule to auto-resolve the noisy category" isn't a hypothetical failure mode — it's the default outcome. I've seen that happen to an on-call rotation in six weeks, not six months.
The three-option framing — blocking, last known good, or serve-with-a-contested-marker — is the sharpest addition to this thread so far, because it names something the "flag to human" answer quietly skips past. Most teams that reach for human triage aren't actually picking a policy for those seven hours. They're picking whatever their database already does by default and calling it done. Serve-contested-but-marked is the cheapest to ship, so it's probably what most systems end up with. My guess is it's also the one most likely to get ignored downstream — an agent pulling a claim at 2am to make a decision isn't necessarily built to check whether that claim is wearing an asterisk.
The TTL point connects straight to what Ali raised. A receipt with an honest expiry doesn't just say who claimed something and when. It says when to stop trusting it without anyone having to notice the rot. That's the difference between a system that ages gracefully and one that needs a human to catch it failing.
Good add. This thread is turning into the real article.
@[Tom Smith] The asterisk getting ignored is the part I would build on, because I think it decides whether the marker is a control or just a comment.
A flag a consumer can ignore is a deprecation warning in a log nobody reads. It does not change behaviour, it just gives you something to point at afterwards. If contested state is supposed to matter, it has to change the shape of the response rather than add a field to it. A different status, or a result the caller has to unwrap before it can reach the value. Then ignoring it is no longer silent. It is code somebody had to write on purpose, and it shows up in review. That is roughly the difference between a system that degrades safely and one that degrades quietly, and the quiet kind is what you find yourself reading about in a postmortem.
The other thing I would stop treating as one policy is the choice between your three options. Blocking a read because two agents disagree is usually worse than serving something slightly stale. Letting an irreversible action through on a contested claim is not. So the answer is probably not one of the three at all. It is a mapping from the class of action to the behaviour, which means what the caller is about to do has to travel with the request. Which lands us back at labelling endpoints by risk, a genuinely unglamorous piece of work that keeps turning out to be the prerequisite for the interesting part.
Enjoying this one. Ali's receipt framing did a lot of the work here.
@[Mike Dabydeen] The control-versus-comment framing is the cleanest way I've seen this put. A field is optional to read. A different return type isn't. If unwrapping a contested result is the only way to get the value out, ignoring it stops being a default and becomes a decision, and decisions get reviewed. Fields don't.
Splitting the three options by action class is the right correction to what I wrote. I was treating "how do you handle disagreement" as one policy when it's actually a risk question wearing a technical costume. Serving something a few seconds stale to a dashboard is nothing. Letting a contested claim authorize a refund, a deployment, or a wire transfer is a different category of nothing. Collapsing those into one read policy was the mistake.
Labelling endpoints by risk sounds like the boring homework nobody wants to do before the interesting agent work starts. It's also apparently the thing that decides whether an asterisk means something or means nothing. That tracks with most of the infrastructure problems I've covered over the years. The unglamorous layer is usually the one holding everything else up.
This thread has more real design decisions in it than most vendor briefings I sit through in a month.
Please log in to add a comment.
Please log in to comment on this post.
More Posts
- © 2026 Coder Legion
- Feedback / Bug
- Privacy
- About Us
- Contacts
- You Tube
- Tiktok
- Premium Subscription
- Terms of Service
- Early Builders
That AI fluency comes from direct experience: I was one of the original six members of Google's Bard training team (now Gemini) and currently evaluate Meta's AI Business Assistant. I understand how these models work from the inside, which shapes how I write about them for a technical audience.
I specialize in LLM evaluation, prompt engineering, and RLHF methodologies, and I write about real-world implementation challenges — not theoretical possibilities. I attend major tech conferences to stay close to what developers actually face when deploying AI in production. Show less
More From Tom Smithverified
Related Jobs
- Live-Service Game Master (Spanish Speaking)Pearl Abyss · Full time · Netherlands
- BROADCAST ENGINEER (COMMUNICATIONS SYSTEMS SPECIALIST)State of Illinois · Full time · Springfield, IL
- ServiceNow - ServiceNow IT Operations Management (ITOM) Manager - Tech Cons - Open LocationErnst & Young Oman · Full time · Springfield, IL
Commenters (This Week)
Contribute meaningful comments to climb the leaderboard and earn badges!