ProtocolAI and GrammarAI: Giving AI a Language Software Can Actually Own

ProtocolAI and GrammarAI: Giving AI a Language Software Can Actually Own

●2 ●41 ●80
calendar_today ago • schedule9 min read

There is a particular problem that keeps appearing as I build the architecture for The Singularity Workshop.

It starts with something that sounds almost trivial:

How does software communicate with an AI when the software already knows what its world means?

An LLM is very good at language.

It can interpret a request, generate possibilities, describe relationships, and produce something that looks structurally correct.

But there is a boundary where the application has to take over.

The application needs to know exactly what something refers to.

And then it needs to know how those things may be connected.

That is where two small packages have emerged:

  • ProtocolAI — the WHAT
  • GrammarAI — the HOW

They are deliberately separate.

And neither one is the AI.


The Problem Starts With Probability

An LLM operates in a probabilistic space.

That is not a criticism. It is the nature of the technology.

Given some context, the model produces a candidate response.

That candidate may be exactly what we wanted.

It may also be something new.

Or something that looks familiar but does not actually identify the same object the application already knows.

Imagine an application that already knows three people:

Bob
Jane
Sara

The application has meaning attached to those names.

The model does not own that meaning.

It can produce "Bob".

But the application needs to turn that representation into something deterministic.

That is the boundary ProtocolAI explores.

probabilistic language
        |
        v
     candidate
        |
        v
    ProtocolAI
        |
    +---+---+
    |       |
  known   unknown
    |       |
    v       v
 integer   literal
    |       |
    +---+---+
        |
        v
application-owned identity

The model proposes.

The application decides.

ProtocolAI provides the semantic address space between those two worlds.

Visualization of the string/integer mapping of the ProtocolAi


ProtocolAI: The WHAT

The core idea behind ProtocolAI is surprisingly small.

A protocol owns a vocabulary.

For example:

[1001] People

[2001] bobId  = "Bob"
[2002] janeId = "Jane"
[2003] saraId = "Sara"

Now the application has an address space.

"Bob"   -> [2001]
"Jane"  -> [2002]
"Sara"  -> [2003]

The integer is not the meaning.

It is the address of the meaning.

That distinction is important.

ProtocolAI does not replace the domain object.

It does not become the database.

It does not become the authority that decides what Bob is allowed to do.

It simply establishes a deterministic identity that the application owns.


The Literal Escape Hatch

There is another important detail.

What happens when the model produces something the application does not already know?

Suppose the vocabulary contains:

[2001] Bob
[2002] Jane

and the model produces:

"Amelia"

The protocol should not have to pretend Amelia already exists.

So ProtocolAI preserves the literal.

[2001] [2002] "Amelia"

That gives the host a choice.

The host can decide whether Amelia should be:

  • created;
  • registered;
  • rejected;
  • authorized;
  • deferred;
  • or otherwise handled.

That is a much more interesting boundary than simply forcing everything into a closed enumeration.

Known things become references.

Unknown things remain visible as unknown.


This Is Not Just About Compression

It would be easy to look at integer identities and conclude that this is merely an optimization.

Smaller values.

Smaller payloads.

Faster lookup.

Those may eventually be useful properties.

But that is not the fundamental idea.

The more important separation is:

Meaning
   |
   v
Identity
   |
   v
Representation

Consider:

Meaning:        Bob
Symbol name:    bobId
Symbol ID:      2001
Protocol ID:    1001
Payload form:   [2001]

The application already owns the meaning.

ProtocolAI gives that meaning an explicit address.

That address can then survive across interactions.

Prompt
"Bob"
   |
   v
[2001]
   |
   v
Response
[2001]
   |
   v
Next interaction
[2001]

The model does not have to repeatedly rediscover that two strings refer to the same application-owned identity.

Again, that is an architectural direction rather than a claim that integer identities automatically make an LLM reason better.

The point is that the application has a semantic address space.


And Then I Ran Into the Next Problem

A vocabulary is useful.

But a vocabulary is not a language.

Knowing:

[2001] Bob
[2002] Jane

does not tell us how those identities may be arranged.

We need another layer.

Something has to describe relationships such as:

this kind of thing
    |
    +--> may contain this
    |
    +--> may reference that
    |
    +--> may be followed by another structure

That is where GrammarAI enters.

Visualization of the GrammarAi


GrammarAI: The HOW

GrammarAI takes the identities supplied by ProtocolAI and gives them a structural language.

The distinction is intentionally simple:

ProtocolAI defines the WHAT. GrammarAI defines the HOW.

For example:

ProtocolAI

[1001] People

[2001] Bob
[2002] Jane

GrammarAI can define a structure that references those identities:

[4001] -> [1001:2001]
[4001] -> [1001:2002]

Now we have two different owners.

[1001:2001]
    |
    +--> ProtocolAI owns the meaning

[4001]
    |
    +--> GrammarAI owns the structural position

That separation is deliberate.

GrammarAI does not copy Bob into itself.

It references the protocol that owns Bob.


A Grammar Is a Relationship System

This is the part I find particularly interesting.

A grammar is not another dictionary.

It is a relationship system.

It can say:

Grammar nonterminal
       |
       v
     [4001]
       |
       +----> [1001:2001]
       |
       +----> [1001:2002]

The grammar owns the relationship.

The protocol owns the thing being referenced.

That gives us a clean composition boundary.

Protocol A
    |
    +-- symbols

Protocol B
    |
    +-- symbols

       \       /
        \     /
         \   /
        Grammar
           |
           +-- structure
           +-- relationships
           +-- references

The grammar connects things without becoming the owner of those things.


Why Integer-Backed Structure?

There is another common thread between the two packages.

Both are self-defining.

A GrammarAI definition has its own identity:

[3001] Greeting

It has a start symbol:

start = [4001]

And it has production rules:

[5001] [4001] -> [1001:2001]
[5002] [4001] -> [1001:2002]

That means the grammar itself can become a piece of data.

It can eventually be:

  • described;
  • serialized;
  • compared;
  • versioned;
  • generated;
  • visualized;
  • translated;
  • supplied to another composition layer.

The grammar does not need to know which of those things will happen.

That is the point of keeping it small.


The AI Still Isn't the Authority

This is probably the most important part of the architecture.

The LLM does not become the owner of the protocol.

The LLM does not become the owner of the grammar.

And the LLM does not become the application's execution authority.

The flow is closer to:

              LLM
               |
          probabilistic
            candidate
               |
               v
          ProtocolAI
             WHAT
               |
               v
           GrammarAI
              HOW
               |
               v
        host validation
               |
               v
       application decision
               |
               v
       deterministic state

The model participates.

It does not rule the system.

That distinction becomes increasingly important as AI systems begin interacting with real applications rather than simply generating text.


This Is the Same Boundary I Keep Finding Elsewhere

Consider a reservation system.

User A requests the last available table.

The centralized authority accepts the reservation.

Users B through K arrive almost immediately afterward.

The system does not need to tell User K:

Failure.

That is technically accurate but semantically terrible.

The actual state is:

You cannot acquire the resource right now.

So the system can return something closer to:

Resource unavailable.

You are currently #10 in the waiting queue.
There are 9 requests ahead of you.

That is an inability, not necessarily a failure.

The system has preserved useful information about the state of the world.

And that is the kind of distinction these AI packages are trying to preserve as well.

An LLM can produce a candidate.

ProtocolAI determines whether that candidate corresponds to a known application identity.

GrammarAI determines whether that identity participates in a valid structure.

The host then decides what that structure means operationally.

At every step, we preserve the distinction between:

what was proposed
what is known
what is structurally valid
what is currently possible
what the application actually decides to do

That is much more useful than collapsing everything into true or false.


ProtocolAI and GrammarAI Are Deliberately Not an AI Framework

This is where the package boundaries become important.

ProtocolAI is not:

  • an LLM;
  • an AI provider;
  • a prompt framework;
  • a schema validator;
  • a database;
  • a grammar compiler;
  • a command executor;
  • a GUI framework;
  • a MicroBundle host.

GrammarAI is not:

  • an LLM client;
  • model selection;
  • tokenization;
  • inference;
  • prompt transport;
  • protocol vocabulary ownership;
  • MicroBundle hosting;
  • REST/OpenAPI;
  • GUI manifestation;
  • tool execution.

That list may look like a lot of things to leave out.

That is intentional.

The packages are supposed to provide semantic infrastructure, not become another giant AI framework.


The Exchange Layer Comes Later

ProtocolAI already describes an interesting future boundary for AI interaction.

The same semantic exchange should eventually work through either:

             AI Exchange
                 |
        +--------+--------+
        |                 |
    Clipboard         Provider
        |              Adapter
        |                 |
        +--------+--------+
                 |
                LLM

Clipboard interaction is not treated as a second protocol.

It is simply another transport.

A developer could extract:

Protocol
Grammar
Context
Request

copy that into an LLM, receive a response, and paste it back.

A connected provider could perform the same semantic exchange programmatically.

The transport changes.

The semantic boundary does not.

That is an important distinction.


Context Is Not Authority

An AI exchange might eventually contain a substantial amount of information:

current Experience
selected objects
available commands
protocol descriptions
grammar definitions
current state
previous interactions
host capabilities

But none of that means the model has authority.

The intended boundary remains:

MODEL OUTPUT
     |
     v
parse / validate
     |
     v
ProtocolAI resolution
     |
     v
host policy
     |
 +---+---+---+---+
 |   |   |   |   |
accept reject clarify create execute

The model proposes.

The host decides.

ProtocolAI makes owned meaning addressable.

GrammarAI makes the relationships between those meanings describable.


And This Is Where the Larger Workshop Architecture Appears

The two packages are small individually.

But together they establish another layer in the larger architecture I have been building.

Domain meaning
      |
      v
ProtocolAI
    WHAT
      |
      v
GrammarAI
    HOW
      |
      v
AI / provider adapter
      |
      v
Host
      |
      v
FSM / GUI / Experience

And eventually, those capabilities can participate in the MicroBundle architecture.

That means the AI layer does not have to become a special universe disconnected from the rest of the system.

It can become another collection of capabilities that the host can compose.


The Interesting Part Is the Separation

There is a temptation when building AI software to put everything in one place.

The model client.

The prompts.

The schemas.

The grammar.

The commands.

The domain objects.

The execution engine.

The UI.

The credentials.

The networking.

Eventually you have an enormous framework that knows everything about everything.

I am deliberately going in the other direction.

ProtocolAI knows about semantic identity.

GrammarAI knows about semantic structure.

The host knows about execution.

The provider adapter knows about transport.

The application knows about meaning and authority.

That makes the pieces composable.


Where the Alpha Actually Stands

These packages are intentionally early.

ProtocolAI 0.1.0-alpha.2 currently establishes:

  • self-defining integer-backed vocabularies;
  • protocol-qualified identity;
  • known-value encoding;
  • literal preservation;
  • deterministic descriptions;
  • ordered mixed payloads;
  • payload validation;
  • mixed reference/literal resolution.

GrammarAI 0.1.0-alpha.1 currently establishes:

  • grammar identity;
  • integer start symbols;
  • integer nonterminals;
  • ordered production rules;
  • external ProtocolAI references;
  • deterministic self-description;
  • validation of referenced nonterminals.

Neither package currently implements the entire AI pipeline.

And that is okay.

The unanswered questions are actually interesting architectural questions.

For example:

Who allocates new symbol IDs?

How do identities persist?

How are protocols versioned?

How are protocols negotiated?

How are grammars versioned?

How are recursive productions handled?

How does an abstract grammar become a provider-specific constraint?

How do we validate generated model output?

How does a semantic artifact become executable behavior?

Those questions should be solved by the architecture.

They should not be hidden inside convenience APIs simply because the API would look nicer.


The Bigger Experiment

What I am really exploring here is a progression.

human / domain language
          |
          v
    self-defining lexicon
          |
          v
     integer identity
          |
          v
    grammar / structure
          |
          v
       protocol
          |
          v
     tool execution
          |
          v
       experience

ProtocolAI is the first semantic compression boundary.

GrammarAI is the next structural boundary.

Neither is trying to replace human language.

Neither is trying to make an LLM deterministic.

Instead, they give software something it has always needed:

a precise address space for meanings that the software itself owns.

And once those meanings have addresses, and those addresses have structure, we have something much more interesting than an AI prompt.

We have the beginning of a language that the application can actually own.


The Current Workshop Stack

The current direction now looks something like this:

                    LLM
                     |
              probability
                     |
                     v
               ProtocolAI
                  WHAT
                     |
                     v
                GrammarAI
                   HOW
                     |
                     v
             AI / Host Adapter
                     |
                     v
                  Host
                     |
                     v
                FSM_COS
                     |
              MicroBundles
                     |
                     v
              RuntimeAssembly
                     |
                     v
            GUI / Experience
                     |
                     v
             Digital Reality

The interesting part is that the model remains at the edge.

The application owns the center.

That is where I want the architecture to live.

Resources



Support The Singularity Workshop

If this work is useful to you, and you'd like to help keep The Singularity Workshop moving, you can support its continued development:

Every bit of support helps provide more time and resources for building, testing, documenting, and publishing the architecture.


The Singularity Workshop

The Singularity Workshop — Tools for the curious, the bold, and the systemically inclined.
Because state shouldn't be a mess.

🔥 Join developers growing publicly
Share your knowledge, build in public, and grow your developer presence with a global community.

More Posts

Cisco's Amy Chang: A Model's "Passport" Doesn't Tell You Where It Actually Came From

Tom Smithverified - Aug 27

The Zero-Net-Loss Fleet & The Mercenary Squad: A Live AI Economy

DEVPlank - Aug 4

Your App Feels Smart, So Why Do Users Still Leave?

kajolshah - Feb 2

I’m a Senior Dev and I’ve Forgotten How to Think Without a Prompt

Karol Modelski - Mar 19

I Built a Serializer That Refuses to Own Your Data

The Singularity Workshop - Sep 28
chevron_left
3k Points • 123 Badges
37Posts
19Comments
15Connections
Architecting the Future of Computation through Relentless Optimization.

The Singularity Workshop is... Show more

Related Jobs

View all jobs →

Commenters (This Week)

4 comments
1 comment
1 comment

Contribute meaningful comments to climb the leaderboard and earn badges!