Your AI Coding Agent Is Paying to Read Code It Doesn't Need
AI coding agents have become surprisingly good at changing code.
But we are still giving them files as if they were human developers sitting in front of an editor.
And that is expensive.
Imagine your agent needs to rename one identifier inside a 2,000-line file.
The actual change might be:
oldName → newName
A few tokens of useful information.
Yet the typical agent workflow looks something like this:
read 2,000 lines
↓
reason about one relevant line
↓
edit the file
↓
read the file or diff again
↓
continue carrying much of that context
The agent may consume tens of thousands of tokens to perform an operation whose semantic payload is only a handful of tokens.
That feels backwards.
Read more: https://github.com/francescobianco/rap/
What if the agent never opened the file?
This is the idea behind RAP — Runtime Apply Patch.
Instead of asking an agent to read a file, locate something, rewrite it, and verify the result, RAP lets the agent address the code directly by its content.
For example:
rap s app.go 'oldName' 'newName'
RAP answers:
updated app.go (lines 1697 -> 1699, bytes +133, replacements 1)
That's it.
The agent did not need to load 2,000 lines into its context.
It already knew what it wanted to change.
RAP simply found it and performed the operation.
The principle can be summarized in three words:
Lines, not files.
Files are a human abstraction
Files are extremely useful for humans.
We open app.go, scroll around, understand the surrounding code, make a change, save it, and inspect the result.
But an AI agent does not necessarily need to reproduce that workflow.
If an agent already knows:
OLD = "oldName"
NEW = "newName"
why should it pay the context cost of reading another 1,999 lines?
The information required to describe the mutation and the information required to store the program are two very different things.
RAP tries to keep them separate.
The filesystem stores the program.
The model describes the mutation.
RAP connects the two.
Content becomes the address
Traditional editing often addresses code through coordinates:
file → line → column
But coordinates are fragile.
Insert ten lines above the target and the address changes.
RAP instead addresses code by content:
file → known text → replacement
The interesting consequence is that the agent often already possesses the address.
Maybe the target came from a previous grep.
Maybe it appeared in an error message.
Maybe the user provided it.
Maybe the agent wrote that code itself thirty seconds ago.
There is no reason to reload the entire file merely to rediscover information already present in the agent's context.
What if the target is ambiguous?
Then RAP stops.
$ rap s app.go 'count' 'total'
rap: OLD matched 3 times; use -all or a more specific OLD
Nothing gets written.
This property matters more than it might initially appear.
Shell tools can be extremely efficient for agents, but commands such as sed can also silently perform a perfectly valid operation on the wrong piece of text.
For an autonomous agent, failure is often cheaper than confidence.
A failed command may cost a dozen tokens.
A wrong edit can contaminate several subsequent reasoning steps before the mistake is discovered.
RAP therefore prefers an explicit refusal over an ambiguous mutation.
But sometimes the agent really does need context
Of course.
The argument is not that agents should never read source code.
The argument is that reading an entire file should not be the default unit of interaction.
RAP provides smaller operations for the questions agents actually ask.
Want to know whether a target is unique?
rap m FILE 'text'
Result:
matches: 1
Want to inspect only the affected area?
rap preview -n FILE 40 60 -- s OLD NEW
The agent receives a small window instead of the whole building.
Need to pass ugly text containing quotes, JSON, shell characters, or newlines?
rap q -token @file
RAP turns it into a safe token representation.
And after an edit, the command itself returns a receipt describing what happened.
No automatic full-file reread is necessary just to ask:
Did it work?
The token economics are surprisingly large
Consider a 2,000-line source file representing roughly 25,000 tokens.
A conventional agent might:
read file ~25,000 tokens
modify
inspect result ~25,000 tokens
Potentially around 50,000 tokens of context traffic around an edit that may involve only one line.
With RAP, the interaction can look more like:
rap s FILE OLD NEW
Perhaps ~40 input tokens and ~20 output tokens.
Obviously, real workloads vary. Agents sometimes genuinely need broader context, and not every edit can be expressed as a simple substitution.
But that is precisely the point.
Pay for context when reasoning requires context.
Do not pay for it merely because the code happens to live inside a large file.
This becomes more important as agents become more autonomous
Token efficiency is usually discussed as a model problem:
smaller prompts, better compression, larger context windows, cheaper inference.
But part of the problem may actually live outside the model.
Our tools still assume that an intelligent actor interacts with software primarily by opening documents.
Coding agents give us an opportunity to reconsider that assumption.
An agent does not necessarily need an editor.
It needs precise operations over a codebase, small observations about their consequences, and predictable failure modes.
That suggests a different interface between models and source code:
observe less
act precisely
verify cheaply
expand context only when necessary
RAP is a small experiment in that direction.
It does not try to make the model smarter.
It tries to make the model need fewer tokens to do the work it already knows how to do.
Lines, not files
The core idea behind RAP is intentionally simple:
The unit of work should be the information required for the change, not the container in which that information happens to live.
Sometimes that means a file.
Sometimes it means fifty lines.
Sometimes it means five.
And sometimes it means:
oldName → newName
If AI agents are going to execute thousands of filesystem operations autonomously, that distinction starts to matter.
A lot.
RAP is open source and available on GitHub.
Try giving your coding agent RAP and watch how often it can edit code without ever asking to see the whole file.