The Goldfish Brain: How AI fakes a memory it doesn't have, and why the illusion is so expensive

The Goldfish Brain: How AI fakes a memory it doesn't have, and why the illusion is so expensive

calendar_today ago • schedule7 min read
— Originally published at avanish-garg.hashnode.dev

EP:5 The Curiosity Crunch


You've been chatting with an AI for an hour. You brainstormed a product name, told it an embarrassing story from college, and cracked a joke about your landlord. Twenty minutes later, it calls back to that joke at exactly the right moment.

It feels personal. It feels like someone is listening, keeping notes, and getting to know you.

Nobody is. There's no one in there who remembers you.

Every time you press Send, the AI starts from zero. It has no idea who you are, what you said a minute ago, or that a conversation is happening at all. It's the most convincing magic trick in modern tech, and by the end of this episode you'll know how it works.

The Crunch in one sentence: an AI doesn't remember your conversation. Your chat app re-reads the whole thing to it, from the top, every single time.


1. The 50 First Dates problem

Models like ChatGPT, Claude and Gemini are stateless. A model is a giant file of numbers (its "weights") that was frozen when training ended. When you send a message, the model reads your text, produces a reply, and stops. Nothing about your conversation is written back into it. It never learns from the chat in real time, and it doesn't keep a file with your name on it.

If you've seen 50 First Dates, you know the setup. Lucy wakes up every morning with her memory wiped. Henry gets around this by showing her a video that catches her up on everything she's missed.

Your chat app is Henry, and it plays that video every time you speak.

When you type a new message, the app doesn't send only that message. It grabs:

  1. the hidden instructions the company gave the model ("you are a helpful assistant…"),

  2. everything you've said so far,

  3. everything the AI has said so far,

  4. your new message,

and glues them into one long document. That whole document goes to the model. The model reads it in a blink, writes the next reply, and forgets everything again.

Here's roughly what your chat app sends on your third message:

[system]    You are a helpful assistant.
[user]      Help me name my coffee startup.
[assistant] How about "Bean There"?
[user]      Lol, my landlord says "Bean There, Done That" is his catchphrase.
[assistant] Ha! Maybe that's a sign. Want a few variations?
[user]      Can you make it funnier?        <-- the only part that's actually "new"

The model sees all six lines as if for the first time. When it replies to "Can you make it funnier?", it can mention the landlord's catchphrase because the catchphrase is sitting right there in the text, not because it remembered anything.

Try it yourself. Use any AI's API (or a playground) and send only "Can you make it funnier?" with no history. You'll get a confused answer. The model is the same. The only thing missing is the transcript.


2. The conveyor belt (a.k.a. the context window)

That transcript can't be infinite. The maximum amount of text a model can read at once is called its context window.

Context windows are measured in tokens, which are chunks of text, roughly three-quarters of a word each. Early chatbots could handle a few thousand tokens, about a few pages. Modern models handle hundreds of thousands of tokens, and a few reach a million or more, which is about the length of several novels.

That sounds like plenty, but think of the context window as a conveyor belt with a fixed length. Every message you send and every reply you receive goes onto the belt. When the belt is full, something has to fall off the end. Depending on the app, one of three things happens:

What the app does What you experience
Truncates: drops the oldest messages The AI forgets the rule you gave it at the very start
Summarizes: compresses old messages into a short recap The AI remembers the gist but loses exact wording and details
Refuses: tells you the conversation is too long "Please start a new chat"

This is why, deep into a long chat, an AI suddenly breaks a rule you set in message one ("always answer in bullet points"). It didn't get lazy or stop listening. That message fell off the belt. If it isn't in the transcript, then as far as the model is concerned, it never existed.


3. "But ChatGPT does remember my name"

Fair objection. Some products now have a memory feature. They really do seem to know your job, your dog's name and your coding style across completely separate chats.

It's the same trick with better stage management.

The app saves short notes about you in an ordinary database, something like "User is a backend engineer. Has a dog named Biscuit. Prefers concise answers." The next time you start a chat, the app quietly pastes those notes onto the top of the transcript before sending it to the model.

The model still remembers nothing. It's just handed a cheat sheet, and the cheat sheet lives in the app, not in the AI.

That's the real distinction worth knowing:

  • The model is stateless. Same input in, same kind of output out.

  • The app around the model holds all the memory: chat history, saved notes, retrieved documents.

When people say "AI memory," they almost always mean the second thing.


4. The detective's corkboard (self-attention)

Now the puzzle. If the model has to re-read the entire transcript on every message, how does it reply so quickly, and how does it know that "it" in "make it funnier" refers to the landlord joke and not the startup name?

The answer is the breakthrough behind modern AI, self-attention, introduced in the 2017 paper "Attention Is All You Need."

Picture a detective in front of a corkboard covered in photos. Using red string, the detective connects pieces of evidence that relate to each other: this suspect to that location, that receipt to this timestamp.

Attention does the same thing with words. For every token in the transcript, the model asks, "Which other tokens matter most for understanding this one?" and draws a weighted red string to each of them. A thick string means "very relevant," and a hair-thin string means "mostly ignore."

For "Can you make it funnier?", the thickest string from "it" runs to the joke you told earlier. That's how the model works out what "it" means.

Here's the part that surprises people. Models don't read left to right the way you do. When they take in your transcript, they process all the words in parallel, which is exactly what GPUs are built for. That's a big reason the first part of the response (reading your whole history) is so fast.

(One honest footnote: when the model writes its reply, it does go one token at a time, left to right. That's why you see the answer stream onto the screen word by word. Reading is parallel and writing is sequential.)


5. The staggering cost of a PDF

Those red strings come at a price.

If every token connects to every other token, then a transcript with N tokens needs about N × N connections. That's quadratic scaling.

Transcript length Red strings to compute
1,000 tokens ~1 million
2,000 tokens ~4 million (2× the text, 4× the work)
10,000 tokens ~100 million (10× the text, 100× the work)
100,000 tokens ~10 billion

Doubling the document quadruples the work. This is why dropping a 100-page PDF into a chat is such a heavy lift, and why long-context models are expensive to run.

Engineers are fighting back. Real systems don't pay the full bill every time. Here's what they use:

  • KV caching. The model saves the intermediate math for the part of the transcript it already processed. When you send message #31, it only needs to work on what's new, not redo messages #1 to #30. (Providers now even sell "prompt caching" as a discount feature.)

  • FlashAttention and similar kernels. Smarter GPU code that does the same math using far less memory and time.

  • Sparse and sliding-window attention. Some models don't draw a string between every pair of tokens, only the ones that probably matter.

So the honest picture is more nuanced than "AI recomputes everything from scratch." Conceptually it re-reads everything, and in practice clever engineering skips the repeated work. But the underlying cost of long context is real, and it's why providers charge more for longer inputs and why "just give it a million tokens" isn't free.


6. So what? Practical takeaways

Understanding the trick makes you better at using the tool:

  1. Long chats get worse, not better. As the belt fills with old, irrelevant text, quality drops and rules get lost. Start a new chat for a new topic.

  2. Repeat what matters. If a rule is critical, restate it in your latest message. Don't trust that it survived 60 messages.

  3. Front-load a summary. Before a long session, paste in a short "here's who I am and what we're doing" block. You're doing manually what memory features do automatically.

  4. Don't treat "forgot" as "ignored you." Usually the instruction was truncated or diluted, not defied.

  5. Be careful what you paste. Everything in the chat is re-sent to the model on every turn. That includes anything private you pasted in at the start.


The Crunch (key takeaways)

  • No memory in the model. The AI is frozen and stateless. It lives in a permanent present tense.

  • The transcript trick. Your chat app re-sends the entire conversation, plus hidden instructions, with every message.

  • The conveyor belt. The context window is a fixed-size box. When it fills, old messages are dropped or summarized.

  • "Memory" features are sticky notes. The app stores notes about you and pastes them into the prompt. The AI never learns you.

  • Attention connects the dots. Red strings between tokens let the model understand "it" and "that joke," and they're why the cost grows quadratically with length.

  • The illusion is engineered. KV caching and smarter attention make it affordable, but it's still a lot of computation.


The next time an AI nails a detail you mentioned an hour ago, don't be flattered. It didn't remember you.

It just read the whole script again, faster than you can blink.

Until next time, keep crunching!!!


Enjoyed this? Forward it to the friend who says "my AI knows me." And tell us what to crunch next: reply with the topic you're most curious about.

Part 2 of 2 in The Curiosity Crunch
🔥 Join developers growing publicly
Share your knowledge, build in public, and grow your developer presence with a global community.

More Posts

Cisco's Amy Chang: A Model's "Passport" Doesn't Tell You Where It Actually Came From

Tom Smithverified - Aug 27

Your AI Doesn't Just Write Tests. It Runs Them Too.

Kevin Martinez - May 12

The Goldfish Brain: How AI fakes a memory it doesn't have, and why the illusion is so expensive

Avanish Garg - Oct 5

Local-First: The Browser as the Vault

Pocket Portfolio - Apr 20

Sovereign Intelligence: The Complete 25,000 Word Blueprint (Download)

Pocket Portfolio - Apr 1
chevron_left
148 Points • 4 Badges
India
2Posts
0Comments

Related Jobs

View all jobs →

Commenters (This Week)

4 comments
1 comment

Contribute meaningful comments to climb the leaderboard and earn badges!