PrismML’s 2-billion-parameter 1-bit Bonsai model moves vision, memory
and reasoning onto the device itself, and shows why quantization might
be the real key to practical local wearable AI.
A few months ago, I built a small project for myself, just to experiment with things out of curiosity. Smart glasses that could remember things I’d seen and let me search through those memories later, all without touching the internet, completely offline. I used Gemini’s embedding model and a local vector database called Qdrant Edge to make it work. And the whole time I was building it, the hardware just kept pushing back.
Every time I wanted the system to actually think, to reason about what it was looking at instead of just storing it, I needed a model. And every model I tried either needed a cloud API call or ate up more memory than the device I was testing on could handle.
That is the catch with building on-device AI. Storing and searching through data locally (i.e. retrieval) is something you can solve with clean code and clever engineering. But making sense of that data (i.e. reasoning) requires a real model sitting right on the chip, and most models are simply too big.
So, when I saw PrismML’s new Bonsai model running locally on Qualcomm’s Snapdragon AR1 chip, on actual smart glasses, I didn’t read it as another AI announcement. I read it as someone solving the exact problem I was stuck on.
Here is a clear look at what actually changed, and why it matters more than the headline number suggests.
Why “on-device AI” has mostly been a marketing phrase until now
Smart glasses are one of the hardest places to run AI. The software is manageable, but the physical constraints are brutal. A phone can afford a chip that runs warm without anyone noticing. Glasses sit on your face. They can’t get hot, and the battery has to fit inside a frame thin enough that you’d wear it in public.
That’s why most “AI glasses” on the market today are really just a camera and a microphone with a Bluetooth connection to your phone. The phone talks to the cloud, the cloud does the thinking, and the answer gets piped back to a tiny speaker near your ear. It works, but it’s not local intelligence. It’s a remote control for a server somewhere else. And that means no internet, no assistant and slow internet, slow assistant. Send an image to the cloud, and now you’ve also sent your surroundings to a company’s servers.
PrismML’s approach is different because it’s not trying to make the cloud faster. It’s trying to make the model small enough that the cloud becomes optional.

What “1-bit” actually means, in plain words
The Bonsai model PrismML built for smart glasses has 1.7 billion parameters in its language model, plus a smaller 0.3 billion parameter vision component that lets it understand what the camera sees. Together that’s about 2 billion parameters, running entirely on the glasses’ own chip.
Normally, a model that size would never fit on something this small. Here’s why it does.
Every parameter in a neural network is a number. Normally that number carries a lot of precision, something like 16 or 32 bits, so the model can represent fine shades of value. PrismML’s model is “1-bit,” which means each parameter gets squeezed down to one of a few values, often -1, 0, or +1, instead of a long decimal. Microsoft Research described the same underlying idea in its BitNet work back in 2023: restrict weights enough, and you change what kind of math the chip has to do to run the model.
Multiplying by 1 or -1 isn’t really multiplication anymore. It’s just flipping a sign. Multiplying by 0 means skipping the operation entirely. So instead of millions of expensive multiplications, the chip is mostly doing additions and sign flips. Cheaper, faster, and far less memory needed to store the model in the first place.
For the model actually on the glasses, the 1.7B language model, that
compression takes it from a 3.45GB reference size at full precision
down to 0.25GB in 1-bit form, or 0.46GB in ternary. Roughly a 14x
reduction.

None of this is free, and by that I mean, Aggressive quantization usually costs accuracy, and pushing an entire network down to 1-bit has historically hurt reasoning tasks specifically. PrismML trains the constraint in from the start rather than compressing an already-finished model, which is what it says keeps quality close to the original.
For this exact model, PrismML’s own benchmark disclosure states the 1-bit Bonsai 1.7B was tested against a 4-bit Qwen3 1.7B across benchmarks including HumanEval+, MMLU Redux, GSM8K, and GPQA Diamond, and reports the two configurations achieved comparable results. That’s PrismML’s own comparison, not an independently run one, and their disclosure itself notes results may vary by configuration and workload.
One more number, with a caveat attached. PrismML has published per-token energy figures showing roughly a 4 to 5x improvement over 4-bit equivalents, but that’s measured on their 8B model, not this 2B glasses configuration. No energy figure specific to the glasses deployment has been published anywhere I could find, so treat the 8B number as a directional signal for the technique, not a spec for this device.
The part almost nobody mentions: KV cache
Shrinking the weights only solves half of the memory problem. Every token a model generates also gets stored in something called the KV cache, short for key-value cache. It’s what lets the model refer back to everything said earlier without recomputing all of it from scratch. The longer the conversation, the bigger this cache gets, and quantizing the weights doesn’t shrink it.
PrismML publishes KV cache figures for the larger models in its lineup: roughly 64 KiB per token at full precision, and about 18 KiB per token with an optional 4-bit cache. I haven’t found a published KV cache breakdown specific to the 1.7B model or to the 4GB glasses configuration.
What’s confirmed, and what actually matters for this article, is the mechanism: cache size scales with context length regardless of model size. On a chip with 4GB total memory, shared between the vision encoder, the operating system, and this cache, memory for holding onto context is a real constraint that exists separately from the size of the model weights themselves.
What’s actually running on the glasses
This is the part that makes the PrismML announcement more than a chip demo. Three separate jobs are now happening locally, all at once, on hardware powered by Qualcomm’s Snapdragon AR1 Gen 1 platform, the same chip family already running Ray-Ban Meta glasses.
▣ Vision: The 0.3 billion parameter vision piece of Bonsai runs at 4-bit precision, not the full 1-bit compression, because vision needs a bit more precision to stay accurate. It’s what lets the glasses actually understand what’s in front of them, not just capture a photo.
▣ Reasoning: The 1.7 billion parameter language model is the part running at 1-bit. This is what turns “here’s an image” into “here’s an answer about that image”, in real time, without sending anything anywhere.
▣ On working memory: Bonsai 1.7B’s architecture supports up to 32,768 tokens, natively trained at 8,192 and extended 4x using a technique called YaRN.
That last part isn’t a flaw in PrismML’s engineering. It’s an intentional trade-off. A bigger memory window costs compute and power the glasses don’t have to spare, so PrismML built a model that’s excellent at reasoning about what’s right in front of it, right now, and left the longer-term memory problem to whatever gets built on top of it.
Qualcomm’s own testing backs up why this tradeoff was worth making. Running the compressed model brought weight memory down by about 3.83 times and pushed token generation speed up by roughly 2.06 times, compared to a standard version of the model at the same size. PrismML’s marketing claims a 4x reduction, and the measured number came in slightly under that, which is worth knowing if you’re the kind of person who checks the fine print. Either way, the direction of the result is real and it’s large.

Where this still needs help
I want to be honest about something, because I’ve built the piece PrismML’s model is missing, and I know exactly where the seams are.
Even at the model’s full 32,768-token ceiling, that’s one long session, not a week of context and not a memory of what you saw yesterday. That’s a different problem from reasoning, and it’s the one I was solving with Qdrant Edge and Gemini’s embeddings in my own project. A small local database that stores what the glasses saw, makes it searchable, and hands the model just the relevant piece when asked. The model doesn’t need to hold a week in its head. It needs a fast way to look something up, the same way you don’t keep every conversation you’ve ever had in your head, you just know roughly where to go looking for it.

Put a model like Bonsai next to a retrieval layer like that, one handling real-time reasoning, the other handling anything longer than a session, and you get something neither piece does alone. That combination is where I think this category is actually headed, not one giant model trying to do everything, but a small reasoning engine paired with a small memory system, both living entirely on the device.
No consumer smart glasses running this software have shipped. This was a chip demo at Qualcomm’s Snapdragon Summit, not a product on a shelf. The technology is real and measured. The product built on top of it doesn’t exist yet.
What this actually changes
If you’re building anything in this space, the useful takeaway isn’t that quantization is impressive. It’s narrower and more useful than that. Local reasoning on genuinely constrained hardware now has real, measured numbers behind it, not just a research demo. The model question stops being “biggest model I can fit” and becomes “smallest model that’s good enough, plus whatever memory system covers the gap it leaves behind.”
If you’re curious what the retrieval half of this looks like in practice, I wrote up the full build of my offline life memorizer, the piece that handles exactly the memory gap Bonsai leaves open. Go read it if you want the other half of this picture.
References
- PrismML Brings 1-Bit Bonsai Models to AI Smart Glasses Powered by Snapdragon, prismml.com, September 23, 2026
- Bonsai 1.7B model specification, docs.prismml.com
- Bonsai introduction and quantization overview, docs.prismml.com
- Bonsai-1.7B-gguf and Bonsai-1.7B-mlx-1bit model cards, Hugging Face
- Bonsai-demo README (KV cache figures for the wider model family), GitHub, PrismML-Eng
- The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits, Microsoft Research
- Snapdragon AR1 Gen 1 Platform, Qualcomm official product page
Smart Glasses Project reference implementation and GitHub link:
- https://coderlegion.com/26957/building-an-offline-life-memorizer-with-gemini-2-0-qdrant-edge
- https://github.com/satyam671/Life-Memorizer-With-Gemini-Embedding-2-And-Qdrant-Edge
Feel free to connect with me on LinkedIn or via email. I’d love to hear about what you’re building. Follow and subscribe to emails for more such posts.
And if you liked this article, make sure to like this post and repost it with your thoughts so that more people discover it and leave your comments down.
❤