CoderLegion has a lot of developers experimenting with language models, and game characters are one of the most demanding places to use them. This is a practical rundown of the latency and cost techniques that make NPCs driven by language models workable outside a demo.
Players expect a reply within a second or two when they talk to a character. Developers need every one of those replies to cost almost nothing, because a game with thousands of daily players having dozens of conversations each turns per token pricing into a real bill. Both constraints have known solutions, and none of them require a bigger model.
Stream First, Optimize Second
Without streaming, a player waits 2 to 5 seconds staring at an empty dialogue box while the full reply generates. With streaming, perceived latency drops to the time to first token, typically 200 to 800 milliseconds from a cloud API and 50 to 150 milliseconds from a local model on a GPU.
Streaming changes the rest of the pipeline. Safety filters need to run sentence by sentence as text arrives, and if the model returns structured JSON with dialogue, emotion and actions, the parser has to tolerate incomplete objects until the stream closes.
Route Each Conversation to the Right Model
Model tiering does more for cost than any other single change. Story critical characters and pivotal scenes get the best model available, general conversation goes to a smaller and faster one, and ambient chatter comes from the smallest model, cached templates or plain rules.
A top tier cloud model can cost 10 to 50 times more per token than a small one, and the bottom tier can be free when it runs locally. If around 70 percent of interactions land in the lower two tiers, the average cost per exchange falls sharply, and players only notice the difference in moments they were not paying close attention to anyway.
Put Hard Limits on Tokens
Set a maximum input budget covering the system prompt, history, game state and retrieved memories, and when a conversation exceeds it, summarize older turns rather than resending them word for word. Cap output too, since NPC dialogue rarely needs more than 150 to 300 tokens per reply.
Generating ahead fills the remaining gap. If the game can predict the first greeting or the line at a quest stage, generate it in advance and serve it from cache, falling back to live generation for everything else.
Combined, these turn LLM NPCs from an expensive novelty into something a small team can ship. The full set of techniques with cost calculations is in this guide to handling latency and cost for LLM NPCs, and the wider architecture, from character prompts to long term memory, is covered in the overview of LLM powered NPCs.