Token Compression and Prompt Optimization: Five Techniques to Cut LLM Costs Without Losing Quality
Structured Constraints Beat Narrative System Prompts
Most system prompts read like onboarding documents. They explain the persona, narrate the task, and politely request behavior in full sentences. That costs significantly more than a tightly structured equivalent doing the same job.
The fix is to replace narrative instructions with declarative constraints, formatted in a schema-like style. Pipe-delimited key-value pairs, YAML-style constraints, and JSON schema snippets all work; the right choice depends on the model family and how it was trained to parse structured input.
A worked example makes the arithmetic concrete. A narrative instruction of 36 tokens can be replaced by a structured line of roughly 14 tokens, a 60 percent reduction on a single instruction. Treat those figures as illustrative rather than measured benchmark results; your own prompts will vary. Now multiply the effect across a system prompt with a dozen such instructions, and across thousands of API calls per day. The savings compound quietly, and because structured constraints are less ambiguous than prose, output consistency often improves at the same time.

Right-Sizing Few-Shot Examples
Few-shot prompting works. Placing example input-output pairs before the request improves output format consistency, particularly for classification and structured generation. The failure mode is not using too few examples. It is using too many.
Research from Anthropic and academic benchmarks (specific papers not named in this summary) points to diminishing returns beyond three to five examples for most classification and generation tasks. Adding ten more examples rarely improves accuracy, and it often introduces contradictions that confuse the model rather than guide it. Two examples that disagree on edge-case handling are worse than no examples at all.
The practical discipline is to benchmark quality at one, three, and five examples before committing to a larger set. For a sentiment labeling task, a lean three-shot prompt with clearly contrasting cases can often match a larger example set while consuming a fraction of the input tokens. Measure rather than assume, because the crossover point differs by task.
Dynamic Context Trimming: Retrieve Passages, Not Documents
Long documents are where token budgets go to die. Transcripts, legal text, and knowledge base articles routinely run to thousands of tokens, and most of that content is irrelevant to the query at hand.
The technique is retrieval at passage granularity rather than document granularity. Instead of passing the raw document, embed the query and the candidate passages, rank by cosine similarity using sentence embeddings, and pass a filtered list of only the passages that clear a relevance threshold.
The numbers are dramatic in principle. For a 10,000-token knowledge base where only 800 tokens are relevant, context costs can drop by more than 90 percent. That figure is an arithmetic illustration of the ratio, not a benchmark result from a named paper. Quality usually holds or improves, because the model is no longer asked to locate a needle in a haystack it was handed wholesale. The trade-off is engineering overhead: you now maintain an embedding pipeline and a retrieval step, and a retrieval miss becomes a quality failure that a full-document prompt would have avoided. Tune the threshold conservatively.
Prefix Caching: Stop Paying Twice for the Same System Prompt
If your application resends the same system prompt on every request, including persona definitions, tool descriptions, and policy constraints, you are paying full input price for identical tokens over and over.
Several inference providers now store and reuse static prompt prefixes server-side, billing cached tokens at a fraction of standard input pricing. Named implementations include Anthropic’s prompt caching and OpenAI’s automatic prefix caching, though the list is not exhaustive and other providers offer comparable features. The structural requirement is consistent across them: stable content first, dynamic content last. A cacheable prefix cannot contain the user’s query.
There are caveats. Anthropic’s prompt caching engages only when the cached prefix exceeds a minimum token threshold, and the cache must be hit within a defined time window. Both conditions vary by provider and change over time, so check each provider’s current caching documentation before you architect around it. Caching is also a moving target, and behavior that holds today may shift with a model update.
Separating Chain-of-Thought Scratchpads from Final Answers
Chain-of-thought prompting improves reasoning on complex tasks, and it does so by generating a reasoning trace that can run to hundreds of tokens. The problem is billing. That trace frequently appears verbatim in the API response even when the application only needs the final answer, which inflates output token costs on every call.
The fix is to separate the scratchpad from the answer using structured markers. Ask the model to wrap its reasoning in <thinking> tags and its final output in <answer> tags, then parse the response, discard the thinking block, and store and return only the answer content. You still pay for the reasoning tokens, because the model generated them, but you stop paying to carry them through your own storage and downstream processing.
For APIs with native extended thinking or reasoning modes, such as Anthropic’s extended thinking, reasoning tokens may be billed at a different rate and can be suppressed from the response body entirely. Reasoning token accounting differs meaningfully between providers and is revised often, so read the current billing documentation for whichever API you use rather than assuming one provider’s rules apply everywhere.
A Practical Workflow: Audit, Measure, Benchmark
None of these techniques require a rewrite. In fact, the rewrite approach is how optimization projects fail. Incremental, evidence-based changes are more sustainable than overhauling every prompt at once.
Start with one frequently used prompt. Measure its current token count. Apply a single technique from this list. Benchmark output quality against the original, using the same evaluation set you would use for any model change. If quality holds and cost drops, keep it and move to the next prompt. If quality drops, you have learned something specific about your workload at the cost of one afternoon.
This is also where the discipline connects to the broader agentic landscape. Systems that spawn many agents and many calls per task, such as the swarm behavior documented in When AI Agents Talk Behind Our Backs: Inside the 2026 Agent Collusion Incidents, multiply every inefficiency in their prompts. Even a small per-call token saving becomes substantial across a fleet of agents making many calls.
Leaner Prompts, Better Models, Lower Bills
The techniques here, structured constraints, right-sized few-shot sets, dynamic context trimming, prefix caching, and scratchpad separation, share a common effect: they reduce the distance between what you mean and what the model reads.
The compounding argument is the strongest one. A large reduction on a single system instruction, a context window trimmed to a fraction of its original size, a cached prefix billed at a fraction of input price, and a discarded reasoning trace all stack. Applied across a production application, they change the economics of what you can build.
The same pattern shows up elsewhere in the field, where efficiency gains and capability gains arrive together rather than trading off. Disciplined method, not scale alone, is what settles hard questions, a lesson that holds whether you are auditing a prompt or reading a contested result. For teams tracking how these shifts land in commercial settings, industry news sources can offer a useful view of how quickly the surrounding market moves.
Pick one prompt this week. Count its tokens. Cut something. Measure what happens.
For more on this, see beyond setup five practices.
One Comment
Comments are closed.