You have a hundred-page report on quarterly sales performance. Your executive reads the first page, glances at the charts, and makes a decision. The executive summary carried the substance. The hundred pages provided backup if anyone wanted to go deeper.
A compressed prompt does the same job for context. You take a long context, extract the essential information, and present the model with a concise version that captures what matters. The model reasons over the summary rather than the full document. The key relationships and facts are preserved. The noise is removed.
Why Compression Matters
Long contexts are expensive. You pay per token. A hundred-page document might contain tens of thousands of tokens. If you can compress that to a single page while preserving the relevant information, you reduce cost and often improve output quality, because the model has less irrelevant content to reason over.
Compression also helps with attention. Models process all tokens in context, though not equally. Irrelevant content in the context dilutes the model’s attention. Removing irrelevant content focuses the model on what matters.
Consider a legal document review task. The documents are hundreds of pages each. You need the model to answer questions about specific contractual terms. If you include the entire document, the model has to search through lengthy background sections, definitions, and boilerplate to find the relevant terms. If you preprocess the document to extract only the key terms and their definitions, the model can reason over a fraction of the tokens with a clearer view of what matters.
The savings compound at scale. If you process a thousand documents per day, compressing each from 50,000 tokens to 5,000 tokens reduces your token costs by a factor of ten and your inference time proportionally.
What Gets Preserved
Effective compression preserves relationships and key facts. The structure of the original matters: what caused what, what is dependent on what, which numbers are important. A compression that strips out causality and leaves only assertions is less useful than one that preserves the logical structure.
Extractive compression selects existing sentences or passages that are most informative. The selected sentences are taken verbatim from the original. This approach is safe: the compressed text is a subset of the original, so nothing is invented. It is less powerful: the summary is limited to sentences that exist in the original.
Abstractive compression generates new text that captures the meaning in fewer tokens. The compressor paraphrases, synthesises, and infers connections that may not be explicit in any single sentence. This achieves higher compression ratios but introduces risk: the compressor’s interpretation of the source may be inaccurate.
Abstractive compression is more powerful but requires careful evaluation. A compressor that misinterprets a source document produces a compressed context that leads the model to wrong conclusions. This is a silent failure: the model reasoning correctly from incorrect context produces confidently wrong answers.
The Compression Evaluation Problem
You cannot easily evaluate whether a compression preserved what matters. You would need to know what matters in the original, which is precisely what you are trying to distill. This makes compression quality hard to verify and easy to overestimate.
One approach is task-based evaluation. Compress the document, ask the model questions that require specific information from the original, and check whether the answers are correct. If the questions are representative of actual use cases, and the answers are correct, the compression is probably adequate. If the answers are wrong, the compression lost something that matters.
This evaluation approach has limitations. You can only test for information you thought to ask about. A compression that passes your test questions but misses other important information will still produce wrong answers on questions you did not anticipate.
The safer approach is conservative compression: only remove content that you are confident does not matter. This limits the compression ratio but reduces the risk of silent failures. As you develop confidence in what your application can tolerate, you can push the compression ratio higher.
Selective Inclusion vs Full Compression
Compression is one approach to long contexts. Another is selective inclusion: instead of compressing the entire context, you decide in advance what to include.
Selective inclusion works when you know what the model needs before you know what the document contains. You design a context that includes only the relevant pieces, assembled from different sources. The model never sees the full document; it only sees what you selected.
This approach requires more upfront work: you have to identify the relevant pieces rather than just running compression. It also requires domain knowledge: someone has to know what pieces matter. But it avoids the compression evaluation problem, because you control exactly what the model sees.
Hybrid approaches combine both. Use selective inclusion for obviously relevant content. Use compression for background material where selective inclusion is too expensive. The compression ratio can be lower for background material because the information density requirement is lower.
RAG with Compression
Retrieval-augmented generation systems often use compression to fit more retrieved documents into the context window. A retrieval pass might return 20 relevant documents; compression can reduce them to 5 while preserving the key information from each.
This is different from compressing a single long document. With multiple retrieved documents, you need to compress each one individually and preserve the distinctions between them. Losing track of which information came from which document is a common failure mode.
Attribution matters in RAG compression. If the model does not know which compressed passage came from which document, it cannot properly attribute its reasoning. Good RAG compression preserves source attribution alongside the compressed content.
Real-World Scenario: The Contract Review
A legal team reviews a 300-page vendor contract. They need to understand: what are the termination clauses, what are the liability limits, what are the payment terms, and are there any unusual provisions.
They build a prompt that includes only the relevant sections, identified by a legal document parsing system. The parser extracts termination, liability, payment, and special provisions sections from the full contract. The model receives a 10-page context instead of a 300-page document.
The model answers the team’s questions. The team verifies the answers against the extracted sections. The answers are accurate because the extraction was accurate and the model had focused context.
If the extraction missed a relevant clause buried in an unexpected section, the model would not catch it. Selective inclusion depends on knowing where relevant content lives.
Real-World Scenario: The Codebase Q&A
An engineer asks: “Which modules handle payment processing and how do they interact with the fraud detection system?” The codebase is millions of lines.
The system does not send the entire codebase in the context. Instead, it retrieves the most relevant files using code search: files that mention payment, files that mention fraud, and files that define interfaces between them. The retrieved code is compressed by removing comments and boilerplate while preserving function signatures and key logic.
The model receives a compressed context of 50 files totalling 10,000 tokens. It traces the payment flow and identifies the fraud detection hooks. The engineer gets an answer that covers the relevant architecture without reading the entire codebase.
The answer is only as good as the retrieval and compression. If the fraud detection system uses an obscure naming convention that the search did not find, its involvement would be missing from the answer.
Real-World Scenario: The Meeting Summarisation
A system summarises hour-long meetings. The transcript is 20,000 tokens. The summary needs to capture key decisions, action items, and questions raised.
Extractive compression selects the most important sentences from the transcript verbatim. A 500-token summary extracts the 20 sentences that contain the most consequential statements.
Abstractive compression generates a new summary that paraphrases and synthesises the key points. A 500-token abstractive summary may capture more nuance but risks misrepresenting what was said.
The choice depends on how the summary will be used. If the summary is for a participant who attended the meeting and just needs a reminder, extractive compression is safer. If the summary is for a participant who did not attend and needs to understand what happened, abstractive compression may be more useful but requires more trust in the compressor.
The Token Budget Mindset
Compression is ultimately a token budget management problem. You have a context window of fixed size. You have content that may be larger than that window. You have to decide what to include and what to sacrifice.
The right token budget depends on the task. If the task requires precision (legal, medical, financial), you need more tokens for accuracy and should compress less. If the task requires only a rough answer, you can compress more aggressively.
Think of it as a budget allocation problem. Each token you spend on background is a token you cannot spend on specifics. The compression ratio should reflect the relative importance of breadth versus depth for your use case.
Compression and the Long Context Problem
Long contexts expose compression’s limits. A model that can handle 200,000 tokens in context is impressive. But a 200,000-token context is not the same as a well-compressed 10,000-token context that contains the right information.
Models have to attend to all tokens in context, though not equally. A 200,000-token context with 90% irrelevant information may produce worse results than a 10,000-token context with 90% relevant information. Compression that removes “irrelevant” content may remove content the model actually needed.
The long context problem is not solved by compression alone. It requires better retrieval, better selection, and better understanding of what “relevant” means for the task.
The Chunking Problem
When compressing long documents, how you chunk matters. A document has structure: sections, paragraphs, sentences. If you compress by chunking arbitrarily, you may break relationships that matter.
A contract clause that refers to definitions in section 2 needs section 2 in context when you compress. If you compress each section independently and drop the cross-references, the compressed context loses the relationship.
Chunk boundaries should respect document structure. Section-level chunking works better than arbitrary token-count chunking. The cost is that chunks may be uneven in size, but the semantic coherence is preserved.
Compression and Reasoning Quality
Compression can degrade reasoning quality in subtle ways. A model reasoning over compressed context may reach correct conclusions for the wrong reasons, because the compression altered the logical flow.
Consider a document that presents an argument: premise A leads to premise B, which leads to conclusion C. If compression removes premise B as “redundant,” the remaining text says A and C with an implicit connection. The model may infer a different connection than the original author intended.
This is different from losing factual information. The facts are preserved but the reasoning path is altered. The model reaches the right answer based on different reasoning. This can be dangerous when the reasoning matters, not just the conclusion.
Testing for reasoning degradation requires evaluating reasoning chains, not just conclusions. Ask the model to explain its reasoning and compare against the original argument structure.
Real-World Scenario: The Research Paper Analysis
A research team analyses hundreds of papers to identify whether a specific hypothesis has been tested. They retrieve papers from a corpus, compress each to a summary, and ask the model whether any paper tests the hypothesis.
Extractive compression works well: it selects the paper sections that discuss methodology and results, verbatim. The model reads these selections and identifies the relevant papers.
But the compression loses nuance about the methodology. The paper that tested the hypothesis used a different population than the team is interested in. The compressed summary mentions the hypothesis test but not the population difference. The model incorrectly concludes the hypothesis has been tested in the relevant population.
The failure was not in retrieval, which found the right paper. It was in compression, which lost the population detail. This detail would have been caught by reading the full paper, but the compression summary did not include it.
The Summary Length Calibration Problem
How long should a summary be? Too short and you lose important details. Too long and you have not compressed enough and you still have the original context window problem.
The right length depends on the compression ratio that preserves task-relevant information. This ratio varies by document type and by task. A news article can be compressed to 10% of original length. A legal contract may need to retain 50% to preserve all relevant details.
Calibrating the right length requires testing with your specific documents and tasks. Start with aggressive compression and back off where quality degrades.
Decision Rules
Use prompt compression when:
- Your context is long and much of it is not relevant to the task
- Token cost and latency are constraints
- You have verified that the compressor preserves task-relevant information
- The compression ratio does not exceed the threshold where accuracy degrades
- You are using extractive compression or have validated abstractive compression quality
- You can preserve attribution when multiple sources are involved
- Chunk boundaries respect document structure
- Reasoning chains are preserved, not just facts
Do not use prompt compression when:
- The full context is short enough that compression gains are minimal
- Task-relevant information is distributed throughout the context and hard to identify in isolation
- The compression introduces inaccuracies that are worse than the noise the compression was meant to remove
- You cannot evaluate whether the compression preserved what matters
- Selective inclusion is a better fit for your use case
- Your task requires high precision and compression risks losing details that matter
- Reasoning chains are critical and compression disrupts them silently
The executive summary works because someone understood what the executive needed to decide. Prompt compression works for the same reason: know what the model needs to know, and compress to that. The executive who reads a summary that misrepresents the argument has been misled, not informed.