Henry Ford did not try to build one car at a time by hand. He designed an assembly line where each station performs one operation and the product moves between stations. The line is optimised for throughput, not for flexibility. Every design decision serves the goal of producing more cars per hour at lower cost per car.
AI inference is being subjected to the same optimisation pressure. Inference is the process of running a trained model to produce outputs. For years, the focus was training: getting the model to learn. Now the focus is shifting to inference: getting the model to produce outputs efficiently and cheaply at scale.
The Batching Opportunity
A model running on a GPU can process multiple inputs simultaneously. If you send one request at a time, you pay the full cost of running the model for one output. If you batch a hundred requests together, you amortise the fixed cost of the GPU operation across a hundred outputs. The per-request cost drops.
Batching works when you have enough concurrent requests to fill the batch and when the latency budget allows waiting for the batch to fill. Real-time applications may not tolerate the wait. A customer service bot that needs to respond within two seconds cannot wait for a batch to fill. A batch processing pipeline that generates reports overnight can wait.
The batch size matters. A batch of 10 gets some amortisation benefit. A batch of 1000 gets more. The optimal batch size depends on your request volume and your latency requirements. Higher volume enables larger batches. Tighter latency requires smaller batches or faster batch-filling mechanisms.
Dynamic batching addresses some of this. Instead of fixed batch sizes, you set a maximum latency budget and fill the batch as requests arrive until the budget is exhausted. This balances throughput against latency. The engineering complexity is real: you need to handle variable batch sizes, monitor batch wait times, and tune the latency budget to your users’ tolerance.
Caching for Repeated Patterns
Many requests are similar. A customer service bot receives variations of the same questions. An internal assistant receives repeated queries about the same policies. If you cache the model’s response to a query, subsequent identical or near-identical queries can be served from cache at a fraction of the inference cost.
Cache hit rate depends on query distribution. High repetition means high cache value. High uniqueness means the cache overhead exceeds its benefit. A system where every user asks a completely different question in completely different words will not benefit from caching. A system where users ask variations of a finite set of questions will.
Cache invalidation is the hard part. When the underlying information changes, cached responses may become stale. If the policy document that answers a question is updated, the cached answer may now be wrong. Cache invalidation strategies range from simple time-based expiration to deliberate invalidation on content changes, depending on how critical cache freshness is.
Approximate caching can help when exact caching is too expensive. If you can define a similarity function that groups similar queries, you can cache responses for query clusters rather than exact queries. This increases cache hit rates at the cost of serving some responses that are close but not exact matches.
Quantisation
A model stores its parameters in full precision numbers. Full precision means high accuracy but also high memory and high compute cost. Quantisation reduces the precision of those numbers, typically from 32-bit floating point to 8-bit integers. The accuracy loss is often small. The speed and cost improvements can be substantial.
Quantisation works because neural networks are robust to small amounts of noise in their parameters. The exact precision of individual parameters matters less than the overall pattern they encode. You can round 3.14159 to 3 and the network still works, because the network learned to use the pattern rather than the exact value.
Different quantisation approaches trade off accuracy against cost. INT8 quantisation is common: parameters are stored as 8-bit integers rather than 32-bit floats. This halves memory usage and often enables faster inference on hardware that supports integer operations. The accuracy loss is typically small for most tasks.
Lower precision options like INT4 exist and achieve higher compression and speed but with larger accuracy losses. These are more appropriate for larger models where the accuracy loss may be acceptable in exchange for the ability to run at all.
The Optimisation Hierarchy
Not all inference optimisations are equal. Some optimisations improve throughput without degrading quality. Some trade quality for speed. The right optimisation depends on your constraints.
Batching improves throughput without changing the model’s computation. You get more outputs per unit time with the same model. The quality is unchanged. Start here if batching is feasible.
Caching improves effective throughput for repeated requests. If your request distribution has enough repetition, caching can dramatically reduce effective inference costs. The quality is unchanged for cache hits and may be slightly degraded for approximate cache hits.
Quantisation reduces the compute required per inference. The model computation changes; the output quality may degrade slightly. Test carefully: some tasks are more sensitive to quantisation than others.
Architecture changes modify the model itself. Smaller models with fewer parameters require less compute but may have lower quality. Distillation trains a smaller model to imitate a larger one. These are larger engineering investments with larger quality trade-offs.
Speculative Decoding
Speculative decoding is a newer optimisation technique. A small “draft” model generates candidate tokens quickly. A larger “verifier” model checks each candidate and accepts or rejects it. When candidates are accepted, you get the draft model’s speed with the verifier model’s quality.
This works best when the draft model and verifier model agree most of the time. If the draft model is much smaller and much less capable, it will be rejected frequently, and you may end up slower than just running the verifier model directly.
Speculative decoding is most valuable when the latency improvement from the draft model exceeds the overhead of verification. For high-throughput batch processing where latency matters less than throughput, standard inference may be better.
Real-World Scenario: The Customer Service Queue
A customer service operation handles 10,000 inquiries per day. Most inquiries fall into 50 categories with standard responses. The system uses caching: when an inquiry matches a cached category, the cached response is served in 50ms. When there is a cache miss, inference takes 2 seconds.
The cache hit rate is 80%. 8,000 inquiries are served from cache at 50ms each. 2,000 inquiries require inference at 2 seconds each. The average cost per inquiry drops by 60% compared to running inference for every request.
The cache hit rate depends on how well the inquiry categories capture the actual distribution. If a new product launches and inquiries shift to new topics, cache hit rates drop until the system adapts or the cache is invalidated.
Real-World Scenario: The Report Generator
A financial institution generates daily reports from market data. The report generation uses a model to summarise 50 pages of market activity into a 2-page executive summary. The generation runs overnight in batch mode.
Without optimisation, each report takes 30 seconds of inference time. With batching across 200 reports, the per-report time drops to 5 seconds. Quantisation further reduces it to 3 seconds per report. The total batch processing time drops from 100 minutes to 10 minutes.
The institution evaluates whether quantisation degrades report quality. Human reviewers compare quantised and non-quantised reports on a sample. The reviewers cannot distinguish them. The quantisation is deployed.
Real-World Scenario: The Real-Time Translator
A video conferencing platform adds real-time translation. The user speaks in English and sees Spanish subtitles appear with less than 500ms delay. The latency requirement is strict: anything above 500ms feels unnatural to users.
The system uses speculative decoding with a small draft model and large verifier model. The draft model generates candidate translations quickly. The verifier model checks them and produces the final output. The pipeline achieves 400ms latency on average.
When the draft and verifier disagree frequently (complex sentences, ambiguous phrasing), the latency spikes above 500ms. The system falls back to direct verifier inference for those cases, accepting higher latency to maintain quality.
The Cost-Quality Frontier
Every optimisation moves you along a cost-quality frontier. Lower cost usually means lower quality. The question is whether the quality loss is acceptable for your use case.
This frontier is not the same for all tasks. A task that requires precise factual accuracy may suffer more from quantisation than a task that generates creative content. A task where occasional errors are harmless may tolerate more optimisation than a task where errors have consequences.
Understanding where your task falls on this frontier requires measurement. Test each optimisation against your specific quality requirements. Do not assume that because one team reported success with an optimisation, it will work for your task.
The Inference Optimisation Technical Debt
Inference optimisations accumulate technical debt. Caching strategies that worked at small scale fail at large scale. Quantisation that was safe at one model version breaks at the next. Batching logic that assumed certain request distributions breaks when distributions change.
This debt is invisible until it causes problems. A cache that was serving 10% of requests at small scale may serve 60% at large scale, and the hit rate explosion may expose bugs that only manifest at scale.
Managing inference optimisation debt requires continuous monitoring. Watch for changes in optimisation effectiveness metrics. When caching hit rates change dramatically, investigate why. When quantisation behaviour changes after a model update, re-evaluate the optimisation.
The Hardware Utilisation Problem
GPUs are expensive. A model running on a GPU may utilise only a fraction of the hardware’s capacity. The GPU sits idle waiting for data, waiting for memory transfers, waiting for the CPU to prepare the next batch.
Low utilisation means you are paying for hardware you are not using. Improving utilisation means reducing the gaps in the processing pipeline where the GPU has nothing to do.
This is a systems engineering problem as much as a model problem. Memory layout, data transfer scheduling, batch construction: these all affect GPU utilisation. Optimising for utilisation requires profiling the actual hardware behaviour, not just the logical inference pipeline.
Inference at the Edge
Running inference on edge devices (phones, IoT devices) has different constraints than running inference in the cloud. Edge devices have limited compute and memory. They may have intermittent connectivity. They have power constraints that cloud servers do not.
Edge inference enables offline operation and reduces latency by processing locally. It also raises privacy concerns: data does not leave the device.
The optimisations for edge inference are more aggressive than for cloud inference. Quantisation to INT4 or lower is common. Model distillation to smaller architectures is often necessary. The quality trade-offs are larger, but the use cases (always-on, offline, private) may justify them.
Real-World Scenario: The Mobile Keyboard
A mobile keyboard app adds next-word prediction. The model runs on the phone, not in the cloud. The user expects predictions to appear instantly, with no network latency.
The model is quantised to INT4 to fit in mobile memory. It is a distilled version of a larger model, sacrificing some accuracy for the ability to run locally. The latency is 20ms per prediction, which feels instant to users.
The quality is slightly worse than the cloud model would be. But the latency advantage and privacy benefits outweigh the quality difference for this use case.
Real-World Scenario: The Voice Assistant
A voice assistant runs wake-word detection locally on the device. The full speech recognition and response generation runs in the cloud. The local wake-word detection is a small model that runs continuously, consuming minimal power.
The local model is quantised to minimise power consumption. It is not the highest-quality wake-word detector possible; it is the highest-quality detector that can run continuously on a battery-powered device.
When the wake word is detected, the system streams audio to the cloud for full processing. The edge processing handles the simple case (wake-word detection) locally, reserving cloud resources for the complex case (speech recognition and generation).
The Optimisation Measurement Problem
Measuring optimisation effectiveness is harder than it sounds. You want to know whether the optimisation improved throughput, reduced latency, or lowered cost without degrading quality.
Throughput is easy to measure: requests per second before and after the optimisation. Latency is easy to measure: time per request. Cost per request is harder: it depends on hardware utilisation, which depends on workload characteristics that vary over time.
Quality is the hardest to measure. For some tasks, there is a clear quality metric: accuracy on a test set, BLEU score on a translation task. For other tasks, quality is subjective: does the output sound natural? Is the response helpful?
Establishing quality metrics before optimising is essential. If you do not know what quality means for your task, you cannot know whether the optimisation degraded it.
Decision Rules
Optimise inference when:
- You are serving at scale and per-request cost matters
- Latency requirements allow for batching or caching strategies
- You have verified that optimisation does not degrade output quality below your threshold
- The engineering cost of optimisation is less than the cost savings it delivers
- Your request distribution has enough repetition for caching to help
- Your users’ experience would benefit from lower latency
- You have established quality metrics that you can measure reliably
- You can monitor optimisation effectiveness over time
Accept higher inference costs when:
- Your volume is low and optimisation overhead exceeds savings
- Your application requires real-time responses that cannot tolerate batching latency
- Output quality is highly sensitive to model precision
- You are still in the exploration phase and have not locked down your inference requirements
- The engineering capacity for optimisation is not available
- You cannot measure quality degradation reliably
- Your hardware utilisation is already high and further optimisation would yield minimal gains
The assembly line made cars affordable by making them fast. Inference optimisation makes AI affordable by making it efficient. But efficiency is not free; it comes with complexity that has to be managed.