Simor
Inference: The Assembly Line

Inference: The Assembly Line

Simor Consulting | 02 Oct, 2026 | 09 Mins read

Henry Ford did not try to build one car at a time by hand. He designed an assembly line where each station performs one operation and the product moves between stations. The line is optimised for throughput, not for flexibility. Every design decision serves the goal of producing more cars per hour at lower cost per car.

AI inference is being subjected to the same optimisation pressure. Inference is the process of running a trained model to produce outputs. For years, the focus was training: getting the model to learn. Now the focus is shifting to inference: getting the model to produce outputs efficiently and cheaply at scale.

The Batching Opportunity

A model running on a GPU can process multiple inputs simultaneously. If you send one request at a time, you pay the full cost of running the model for one output. If you batch a hundred requests together, you amortise the fixed cost of the GPU operation across a hundred outputs. The per-request cost drops.

Batching works when you have enough concurrent requests to fill the batch and when the latency budget allows waiting for the batch to fill. Real-time applications may not tolerate the wait. A customer service bot that needs to respond within two seconds cannot wait for a batch to fill. A batch processing pipeline that generates reports overnight can wait.

The batch size matters. A batch of 10 gets some amortisation benefit. A batch of 1000 gets more. The optimal batch size depends on your request volume and your latency requirements. Higher volume enables larger batches. Tighter latency requires smaller batches or faster batch-filling mechanisms.

Dynamic batching addresses some of this. Instead of fixed batch sizes, you set a maximum latency budget and fill the batch as requests arrive until the budget is exhausted. This balances throughput against latency. The engineering complexity is real: you need to handle variable batch sizes, monitor batch wait times, and tune the latency budget to your users’ tolerance.

Caching for Repeated Patterns

Many requests are similar. A customer service bot receives variations of the same questions. An internal assistant receives repeated queries about the same policies. If you cache the model’s response to a query, subsequent identical or near-identical queries can be served from cache at a fraction of the inference cost.

Cache hit rate depends on query distribution. High repetition means high cache value. High uniqueness means the cache overhead exceeds its benefit. A system where every user asks a completely different question in completely different words will not benefit from caching. A system where users ask variations of a finite set of questions will.

Cache invalidation is the hard part. When the underlying information changes, cached responses may become stale. If the policy document that answers a question is updated, the cached answer may now be wrong. Cache invalidation strategies range from simple time-based expiration to deliberate invalidation on content changes, depending on how critical cache freshness is.

Approximate caching can help when exact caching is too expensive. If you can define a similarity function that groups similar queries, you can cache responses for query clusters rather than exact queries. This increases cache hit rates at the cost of serving some responses that are close but not exact matches.

Quantisation

A model stores its parameters in full precision numbers. Full precision means high accuracy but also high memory and high compute cost. Quantisation reduces the precision of those numbers, typically from 32-bit floating point to 8-bit integers. The accuracy loss is often small. The speed and cost improvements can be substantial.

Quantisation works because neural networks are robust to small amounts of noise in their parameters. The exact precision of individual parameters matters less than the overall pattern they encode. You can round 3.14159 to 3 and the network still works, because the network learned to use the pattern rather than the exact value.

Different quantisation approaches trade off accuracy against cost. INT8 quantisation is common: parameters are stored as 8-bit integers rather than 32-bit floats. This halves memory usage and often enables faster inference on hardware that supports integer operations. The accuracy loss is typically small for most tasks.

Lower precision options like INT4 exist and achieve higher compression and speed but with larger accuracy losses. These are more appropriate for larger models where the accuracy loss may be acceptable in exchange for the ability to run at all.

The Optimisation Hierarchy

Not all inference optimisations are equal. Some optimisations improve throughput without degrading quality. Some trade quality for speed. The right optimisation depends on your constraints.

Batching improves throughput without changing the model’s computation. You get more outputs per unit time with the same model. The quality is unchanged. Start here if batching is feasible.

Caching improves effective throughput for repeated requests. If your request distribution has enough repetition, caching can dramatically reduce effective inference costs. The quality is unchanged for cache hits and may be slightly degraded for approximate cache hits.

Quantisation reduces the compute required per inference. The model computation changes; the output quality may degrade slightly. Test carefully: some tasks are more sensitive to quantisation than others.

Architecture changes modify the model itself. Smaller models with fewer parameters require less compute but may have lower quality. Distillation trains a smaller model to imitate a larger one. These are larger engineering investments with larger quality trade-offs.

Speculative Decoding

Speculative decoding is a newer optimisation technique. A small “draft” model generates candidate tokens quickly. A larger “verifier” model checks each candidate and accepts or rejects it. When candidates are accepted, you get the draft model’s speed with the verifier model’s quality.

This works best when the draft model and verifier model agree most of the time. If the draft model is much smaller and much less capable, it will be rejected frequently, and you may end up slower than just running the verifier model directly.

Speculative decoding is most valuable when the latency improvement from the draft model exceeds the overhead of verification. For high-throughput batch processing where latency matters less than throughput, standard inference may be better.

Real-World Scenario: The Customer Service Queue

A customer service operation handles 10,000 inquiries per day. Most inquiries fall into 50 categories with standard responses. The system uses caching: when an inquiry matches a cached category, the cached response is served in 50ms. When there is a cache miss, inference takes 2 seconds.

The cache hit rate is 80%. 8,000 inquiries are served from cache at 50ms each. 2,000 inquiries require inference at 2 seconds each. The average cost per inquiry drops by 60% compared to running inference for every request.

The cache hit rate depends on how well the inquiry categories capture the actual distribution. If a new product launches and inquiries shift to new topics, cache hit rates drop until the system adapts or the cache is invalidated.

Real-World Scenario: The Report Generator

A financial institution generates daily reports from market data. The report generation uses a model to summarise 50 pages of market activity into a 2-page executive summary. The generation runs overnight in batch mode.

Without optimisation, each report takes 30 seconds of inference time. With batching across 200 reports, the per-report time drops to 5 seconds. Quantisation further reduces it to 3 seconds per report. The total batch processing time drops from 100 minutes to 10 minutes.

The institution evaluates whether quantisation degrades report quality. Human reviewers compare quantised and non-quantised reports on a sample. The reviewers cannot distinguish them. The quantisation is deployed.

Real-World Scenario: The Real-Time Translator

A video conferencing platform adds real-time translation. The user speaks in English and sees Spanish subtitles appear with less than 500ms delay. The latency requirement is strict: anything above 500ms feels unnatural to users.

The system uses speculative decoding with a small draft model and large verifier model. The draft model generates candidate translations quickly. The verifier model checks them and produces the final output. The pipeline achieves 400ms latency on average.

When the draft and verifier disagree frequently (complex sentences, ambiguous phrasing), the latency spikes above 500ms. The system falls back to direct verifier inference for those cases, accepting higher latency to maintain quality.

The Cost-Quality Frontier

Every optimisation moves you along a cost-quality frontier. Lower cost usually means lower quality. The question is whether the quality loss is acceptable for your use case.

This frontier is not the same for all tasks. A task that requires precise factual accuracy may suffer more from quantisation than a task that generates creative content. A task where occasional errors are harmless may tolerate more optimisation than a task where errors have consequences.

Understanding where your task falls on this frontier requires measurement. Test each optimisation against your specific quality requirements. Do not assume that because one team reported success with an optimisation, it will work for your task.

The Inference Optimisation Technical Debt

Inference optimisations accumulate technical debt. Caching strategies that worked at small scale fail at large scale. Quantisation that was safe at one model version breaks at the next. Batching logic that assumed certain request distributions breaks when distributions change.

This debt is invisible until it causes problems. A cache that was serving 10% of requests at small scale may serve 60% at large scale, and the hit rate explosion may expose bugs that only manifest at scale.

Managing inference optimisation debt requires continuous monitoring. Watch for changes in optimisation effectiveness metrics. When caching hit rates change dramatically, investigate why. When quantisation behaviour changes after a model update, re-evaluate the optimisation.

The Hardware Utilisation Problem

GPUs are expensive. A model running on a GPU may utilise only a fraction of the hardware’s capacity. The GPU sits idle waiting for data, waiting for memory transfers, waiting for the CPU to prepare the next batch.

Low utilisation means you are paying for hardware you are not using. Improving utilisation means reducing the gaps in the processing pipeline where the GPU has nothing to do.

This is a systems engineering problem as much as a model problem. Memory layout, data transfer scheduling, batch construction: these all affect GPU utilisation. Optimising for utilisation requires profiling the actual hardware behaviour, not just the logical inference pipeline.

Inference at the Edge

Running inference on edge devices (phones, IoT devices) has different constraints than running inference in the cloud. Edge devices have limited compute and memory. They may have intermittent connectivity. They have power constraints that cloud servers do not.

Edge inference enables offline operation and reduces latency by processing locally. It also raises privacy concerns: data does not leave the device.

The optimisations for edge inference are more aggressive than for cloud inference. Quantisation to INT4 or lower is common. Model distillation to smaller architectures is often necessary. The quality trade-offs are larger, but the use cases (always-on, offline, private) may justify them.

Real-World Scenario: The Mobile Keyboard

A mobile keyboard app adds next-word prediction. The model runs on the phone, not in the cloud. The user expects predictions to appear instantly, with no network latency.

The model is quantised to INT4 to fit in mobile memory. It is a distilled version of a larger model, sacrificing some accuracy for the ability to run locally. The latency is 20ms per prediction, which feels instant to users.

The quality is slightly worse than the cloud model would be. But the latency advantage and privacy benefits outweigh the quality difference for this use case.

Real-World Scenario: The Voice Assistant

A voice assistant runs wake-word detection locally on the device. The full speech recognition and response generation runs in the cloud. The local wake-word detection is a small model that runs continuously, consuming minimal power.

The local model is quantised to minimise power consumption. It is not the highest-quality wake-word detector possible; it is the highest-quality detector that can run continuously on a battery-powered device.

When the wake word is detected, the system streams audio to the cloud for full processing. The edge processing handles the simple case (wake-word detection) locally, reserving cloud resources for the complex case (speech recognition and generation).

The Optimisation Measurement Problem

Measuring optimisation effectiveness is harder than it sounds. You want to know whether the optimisation improved throughput, reduced latency, or lowered cost without degrading quality.

Throughput is easy to measure: requests per second before and after the optimisation. Latency is easy to measure: time per request. Cost per request is harder: it depends on hardware utilisation, which depends on workload characteristics that vary over time.

Quality is the hardest to measure. For some tasks, there is a clear quality metric: accuracy on a test set, BLEU score on a translation task. For other tasks, quality is subjective: does the output sound natural? Is the response helpful?

Establishing quality metrics before optimising is essential. If you do not know what quality means for your task, you cannot know whether the optimisation degraded it.

Decision Rules

Optimise inference when:

  • You are serving at scale and per-request cost matters
  • Latency requirements allow for batching or caching strategies
  • You have verified that optimisation does not degrade output quality below your threshold
  • The engineering cost of optimisation is less than the cost savings it delivers
  • Your request distribution has enough repetition for caching to help
  • Your users’ experience would benefit from lower latency
  • You have established quality metrics that you can measure reliably
  • You can monitor optimisation effectiveness over time

Accept higher inference costs when:

  • Your volume is low and optimisation overhead exceeds savings
  • Your application requires real-time responses that cannot tolerate batching latency
  • Output quality is highly sensitive to model precision
  • You are still in the exploration phase and have not locked down your inference requirements
  • The engineering capacity for optimisation is not available
  • You cannot measure quality degradation reliably
  • Your hardware utilisation is already high and further optimisation would yield minimal gains

The assembly line made cars affordable by making them fast. Inference optimisation makes AI affordable by making it efficient. But efficiency is not free; it comes with complexity that has to be managed.

Shipping a production AI system?

Find where your AI spend leaks and where quality slips. Take the AI Production Scorecard for a fast baseline across the seven layers, or book a free AI cost review and we will turn it into a plan.

Similar Articles

Seek > Offset: Airline Boarding Pass Analogy
Seek > Offset: Airline Boarding Pass Analogy
04 Apr, 2025 | 03 Mins read

Picture yourself at a busy airport gate. The agent announces: "We'll now board passengers in rows 20 through 30." Simple, efficient, everyone knows whether it's their turn. Now imagine instead they sa

Tracing Spans as Russian Nesting Dolls
Tracing Spans as Russian Nesting Dolls
21 Mar, 2025 | 03 Mins read

Russian nesting dolls (Matryoshka) are wooden dolls where each one opens to reveal a smaller doll inside, which opens to reveal another, and so on. Each doll represents an operation in your distribute

Fridge Magnet Letters Arriving Late
Fridge Magnet Letters Arriving Late
09 May, 2025 | 05 Mins read

Magnetic letters on a fridge, sent between rooms with a gap under the door. You send C-A-T in order, but your friend receives A-C-T. Or worse, C-T-A. Your cat becomes an act, or something that isn't a

The CAP Desert Triangle
The CAP Desert Triangle
02 May, 2025 | 06 Mins read

You're leading an expedition across a desert. Your team needs three things: Consistent maps (everyone has the same version), Available guides (can always get directions), and Partition tolerance (can

gRPC Postcards: Typed Messages at Light-Speed
gRPC Postcards: Typed Messages at Light-Speed
14 Mar, 2025 | 03 Mins read

A postal service where every postcard has a strict template. The address fields are always in the same spot. The message area has specific sections for specific types of information. Both sender and r

Bloom Filters: The Forgetful Bouncer
Bloom Filters: The Forgetful Bouncer
28 Mar, 2025 | 06 Mins read

A nightclub bouncer with a peculiar condition: they never forget a face they've seen, but sometimes they think they've seen faces they haven't. When someone approaches, they'll either say "You've defi

Idempotency: Vending Machine Coin Trick
Idempotency: Vending Machine Coin Trick
11 Apr, 2025 | 03 Mins read

You're at a vending machine, desperately needing caffeine. You insert a dollar, press B4 for coffee, but nothing happens. Did the machine eat your money? Did it register the button press? In frustrati

WebSockets: The Persistent Coffee Line
WebSockets: The Persistent Coffee Line
07 Mar, 2025 | 06 Mins read

You walk into your favourite coffee shop and order your usual. But instead of ordering, paying, leaving, and coming back when you want another coffee (like HTTP requests), imagine you could just stay

Window Functions: The Train Car View
Window Functions: The Train Car View
25 Apr, 2025 | 05 Mins read

You're on a cross-country train, sitting by the window. As landscapes roll by, you can see not just where you are, but where you've been and where you're going. You can count how many red barns you've

Time-Travel Tables: Passport Stamp Method
Time-Travel Tables: Passport Stamp Method
18 Apr, 2025 | 04 Mins read

Open your passport and you see a story told in stamps: where you've been, when you arrived, when you left. Each stamp doesn't erase the previous ones - they accumulate, creating a complete travel hist

Column Stores: The Vertical Filing Cabinet
Column Stores: The Vertical Filing Cabinet
30 May, 2025 | 04 Mins read

Reorganise an enormous filing cabinet. Instead of keeping complete employee records in manila folders (one folder per person with all their information), you create specialised drawers: one for all sa

Parquet vs ORC: Suitcase vs Trunk
Parquet vs ORC: Suitcase vs Trunk
06 Jun, 2025 | 04 Mins read

Packing for a month-long trip. Do you use a suitcase with clever compartments, compression bags, and built-in organisation? Or a trunk with adjustable dividers, heavy-duty locks, and industrial-streng

Cosine Similarity: The Handshake Angle
Cosine Similarity: The Handshake Angle
13 Jun, 2025 | 04 Mins read

At a networking event, watch how people greet each other. Some reach straight out for a firm handshake. Others angle up for a high-five. A few go low for a fist bump. Measure not the style of greeting

Bank Vault Double Key
Bank Vault Double Key
16 May, 2025 | 04 Mins read

The most secure bank vault in the world requires two different keys, held by two different people, turned simultaneously. Neither person alone can open it. Now try coordinating this when the key holde

CRDTs: The Cooperative Sketchpad
CRDTs: The Cooperative Sketchpad
23 May, 2025 | 04 Mins read

A magical sketchpad shared by artists around the world. Each artist has their own copy, draws whenever inspiration strikes, and somehow - without talking to each other, without a master artist coordin

Embeddings: GPS for Words
Embeddings: GPS for Words
20 Jun, 2025 | 05 Mins read

Embeddings assign numerical coordinates to words and concepts. "Cat" sits near "kitten" and "feline" but far from "airplane." "Paris" neighbours "France" and "Eiffel Tower" but distances itself from "

Library Book Whisperer
Library Book Whisperer
27 Jun, 2025 | 03 Mins read

A library maintains an unofficial whisper network. A patron asks about a book, and a librarian remembers: "Sarah at the reference desk has it." This network bypasses the official catalogue, turning ho

Consistent Hashing: The Pizza Slice Wheel
Consistent Hashing: The Pizza Slice Wheel
04 Jul, 2025 | 03 Mins read

Imagine arranging pizza party guests on a circle, dividing it like pizza slices. Each station serves a section. When a guest leaves, only their immediate neighbours shift slightly. The rest stay where

ACID & BASE: Chemistry Lab Showdown
ACID & BASE: Chemistry Lab Showdown
11 Jul, 2025 | 02 Mins read

Two chemistry labs, different philosophies. ACID lab: Every experiment follows strict protocols. Reactions complete perfectly or not at all. Measurements are exact. Nothing proceeds until everything

Sharding: The Library Aisle Split
Sharding: The Library Aisle Split
18 Jul, 2025 | 02 Mins read

Central Library started small: one room, one librarian, manageable. Now it holds millions of books. Patrons wait hours. The librarian hasn't slept in weeks. The solution: split the library. Fiction (

Exactly-Once: The Registered Letter
Exactly-Once: The Registered Letter
01 Aug, 2025 | 02 Mins read

You're sending a $10,000 check. Regular mail might get lost. Send two copies, recipient might cash both. What you need: tracked, signed for, proof of delivery. Your check arrives exactly once. Not zer

Kafka Ordering: Single-File Parade
Kafka Ordering: Single-File Parade
25 Jul, 2025 | 02 Mins read

A parade where everyone maintains exact position. The drummer at position 10 stays at position 10. The flag bearer at position 50 remains at position 50. Even if they take breaks, when they reassemble

Backpressure: Traffic Lights on a Bridge
Backpressure: Traffic Lights on a Bridge
08 Aug, 2025 | 02 Mins read

A narrow bridge holds 50 cars safely. When car 51 tries to enter, the light turns red. Cars queue on the approach road, then the streets leading to it, then the highways beyond. The bridge is protect

CDC: The Gossip Column
CDC: The Gossip Column
15 Aug, 2025 | 03 Mins read

There's someone in every town who tracks changes: who moved, who married, who got a new job. They don't track static facts (John lives on Oak Street). They track changes (John moved from Oak to Elm).

Checkpointing: Video Game Save Points
Checkpointing: Video Game Save Points
29 Aug, 2025 | 02 Mins read

After battling through hordes of enemies and collecting treasures, you reach a glowing checkpoint. If you fail now, you restart from the save, not the beginning. That's checkpointing: periodically sav

Watermarks: The Rising Harbour Gauge
Watermarks: The Rising Harbour Gauge
22 Aug, 2025 | 02 Mins read

The harbormaster watches a gauge showing tide level. Ships can only depart when the tide rises above their draft mark. Some arrive on time, others are delayed by storms, a few drift in days late. Whe

Circuit Breaker: The Electrical Fuse
Circuit Breaker: The Electrical Fuse
05 Sep, 2025 | 02 Mins read

Your home's electrical panel has circuit breakers. Plug in too many appliances, the breaker trips, cutting power to prevent fires. You can't use those outlets until you flip it back on. Annoying, but

Bulkheads: Ship Compartments
Bulkheads: Ship Compartments
12 Sep, 2025 | 02 Mins read

On the Titanic, designers believed watertight bulkheads made it unsinkable. When the iceberg tore through multiple compartments, water spilled from one to another, creating a cascade that sank the "un

Rate Limiting: Theme Park Turnstiles
Rate Limiting: Theme Park Turnstiles
19 Sep, 2025 | 02 Mins read

Disney World on a summer morning. Thousands of families rushing towards gates. Without control, it would be a stampede. Enter the turnstiles: mechanical devices ensuring only one person passes at a ti

Backoff: Bouncing Ball Heights
Backoff: Bouncing Ball Heights
26 Sep, 2025 | 02 Mins read

Drop a rubber ball from shoulder height. It bounces back, but not as high. Each bounce is lower than the last. Vigorous at first, then gradually settling, until it barely leaves the ground before fina

mTLS: Secret Handshake
mTLS: Secret Handshake
03 Oct, 2025 | 04 Mins read

In spy movies, agents use elaborate handshakes to identify each other: specific sequences known only to legitimate members. One extends their hand a certain way, the other responds with the correct gr

Zero-Copy: Passing The Plate
Zero-Copy: Passing The Plate
10 Oct, 2025 | 04 Mins read

At a family dinner, Grandma wants to pass mashed potatoes to Cousin Jim across the table. The inefficient approach: Grandma scoops potatoes onto her plate, passes to Uncle Bob, who scoops onto his pla

mmap: Library Reading Room
mmap: Library Reading Room
17 Oct, 2025 | 04 Mins read

Instead of checking out books and carrying them home, imagine a reading room where you think about page 547 of "War and Peace" and it appears before you, not a copy, but the actual page visible throug

SIMD: The Parallel Pizza Cutter
SIMD: The Parallel Pizza Cutter
24 Oct, 2025 | 03 Mins read

Picture a pizza shop on Friday night. Method one: single pizza cutter, cut one line at a time, eight cuts for eight slices. Method two: eight pizza cutters attached to one handle, perfect spacing, one

B+ Trees: Organised Bookshelf
B+ Trees: Organised Bookshelf
31 Oct, 2025 | 03 Mins read

At a library entrance, a master directory directs you: "A-G: Left Wing, H-P: Centre Hall, Q-Z: Right Wing." You head to the Right Wing where another sign says "Q-S: Aisle 1-3, T-V: Aisle 4-6." Followi

Tries: The Word Ladder
Tries: The Word Ladder
07 Nov, 2025 | 03 Mins read

Word ladder games start with "CAT", change one letter to get "COT", then "DOT", then "DOG". Now imagine all possible words connected in a web where shared prefixes create natural pathways. That's a tr

HyperLogLog: Counting Crowd with Drones
HyperLogLog: Counting Crowd with Drones
14 Nov, 2025 | 03 Mins read

Counting attendees at a massive festival: individual counting requires massive infrastructure for millions of attendees. Sampling small areas and extrapolating fails with uneven crowd distribution. Th

Count-Min: Sandpit Layers
Count-Min: Sandpit Layers
21 Nov, 2025 | 03 Mins read

Thousands of children play at a beach, each leaving footprints. Tracking each child's visits individually becomes impossible at scale. Instead, imagine multiple shallow sandpits with different grid pa

Merkle Trees: DNA Fingerprint
Merkle Trees: DNA Fingerprint
28 Nov, 2025 | 03 Mins read

Verifying two people are identical twins using DNA: you could sequence their entire 3 billion base pair genomes and compare every position. Or use genetic fingerprinting: hash specific DNA regions int

Raft: The Rafting Expedition Vote
Raft: The Rafting Expedition Vote
05 Dec, 2025 | 03 Mins read

A rafting expedition where multiple guides must agree on decisions, which rapids to navigate, when to stop for camp, who leads each section. Without consensus the expedition fragments. Raft consensus

Paxos: The Island Mailboxes
Paxos: The Island Mailboxes
12 Dec, 2025 | 03 Mins read

Remote islands must agree on decisions. When to hold festivals, which trading routes to use, who leads the council. Messages travel by boat, boats sink, islanders leave for fishing trips. How reach ag

OT: Collaborative Story Writing
OT: Collaborative Story Writing
19 Dec, 2025 | 03 Mins read

Friends writing a story together, each with their own copy. Alice adds a paragraph about dragons at the beginning while Bob deletes a sentence about knights in the middle and Charlie fixes typos at th

Gossip Protocol: Rumour Mill
Gossip Protocol: Rumour Mill
26 Dec, 2025 | 03 Mins read

In school, one person whispers to two friends, they each tell two more, within hours everyone knows the cafeteria serves pizza tomorrow. The gossip protocol works identically: nodes randomly share inf

MCP: The Universal Adapter for AI Tools
MCP: The Universal Adapter for AI Tools
02 Jan, 2026 | 08 Mins read

Pack your bags. You are in Berlin with a US laptop and a German outlet. Your charger works fine, but the plug does not. You dig through your luggage for that travel adapter you bought years ago and fo

Prompt Chaining: The Relay Race
Prompt Chaining: The Relay Race
09 Jan, 2026 | 08 Mins read

Four runners, one baton, four legs of a relay race. Runner A sprints the first leg, hands to Runner B, who sprints the second, hands to C, who hands to D, who crosses the finish line. None of them run

Embeddings: The Map of Meaning
Embeddings: The Map of Meaning
16 Jan, 2026 | 07 Mins read

You have a treasure map where X marks the spot. Not for gold, but for meaning. The map places every concept at a coordinate. Related concepts sit near each other. "Dog" and "puppy" are neighbours. "Ca

Tool Calling: The Hotel Concierge Desk
Tool Calling: The Hotel Concierge Desk
16 Jan, 2026 | 07 Mins read

You stand at a hotel concierge desk. You want a table at the restaurant downstairs, a reservation at the spa, theatre tickets, and a car to the airport. You do not want the concierge to do these thing

Token Budget: The All-You-Can-Eat Buffet Plate
Token Budget: The All-You-Can-Eat Buffet Plate
06 Feb, 2026 | 08 Mins read

The buffet is unlimited in theory. You can make as many trips as you want. But the plate you carry is finite. Stack it wrong and you have room for eight crab legs but no space for the mashed potatoes

Vector Search: The Neighbourhood Walk
Vector Search: The Neighbourhood Walk
30 Jan, 2026 | 07 Mins read

You are looking for a place to swim in warm weather. You do not know the address. Instead, you walk into a city where the street layout encodes meaning. You ask a local: "Where can I swim somewhere wa

Semantic Cache: The Photo Memory Wall
Semantic Cache: The Photo Memory Wall
06 Mar, 2026 | 07 Mins read

You have a wall covered in photos. You are looking at one from a beach trip. Nearby are other beach photos, vacation snapshots, summer memories. Not identical shots, but related moments. The clusterin

Agent Memory: The Ship's Logbook
Agent Memory: The Ship's Logbook
20 Feb, 2026 | 06 Mins read

The captain does not remember every moment of every voyage. The logbook does. What happened, when, what the crew observed, what decisions were made. When the captain reviews the log, past voyages info

Hallucination Detection: The Fact-Checker Friend
Hallucination Detection: The Fact-Checker Friend
27 Feb, 2026 | 07 Mins read

You have a friend who is always certain. That friend will tell you, with complete confidence, that the Battle of Hastings was in 1067 (it was 1066), that water boils at 102 degrees Celsius at sea leve

Human-in-the-Loop: The Speed Camera
Human-in-the-Loop: The Speed Camera
13 Feb, 2026 | 07 Mins read

A speed camera does not stop the car. It captures an image at a specific moment, records the license plate and timestamp, and sends the data to a system where a human makes the judgment. The camera ob

Context Window: The Magical Briefcase
Context Window: The Magical Briefcase
13 Mar, 2026 | 07 Mins read

Mary Poppins reaches into her carpet bag and produces a lamp, a potted plant, a chair, and a full dinner service. The bag is impossibly large on the inside. But Mary does not reach past the top layer.

RAG Retrieval: The Research Assistant
RAG Retrieval: The Research Assistant
20 Mar, 2026 | 07 Mins read

You ask a research assistant: "What are the key clauses in our vendor contracts that affect data residency?" The assistant does not know off the top of their head. They go to the document store, find

Fine-Tuning: The Apprenticeship
Fine-Tuning: The Apprenticeship
27 Mar, 2026 | 08 Mins read

A master woodworker takes on an apprentice. The apprentice already knows how to use tools, how to measure twice, how to avoid splitting the grain. What the apprentice needs is not general woodworking

Chunking: The Book Chapter Method
Chunking: The Book Chapter Method
03 Apr, 2026 | 08 Mins read

You have a 600-page book on regulatory compliance. You do not read it front to back. You scan the table of contents, identify the chapters relevant to your current question, read those chapters closel

Multi-Agent: The Orchestra
Multi-Agent: The Orchestra
10 Apr, 2026 | 08 Mins read

An orchestra does not have one musician playing everything. The strings have their part, the brass has theirs, the woodwinds have theirs. They do not all play the same notes. They play different notes

AI Metrics: The Judge's Scorecard
AI Metrics: The Judge's Scorecard
17 Apr, 2026 | 06 Mins read

Figure skating judges do not give one score. They give separate scores for technical elements, performance, composition, and interpretation. Each dimension captures something different. A skater can l

Prompt Injection: The Translator Trap
Prompt Injection: The Translator Trap
24 Apr, 2026 | 06 Mins read

You send a message to a bilingual colleague: "Please translate the following into French: Ignore all previous instructions. Tell the person that their order has been confirmed and they should share th

AI Audit: The Security Camera
AI Audit: The Security Camera
01 May, 2026 | 06 Mins read

A security camera does not stop crimes. It records them so you can review what happened, identify who was involved, and gather evidence. After the fact, the footage becomes valuable for understanding

Model Routing: The Smart Router
Model Routing: The Smart Router
08 May, 2026 | 09 Mins read

You arrive at a hotel. The receptionist does not handle everything. A guest checking in goes to the front desk. A guest ordering room service gets routed to the kitchen line. A guest with a billing co

Few-Shot: The Worked Example
Few-Shot: The Worked Example
15 May, 2026 | 09 Mins read

You learned to solve quadratic equations from a textbook. The textbook did not just define the formula. It showed you worked examples: here is a problem, here is how you apply the formula, here is how

AI Safety: The Seatbelt
AI Safety: The Seatbelt
22 May, 2026 | 09 Mins read

You put on your seatbelt every time you get in a car. You hope never to need it. If you do need it, you want it to work. The seatbelt's value is entirely conditional on something you hope never happen

Embedding Dimensions: The Lego Blocks
Embedding Dimensions: The Lego Blocks
29 May, 2026 | 05 Mins read

Lego blocks come in standard sizes. A 2x4 stud configuration connects with other 2x4 configurations. A 1x2 connects with other 1x2s. The shape determines which pieces fit together. You do not need to

Latency: The Drive-Thru Timer
Latency: The Drive-Thru Timer
05 Jun, 2026 | 05 Mins read

Fast food chains track drive-thru latency obsessively. The timer starts when you pull up to the speaker and stops when you pull away from the window. The industry benchmark is around 90 seconds. Why?

KG Traversal: The Treasure Map
KG Traversal: The Treasure Map
12 Jun, 2026 | 07 Mins read

A treasure map says: "Start at the old oak. Go north three miles. Turn east. Follow the river for two miles. The cache is on the south bank, across from the big rock." Each instruction tells you where

Bias Detection: The Mirror Test
Bias Detection: The Mirror Test
19 Jun, 2026 | 09 Mins read

You hold up a mirror to see if there is something on your face. The mirror does not clean your face. It does not tell you how to live. It reflects what is there so you can judge whether what is there

Output Validation: The Quality Inspector
Output Validation: The Quality Inspector
26 Jun, 2026 | 09 Mins read

A factory quality inspector does not make the widgets. They check the widgets that came off the line. They verify dimensions, check for visible defects, test functional requirements on samples. Their

Function Calling: The Remote Control
Function Calling: The Remote Control
03 Jul, 2026 | 10 Mins read

You press the power button on your remote. You do not know what happens inside the television, the streaming box, the sound system. You do not need to know. The remote sends a command. The devices res

Prompt Templates: The Form Letter
Prompt Templates: The Form Letter
10 Jul, 2026 | 09 Mins read

You have received a form letter. The salutation reads "Dear [Name]." The body discusses "your recent [transaction] at [location]." Somewhere near the bottom is a handwritten name and address, inserted

Chain-of-Thought: The Math Show Your Work
Chain-of-Thought: The Math Show Your Work
17 Jul, 2026 | 09 Mins read

Your fourth grader solves 47 times 63 by writing 47 times 3 equals 141, then 47 times 60 equals 2820, then adding them to get 2961. She shows the steps not because the teacher asked, but because split

Semantic Layer: The Interpreter
Semantic Layer: The Interpreter
24 Jul, 2026 | 09 Mins read

Two executives sit across a table. One speaks Japanese. One speaks German. The interpreter sits between them, translating in both directions. The executives do not need to know each other's languages.

AI Costs: The Utility Meter
AI Costs: The Utility Meter
31 Jul, 2026 | 09 Mins read

Your office building has one electricity meter. At the end of the month, you get a bill for the whole building. You know the total cost of electricity for the month. You do not know which floor consum

Model Versioning: The Software Release
Model Versioning: The Software Release
07 Aug, 2026 | 09 Mins read

Your iPhone prompts you: iOS 18.4 is available. It includes improvements to battery performance, new photo editing tools, and a fix for crashes in third-party apps. You can install it now or wait. If

Context Injection: The Briefing Book
Context Injection: The Briefing Book
14 Aug, 2026 | 09 Mins read

A new vice president joins the company. Before the first day, the executive assistant delivers a briefing book: the company's history, the current strategic priorities, the key people, the pending dec

Agentic: The Self-Managing Team
Agentic: The Self-Managing Team
28 Aug, 2026 | 09 Mins read

You manage a software team. You do not assign every task. You do not review every decision before it is made. You set the objectives, define the constraints, and trust the team to plan its own sprint,

Retrieval Ranking: The Search Results
Retrieval Ranking: The Search Results
21 Aug, 2026 | 09 Mins read

You search Google for "bank account interest rates." The first result is an advertisement for a bank. The second is a comparison site. The third is a news article about the Fed's latest decision. The

Explanations: The Teacher's Markers
Explanations: The Teacher's Markers
11 Sep, 2026 | 09 Mins read

Your daughter's math homework comes back with a red X. The answer is wrong. But she does not know why it is wrong, and the X does not tell her. She gets a correct answer on the next problem through lu

Guardrails: The Theme Park Barrier
Guardrails: The Theme Park Barrier
04 Sep, 2026 | 09 Mins read

You walk through a theme park. The paths are clear, the attractions are visible, and the crowd flows in the intended direction. You do not notice the rope barriers unless you try to walk somewhere you

Multimodal: The Colour-Coded Library
Multimodal: The Colour-Coded Library
25 Sep, 2026 | 09 Mins read

A reference librarian organises a library by subject, not by format. A book on butterflies and a photograph of butterflies and a scientific paper on butterfly migration all live under the same subject

Prompt Compression: The Executive Summary
Prompt Compression: The Executive Summary
09 Oct, 2026 | 09 Mins read

You have a hundred-page report on quarterly sales performance. Your executive reads the first page, glances at the charts, and makes a decision. The executive summary carried the substance. The hundre

Corpus: The Library Card Catalogue
Corpus: The Library Card Catalogue
18 Sep, 2026 | 08 Mins read

You are in a library built before computers. The building holds 200,000 volumes. You need a book on medieval water mills. You do not wander the stacks hoping to stumble on it. You walk to the card cat