Simor
Multi-Agent Failure Modes: What Breaks When Agents Call Agents

Multi-Agent Failure Modes: What Breaks When Agents Call Agents

Simor Consulting | 24 Jun, 2026 | 12 Mins read

A single agent can already fail in difficult ways: a tool times out after applying a change, retrieved evidence is stale, or an answer invents a policy. With one decision loop, there are fewer handoffs to reconstruct. That does not make the system safe, but it gives an incident responder a smaller execution path to investigate.

Multi-agent systems add handoffs that can amplify these failures. The failure of one agent can cascade through the system in ways that are difficult to predict and harder to debug. When agent A calls agent B to complete a subtask, agent A’s correctness now depends on agent B’s correctness, which may depend on agent C’s correctness. A failure at any point in this chain propagates upward.

The failure may be obvious: agent B returns an error. Or it may be subtle: agent B returns a plausible but incorrect result that agent A incorporates into its reasoning without detecting the error. The subtle failures are the dangerous ones because they produce confident, wrong outputs that propagate through the system as verified facts.

Cascading Hallucination

The most dangerous multi-agent failure mode is cascading hallucination. Agent B hallucinates a fact. Agent A receives the hallucinated fact, treats it as ground truth because it came from another agent, and builds further reasoning on top of it. The hallucination is now embedded in agent A’s context and will be presented to the user with the confidence of a verified fact.

This is worse than single-agent hallucination for two reasons. First, the user may trust multi-agent output more than single-agent output. The assumption is that multiple agents checking each other’s work should reduce errors. In practice, without explicit verification mechanisms, multiple agents can amplify errors rather than catching them. Second, the hallucination is harder to trace. In a single-agent system, you can see the hallucination in the agent’s output. In a multi-agent system, the hallucination is buried in an intermediate agent’s response that was consumed by the calling agent and is not directly visible in the final output.

A Concrete Scenario

Consider a customer support system with three agents. Agent A is the orchestrator that receives the customer query. Agent B is a knowledge agent that searches documentation and returns facts. Agent C is a policy agent that checks company policies and returns guidance.

A customer asks about their refund eligibility. Agent A delegates to Agent B to find the refund policy. Agent B searches its knowledge base and hallucinates that the policy was updated last week to extend the refund window from 30 to 60 days. This is wrong. The policy has not changed. But Agent B presents it confidently with a plausible-sounding effective date.

Agent A receives this information and delegates to Agent C to check whether this customer qualifies under the 60-day window. Agent C confirms that the customer qualifies because they are within 45 days of purchase. Agent A tells the customer they are eligible for a refund under the extended 60-day policy.

The customer contacts the company expecting a refund under a policy that does not exist. The hallucination travelled through three agents, each of which trusted the previous agent’s output, and emerged as a confident, specific, wrong answer.

Mitigation: Inter-Agent Verification

The mitigation is risk-based source verification, not agreement between agents. When agent A receives a result from agent B, it should carry the source reference, the relevant passage or structured value, and the version or retrieval time needed to check it. Another agent can help inspect the evidence, but repeating the same unsupported claim is not independent confirmation.

Choose checks by the consequence of being wrong, the authority and freshness of the source, and whether the proposed action can be reversed. In the refund example, verify the eligibility rule against the authoritative policy and the purchase date against the account system before promising or issuing a refund. If either source is unavailable or conflicts with the claim, withhold the eligibility decision and route it for review. For a low-impact internal summary, sampling non-critical claims may be sufficient; requirements that affect permissions, money, or customer commitments need explicit checks.

An agent’s self-reported confidence is not permission to skip those checks. A confident answer can still be wrong, and several agents can share the same mistaken premise. Record unresolved claims and failed source checks in the handoff so the next agent cannot silently promote them to verified facts. Verification effort belongs in an explicit risk policy, not in a model’s assessment of its own certainty.

Context Window Exhaustion

Each agent in a multi-agent system has its own context window. When agent A calls agent B, agent A must pass enough context for agent B to do its work. If agent B then calls agent C, the context must be further compressed or truncated. By the time you reach three or four levels of delegation, the context available to the deepest agent may be insufficient for accurate work.

Consider a scenario where agent A is orchestrating a complex research task. It has accumulated 80,000 tokens of context from previous steps. It delegates to agent B to analyse a specific document. Agent B receives the document (20,000 tokens) plus a summary of agent A’s context (10,000 tokens). Agent B then delegates to agent C to extract key entities from a section of the document. Agent C receives a compressed version of the document section (5,000 tokens) plus a summary of agent B’s reasoning (3,000 tokens).

Agent C misses entities that were in the full document but not in the compressed version. Agent B’s analysis is incomplete because it was based on agent C’s incomplete extraction. Agent A’s final output is wrong, but the error originated three levels deep in the delegation chain, in a context compression step that is not visible in the final trace.

Mitigation: Context Management Discipline

Each agent should pass the minimum context needed for the sub-agent to do its work, not the entire conversation history. Critical context like specific constraints, quality requirements, and known facts should be explicitly included. Background context that is nice to have but not essential should be omitted to preserve context window space.

Hierarchical summarisation can help. Agent A provides agent B with a summary of the relevant context plus the specific task. Agent B provides agent C with an even more focused summary. Neither summarisation nor truncation guarantees that the task’s requirements survive. Keep authoritative evidence accessible by reference, and check each handoff against the requirements rather than assuming a fluent summary is complete.

Context pinning is another technique. Identify the most critical pieces of context (the user’s original request, key constraints, known facts) and pin them at every level of delegation. Pinned context is always included, even if it means reducing other context. This ensures that the most important information survives the delegation chain.

Monitoring Context Degradation

Check retained requirements and evidence at each delegation level. Before delegating, identify the constraints that determine whether the result is usable: tenant scope, policy version, relevant dates, allowed actions, and unresolved questions. After summarisation, verify that each required item is still present or retrievable and that the receiving agent’s result respects it. Keep the original evidence available for that comparison.

In the refund scenario, a summary that retains the customer’s purchase date but drops a policy exception is incomplete even if it retains nearly all the original tokens. A much shorter summary may be adequate if it preserves the applicable rule, exception, and source reference. Token counts measure context size and cost, not the proportion of meaning retained.

Use incident-derived evaluation cases that deliberately omit or contradict one required item. Check whether the receiving agent requests the missing evidence or stops the decision, rather than filling the gap with a guess. Track missing requirements, unsupported claims, and incorrect decisions separately from token savings. If a critical requirement disappears, repair the handoff or reduce delegation depth before continuing.

Dead Loops and Infinite Delegation

Multi-agent systems can enter loops where agent A calls agent B, agent B calls agent C, and agent C calls agent A. Without loop detection, the agents keep delegating to each other indefinitely, consuming tokens and adding latency without making progress.

A delegation stack can detect explicit cycles when the orchestrator tracks stable task identities. If a task is already active in the chain, reject the repeated delegation and return a structured stop reason. Do not force the calling agent to guess an answer. Changed wording can conceal the same task, so cycle detection needs independent execution limits as well.

The harder problem is implicit loops or infinite delegation behaviour. Agent A is unsure about a decision and delegates to agent B for a second opinion. Agent B is also unsure and delegates to agent C. Agent C is also unsure and delegates back to agent A, not because of a circular dependency but because all three agents have the same uncertainty threshold and none of them can resolve the uncertainty.

Mitigation: Delegation Depth Limits

Set a maximum depth for the delegation chain. At the limit, return control with verified partial results and a stop reason, or escalate for human review. Reaching the limit is not a reason to manufacture an answer or execute an unverified action.

Choose the depth limit from the task graph and test it against representative workloads. Depth alone does not bound cost: a shallow agent can repeatedly call siblings or retry tools. The orchestration layer also needs a shared per-task call budget and deadline, with bounded retries that consume the same budget. Reserve capacity before dispatch so parallel branches cannot each spend the remaining allowance.

Use missing evidence or a failed check to decide when to stop or request help. An agent’s confidence must not override the source-verification policy. Repeated delegation without new evidence should return an unresolved result, not another request for the same opinion.

Loop Detection at the Orchestration Layer

The orchestrator agent should track delegation patterns and detect loops before they consume significant resources. If the orchestrator sees that agent B has been called three times in the last ten seconds with similar tasks, it should stop delegating to agent B and handle the task differently. This is rate-limiting applied to delegation.

The orchestrator should also monitor token consumption per task. If a task has consumed more tokens than expected without producing a result, it should abort the delegation chain and return a partial result with an explanation. Token-based abort prevents runaway delegation from generating unexpected costs.

Inconsistent State

When multiple agents operate on shared state, they can create inconsistent state. Agent A reads a customer record, decides to update the address, and sends the update. Agent B reads the same customer record before agent A’s update is applied, decides to update the phone number, and sends its update. Depending on the update mechanism, one update may overwrite the other.

This is the classic concurrent modification problem from distributed systems. It applies to multi-agent systems whenever agents share state through external systems like databases or APIs. The problem is worse in multi-agent systems because agents may not be aware that other agents are operating on the same state.

Mitigation: Optimistic Concurrency Control

Each agent reads the state with a version identifier enforced by the storage or API layer. When updating, the agent includes that version as a precondition. The server must check the version and apply the write atomically; a separate read followed by an unconditional write still has a race. If the version has changed, reject the update so the caller can re-read and reconsider. For HTTP APIs, a strong ETag with If-Match is one way to express that precondition.

The retry logic adds complexity to the agent. The agent must detect the rejection, re-read the current state, re-evaluate its decision based on the updated state, and retry the update. If the state has changed significantly, the agent’s original decision may no longer be valid. The agent must handle this gracefully rather than blindly retrying.

Worked Example: Two Updates, One Stale Version

Consider a fictional support system where two agents read customer record version 17. The address agent proposes a new delivery address. The contact agent proposes a new phone number, but its write includes the whole record, including the old address. Without a precondition, the contact agent can silently restore the old address after the address agent has updated it.

Assume this API supplies strong ETags and requires If-Match for updates. Both agents initially receive ETag "17". The server accepts the address update with If-Match: "17" and advances the record to version 18. When the contact update arrives with the same precondition, the server returns 412 Precondition Failed without applying it. The address remains intact. This is the behaviour defined for a failed If-Match precondition in HTTP Semantics, RFC 9110.

The contact agent now reads version 18, verifies that the phone change is still authorised, and submits only the intended change against that version. If the server accepts it, version 19 contains both the new address and the new phone number. If the precondition fails again, the retry consumes the original task’s budget; exhausting that budget stops the write and returns a conflict for review. These version numbers illustrate the sequence, not a required ETag format.

A non-production test should force both agents to read version 17 before either writes. Assert that the stale write is rejected, that the rejection leaves the address unchanged, and that a bounded retry preserves both intended changes. Also force repeated conflicts and check that the stop path makes no further write. Save the precondition, server outcome, and resulting version in the trace without logging customer PII. A plausible final answer from the agents is not evidence that the record stayed consistent.

This example prevents a lost update to one record. It does not make a multi-record workflow transactional or solve a timeout after an external side effect. Those need their own transaction, reconciliation, and idempotency controls.

State Ownership

Another mitigation is to designate a single agent as the owner of each piece of state. If agent A owns customer records and agent B needs to update a customer, agent B delegates the update to agent A rather than performing it directly. This serialises access through the owning agent.

State ownership prevents concurrent modification only if the owner serialises writes and all writers use that path. A shared agent name alone provides no locking. Enforced ownership may create bottlenecks. If agent A is the sole owner of customer records and ten agents need to update customers simultaneously, agent A becomes a serialisation point. The throughput is limited by agent A’s processing speed.

The trade-off depends on your concurrency requirements. If concurrent modifications are rare, optimistic concurrency control is simpler and higher-throughput. If concurrent modifications are common, state ownership provides stronger consistency guarantees at the cost of throughput.

Idempotent Operations

Design tool operations to be idempotent. If agent A sends the same update twice (because it retried after a timeout), the second update should have no effect. Idempotency requires the tool to detect duplicate requests, usually through a request ID or idempotency key.

Idempotency is essential for reliable multi-agent systems because agents will retry failed operations. Without idempotency, a retry can create duplicate records, send duplicate emails, or apply duplicate charges. The tool layer must handle idempotency, not the agent layer, because the agent does not know whether the tool received the first request.

Cascading Latency

Each agent call adds latency. Agent A waits for agent B, which waits for agent C. The total latency is the sum of all agent call latencies plus network overhead. In a chain of four agents, each taking two seconds, the total latency is eight seconds before the user receives a response.

Latency compounds non-linearly in practice because each agent may make multiple tool calls, and the tool calls add their own latency. A multi-agent system where each agent makes two tool calls at 500ms each, in a chain of three agents, produces a total latency of at least three seconds from tool calls alone, plus model inference time at each level.

Parallel Execution

The primary mitigation is parallel execution of independent subtasks. If agent A needs results from both agent B and agent C, and B and C are independent, call them in parallel rather than sequentially. This reduces the total latency to the maximum of the individual latencies rather than their sum.

Parallel execution requires the orchestrator to identify independence. If agent B’s task depends on agent C’s output, they cannot be parallelised. If they are independent, parallelisation cuts latency roughly in half for two parallel calls, in thirds for three, and so on.

Timeout Enforcement

Each agent call should have a timeout. If the sub-agent does not respond within the timeout, the calling agent should handle the timeout gracefully. Three options exist. Retry the call (appropriate for transient failures). Use a degraded fallback like a cached result or a simplified heuristic (appropriate when some answer is better than no answer). Return an incomplete result with a clear explanation that some subtasks timed out.

The timeout should be set based on the expected latency of the sub-agent’s task plus a buffer for variance. A simple retrieval task might have a five-second timeout. A complex analysis task might have a thirty-second timeout. Do not set a single global timeout for all agent calls because different tasks have different latency profiles.

Circuit Breaking

If a sub-agent consistently fails or times out, stop calling it. The circuit-breaker pattern applies to multi-agent delegation: if agent B fails three times in a row, open the circuit and stop delegating to agent B. Route around the failure by using a fallback or reducing the scope of the task.

The circuit breaker should have a half-open state where it periodically retries the failed agent to see if it has recovered. If the retry succeeds, close the circuit and resume normal delegation. If it fails, keep the circuit open.

Failure Mode Summary

Cascading hallucination requires risk-based checks against authoritative sources. Context window exhaustion requires retained-requirement and evidence checks, supported by context pinning and careful summarisation. Dead loops require delegation depth limits, shared execution budgets, and loop detection. Inconsistent state requires optimistic concurrency control, state ownership, or idempotent operations. Cascading latency requires parallel execution, timeout enforcement, and circuit breaking.

The common thread is that multi-agent systems need the same distributed-systems patterns that microservices need: circuit breakers, timeouts, retries, idempotency, and eventual consistency. The models are non-deterministic, which adds a layer of complexity that deterministic microservices do not have. But the architectural patterns for handling failure, latency, and consistency are the same patterns that distributed systems engineers have been applying for decades.

Start with a single agent and add multi-agent delegation only when the task genuinely requires it. Every level of delegation adds failure modes, latency, and debugging complexity. The question is not whether multi-agent architectures are possible but whether the delegation justifies the operational overhead for your specific use case. Compare the simpler design against the delegated one on incident-derived cases, including failed source checks, missing requirements, budget exhaustion, and conflicting writes. Add delegation only when the measured benefit justifies those extra failure paths.

Can your agents stop before a failure spreads?

Use the free, anonymous AI Production Scorecard to review Guardrails and Budget governance alongside the other five control layers. Get full results immediately, without an email address. The scores suggest review priorities; they do not verify that your agents are safe.

Similar Articles

AI Agent Platforms Compared: CrewAI, AutoGen, and LangGraph for Mid-Market Operations
AI Agent Platforms Compared: CrewAI, AutoGen, and LangGraph for Mid-Market Operations
10 Jul, 2026 | 10 Mins read

You have signed off on an AI initiative. Your team has a real workflow in mind. Say, triaging inbound operations tickets, drafting first-pass vendor reviews, or reconciling exception cases across thre

How to Measure AI Consulting ROI: A Practitioner's Guide
10 Jul, 2026 | 06 Mins read

You approved the AI initiative. You hired the consultants. Six months later, the CFO is asking what you got for the spend. If your answer is a slide deck of demos and "strategic enablement," the budge

Practical AI Governance for Mid-Market Companies: A Framework That Won't Break Your Team
10 Jul, 2026 | 08 Mins read

A mid-market operations director told us recently that her legal team had forwarded a 47-page enterprise AI governance policy and asked her to "just adapt it." She has a five-person analytics team, tw

Practical LLM Evaluation Metrics Beyond Vibes: Building a Repeatable Scoring Pipeline
Practical LLM Evaluation Metrics Beyond Vibes: Building a Repeatable Scoring Pipeline
10 Jul, 2026 | 11 Mins read

The demo looked great. The model summarised the document cleanly, answered the test question correctly, and produced prose that read well enough to ship. Two weeks later it is in production, and the c

Lightweight MLOps for Mid-Market Teams: Ship Models Without a Platform Engineering Org
Lightweight MLOps for Mid-Market Teams: Ship Models Without a Platform Engineering Org
10 Jul, 2026 | 11 Mins read

A head of ML at a 120-person company told us recently that his team had spent nine months trying to stand up a "proper MLOps platform." They had evaluated three orchestration tools, designed a feature

Building AI-Ready Data Pipelines: Key Architecture Considerations
Building AI-Ready Data Pipelines: Key Architecture Considerations
04 Mar, 2025 | 02 Mins read

Data pipelines built for business intelligence often fail when supporting AI workloads. The root cause is usually architectural: BI pipelines assume bounded, relatively static datasets, while AI syste

The Modern Data Stack for AI Readiness: Architecture and Implementation
The Modern Data Stack for AI Readiness: Architecture and Implementation
28 Jan, 2025 | 03 Mins read

Existing data infrastructure often cannot support ML workflows. The modern data stack offers a foundation, but it requires adaptation to become AI-ready. This article covers building a data architectu

Model Context Protocol: The USB-C Moment for AI Tooling
Model Context Protocol: The USB-C Moment for AI Tooling
16 Jul, 2026 | 21 Mins read

Every AI agent system eventually faces the same problem. You have built a capable language model. You want it to interact with your tools, your data, your APIs. So you write a custom integration layer

The AI Model Registry: Managing Model Versions, Lineage, and Governance
The AI Model Registry: Managing Model Versions, Lineage, and Governance
09 Sep, 2026 | 20 Mins read

When a model stops working correctly in production, the first question is always the same: what changed? Which version of the model is currently deployed? What training data was used? What evaluation

Fine-Tuning vs RAG vs Prompt Engineering: Decision Framework
Fine-Tuning vs RAG vs Prompt Engineering: Decision Framework
29 Aug, 2026 | 14 Mins read

Teams new to applied AI often fixate on which foundation model to use. The more important decision is how to shape the model's behaviour for your specific task. The three primary levers are prompt eng

Building an Eval Harness That Ships With Every Release
Building an Eval Harness That Ships With Every Release
18 Jun, 2026 | 10 Mins read

A fintech company shipped a prompt update to their underwriting assistant on a Friday afternoon. The update improved response quality on three of four test cases. On Monday, the risk team reported tha

Model Gateway Patterns: When to Route, When to Fail Over
Model Gateway Patterns: When to Route, When to Fail Over
20 Jun, 2026 | 11 Mins read

The first time your model provider has an outage at 2 AM and your entire application goes dark, you learn something important about architectural dependencies. The second time it happens, you start bu

Tool Governance for MCP: Scoping Permissions Before They Drift
Tool Governance for MCP: Scoping Permissions Before They Drift
21 Jun, 2026 | 10 Mins read

When an AI agent can call external tools, the security boundary shifts from the model to the tool layer. The model generates a request to call a tool. The tool executes against real systems, reading d

AI Observability Beyond Logging: Trace Replay, Incident Forensics, and Cost Attribution
AI Observability Beyond Logging: Trace Replay, Incident Forensics, and Cost Attribution
22 Jun, 2026 | 11 Mins read

Traditional application observability focuses on three signals: request latency, error rates, and resource utilisation. If the request returns a 200 in under two hundred milliseconds, the system is he

MCP in Production: Registry, Auth, and Permission Models
MCP in Production: Registry, Auth, and Permission Models
23 Jun, 2026 | 11 Mins read

The Model Context Protocol gives AI agents a standardised way to discover and invoke external tools. In development, MCP works well with a local server running on localhost and a handful of tools. The

Agent Guardrails: Containing What an Agent Can Do in Production
Agent Guardrails: Containing What an Agent Can Do in Production
25 Jun, 2026 | 09 Mins read

Input guardrails check whether a user prompt is safe. Output guardrails check whether a model response is appropriate. Agent guardrails check whether the actions an agent takes are within bounds. Thes

From Single-User to Multi-User: Ten Tenancy Controls for Production AI
From Single-User to Multi-User: Ten Tenancy Controls for Production AI
26 Jun, 2026 | 11 Mins read

An AI application built for a single user has no tenancy concerns. The user is the user. There is no data isolation problem because there is only one data set. There is no cost attribution problem bec

A2A and MCP: How Agent-to-Agent Protocol Fits the Control Layer Model
A2A and MCP: How Agent-to-Agent Protocol Fits the Control Layer Model
28 Jun, 2026 | 09 Mins read

Google announced the Agent-to-Agent protocol, A2A, as a standard for how AI agents communicate with each other. This sits alongside the Model Context Protocol, MCP, which standardises how agents acces

OpenAI vs Anthropic vs Google: Model Provider Failover Strategies
OpenAI vs Anthropic vs Google: Model Provider Failover Strategies
29 Jun, 2026 | 10 Mins read

Every major model provider has had outages. OpenAI has gone down during peak hours. Anthropic has experienced degraded performance. Google Gemini has had API issues. If your application depends on a s

AI Rollback Patterns: When to Roll Back a Prompt, a Model, or the Whole Release
AI Rollback Patterns: When to Roll Back a Prompt, a Model, or the Whole Release
27 Jun, 2026 | 11 Mins read

Software rollbacks are well-understood. You deploy a new version, detect an issue, and roll back to the previous version. The rollback is atomic: the entire application reverts to the previous state.

AI Middleware: The Missing Abstraction Between Your App and the Model
AI Middleware: The Missing Abstraction Between Your App and the Model
30 Jun, 2026 | 09 Mins read

When web applications needed to talk to databases, the industry created ORMs and connection pools. When microservices needed to talk to each other, the industry created API gateways and service meshes

Prompt Versioning in Git: Prompts as Code, Not Configuration
Prompt Versioning in Git: Prompts as Code, Not Configuration
01 Jul, 2026 | 10 Mins read

Prompts are the most frequently changed component of an AI application. They are updated to fix edge cases, improve output quality, accommodate new use cases, and adapt to model behaviour changes. Des

How a retailer reduced inference latency 90% with feature store caching
How a retailer reduced inference latency 90% with feature store caching
21 Apr, 2026 | 04 Mins read

A mid-market e-commerce retailer with roughly $200M in annual revenue had invested eighteen months building a product recommendation engine. The models were accurate. Offline evaluation showed meaning

The 7-step vector database selection checklist
The 7-step vector database selection checklist
26 Apr, 2026 | 06 Mins read

Most vector database selection failures come down to one mistake: picking the technology before mapping the workload. Teams benchmark embedding search speed on a curated dataset, pick the fastest opti

The open-source LLM landscape just shifted: again
The open-source LLM landscape just shifted: again
02 May, 2026 | 03 Mins read

Three releases in the last six weeks have redrawn the open-source LLM map. Meta shipped Llama 4 with a mixture-of-experts architecture that narrows the gap with proprietary frontier models. Mistral re

Build vs buy: a decision tree for AI infrastructure
Build vs buy: a decision tree for AI infrastructure
03 May, 2026 | 06 Mins read

Every AI infrastructure team eventually faces the same argument. One faction wants to build a custom solution because the commercial options do not handle their specific requirements. The other factio

Why every cloud provider launched an AI operating system this year
Why every cloud provider launched an AI operating system this year
09 May, 2026 | 03 Mins read

AWS announced Bedrock Studio. Google shipped Vertex AI Platform as a unified surface. Azure consolidated its AI offerings under a single "AI Foundry" brand. Databricks, Snowflake, and even Cloudflare

The vector database that couldn't scale, and what we did instead
The vector database that couldn't scale, and what we did instead
12 May, 2026 | 05 Mins read

A media company with a library of twelve million articles, transcripts, and research documents had built a semantic search system on a managed vector database. The system was designed to let journalis

LLM evaluation platforms compared: LangSmith, Braintrust, Patronus
LLM evaluation platforms compared: LangSmith, Braintrust, Patronus
14 May, 2026 | 06 Mins read

Building an LLM application is the easy part. Knowing whether it works (whether it still works after you change a prompt, swap a model, or add a tool) is the hard part. LLM evaluation platforms exist

The A2A protocol and what it means for enterprise AI
The A2A protocol and what it means for enterprise AI
16 May, 2026 | 03 Mins read

Google published the Agent-to-Agent (A2A) protocol specification in late 2025 and, as of this quarter, has secured endorsement from over fifty technology companies including Salesforce, SAP, ServiceNo

Building an AI operating system for a 10,000-person company
Building an AI operating system for a 10,000-person company
19 May, 2026 | 05 Mins read

A diversified industrial company with 10,000 employees across manufacturing, logistics, and field services had accumulated forty-seven separate AI projects over three years. Each business unit had bui

A cost optimisation framework for LLM inference
A cost optimisation framework for LLM inference
24 May, 2026 | 06 Mins read

LLM inference costs follow a pattern that catches teams off guard. The first prototype costs almost nothing: a few hundred dollars a month during development. The pilot scales to a few thousand. Produ

AI spending is up 300%: where is it actually going?
AI spending is up 300%: where is it actually going?
27 May, 2026 | 03 Mins read

Enterprise AI spending increased roughly 300% year-over-year according to multiple industry surveys released this quarter. The headline number gets attention, but the breakdown is where the actionable

The observability stack: Datadog vs Grafana vs Monte Carlo
The observability stack: Datadog vs Grafana vs Monte Carlo
28 May, 2026 | 07 Mins read

Observability is not one problem. It is three. Infrastructure observability watches your servers, containers, and network. Application observability watches your code, APIs, and user-facing behaviour.

RAG frameworks head-to-head: LlamaIndex vs Haystack vs Semantic Kernel
RAG frameworks head-to-head: LlamaIndex vs Haystack vs Semantic Kernel
04 Jun, 2026 | 05 Mins read

Retrieval-augmented generation is simple in theory: retrieve relevant documents, stuff them into a prompt, get a grounded answer. In practice, the retrieval step is where most RAG applications fail. T

Designing guardrails: a practical architecture guide
Designing guardrails: a practical architecture guide
21 Jun, 2026 | 06 Mins read

The guardrail problem in AI is a tension between two failure modes. Too few guardrails and the system produces harmful, inaccurate, or brand-damaging outputs. Too many guardrails and the system refuse

When your AI vendor goes bankrupt: surviving platform lock-in
When your AI vendor goes bankrupt: surviving platform lock-in
23 Jun, 2026 | 05 Mins read

A healthcare analytics company received notice on a Tuesday afternoon that their primary AI infrastructure vendor was filing for Chapter 7 bankruptcy. The platform hosted their patient risk stratifica

Real-time fraud detection: from proof-of-concept to production in 90 days
Real-time fraud detection: from proof-of-concept to production in 90 days
30 Jun, 2026 | 05 Mins read

A payment processor handling twelve million transactions per day had a fraud detection system that was accurate but slow. The system reviewed transactions in batch, four times per day. A fraudulent tr

Graph databases for AI: Neo4j vs Amazon Neptune vs ArangoDB
Graph databases for AI: Neo4j vs Amazon Neptune vs ArangoDB
02 Jul, 2026 | 05 Mins read

Graph databases went from niche to essential as AI applications discovered that relationships matter. RAG applications that only search by vector similarity miss the connections between entities. Reco

The hidden environmental cost of your RAG pipeline
The hidden environmental cost of your RAG pipeline
04 Jul, 2026 | 03 Mins read

Retrieval-augmented generation is the default architecture for enterprise AI applications that need to ground model outputs in organisational data. The standard RAG pipeline ingests documents, chunks

Synthetic data tools: Gretel, Mostly AI, Tonic
Synthetic data tools: Gretel, Mostly AI, Tonic
09 Jul, 2026 | 05 Mins read

Real data is expensive, restricted, and often unusable. Privacy regulations block access to customer records. Data sharing agreements prevent using production data in development environments. Class i

Agentic AI in production: hype vs reality check
Agentic AI in production: hype vs reality check
18 Jul, 2026 | 03 Mins read

Agentic AI (systems where language models plan, execute multi-step tasks, and use tools autonomously) is the dominant topic at every AI conference, vendor pitch, and engineering blog. The hype is inte

Capacity planning for vector databases
Capacity planning for vector databases
19 Jul, 2026 | 07 Mins read

Vector database capacity planning fails in predictable ways. Teams estimate storage based on vector count alone and discover at 60% capacity that memory consumption is growing faster than disk because

Prompt management tools: PromptLayer, Humanloop, Promptfoo
Prompt management tools: PromptLayer, Humanloop, Promptfoo
22 Jul, 2026 | 05 Mins read

Prompts are code. They have versions, they break when changed carelessly, and they need testing. Yet most teams manage prompts as string literals in source files or as unversioned entries in a databas

The $100B AI infrastructure buildout: who benefits?
The $100B AI infrastructure buildout: who benefits?
25 Jul, 2026 | 03 Mins read

The combined AI infrastructure capital expenditure of the four largest cloud providers exceeded $100 billion in the trailing twelve months. Microsoft, Google, Amazon, and Meta are building data centre

Setting up a model registry: the minimal viable approach
Setting up a model registry: the minimal viable approach
02 Aug, 2026 | 06 Mins read

A model registry is the version control system for your trained models. Without one, teams track model versions by filename, store artifacts in ad-hoc cloud storage locations, and discover which model

Privacy-preserving computation: differential privacy tools compared
Privacy-preserving computation: differential privacy tools compared
06 Aug, 2026 | 06 Mins read

Publishing aggregate statistics about a dataset sounds safe. The average salary in a department. The number of users in a geographic region. The distribution of query types in a search engine. But agg

Scaling a recommendation engine from 1K to 10M users
Scaling a recommendation engine from 1K to 10M users
11 Aug, 2026 | 06 Mins read

A video streaming platform grew from 1,000 beta users to 10 million subscribers over thirty months. Their recommendation system was rebuilt three times during this period. Each rebuild was triggered n

How to run an AI architecture review
How to run an AI architecture review
12 Aug, 2026 | 07 Mins read

An architecture review for an AI system catches design flaws at the cheapest possible stage: before implementation. A data pipeline that cannot handle the expected volume, a model serving architecture

MCP server ecosystem: what's production-ready in 2026?
MCP server ecosystem: what's production-ready in 2026?
13 Aug, 2026 | 05 Mins read

The Model Context Protocol (MCP) was released in late 2024 as a standardised way for AI models to interact with external tools and data sources. By mid-2026, the server ecosystem has grown to hundreds

The observability maturity model for AI systems
The observability maturity model for AI systems
16 Aug, 2026 | 07 Mins read

Most AI systems in production operate with observability that was designed for traditional software. Teams monitor CPU, memory, network, and error rates. These metrics tell you whether the server is r

Agent frameworks compared: LangGraph vs CrewAI vs AutoGen
Agent frameworks compared: LangGraph vs CrewAI vs AutoGen
20 Aug, 2026 | 06 Mins read

Single-agent applications (one LLM, one set of tools, one task) are straightforward to build and debug. The agent receives input, calls tools, produces output. When multi-step reasoning or collaborati

Embedding models compared: OpenAI, Cohere, Voyage, and open-source options
Embedding models compared: OpenAI, Cohere, Voyage, and open-source options
27 Aug, 2026 | 04 Mins read

Choosing an embedding model is one of the first decisions you make when building a retrieval-augmented generation system, and it is one of the hardest to reverse. The model you pick determines your ve

LLM cost calculator: estimating spend before you deploy
LLM cost calculator: estimating spend before you deploy
30 Aug, 2026 | 05 Mins read

Teams approve LLM projects based on per-query cost estimates, then get blindsided by the actual invoice. The gap between estimate and reality is not a rounding error. It is a structural problem: the e

The consolidation wave: 5 AI acquisitions that reshaped the market this quarter
The consolidation wave: 5 AI acquisitions that reshaped the market this quarter
02 Sep, 2026 | 04 Mins read

The acquisition wave in AI this quarter was not random. Five deals, each above the billion-dollar threshold, closed within weeks of each other, and they share a common logic: the companies being acqui

Why enterprises are repatriating from managed AI services
Why enterprises are repatriating from managed AI services
05 Sep, 2026 | 04 Mins read

A quiet but significant trend has emerged over the past two quarters: enterprises are moving AI workloads off managed services and back onto infrastructure they control. The pattern is not universal,

The invisible labour of maintaining AI systems in production
The invisible labour of maintaining AI systems in production
07 Sep, 2026 | 04 Mins read

Every AI demo is impressive. Every AI production system is a maintenance burden. The distance between those two statements is where most AI initiatives quietly fail. The demo shows a model producing

Text-to-SQL tools in 2026: which ones actually work?
Text-to-SQL tools in 2026: which ones actually work?
10 Sep, 2026 | 05 Mins read

Text-to-SQL has been promised for a decade. Every year, a new tool claims to convert natural language to production-ready SQL. Every year, the demos look impressive and the production deployments disa

The LLM cost optimisation playbook: 12 techniques that actually save money
The LLM cost optimisation playbook: 12 techniques that actually save money
13 Sep, 2026 | 04 Mins read

LLM costs are easy to start and hard to control. A team ships a feature that calls GPT-4, the feature works, users like it, and the invoice climbs 15 percent month over month. The cost is not a proble

Lessons from manufacturing quality control for AI system reliability
Lessons from manufacturing quality control for AI system reliability
14 Sep, 2026 | 04 Mins read

Manufacturing figured out quality control decades ago. AI is still learning the lesson the hard way. When a car leaves the factory with a defect, the manufacturer does not shrug and say "models are p

When the CDO and CTO disagreed on AI strategy, and what happened
When the CDO and CTO disagreed on AI strategy, and what happened
15 Sep, 2026 | 06 Mins read

At a mid-market insurance company with eight thousand employees, the Chief Data Officer and the Chief Technology Officer had fundamentally different views on how AI should be adopted. The CDO believed

Document intelligence platforms: AWS Textract vs Azure AI Doc Intelligence vs Google DocAI
Document intelligence platforms: AWS Textract vs Azure AI Doc Intelligence vs Google DocAI
17 Sep, 2026 | 05 Mins read

Every enterprise processes documents. Invoices, contracts, forms, receipts, medical records, insurance claims. The volume is measured in millions of pages per month for large organisations. The questi

AI chip wars: NVIDIA, AMD, Intel, and custom silicon: who wins?
AI chip wars: NVIDIA, AMD, Intel, and custom silicon: who wins?
19 Sep, 2026 | 04 Mins read

NVIDIA still dominates AI inference and training hardware, but the dominance is no longer absolute in the way it was 18 months ago. AMD has shipped competitive alternatives at lower price points. Inte

An insurance firm's journey from PDF extraction to automated underwriting
An insurance firm's journey from PDF extraction to automated underwriting
29 Sep, 2026 | 06 Mins read

A specialty insurance firm underwriting commercial property policies received submission packets as PDF documents. Each packet contained an ACORD application, loss runs from prior carriers, a statemen

Low-code AI platforms: worth it for data teams?
Low-code AI platforms: worth it for data teams?
30 Sep, 2026 | 05 Mins read

Low-code AI platforms promise to put machine learning in the hands of people who cannot write Python. The pitch is compelling: connect your data, configure a pipeline with drag-and-drop components, de

What jazz improvisation teaches us about multi-agent coordination
What jazz improvisation teaches us about multi-agent coordination
05 Oct, 2026 | 04 Mins read

The typical approach to multi-agent AI systems is choreographed. A central orchestrator assigns tasks, sequences handoffs, and controls the flow. Agent A finishes, passes to Agent B, which passes to A

How we reduced cloud data spend 40% without cutting features
How we reduced cloud data spend 40% without cutting features
06 Oct, 2026 | 06 Mins read

A media analytics company running its entire data platform on AWS was spending $480,000 per month on cloud infrastructure. The bill had grown organically over three years as the platform expanded from

The carbon footprint of training frontier models: what the latest research shows
The carbon footprint of training frontier models: what the latest research shows
10 Oct, 2026 | 04 Mins read

The energy consumption numbers for training frontier AI models have crossed a threshold that makes them difficult to ignore. Training a single large language model now consumes between 50 and 100 giga

Zero-trust architecture for data pipelines: a practical guide
Zero-trust architecture for data pipelines: a practical guide
11 Oct, 2026 | 04 Mins read

Data pipelines are soft targets. They move sensitive data across network boundaries, authenticate with service accounts that have broad permissions, and log enough information to reconstruct entire da

LLM gateway comparison: LiteLLM, Portkey, Martian
LLM gateway comparison: LiteLLM, Portkey, Martian
29 Jun, 2026 | 07 Mins read

A production AI application calls multiple LLM providers. The primary model is GPT-4o for complex reasoning, but simple classification tasks use Claude Haiku for cost savings, and the fallback for rat

The Rise of GPU Databases for AI Workloads
The Rise of GPU Databases for AI Workloads
22 Jan, 2024 | 03 Mins read

Traditional relational database management systems were designed for an era of megabyte-scale datasets and batch reporting. AI workloads demand processing terabyte-scale datasets with complex analytic

Vector Databases: The Missing Piece in Your AI Infrastructure
Vector Databases: The Missing Piece in Your AI Infrastructure
12 Jan, 2024 | 02 Mins read

Vector databases index and query high-dimensional vector embeddings. Unlike traditional databases that excel at exact matches, vector databases enable similarity search: finding items conceptually clo

Designing the Enterprise Knowledge Layer: Beyond RAG
Designing the Enterprise Knowledge Layer: Beyond RAG
16 Jan, 2026 | 14 Mins read

Most teams implement retrieval-augmented generation and call it a knowledge layer. Give the model access to a vector database, stuff in some documents, and ship. This approach works for demos. It fall

AI Agent Orchestration Patterns: From Chaining to Multi-Agent Systems
AI Agent Orchestration Patterns: From Chaining to Multi-Agent Systems
27 Jan, 2026 | 13 Mins read

A software debugging agent receives a bug report. It needs to search code, understand the error, propose a fix, write tests, and summarise for the developer. None of these steps are independent. Each

Feature Stores for AI: The Missing MLOps Component Reaching Maturity
Feature Stores for AI: The Missing MLOps Component Reaching Maturity
12 Mar, 2026 | 11 Mins read

A recommendation system team built their tenth model. Each model required feature engineering. Each feature engineering project started by copying code from the previous project, then modifying it for

Tool Calling and Function Calling: Connecting AI to Enterprise Systems
Tool Calling and Function Calling: Connecting AI to Enterprise Systems
28 Mar, 2026 | 14 Mins read

A language model that only generates text is not enough for most enterprise problems. The real value emerges when an AI system can look up your customer record, check inventory levels across warehouse

Case Study: Multi-Agent System for Supply Chain Optimisation
Case Study: Multi-Agent System for Supply Chain Optimisation
13 Jun, 2026 | 12 Mins read

A mid-size automotive parts manufacturer with operations spanning 15 countries and relationships with over 200 suppliers faced a supply chain coordination problem that was consuming too much of their

AI Infrastructure for Legacy Systems: Modernising 20-Year-Old ERPs with AI
AI Infrastructure for Legacy Systems: Modernising 20-Year-Old ERPs with AI
18 Feb, 2026 | 13 Mins read

A manufacturing company runs their operations on an ERP system installed in 2004. The vendor still supports it. The team knows how to maintain it. The integrations are stable. It works. The problem i

The AI Data Pipeline: Special Considerations for Unstructured and Structured Data
The AI Data Pipeline: Special Considerations for Unstructured and Structured Data
11 May, 2026 | 13 Mins read

Data pipelines for AI are not the same as data pipelines for traditional software systems. The outputs are different. The failure modes are different. The tolerance for data quality issues is differen

AI Observability: Monitoring Hallucinations, Latency, and Cost in Production
AI Observability: Monitoring Hallucinations, Latency, and Cost in Production
30 Apr, 2026 | 09 Mins read

Traditional software monitoring tracks CPU utilisation, memory consumption, request rates, and error counts. These metrics tell you whether your service is running and whether it is handling load. The

Semantic Caching for AI: Reducing Latency and Cost with Meaning-Based Retrieval
Semantic Caching for AI: Reducing Latency and Cost with Meaning-Based Retrieval
19 May, 2026 | 07 Mins read

Every repeated question your AI system answers is money spent and latency incurred that you did not need to. If a thousand users ask the same question in a week, running it through the language model

Evaluating LLM Providers for Enterprise: A Framework Beyond Benchmark
Evaluating LLM Providers for Enterprise: A Framework Beyond Benchmark
08 Apr, 2026 | 10 Mins read

Benchmark scores tell you how a model performs on problems that someone else chose. Your enterprise systems present different problems: your proprietary terminology, your specific data distributions,

RAG vs Fine-Tuning: Choosing the Right Approach for Your Use Case
RAG vs Fine-Tuning: Choosing the Right Approach for Your Use Case
10 Jul, 2026 | 09 Mins read

Your team has a real use case. Maybe it is a support assistant that answers from your knowledge base, a contracts reviewer that applies your house clause library, or an ops copilot that understands yo

Choosing a Vector Database for Production AI Applications
Choosing a Vector Database for Production AI Applications
10 Jul, 2026 | 12 Mins read

You have a retrieval-augmented generation proof of concept that works on a laptop. The embeddings are in a CSV file, the search is brute force, and the demo impresses the steering committee. Now someo