Simor
The ML model that predicted churn but couldn't explain why

The ML model that predicted churn but couldn't explain why

Simor Consulting | 22 Sep, 2026 | 06 Mins read

A subscription media company with 1.2 million subscribers built a machine learning model to predict churn. The model worked. It identified at-risk subscribers with seventy-nine percent precision and eighty-three percent recall, which was significantly better than the rule-based system it replaced: a simple heuristic that flagged subscribers who had not logged in for thirty days.

The problem was not the model’s accuracy. The problem was that the retention team would not use it. The team consisted of twelve people who called at-risk subscribers and offered incentives to stay. They had been doing this job for years and had strong intuitions about why subscribers churned: content fatigue, price sensitivity, life events, competitor offers. When the model flagged a subscriber as high-risk, the retention specialist needed to know why in order to craft the right intervention. A generic “the model says this person will churn” was not actionable. The specialist needed to know: is this person price-sensitive, or are they bored with the content?

The model could not answer that question. It was a gradient-boosted ensemble with two hundred features. The feature importance rankings showed which features were most predictive across the entire population (login frequency, content diversity score, billing cycle) but these global rankings did not explain any individual prediction. A subscriber might be flagged as high-risk despite having high login frequency because the combination of three other features (declining content diversity, recent support ticket, and mid-contract billing) pushed the prediction over the threshold. The model knew this. The retention specialist did not.

The Explainability Gap

The company had tried two approaches to explainability before engaging us. Both failed for different reasons.

The first attempt was LIME, Local Interpretable Model-agnostic Explanations. LIME generates a local linear approximation of the model’s behaviour around a specific prediction. For each flagged subscriber, LIME produced a ranked list of features that contributed to the prediction. The retention team found these lists confusing. The features were technical: “days since last content interaction weighted by content type diversity coefficient”, and the contribution scores were abstract numbers that did not map to the retention specialist’s mental model of subscriber behaviour.

The second attempt was to build a simpler model, a decision tree, that the team could understand. The decision tree was interpretable. Each prediction came with a clear path: “subscriber has not logged in for 14 days AND has a content diversity score below 0.3 AND is within 60 days of renewal.” The problem was accuracy. The decision tree achieved sixty-four percent precision, fifteen percentage points worse than the gradient-boosted model. The retention team understood the predictions but the predictions were wrong too often to be useful.

The company was stuck in a trade-off that is common in applied ML: accuracy versus interpretability. The accurate model was a black box. The interpretable model was inaccurate. They needed both.

The Explanation Layer

The solution was not to make the model interpretable. It was to build a separate system that translated model outputs into the language the retention team used.

We started by interviewing the retention specialists. We asked them to describe how they decided what to say to a subscriber during a retention call. Their answers revealed a taxonomy of churn drivers that they used implicitly but had never documented:

  • Content fatigue: The subscriber has consumed most of the available content in their interest areas and finds nothing new.
  • Price objection: The subscriber has compared competitor pricing and believes the service is overpriced.
  • Usage decline: The subscriber’s engagement has gradually decreased, often due to a life change (new job, new baby, relocation).
  • Service issue: The subscriber had a negative support experience and lost trust.
  • Competitor switch: The subscriber has already signed up for a competitor and is letting the subscription lapse.

Each driver had a different intervention. Content fatigue required surfacing undiscovered content. Price objection required a discount or value framing. Usage decline required a gentle re-engagement, not a hard sell. Service issue required an apology and a resolution. Competitor switch required understanding what the competitor offered and whether it could be matched.

The explanation layer mapped model features to churn drivers. A subscriber with declining content diversity and no new category exploration in sixty days was mapped to content fatigue. A subscriber who had searched the cancellation page three times in two weeks was mapped to price objection. A subscriber with a recent support ticket and a negative satisfaction survey was mapped to service issue.

This diagram requires JavaScript.

Enable JavaScript in your browser to use this feature.

The driver classification was not part of the ML model. It was a rule-based layer that operated on the model’s input features: the same features the model used to predict churn, but interpreted through the lens of the retention team’s domain knowledge. The rule-based layer did not need to be as accurate as the ML model. It needed to be plausible. If the model predicted churn with seventy-nine percent precision and the driver classification was correct seventy percent of the time, the retention specialist received the right intervention recommendation fifty-five percent of the time (0.79 x 0.70). This was significantly better than the specialist’s baseline intuition, which internal testing showed was correct approximately forty percent of the time.

The Dashboard

The retention specialist’s dashboard showed three things for each flagged subscriber: the churn risk score from the model, the likely churn driver from the explanation layer, and the recommended intervention. The specialist could accept the recommendation, choose a different intervention, or mark the subscriber as uncallable (wrong number, deceased, already cancelled).

The dashboard also showed the specialist’s override history. When a specialist consistently chose a different intervention than the recommendation for a specific driver, the system learned from the pattern. If a specialist always chose a discount for subscribers that the system classified as content fatigue, the system asked: is this specialist’s local knowledge revealing a pattern the rules missed? The override patterns were reviewed monthly and used to update the classification rules.

This feedback loop was the most valuable component of the system. The ML model predicted churn. The explanation layer translated predictions to drivers. The retention specialists applied interventions. The outcome (did the subscriber stay or leave) was tracked. The override patterns and outcome data fed back into both the model and the explanation layer. Over twelve months, the model improved from seventy-nine to eighty-two percent precision, and the driver classification accuracy improved from seventy to seventy-six percent.

What Explainability Actually Costs

Building the explanation layer took six weeks: roughly the same time as building the ML model itself. This is the hidden cost of explainability. The model is half the project. The explanation system is the other half.

The ongoing maintenance cost was also non-trivial. The churn driver taxonomy had to be updated when subscriber behaviour changed. When the company launched a new content format, a new driver category, format incompatibility, emerged. Subscribers who preferred long-form articles churned when the platform shifted to short-form video. The existing taxonomy did not have a category for this, so the explanation layer mapped these subscribers to content fatigue, which led to the wrong intervention (surfacing more short-form video to people who did not want short-form video). Adding the new driver category required interviewing specialists, defining classification rules, and testing the rules against known cases.

The rule-based classification was also brittle in edge cases. When a subscriber exhibited signals for two drivers simultaneously, declining engagement and a price comparison search, the classification chose the dominant signal. This was correct most of the time but wrong when the two signals interacted. A subscriber who was both bored with content and price-shopping needed a different intervention than either condition alone. The system handled these cases by flagging subscribers with multi-driver signals for senior specialist review, which added a manual step but avoided the worst outcomes.

The Outcome

After twelve months, the retention team’s save rate, the percentage of at-risk subscribers who were retained after intervention, increased from thirty-one percent to forty-four percent. The improvement was attributable to two factors: better targeting (the model identified at-risk subscribers earlier than the rule-based predecessor) and better intervention selection (the explanation layer matched interventions to drivers more accurately than the specialists’ intuition alone).

The retention specialists reported higher job satisfaction. Before the system, they described their work as “cold-calling people and guessing what to say.” After the system, they described it as “having a conversation with someone whose situation I understand.” The shift from guessing to understanding was the practical impact of explainability.

The Rule

If your model’s consumers need to act on its predictions, the model is only half the project. The other half is translating predictions into the language and mental model of the people who will use them. This translation does not require making the model itself interpretable. It requires building a bridge between the model’s feature space and the user’s decision space.

The bridge is custom for every organisation because every organisation has its own vocabulary for the decisions it makes. No off-the-shelf explainability tool can provide this because the tool does not know your retention team’s taxonomy of churn drivers, your clinical team’s diagnostic framework, or your fraud team’s investigation workflow. Build the translation layer. Budget for it. Maintain it. Without it, even an accurate model will sit on a shelf.

Shipping a production AI system?

Find where your AI spend leaks and where quality slips. Take the AI Production Scorecard for a fast baseline across the seven layers, or book a free AI cost review and we will turn it into a plan.

Similar Articles

The AI Model Registry: Managing Model Versions, Lineage, and Governance
The AI Model Registry: Managing Model Versions, Lineage, and Governance
09 Sep, 2026 | 20 Mins read

When a model stops working correctly in production, the first question is always the same: what changed? Which version of the model is currently deployed? What training data was used? What evaluation

MLOps vs DataOps: Understanding the Differences and Overlaps
MLOps vs DataOps: Understanding the Differences and Overlaps
08 Feb, 2024 | 03 Mins read

DataOps and MLOps both aim to improve reliability and efficiency in data-centric workflows, but they address different parts of the data science lifecycle. Understanding their boundaries helps organis

How a retailer reduced inference latency 90% with feature store caching
How a retailer reduced inference latency 90% with feature store caching
21 Apr, 2026 | 04 Mins read

A mid-market e-commerce retailer with roughly $200M in annual revenue had invested eighteen months building a product recommendation engine. The models were accurate. Offline evaluation showed meaning

The data pipeline that cost $50K/month, and the audit that found why
The data pipeline that cost $50K/month, and the audit that found why
22 Apr, 2026 | 04 Mins read

A financial services firm running analytics on trade settlement data came to us with a specific complaint: their cloud data platform cost had tripled in eighteen months, and nobody could explain why.

Migrating from batch to streaming: a 6-month journey
Migrating from batch to streaming: a 6-month journey
28 Apr, 2026 | 05 Mins read

A logistics company processing two million shipments per day ran their entire operational reporting stack on nightly batch ETL. Every morning at 6 AM, operations managers reviewed dashboards built on

When RAG failed: a knowledge retrieval project post-mortem
When RAG failed: a knowledge retrieval project post-mortem
29 Apr, 2026 | 05 Mins read

A legal technology company had invested six months building a retrieval-augmented generation system to help contract attorneys find relevant precedent clauses across a corpus of 180,000 executed agree

From 3-hour dashboards to 3-minute insights: a BI modernisation story
From 3-hour dashboards to 3-minute insights: a BI modernisation story
05 May, 2026 | 05 Mins read

A manufacturing company with facilities in twelve countries ran its operational reporting on a traditional BI stack: a data warehouse, an ETL pipeline, and a dashboard tool that had been deployed six

The vector database that couldn't scale, and what we did instead
The vector database that couldn't scale, and what we did instead
12 May, 2026 | 05 Mins read

A media company with a library of twelve million articles, transcripts, and research documents had built a semantic search system on a managed vector database. The system was designed to let journalis

Building an AI operating system for a 10,000-person company
Building an AI operating system for a 10,000-person company
19 May, 2026 | 05 Mins read

A diversified industrial company with 10,000 employees across manufacturing, logistics, and field services had accumulated forty-seven separate AI projects over three years. Each business unit had bui

Feature store comparison: Feast, Tecton, Hopsworks
Feature store comparison: Feast, Tecton, Hopsworks
20 May, 2026 | 05 Mins read

Feature stores solve a specific problem: the features you use to train a model must be the same features you use to serve it. When the training pipeline computes features differently than the serving

How we killed our ETL pipeline (and productivity went up)
How we killed our ETL pipeline (and productivity went up)
26 May, 2026 | 05 Mins read

A B2B SaaS company running a customer success platform had a data pipeline that consumed sixty percent of the data engineering team's time. Not feature work. Not analytics. Pipeline maintenance. The p

A compliance-first AI rollout in financial services
A compliance-first AI rollout in financial services
03 Jun, 2026 | 05 Mins read

A regional bank with $12 billion in assets wanted to use machine learning to improve its commercial loan underwriting process. The existing process was manual, relying on credit analysts who spent fou

The $2M model that never made it to production
The $2M model that never made it to production
09 Jun, 2026 | 05 Mins read

A retail chain with 400 stores spent two years and $2.1 million building an inventory optimisation model. The model was technically excellent. It reduced predicted stockouts by thirty-two percent and

Data mesh in practice: year 2 retrospective
Data mesh in practice: year 2 retrospective
16 Jun, 2026 | 05 Mins read

An insurance company with $400 million in premium volume adopted data mesh two years ago. The central data team had become a bottleneck. Every business unit (claims, underwriting, actuarial, and distr

Model serving: vLLM, TGI, Triton: which fits your stack?
Model serving: vLLM, TGI, Triton: which fits your stack?
18 Jun, 2026 | 05 Mins read

Serving a language model in production is an infrastructure problem, not a model problem. The model weights are the same regardless of how you serve them. What differs is throughput (how many requests

When your AI vendor goes bankrupt: surviving platform lock-in
When your AI vendor goes bankrupt: surviving platform lock-in
23 Jun, 2026 | 05 Mins read

A healthcare analytics company received notice on a Tuesday afternoon that their primary AI infrastructure vendor was filing for Chapter 7 bankruptcy. The platform hosted their patient risk stratifica

CI/CD for ML: MLflow vs Weights & Biases vs Neptune
CI/CD for ML: MLflow vs Weights & Biases vs Neptune
25 Jun, 2026 | 05 Mins read

Machine learning teams face a version control problem that Git does not solve. Git tracks code changes, but ML experiments change more than code. They change hyperparameters, datasets, model architect

Real-time fraud detection: from proof-of-concept to production in 90 days
Real-time fraud detection: from proof-of-concept to production in 90 days
30 Jun, 2026 | 05 Mins read

A payment processor handling twelve million transactions per day had a fraud detection system that was accurate but slow. The system reviewed transactions in batch, four times per day. A fraudulent tr

Consolidating 47 data sources into one knowledge layer
Consolidating 47 data sources into one knowledge layer
01 Jul, 2026 | 05 Mins read

A global professional services firm with 8,000 consultants maintained institutional knowledge across forty-seven separate systems. Project proposals lived in a document management system. Client engag

The GDPR audit that reshaped our entire ML pipeline
The GDPR audit that reshaped our entire ML pipeline
07 Jul, 2026 | 05 Mins read

A European fintech with twelve million customers received a GDPR audit notice from their national data protection authority. The audit focused on the company's machine learning pipeline, which powered

How a healthcare org deployed LLMs without violating HIPAA
How a healthcare org deployed LLMs without violating HIPAA
14 Jul, 2026 | 05 Mins read

A hospital system with twelve facilities and 14,000 clinical staff wanted to use large language models to assist with clinical documentation. Physicians spent an average of two hours per day on docume

Legacy mainframe to cloud-native: the data migration they said was impossible
Legacy mainframe to cloud-native: the data migration they said was impossible
21 Jul, 2026 | 06 Mins read

An insurance company running on an IBM mainframe had accumulated forty years of policy data in VSAM files and DB2 tables. The mainframe processed 600,000 transactions per day across policy administrat

Building trust in AI recommendations: the change management story
Building trust in AI recommendations: the change management story
28 Jul, 2026 | 06 Mins read

A consumer goods company built an AI system that recommended reorder quantities for 12,000 SKUs across 340 distribution points. The system optimised for a multi-objective function that balanced invent

Scaling Machine Learning Infrastructure: From POC to Production
Scaling Machine Learning Infrastructure: From POC to Production
10 May, 2024 | 04 Mins read

# Scaling Machine Learning Infrastructure: From POC to Production Moving a machine learning model from notebook to production exposes gaps that notebooks hide. Data scientists produce working models

When the model was right but nobody believed it
When the model was right but nobody believed it
04 Aug, 2026 | 05 Mins read

An agriculture technology company built a crop yield prediction model that combined satellite imagery, soil sensor data, weather forecasts, and historical yield records. The model predicted per-field

Scaling a recommendation engine from 1K to 10M users
Scaling a recommendation engine from 1K to 10M users
11 Aug, 2026 | 06 Mins read

A video streaming platform grew from 1,000 beta users to 10 million subscribers over thirty months. Their recommendation system was rebuilt three times during this period. Each rebuild was triggered n

Real-time pricing engine: from batch overnight to sub-second
Real-time pricing engine: from batch overnight to sub-second
19 Aug, 2026 | 05 Mins read

An online travel agency processed 2.3 million flight searches per day. Each search triggered a pricing computation that determined the displayed fare for every matching itinerary. The pricing computat

The data catalogue project that actually stuck: 18 months later
The data catalogue project that actually stuck: 18 months later
25 Aug, 2026 | 07 Mins read

Most data catalogue projects die within six months. The tool gets purchased, a team populates it with metadata for a few hundred tables, enthusiasm fades, and twelve months later the catalogue is a st

Replatforming a decade of analytics from Oracle to Snowflake
Replatforming a decade of analytics from Oracle to Snowflake
01 Sep, 2026 | 07 Mins read

Ten years of analytics built on Oracle means ten years of accumulated PL/SQL, materialised views, database links, stored procedures, and ETL jobs that nobody fully understands. The schema has four hun

Deploying ML Models on Kubernetes: Best Practices
Deploying ML Models on Kubernetes: Best Practices
06 May, 2024 | 03 Mins read

# Deploying ML Models on Kubernetes: Best Practices ML models in production need orchestration, scaling, and monitoring infrastructure. Kubernetes provides these capabilities, though the learning cur

How a logistics company predicted delivery failures before they happened
How a logistics company predicted delivery failures before they happened
08 Sep, 2026 | 06 Mins read

A regional logistics company running three thousand deliveries per day across a six-state territory had a late-delivery rate of fourteen percent. The cost of a late delivery was not just the apology.

When the CDO and CTO disagreed on AI strategy, and what happened
When the CDO and CTO disagreed on AI strategy, and what happened
15 Sep, 2026 | 06 Mins read

At a mid-market insurance company with eight thousand employees, the Chief Data Officer and the Chief Technology Officer had fundamentally different views on how AI should be adopted. The CDO believed

Building a customer 360 from 12 disconnected CRM systems
Building a customer 360 from 12 disconnected CRM systems
16 Sep, 2026 | 07 Mins read

A healthcare conglomerate grew through acquisition for fifteen years. Each acquisition brought its own CRM. Salesforce in three divisions. Microsoft Dynamics in two. HubSpot in one. A custom-built CRM

An insurance firm's journey from PDF extraction to automated underwriting
An insurance firm's journey from PDF extraction to automated underwriting
29 Sep, 2026 | 06 Mins read

A specialty insurance firm underwriting commercial property policies received submission packets as PDF documents. Each packet contained an ACORD application, loss runs from prior carriers, a statemen

How we reduced cloud data spend 40% without cutting features
How we reduced cloud data spend 40% without cutting features
06 Oct, 2026 | 06 Mins read

A media analytics company running its entire data platform on AWS was spending $480,000 per month on cloud infrastructure. The bill had grown organically over three years as the platform expanded from

Container orchestration for ML: K8s vs ECS vs Fly.io
Container orchestration for ML: K8s vs ECS vs Fly.io
08 Oct, 2026 | 05 Mins read

Running a model in a Jupyter notebook is trivial. Running a model that serves 500 predictions per second with 99.9% uptime, auto-scales with traffic, recovers from node failures, and costs less than $

Incremental ML: Continuous Learning Systems
Incremental ML: Continuous Learning Systems
12 Jul, 2024 | 11 Mins read

Traditional ML trains on historical data, deploys, and waits until performance degrades. This fails in dynamic environments where data patterns evolve. Incremental ML continuously updates models as ne

Serverless Machine Learning: Patterns with AWS Lambda, GCP Cloud Run & Azure Functions
Serverless Machine Learning: Patterns with AWS Lambda, GCP Cloud Run & Azure Functions
18 Jul, 2025 | 05 Mins read

A social media analytics company watched their Kubernetes cluster fail to handle traffic spikes from trending topics. The cluster would scale from 50 to 500 pods in minutes, but not fast enough to pre

AI Observability: Monitoring Drift, Data Quality & Model Performance
AI Observability: Monitoring Drift, Data Quality & Model Performance
12 Sep, 2025 | 02 Mins read

An insurance company's premium pricing model had been quietly going haywire for two weeks. Young drivers in high-risk areas were getting bargain prices while safe drivers faced astronomical quotes. By

Case Study: End-to-End RAG Platform for Customer Support
Case Study: End-to-End RAG Platform for Customer Support
05 Dec, 2025 | 05 Mins read

A SaaS company with 200 support agents and 10,000+ knowledge base articles had an 18-hour average response time and 23% first-contact resolution. Their largest enterprise client threatened to cancel a

Case Study: Building a Production AI Knowledge Layer for Financial Services
Case Study: Building a Production AI Knowledge Layer for Financial Services
01 Mar, 2026 | 10 Mins read

A regional bank's investment research team spent 60% of their time gathering information and 40% doing analysis. Analysts had to search through regulatory filings, internal research memos, market data

Case Study: Multi-Agent System for Supply Chain Optimisation
Case Study: Multi-Agent System for Supply Chain Optimisation
13 Jun, 2026 | 12 Mins read

A mid-size automotive parts manufacturer with operations spanning 15 countries and relationships with over 200 suppliers faced a supply chain coordination problem that was consuming too much of their