A subscription media company with 1.2 million subscribers built a machine learning model to predict churn. The model worked. It identified at-risk subscribers with seventy-nine percent precision and eighty-three percent recall, which was significantly better than the rule-based system it replaced: a simple heuristic that flagged subscribers who had not logged in for thirty days.
The problem was not the model’s accuracy. The problem was that the retention team would not use it. The team consisted of twelve people who called at-risk subscribers and offered incentives to stay. They had been doing this job for years and had strong intuitions about why subscribers churned: content fatigue, price sensitivity, life events, competitor offers. When the model flagged a subscriber as high-risk, the retention specialist needed to know why in order to craft the right intervention. A generic “the model says this person will churn” was not actionable. The specialist needed to know: is this person price-sensitive, or are they bored with the content?
The model could not answer that question. It was a gradient-boosted ensemble with two hundred features. The feature importance rankings showed which features were most predictive across the entire population (login frequency, content diversity score, billing cycle) but these global rankings did not explain any individual prediction. A subscriber might be flagged as high-risk despite having high login frequency because the combination of three other features (declining content diversity, recent support ticket, and mid-contract billing) pushed the prediction over the threshold. The model knew this. The retention specialist did not.
The Explainability Gap
The company had tried two approaches to explainability before engaging us. Both failed for different reasons.
The first attempt was LIME, Local Interpretable Model-agnostic Explanations. LIME generates a local linear approximation of the model’s behaviour around a specific prediction. For each flagged subscriber, LIME produced a ranked list of features that contributed to the prediction. The retention team found these lists confusing. The features were technical: “days since last content interaction weighted by content type diversity coefficient”, and the contribution scores were abstract numbers that did not map to the retention specialist’s mental model of subscriber behaviour.
The second attempt was to build a simpler model, a decision tree, that the team could understand. The decision tree was interpretable. Each prediction came with a clear path: “subscriber has not logged in for 14 days AND has a content diversity score below 0.3 AND is within 60 days of renewal.” The problem was accuracy. The decision tree achieved sixty-four percent precision, fifteen percentage points worse than the gradient-boosted model. The retention team understood the predictions but the predictions were wrong too often to be useful.
The company was stuck in a trade-off that is common in applied ML: accuracy versus interpretability. The accurate model was a black box. The interpretable model was inaccurate. They needed both.
The Explanation Layer
The solution was not to make the model interpretable. It was to build a separate system that translated model outputs into the language the retention team used.
We started by interviewing the retention specialists. We asked them to describe how they decided what to say to a subscriber during a retention call. Their answers revealed a taxonomy of churn drivers that they used implicitly but had never documented:
- Content fatigue: The subscriber has consumed most of the available content in their interest areas and finds nothing new.
- Price objection: The subscriber has compared competitor pricing and believes the service is overpriced.
- Usage decline: The subscriber’s engagement has gradually decreased, often due to a life change (new job, new baby, relocation).
- Service issue: The subscriber had a negative support experience and lost trust.
- Competitor switch: The subscriber has already signed up for a competitor and is letting the subscription lapse.
Each driver had a different intervention. Content fatigue required surfacing undiscovered content. Price objection required a discount or value framing. Usage decline required a gentle re-engagement, not a hard sell. Service issue required an apology and a resolution. Competitor switch required understanding what the competitor offered and whether it could be matched.
The explanation layer mapped model features to churn drivers. A subscriber with declining content diversity and no new category exploration in sixty days was mapped to content fatigue. A subscriber who had searched the cancellation page three times in two weeks was mapped to price objection. A subscriber with a recent support ticket and a negative satisfaction survey was mapped to service issue.
This diagram requires JavaScript.
Enable JavaScript in your browser to use this feature.
The driver classification was not part of the ML model. It was a rule-based layer that operated on the model’s input features: the same features the model used to predict churn, but interpreted through the lens of the retention team’s domain knowledge. The rule-based layer did not need to be as accurate as the ML model. It needed to be plausible. If the model predicted churn with seventy-nine percent precision and the driver classification was correct seventy percent of the time, the retention specialist received the right intervention recommendation fifty-five percent of the time (0.79 x 0.70). This was significantly better than the specialist’s baseline intuition, which internal testing showed was correct approximately forty percent of the time.
The Dashboard
The retention specialist’s dashboard showed three things for each flagged subscriber: the churn risk score from the model, the likely churn driver from the explanation layer, and the recommended intervention. The specialist could accept the recommendation, choose a different intervention, or mark the subscriber as uncallable (wrong number, deceased, already cancelled).
The dashboard also showed the specialist’s override history. When a specialist consistently chose a different intervention than the recommendation for a specific driver, the system learned from the pattern. If a specialist always chose a discount for subscribers that the system classified as content fatigue, the system asked: is this specialist’s local knowledge revealing a pattern the rules missed? The override patterns were reviewed monthly and used to update the classification rules.
This feedback loop was the most valuable component of the system. The ML model predicted churn. The explanation layer translated predictions to drivers. The retention specialists applied interventions. The outcome (did the subscriber stay or leave) was tracked. The override patterns and outcome data fed back into both the model and the explanation layer. Over twelve months, the model improved from seventy-nine to eighty-two percent precision, and the driver classification accuracy improved from seventy to seventy-six percent.
What Explainability Actually Costs
Building the explanation layer took six weeks: roughly the same time as building the ML model itself. This is the hidden cost of explainability. The model is half the project. The explanation system is the other half.
The ongoing maintenance cost was also non-trivial. The churn driver taxonomy had to be updated when subscriber behaviour changed. When the company launched a new content format, a new driver category, format incompatibility, emerged. Subscribers who preferred long-form articles churned when the platform shifted to short-form video. The existing taxonomy did not have a category for this, so the explanation layer mapped these subscribers to content fatigue, which led to the wrong intervention (surfacing more short-form video to people who did not want short-form video). Adding the new driver category required interviewing specialists, defining classification rules, and testing the rules against known cases.
The rule-based classification was also brittle in edge cases. When a subscriber exhibited signals for two drivers simultaneously, declining engagement and a price comparison search, the classification chose the dominant signal. This was correct most of the time but wrong when the two signals interacted. A subscriber who was both bored with content and price-shopping needed a different intervention than either condition alone. The system handled these cases by flagging subscribers with multi-driver signals for senior specialist review, which added a manual step but avoided the worst outcomes.
The Outcome
After twelve months, the retention team’s save rate, the percentage of at-risk subscribers who were retained after intervention, increased from thirty-one percent to forty-four percent. The improvement was attributable to two factors: better targeting (the model identified at-risk subscribers earlier than the rule-based predecessor) and better intervention selection (the explanation layer matched interventions to drivers more accurately than the specialists’ intuition alone).
The retention specialists reported higher job satisfaction. Before the system, they described their work as “cold-calling people and guessing what to say.” After the system, they described it as “having a conversation with someone whose situation I understand.” The shift from guessing to understanding was the practical impact of explainability.
The Rule
If your model’s consumers need to act on its predictions, the model is only half the project. The other half is translating predictions into the language and mental model of the people who will use them. This translation does not require making the model itself interpretable. It requires building a bridge between the model’s feature space and the user’s decision space.
The bridge is custom for every organisation because every organisation has its own vocabulary for the decisions it makes. No off-the-shelf explainability tool can provide this because the tool does not know your retention team’s taxonomy of churn drivers, your clinical team’s diagnostic framework, or your fraud team’s investigation workflow. Build the translation layer. Budget for it. Maintain it. Without it, even an accurate model will sit on a shelf.