Aayush Damani
PredictiveGroup project · Data Wrangling and Visualisation, Imperial Business School

Customer Segmentation & Heterogeneous Churn Dynamics

K-Prototypes clustering + segment-specific regression on 5,329 e-commerce customers

Cluster AnalysisLogistic RegressionPredictive ModelingRCustomer Segmentation

In most customer segments, ordering more predicts a higher chance of churning. In the highest-risk segment, it's the opposite: ordering more predicts they'll stay.

Most companies build one churn model for their whole customer base and apply one retention strategy to everyone. This project asked whether that's actually safe: do different types of customers leave for different reasons?

Starting with 5,329 e-commerce customers, of whom about 17% had churned, we first grouped customers by how they actually behave (how long they'd been around, what they bought, how often, how much they used coupons) rather than by who they are demographically. Four distinct groups fell out of that.

The gap between groups was large. The riskiest group churned at 27.9%, more than three times the safest group's 8.5%. The single biggest driver was simply how new someone was: customers in their first three months churned at 41.9%, while customers past two years churned at essentially zero. Practically, this means retention money is almost always best spent on the first few months of a customer's life.

The more surprising result came from modelling each group separately instead of all together. In most groups, ordering more and using more coupons predicted a higher chance of leaving, which fits a picture of people who show up for a promotion and disappear afterwards. But in the newest, highest-risk group, that flipped: there, more orders predicted people were more likely to stay. Same behaviour, opposite meaning, depending on who's doing it. A single company-wide model averages those two opposite effects together and effectively hides both.

Why it matters

A customer ordering a lot is usually read as a happy customer. In this data, that was often wrong: for most people, a burst of orders meant they were chasing discounts and would leave once the discounts stopped. So a business that rewards its heaviest orderers with more coupons could be spending money on exactly the customers it's about to lose anyway, while ignoring the newer customers where more orders genuinely does mean they're sticking around.

Churn rate by segment. The highest-risk group churns 3.3x more than the safest one, and behaves differently, not just worse.

Churn is front-loaded: 41.9% in the first three months against a 16.8% overall baseline, falling to zero past two years. Retention spend belongs early.

Dataset, tools and how it was done+

Dataset: E-commerce Customer Churn Analysis and Prediction dataset (Kaggle, Ankit Verma, 2021)

Tools: R · K-Prototypes clustering · clustMixType · Logistic regression

  • 5,329 customers, 20 variables, from the Kaggle E-commerce Customer Churn dataset
  • K-Prototypes clustering (mixed numeric and categorical features) identified 4 behaviourally distinct segments
  • Cluster 3 ("High-Risk New Customers"): 27.9% churn, 5.25-month average tenure, concentrated in one-off mobile phone purchases
  • Cluster 4 ("Loyal Customers"): 8.5% churn, 16.6-month average tenure, concentrated in repeat grocery purchases
  • Segment-specific logistic regressions outperformed a single global model: AUC 0.863 to 0.872, F1 0.543 to 0.556, recall 0.432 to 0.463
  • Order count and coupon usage predict higher churn in most segments but reverse sign in Cluster 3