6.3.8. Cell2Cell Customer Churn#
6.3.8.1. Summary#
Customer data of the US wireless operator Cell2Cell, released for the churn modelling tournament of the Teradata Center for Customer Relationship Management at Duke University. It is the “Duke” data that runs through the profit-driven churn literature [1] [2] [3]. Each row is a subscriber, described by usage, call quality, handset, household and retention-contact variables, with a label indicating whether they churned.
The Duke center no longer distributes the data, so Empulse downloads the cell2celltrain.csv
file from a public GitHub mirror, pinned to a fixed commit. The 156 customers without a
MonthlyRevenue are dropped.
Classes |
2 |
Churners |
14641 |
Non-churners |
36250 |
Samples |
50891 |
Features |
56 |
6.3.8.2. Using the Dataset#
The dataset is fetched through fetch_cell2cell. It is downloaded on first
use and cached under ~/empulse_data (override with $EMPULSE_DATA_HOME or the
data_home argument), so later calls work offline.
It returns a Dataset object with the following attributes:
data: the feature matrixtarget: the target vectorcost_matrix: aCostMatrixwith default values pre-filledinstance_costs: a dict of per-instance cost drivers ('monthly_revenue')feature_names: the feature namestarget_names: the target namesDESCR: the full description of the dataset
import pandas as pd
from empulse.datasets import fetch_cell2cell
dataset = fetch_cell2cell(backend=pd)
X, y = dataset.data, dataset.target
The backend argument selects the dataframe library used for data and target.
Pass the module itself — backend=pd for pandas or backend=pl for polars.
A few numeric features have missing values, and the categorical features need encoding. Pass the
cost matrix to the model as a Metric loss, and hand it the monthly
revenue at fit time:
from empulse.metrics import Metric, Cost
from empulse.models import CSLogitClassifier
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline, make_pipeline
from sklearn.preprocessing import StandardScaler, TargetEncoder
numeric = X.select_dtypes(include=['number']).columns
categorical = X.select_dtypes(exclude=['number']).columns
pipeline = Pipeline([
('preprocessor', ColumnTransformer([
('num', make_pipeline(SimpleImputer(strategy='median'), StandardScaler()), numeric),
('cat', make_pipeline(
SimpleImputer(strategy='constant', fill_value='missing'),
TargetEncoder(),
), categorical),
])),
('model', CSLogitClassifier(loss=Metric(dataset.cost_matrix, Cost()))),
])
pipeline.fit(X, y, model__monthly_revenue=dataset.instance_costs['monthly_revenue'])
6.3.8.3. Cost Matrix#
The cost matrix is the churn retention matrix of Verbraken et al. [4], the framing every paper above uses for this data. Contacting a customer costs a fixed amount \(f\), whether or not they accept. A contacted churner accepts the retention offer with probability \(\gamma\), in which case their value is retained minus the incentive, a fraction \(d\) of that value. A churner who is not contacted simply leaves, which costs the campaign nothing.
Actual churner \(y_i = 1\) |
Actual non-churner \(y_i = 0\) |
|
Predicted churner \(\hat{y}_i = 1\) |
|
|
Predicted non-churner \(\hat{y}_i = 0\) |
|
|
The papers use a single average lifetime value of 200 for every customer [1]. This version of the data records each customer’s monthly revenue, so Empulse makes the lifetime value instance-dependent by counting \(m\) months of it:
The 3 customers with a negative revenue (net credits) get a lifetime value of 0: a customer who costs money has nothing to retain.
The symbolic parameters carry these defaults, and can be overridden by passing their alias:
Alias |
Default |
Meaning |
|---|---|---|
|
12 |
Months of revenue counted as lifetime value |
|
0.3 |
Probability a contacted churner accepts the offer, the mean of the Beta(6, 14) distribution used by the expected maximum profit measure |
|
0.05 |
Retention incentive as a fraction of CLV, as an incentive of 10 on a CLV of 200 in Verbeke et al. [1] |
|
1 |
Cost of contacting a customer, as an absolute amount |
y_score = pipeline.predict_proba(X)[:, 1]
cost = Metric(dataset.cost_matrix, Cost())
default_cost = cost(y, y_score, **dataset.instance_costs)
longer_lifetime = cost(y, y_score, clv_months=24, **dataset.instance_costs)
6.3.8.4. Data Description#
The mirror does not document the variables individually; they fall into these groups.
Features |
Description |
|---|---|
|
Revenue and usage, and their recent percentage change. |
|
Call volumes and call quality |
|
Account tenure, number of subscriptions and service area |
|
Handset history and the current handset |
|
Household demographics and lifestyle |
|
Credit rating and adjustments to it |
|
Contacts with the retention team, and referrals |
churn (target) |
Whether the customer churned (1 = yes, 0 = no) |