6.3.9. KDD Cup 2009 Orange Customer Churn#
6.3.9.1. Summary#
Customer relationship management data from the French telecom operator Orange, released for the ACM KDD Cup 2009 [1]. Each row is a customer, and the task is to predict whether they will switch provider. This is the small version of the competition data, hosted on OpenML.
All 230 variables are anonymised, and none of them is a revenue or customer value, so the cost of a retention campaign cannot differ between customers. The dataset therefore uses the churn cost matrix with the average customer lifetime value that Verbeke et al. [2] calibrated for the telecom sector, the same setting under which the data has been benchmarked in the profit-driven churn literature [2] [3].
Classes |
2 |
Churners |
3672 |
Non-churners |
46328 |
Samples |
50000 |
Features |
230 |
6.3.9.2. Using the Dataset#
The dataset is fetched through fetch_kddcup09_churn. It is downloaded
from OpenML on first use and cached under ~/empulse_data (override with $EMPULSE_DATA_HOME
or the data_home argument), so later calls work offline.
It returns a Dataset object with the following attributes:
data: the feature matrixtarget: the target vectorcost_matrix: aCostMatrixwith default values pre-filledinstance_costs: a dict of per-instance cost drivers ('clv', 200 for every customer)feature_names: the feature namestarget_names: the target namesDESCR: the full description of the dataset
import pandas as pd
from empulse.datasets import fetch_kddcup09_churn
dataset = fetch_kddcup09_churn(backend=pd)
X, y = dataset.data, dataset.target
The backend argument selects the dataframe library used for data and target.
Pass the module itself — backend=pd for pandas or backend=pl for polars.
Many variables are mostly missing and the nominal ones have thousands of levels, so the data needs imputation and an encoding that copes with high cardinality before a linear model can use it:
from empulse.metrics import Metric, Cost
from empulse.models import CSLogitClassifier
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline, make_pipeline
from sklearn.preprocessing import StandardScaler, TargetEncoder
numeric = X.select_dtypes(include=['number']).columns
nominal = X.select_dtypes(exclude=['number']).columns
pipeline = Pipeline([
('preprocessor', ColumnTransformer([
('num', make_pipeline(SimpleImputer(keep_empty_features=True), StandardScaler()), numeric),
('cat', make_pipeline(
SimpleImputer(strategy='constant', fill_value='missing', keep_empty_features=True),
TargetEncoder(),
), nominal),
])),
('model', CSLogitClassifier(loss=Metric(dataset.cost_matrix, Cost()))),
])
pipeline.fit(X, y, model__clv=dataset.instance_costs['clv'])
The defaults of empc_score are the same parameters (a CLV of 200, an
incentive of 10 and a contact cost of 1), with the acceptance rate drawn from Beta(6, 14) instead
of fixed at its mean. The expected maximum profit measure of the literature therefore needs no
extra arguments:
from empulse.metrics import empc_score
y_score = pipeline.predict_proba(X)[:, 1]
expected_profit = empc_score(y, y_score)
6.3.9.3. Cost Matrix#
Contacting a customer costs a fixed amount \(f\), whether or not they accept. A contacted churner accepts the retention offer with probability \(\gamma\), in which case their value is retained minus the incentive, a fraction \(d\) of that value. A churner who is not contacted simply leaves, which costs the campaign nothing [4].
Actual churner \(y_i = 1\) |
Actual non-churner \(y_i = 0\) |
|
Predicted churner \(\hat{y}_i = 1\) |
|
|
Predicted non-churner \(\hat{y}_i = 0\) |
|
|
\(CLV_i\) is 200 for every customer. The symbolic parameters carry these defaults, which are those of Verbeke et al. [2], and can be overridden by passing their alias:
Alias |
Default |
Meaning |
|---|---|---|
|
0.3 |
Probability a contacted churner accepts the offer, the mean of the Beta(6, 14) distribution used by the expected maximum profit measure |
|
0.05 |
Retention incentive as a fraction of CLV, an incentive of 10 on a CLV of 200 |
|
1 |
Cost of contacting a customer, as an absolute amount |
cost = Metric(dataset.cost_matrix, Cost())
generous_offer = cost(
y,
y_score,
accept_rate=0.5,
incentive_fraction=0.10,
clv=dataset.instance_costs['clv'],
)
6.3.9.4. Data Description#
The variables are anonymised by Orange and cannot be interpreted.
Feature |
Description |
|---|---|
|
Numeric variables. Many are mostly missing. |
|
Nominal variables, some with thousands of levels. |
|
Numeric, but missing for every customer (as is the nominal |