6.3.9. KDD Cup 2009 Orange Customer Churn#

6.3.9.1. Summary#

Customer relationship management data from the French telecom operator Orange, released for the ACM KDD Cup 2009 [1]. Each row is a customer, and the task is to predict whether they will switch provider. This is the small version of the competition data, hosted on OpenML.

All 230 variables are anonymised, and none of them is a revenue or customer value, so the cost of a retention campaign cannot differ between customers. The dataset therefore uses the churn cost matrix with the average customer lifetime value that Verbeke et al. [2] calibrated for the telecom sector, the same setting under which the data has been benchmarked in the profit-driven churn literature [2] [3].

Classes

2

Churners

3672

Non-churners

46328

Samples

50000

Features

230

6.3.9.2. Using the Dataset#

The dataset is fetched through fetch_kddcup09_churn. It is downloaded from OpenML on first use and cached under ~/empulse_data (override with $EMPULSE_DATA_HOME or the data_home argument), so later calls work offline.

It returns a Dataset object with the following attributes:

  • data: the feature matrix

  • target: the target vector

  • cost_matrix: a CostMatrix with default values pre-filled

  • instance_costs: a dict of per-instance cost drivers ('clv', 200 for every customer)

  • feature_names: the feature names

  • target_names: the target names

  • DESCR: the full description of the dataset

import pandas as pd
from empulse.datasets import fetch_kddcup09_churn

dataset = fetch_kddcup09_churn(backend=pd)
X, y = dataset.data, dataset.target

The backend argument selects the dataframe library used for data and target. Pass the module itself — backend=pd for pandas or backend=pl for polars.

Many variables are mostly missing and the nominal ones have thousands of levels, so the data needs imputation and an encoding that copes with high cardinality before a linear model can use it:

from empulse.metrics import Metric, Cost
from empulse.models import CSLogitClassifier
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline, make_pipeline
from sklearn.preprocessing import StandardScaler, TargetEncoder

numeric = X.select_dtypes(include=['number']).columns
nominal = X.select_dtypes(exclude=['number']).columns

pipeline = Pipeline([
    ('preprocessor', ColumnTransformer([
        ('num', make_pipeline(SimpleImputer(keep_empty_features=True), StandardScaler()), numeric),
        ('cat', make_pipeline(
            SimpleImputer(strategy='constant', fill_value='missing', keep_empty_features=True),
            TargetEncoder(),
        ), nominal),
    ])),
    ('model', CSLogitClassifier(loss=Metric(dataset.cost_matrix, Cost()))),
])
pipeline.fit(X, y, model__clv=dataset.instance_costs['clv'])

The defaults of empc_score are the same parameters (a CLV of 200, an incentive of 10 and a contact cost of 1), with the acceptance rate drawn from Beta(6, 14) instead of fixed at its mean. The expected maximum profit measure of the literature therefore needs no extra arguments:

from empulse.metrics import empc_score

y_score = pipeline.predict_proba(X)[:, 1]
expected_profit = empc_score(y, y_score)

6.3.9.3. Cost Matrix#

Contacting a customer costs a fixed amount \(f\), whether or not they accept. A contacted churner accepts the retention offer with probability \(\gamma\), in which case their value is retained minus the incentive, a fraction \(d\) of that value. A churner who is not contacted simply leaves, which costs the campaign nothing [4].

Actual churner \(y_i = 1\)

Actual non-churner \(y_i = 0\)

Predicted churner \(\hat{y}_i = 1\)

tp_benefit \(= \gamma (CLV_i - d \cdot CLV_i - f) - (1-\gamma) f\)

fp_cost \(= d \cdot CLV_i + f\)

Predicted non-churner \(\hat{y}_i = 0\)

fn_cost \(= 0\)

tn_benefit \(= 0\)

\(CLV_i\) is 200 for every customer. The symbolic parameters carry these defaults, which are those of Verbeke et al. [2], and can be overridden by passing their alias:

Alias

Default

Meaning

accept_rate (\(\gamma\))

0.3

Probability a contacted churner accepts the offer, the mean of the Beta(6, 14) distribution used by the expected maximum profit measure

incentive_fraction (\(d\))

0.05

Retention incentive as a fraction of CLV, an incentive of 10 on a CLV of 200

contact_cost (\(f\))

1

Cost of contacting a customer, as an absolute amount

cost = Metric(dataset.cost_matrix, Cost())
generous_offer = cost(
    y,
    y_score,
    accept_rate=0.5,
    incentive_fraction=0.10,
    clv=dataset.instance_costs['clv'],
)

6.3.9.4. Data Description#

The variables are anonymised by Orange and cannot be interpreted.

Feature

Description

var1 – var190

Numeric variables. Many are mostly missing.

var191 – var229

Nominal variables, some with thousands of levels.

var230

Numeric, but missing for every customer (as is the nominal var209).

6.3.9.5. References#