fetch_kddcup09_churn#

empulse.datasets.fetch_kddcup09_churn(*, backend, data_home=None, download_if_missing=True)[source]#

Fetch the KDD Cup 2009 / Orange Customer Churn dataset from OpenML (binary classification).

The goal is to predict customer churn (target: 1 = churn, 0 = no churn).

The 230 variables are anonymised and include no revenue or customer value, so the cost matrix uses the average customer lifetime value of Verbeke et al. (2012) and is the same for every customer.

For additional information about the dataset, consult the User Guide.

Classes

2

Churners

3672

Non-churners

46328

Samples

50000

Features

230

Parameters:
backendmodule

Dataframe library to use for data and target. Pass the library module directly, e.g. backend=polars or backend=pandas.

data_homestr or Path, optional

Directory used for caching downloaded data. Defaults to ~/empulse_data (or $EMPULSE_DATA_HOME).

download_if_missingbool, default=True

If False, raise an OSError when the data is not cached locally.

Returns:
datasetDataset

instance_costs contains {'clv': array} — the customer lifetime value, 200 for every customer.

Notes

Cost matrix (Verbraken et al. 2013, with the parameters of Verbeke et al. 2012; \(\gamma\) treated as a fixed scalar):

Actual positive \(y_i = 1\)

Actual negative \(y_i = 0\)

Predicted positive \(\hat{y}_i = 1\)

tp_benefit \(= \gamma (CLV_i - d \cdot CLV_i - f) - (1-\gamma) f\)

fp_cost \(= d \cdot CLV_i + f\)

Predicted negative \(\hat{y}_i = 0\)

fn_cost \(= 0\)

tn_cost \(= 0\)

The cost matrix uses symbolic parameters with the following defaults:

  • incentive_fraction (\(d\)) = 0.05 (an incentive of 10 on a CLV of 200)

  • contact_cost (\(f\)) = 1

  • accept_rate (\(\gamma\)) = 0.3 (the mean of the Beta(6, 14) distribution used by the expected maximum profit measure)

To override these defaults, pass the desired values when evaluating the metric:

metric(dataset.target, y_score, accept_rate=0.5, **dataset.instance_costs)

References

[1]

Verbeke, W., Dejaeger, K., Martens, D., Hur, J., & Baesens, B. (2012). New insights into churn prediction in the telecommunication sector: A profit driven data mining approach. European Journal of Operational Research, 218(1), 211–229.

[2]

Verbraken, T., Verbeke, W., & Baesens, B. (2013). A novel profit maximizing metric for measuring classification performance of customer churn prediction models. IEEE Transactions on Knowledge and Data Engineering, 25(5), 961–973.

[3]

Stripling, E., vanden Broucke, S., Antonio, K., Baesens, B., & Snoeck, M. (2018). Profit maximizing logistic model for customer churn prediction using genetic algorithms. Swarm and Evolutionary Computation, 40, 116–130.

Examples

import numpy as np
import pandas as pd
from empulse.datasets import fetch_kddcup09_churn
from empulse.metrics import Metric, Cost

dataset = fetch_kddcup09_churn(backend=pd)

# replace with your own model's predicted probabilities
y_score = np.random.default_rng(0).uniform(size=len(dataset.target))

metric = Metric(dataset.cost_matrix, Cost())
score = metric(dataset.target, y_score, **dataset.instance_costs)