fetch_kddcup09_churn#
- empulse.datasets.fetch_kddcup09_churn(*, backend, data_home=None, download_if_missing=True)[source]#
Fetch the KDD Cup 2009 / Orange Customer Churn dataset from OpenML (binary classification).
The goal is to predict customer churn (target: 1 = churn, 0 = no churn).
The 230 variables are anonymised and include no revenue or customer value, so the cost matrix uses the average customer lifetime value of Verbeke et al. (2012) and is the same for every customer.
For additional information about the dataset, consult the User Guide.
Classes
2
Churners
3672
Non-churners
46328
Samples
50000
Features
230
- Parameters:
- backendmodule
Dataframe library to use for
dataandtarget. Pass the library module directly, e.g.backend=polarsorbackend=pandas.- data_homestr or Path, optional
Directory used for caching downloaded data. Defaults to
~/empulse_data(or$EMPULSE_DATA_HOME).- download_if_missingbool, default=True
If False, raise an
OSErrorwhen the data is not cached locally.
- Returns:
- dataset
Dataset instance_costscontains{'clv': array}— the customer lifetime value, 200 for every customer.
- dataset
Notes
Cost matrix (Verbraken et al. 2013, with the parameters of Verbeke et al. 2012; \(\gamma\) treated as a fixed scalar):
Actual positive \(y_i = 1\)
Actual negative \(y_i = 0\)
Predicted positive \(\hat{y}_i = 1\)
tp_benefit\(= \gamma (CLV_i - d \cdot CLV_i - f) - (1-\gamma) f\)fp_cost\(= d \cdot CLV_i + f\)Predicted negative \(\hat{y}_i = 0\)
fn_cost\(= 0\)tn_cost\(= 0\)The cost matrix uses symbolic parameters with the following defaults:
incentive_fraction(\(d\)) = 0.05 (an incentive of 10 on a CLV of 200)contact_cost(\(f\)) = 1accept_rate(\(\gamma\)) = 0.3 (the mean of the Beta(6, 14) distribution used by the expected maximum profit measure)
To override these defaults, pass the desired values when evaluating the metric:
metric(dataset.target, y_score, accept_rate=0.5, **dataset.instance_costs)
References
[1]Verbeke, W., Dejaeger, K., Martens, D., Hur, J., & Baesens, B. (2012). New insights into churn prediction in the telecommunication sector: A profit driven data mining approach. European Journal of Operational Research, 218(1), 211–229.
[2]Verbraken, T., Verbeke, W., & Baesens, B. (2013). A novel profit maximizing metric for measuring classification performance of customer churn prediction models. IEEE Transactions on Knowledge and Data Engineering, 25(5), 961–973.
[3]Stripling, E., vanden Broucke, S., Antonio, K., Baesens, B., & Snoeck, M. (2018). Profit maximizing logistic model for customer churn prediction using genetic algorithms. Swarm and Evolutionary Computation, 40, 116–130.
Examples
import numpy as np import pandas as pd from empulse.datasets import fetch_kddcup09_churn from empulse.metrics import Metric, Cost dataset = fetch_kddcup09_churn(backend=pd) # replace with your own model's predicted probabilities y_score = np.random.default_rng(0).uniform(size=len(dataset.target)) metric = Metric(dataset.cost_matrix, Cost()) score = metric(dataset.target, y_score, **dataset.instance_costs)