fetch_cell2cell#
- empulse.datasets.fetch_cell2cell(*, backend, data_home=None, download_if_missing=True)[source]#
Fetch the Cell2Cell Customer Churn dataset (binary classification).
The goal is to predict whether a customer will churn or not. The target variable is whether the customer churned: 1 = churned, 0 = active.
This is the “Duke” churn data of the Teradata Center for Customer Relationship Management tournament. The Duke center no longer distributes it, so it is downloaded from a public GitHub mirror pinned to a fixed commit. The 156 customers without a
MonthlyRevenueare dropped.For additional information about the dataset, consult the User Guide.
Classes
2
Churners
14641
Non-churners
36250
Samples
50891
Features
56
- Parameters:
- backendmodule
Dataframe library to use for
dataandtarget. Pass the library module directly, e.g.backend=polarsorbackend=pandas.- data_homestr or Path, optional
Directory used for caching downloaded data. Defaults to
~/empulse_data(or$EMPULSE_DATA_HOME).- download_if_missingbool, default=True
If False, raise an
OSErrorwhen the data is not cached locally.
- Returns:
- dataset
Dataset instance_costscontains{'monthly_revenue': array}— each customer’s average monthly revenue, from which their lifetime value is approximated. The 3 negative revenues are set to 0: a customer who costs money has no value to retain.
- dataset
Notes
Cost matrix (Verbraken et al. 2013, \(\gamma\) treated as a fixed scalar). The data has no customer lifetime value, so it is approximated by \(CLV_i = m \cdot MonthlyRevenue_i\):
Actual positive \(y_i = 1\)
Actual negative \(y_i = 0\)
Predicted positive \(\hat{y}_i = 1\)
tp_benefit\(= \gamma (CLV_i - d \cdot CLV_i - f) - (1-\gamma) f\)fp_cost\(= d \cdot CLV_i + f\)Predicted negative \(\hat{y}_i = 0\)
fn_cost\(= 0\)tn_cost\(= 0\)The cost matrix uses symbolic parameters with the following defaults:
clv_months(\(m\)) = 12 (months of revenue counted as lifetime value)incentive_fraction(\(d\)) = 0.05 (as an incentive of 10 on a CLV of 200 in Verbeke et al. 2012)contact_cost(\(f\)) = 1accept_rate(\(\gamma\)) = 0.3 (the mean of the Beta(6, 14) distribution used by the expected maximum profit measure)
To override these defaults, pass the desired values when evaluating the metric:
metric(dataset.target, y_score, clv_months=24, **dataset.instance_costs)
References
[1]Verbeke, W., Dejaeger, K., Martens, D., Hur, J., & Baesens, B. (2012). New insights into churn prediction in the telecommunication sector: A profit driven data mining approach. European Journal of Operational Research, 218(1), 211–229.
[2]Verbraken, T., Verbeke, W., & Baesens, B. (2013). A novel profit maximizing metric for measuring classification performance of customer churn prediction models. IEEE Transactions on Knowledge and Data Engineering, 25(5), 961–973.
[3]Höppner, S., Stripling, E., Baesens, B., vanden Broucke, S., & Verdonck, T. (2020). Profit driven decision trees for churn prediction. European Journal of Operational Research, 284(3), 920–933.
[4]Maldonado, S., López, J., & Vairetti, C. (2020). Profit-based churn prediction based on Minimax Probability Machines. European Journal of Operational Research, 284(1), 273–284.
Examples
import numpy as np import pandas as pd from empulse.datasets import fetch_cell2cell from empulse.metrics import Metric, Cost dataset = fetch_cell2cell(backend=pd) # replace with your own model's predicted probabilities y_score = np.random.default_rng(0).uniform(size=len(dataset.target)) metric = Metric(dataset.cost_matrix, Cost()) score = metric(dataset.target, y_score, **dataset.instance_costs)