fetch_cell2cell#

empulse.datasets.fetch_cell2cell(*, backend, data_home=None, download_if_missing=True)[source]#

Fetch the Cell2Cell Customer Churn dataset (binary classification).

The goal is to predict whether a customer will churn or not. The target variable is whether the customer churned: 1 = churned, 0 = active.

This is the “Duke” churn data of the Teradata Center for Customer Relationship Management tournament. The Duke center no longer distributes it, so it is downloaded from a public GitHub mirror pinned to a fixed commit. The 156 customers without a MonthlyRevenue are dropped.

For additional information about the dataset, consult the User Guide.

Classes

2

Churners

14641

Non-churners

36250

Samples

50891

Features

56

Parameters:
backendmodule

Dataframe library to use for data and target. Pass the library module directly, e.g. backend=polars or backend=pandas.

data_homestr or Path, optional

Directory used for caching downloaded data. Defaults to ~/empulse_data (or $EMPULSE_DATA_HOME).

download_if_missingbool, default=True

If False, raise an OSError when the data is not cached locally.

Returns:
datasetDataset

instance_costs contains {'monthly_revenue': array} — each customer’s average monthly revenue, from which their lifetime value is approximated. The 3 negative revenues are set to 0: a customer who costs money has no value to retain.

Notes

Cost matrix (Verbraken et al. 2013, \(\gamma\) treated as a fixed scalar). The data has no customer lifetime value, so it is approximated by \(CLV_i = m \cdot MonthlyRevenue_i\):

Actual positive \(y_i = 1\)

Actual negative \(y_i = 0\)

Predicted positive \(\hat{y}_i = 1\)

tp_benefit \(= \gamma (CLV_i - d \cdot CLV_i - f) - (1-\gamma) f\)

fp_cost \(= d \cdot CLV_i + f\)

Predicted negative \(\hat{y}_i = 0\)

fn_cost \(= 0\)

tn_cost \(= 0\)

The cost matrix uses symbolic parameters with the following defaults:

  • clv_months (\(m\)) = 12 (months of revenue counted as lifetime value)

  • incentive_fraction (\(d\)) = 0.05 (as an incentive of 10 on a CLV of 200 in Verbeke et al. 2012)

  • contact_cost (\(f\)) = 1

  • accept_rate (\(\gamma\)) = 0.3 (the mean of the Beta(6, 14) distribution used by the expected maximum profit measure)

To override these defaults, pass the desired values when evaluating the metric:

metric(dataset.target, y_score, clv_months=24, **dataset.instance_costs)

References

[1]

Verbeke, W., Dejaeger, K., Martens, D., Hur, J., & Baesens, B. (2012). New insights into churn prediction in the telecommunication sector: A profit driven data mining approach. European Journal of Operational Research, 218(1), 211–229.

[2]

Verbraken, T., Verbeke, W., & Baesens, B. (2013). A novel profit maximizing metric for measuring classification performance of customer churn prediction models. IEEE Transactions on Knowledge and Data Engineering, 25(5), 961–973.

[3]

Höppner, S., Stripling, E., Baesens, B., vanden Broucke, S., & Verdonck, T. (2020). Profit driven decision trees for churn prediction. European Journal of Operational Research, 284(3), 920–933.

[4]

Maldonado, S., López, J., & Vairetti, C. (2020). Profit-based churn prediction based on Minimax Probability Machines. European Journal of Operational Research, 284(1), 273–284.

Examples

import numpy as np
import pandas as pd
from empulse.datasets import fetch_cell2cell
from empulse.metrics import Metric, Cost

dataset = fetch_cell2cell(backend=pd)

# replace with your own model's predicted probabilities
y_score = np.random.default_rng(0).uniform(size=len(dataset.target))

metric = Metric(dataset.cost_matrix, Cost())
score = metric(dataset.target, y_score, **dataset.instance_costs)