fetch_kdd98#

empulse.datasets.fetch_kdd98(*, backend, data_home=None, download_if_missing=True)[source]#

Fetch the KDD Cup 1998 Direct Mailing dataset from the UCI KDD Archive (binary classification).

The goal is to predict whether a recipient will donate in response to a direct mailing. Target variable: 1 = donated, 0 = did not donate.

The learning (cup98LRN) and validation (cup98VAL + valtargt) sets are downloaded and combined, as in Vanderschueren et al. (2022). Only the 22 attributes selected by Petrides & Verbeke (2021) are kept. The download is about 75 MB and is cached locally.

For additional information about the dataset, consult the User Guide.

Classes

2

Donors

9716

Non-donors

182063

Samples

191779

Features

22

Parameters:
backendmodule

Dataframe library to use for data and target. Pass the library module directly, e.g. backend=polars or backend=pandas.

data_homestr or Path, optional

Directory used for caching downloaded data. Defaults to ~/empulse_data (or $EMPULSE_DATA_HOME).

download_if_missingbool, default=True

If False, raise an OSError when the data is not cached locally.

Returns:
datasetDataset

instance_costs contains {'amount': array} — the donation amount (TARGET_D), which is lost when a donor is not mailed. It is 0 for non-donors, for whom the false-negative cost never applies.

Notes

Cost matrix (Zadrozny et al. 2003, Vanderschueren et al. 2022):

Actual positive \(y_i = 1\)

Actual negative \(y_i = 0\)

Predicted positive \(\hat{y}_i = 1\)

tp_cost \(= c_f\)

fp_cost \(= c_f\)

Predicted negative \(\hat{y}_i = 0\)

fn_cost \(= amount_i\)

tn_cost \(= 0\)

The cost matrix uses symbolic parameters with the following defaults:

  • contact_cost (\(c_f\)) = 0.68 (cost of one mailing, in dollars)

References

[1]

Zadrozny, B., Langford, J., & Abe, N. (2003). Cost-sensitive learning by cost-proportionate example weighting. In Third IEEE International Conference on Data Mining (pp. 435-442).

[2]

Petrides, G., & Verbeke, W. (2021). Cost-sensitive ensemble learning: a unifying framework. Data Mining and Knowledge Discovery, 1–28.

[3]

Vanderschueren, T., Verdonck, T., Baesens, B., & Verbeke, W. (2022). Predict-then-optimize or predict-and-optimize? An empirical evaluation of cost-sensitive learning strategies. Information Sciences, 594, 400–415.

Examples

import numpy as np
import pandas as pd
from empulse.datasets import fetch_kdd98
from empulse.metrics import Metric, Cost

dataset = fetch_kdd98(backend=pd)

# replace with your own model's predicted probabilities
y_score = np.random.default_rng(0).uniform(size=len(dataset.target))

metric = Metric(dataset.cost_matrix, Cost())
score = metric(dataset.target, y_score, **dataset.instance_costs)