fetch_kdd98#
- empulse.datasets.fetch_kdd98(*, backend, data_home=None, download_if_missing=True)[source]#
Fetch the KDD Cup 1998 Direct Mailing dataset from the UCI KDD Archive (binary classification).
The goal is to predict whether a recipient will donate in response to a direct mailing. Target variable: 1 = donated, 0 = did not donate.
The learning (
cup98LRN) and validation (cup98VAL+valtargt) sets are downloaded and combined, as in Vanderschueren et al. (2022). Only the 22 attributes selected by Petrides & Verbeke (2021) are kept. The download is about 75 MB and is cached locally.For additional information about the dataset, consult the User Guide.
Classes
2
Donors
9716
Non-donors
182063
Samples
191779
Features
22
- Parameters:
- backendmodule
Dataframe library to use for
dataandtarget. Pass the library module directly, e.g.backend=polarsorbackend=pandas.- data_homestr or Path, optional
Directory used for caching downloaded data. Defaults to
~/empulse_data(or$EMPULSE_DATA_HOME).- download_if_missingbool, default=True
If False, raise an
OSErrorwhen the data is not cached locally.
- Returns:
- dataset
Dataset instance_costscontains{'amount': array}— the donation amount (TARGET_D), which is lost when a donor is not mailed. It is 0 for non-donors, for whom the false-negative cost never applies.
- dataset
Notes
Cost matrix (Zadrozny et al. 2003, Vanderschueren et al. 2022):
Actual positive \(y_i = 1\)
Actual negative \(y_i = 0\)
Predicted positive \(\hat{y}_i = 1\)
tp_cost\(= c_f\)fp_cost\(= c_f\)Predicted negative \(\hat{y}_i = 0\)
fn_cost\(= amount_i\)tn_cost\(= 0\)The cost matrix uses symbolic parameters with the following defaults:
contact_cost(\(c_f\)) = 0.68 (cost of one mailing, in dollars)
References
[1]Zadrozny, B., Langford, J., & Abe, N. (2003). Cost-sensitive learning by cost-proportionate example weighting. In Third IEEE International Conference on Data Mining (pp. 435-442).
[2]Petrides, G., & Verbeke, W. (2021). Cost-sensitive ensemble learning: a unifying framework. Data Mining and Knowledge Discovery, 1–28.
[3]Vanderschueren, T., Verdonck, T., Baesens, B., & Verbeke, W. (2022). Predict-then-optimize or predict-and-optimize? An empirical evaluation of cost-sensitive learning strategies. Information Sciences, 594, 400–415.
Examples
import numpy as np import pandas as pd from empulse.datasets import fetch_kdd98 from empulse.metrics import Metric, Cost dataset = fetch_kdd98(backend=pd) # replace with your own model's predicted probabilities y_score = np.random.default_rng(0).uniform(size=len(dataset.target)) metric = Metric(dataset.cost_matrix, Cost()) score = metric(dataset.target, y_score, **dataset.instance_costs)