6.3.10. KDD Cup 1998 Direct Mailing#
6.3.10.1. Summary#
The KDD Cup 1998 data comes from the Paralyzed Veterans of America (PVA), a not-for-profit organisation that raises money through direct mail [1]. It covers all 191,779 “lapsed” donors, who last gave between 13 and 24 months earlier, and who received PVA’s June 1997 renewal mailing. The task is to decide whom to mail: each piece costs $0.68, and only about one in twenty recipients donates.
The competition split the donors into a learning and a validation set. Empulse combines them, as Vanderschueren et al. [2] do, and keeps the 22 donor attributes selected by Petrides & Verbeke [3] out of the original 479. The donation amount itself is not a feature: it is the false-negative cost.
Classes |
2 |
Donors |
9716 |
Non-donors |
182063 |
Samples |
191779 |
Features |
22 |
6.3.10.2. Using the Dataset#
The dataset is fetched through fetch_kdd98. The original files are
downloaded from the UCI KDD Archive on first use (about 75 MB) and cached under ~/empulse_data
(override with $EMPULSE_DATA_HOME or the data_home argument), so later calls work offline.
It returns a Dataset object with the following attributes:
data: the feature matrixtarget: the target vectorcost_matrix: aCostMatrixwith default values pre-filledinstance_costs: a dict of per-instance cost drivers ('amount', the donation)feature_names: the feature namestarget_names: the target namesDESCR: the full description of the dataset
import pandas as pd
from empulse.datasets import fetch_kdd98
dataset = fetch_kdd98(backend=pd)
X, y = dataset.data, dataset.target
The backend argument selects the dataframe library used for data and target.
Pass the module itself — backend=pd for pandas or backend=pl for polars.
The data mixes numeric attributes with flags and dates, and has missing values in both, so it needs imputation and encoding before a linear model can use it:
from empulse.metrics import Metric, Cost
from empulse.models import CSLogitClassifier
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline, make_pipeline
from sklearn.preprocessing import StandardScaler, TargetEncoder
numeric = X.select_dtypes(include=['number']).columns
categorical = X.select_dtypes(exclude=['number']).columns
pipeline = Pipeline([
('preprocessor', ColumnTransformer([
('num', make_pipeline(SimpleImputer(strategy='median'), StandardScaler()), numeric),
('cat', make_pipeline(
SimpleImputer(strategy='constant', fill_value='missing'),
TargetEncoder(),
), categorical),
])),
('model', CSLogitClassifier(loss=Metric(dataset.cost_matrix, Cost()))),
])
pipeline.fit(X, y, model__amount=dataset.instance_costs['amount'])
6.3.10.3. Cost Matrix#
Mailing a donor costs \(c_f\), whether or not they respond. A donor who is not mailed does not donate, which costs the donation \(A_i\) they would have made [2] [4].
Actual donor \(y_i = 1\) |
Actual non-donor \(y_i = 0\) |
|
Predicted donor \(\hat{y}_i = 1\) |
|
|
Predicted non-donor \(\hat{y}_i = 0\) |
|
|
\(A_i\) is the donation in response to the June 1997 mailing (TARGET_D in the original
data). It is only known for donors, and is 0 for non-donors, for whom the false-negative cost
never applies.
The mailing cost is a symbolic parameter with the default \(c_f = 0.68\), the cost per piece
mailed stated in the competition documentation. It is exposed under the alias contact_cost.
Override it by passing the alias when evaluating the metric:
y_score = pipeline.predict_proba(X)[:, 1]
cost = Metric(dataset.cost_matrix, Cost())
default_cost = cost(y, y_score, **dataset.instance_costs)
expensive_mailing = cost(y, y_score, contact_cost=2.0, **dataset.instance_costs)
6.3.10.4. Data Description#
Descriptions follow the competition’s data dictionary; the original column name is given in
brackets. Dates are YYMM codes and are kept as categories.
Feature |
Description |
Type |
|---|---|---|
|
|
categorical |
|
Whether the donor may not be exchanged in list rentals |
categorical |
|
Age of the donor |
numeric |
|
|
categorical |
|
Number of children |
numeric |
|
Household income band |
numeric |
|
|
categorical |
|
Wealth rating |
numeric |
|
|
categorical |
|
Lifetime number of card promotions received |
numeric |
|
Date of the most recent promotion received |
categorical |
|
Lifetime number of gifts to card promotions |
numeric |
|
Amount of the smallest gift to date, in dollars |
numeric |
|
Date of the smallest gift |
categorical |
|
Amount of the largest gift to date, in dollars |
numeric |
|
Date of the largest gift |
categorical |
|
Amount of the most recent gift, in dollars |
numeric |
|
Date of the most recent gift |
categorical |
|
Date of the first gift |
categorical |
|
Date of the second gift |
categorical |
|
Months between the first and second gift |
numeric |
|
Average amount of the gifts to date, in dollars |
numeric |
donated (target) |
Whether the donor responded to the June 1997 mailing (1 = yes, 0 = no) |
binary |