fetch_default_credit_card_clients#

empulse.datasets.fetch_default_credit_card_clients(*, backend, data_home=None, download_if_missing=True)[source]#

Fetch the Default of Credit Card Clients dataset from OpenML (binary classification).

The goal is to predict whether a client will default on their credit card payment. Target variable: 1 = default, 0 = no default.

For additional information about the dataset, consult the User Guide.

Classes

2

Defaulters

6636

Non-defaulters

23364

Samples

30000

Features

23

Parameters:
backendmodule

Dataframe library to use for data and target. Pass the library module directly, e.g. backend=polars or backend=pandas.

data_homestr or Path, optional

Directory used for caching downloaded data. Defaults to ~/empulse_data (or $EMPULSE_DATA_HOME).

download_if_missingbool, default=True

If False, raise an OSError when the data is not cached locally.

Returns:
datasetDataset

instance_costs contains:

  • 'cl': credit limit (limit_bal) per client.

  • 'fp_cost': precomputed FP cost per client following Bahnsen et al. (2014).

Notes

Cost matrix (Bahnsen et al. 2014, Vanderschueren et al. 2022):

Actual positive \(y_i = 1\)

Actual negative \(y_i = 0\)

Predicted positive \(\hat{y}_i = 1\)

tp_cost \(= 0\)

fp_cost (precomputed per client)

Predicted negative \(\hat{y}_i = 0\)

fn_cost \(= Cl_i \cdot L_{gd}\)

tn_cost \(= 0\)

The cost matrix uses symbolic parameters with the following defaults:

  • loss_given_default (\(L_{gd}\)) = 0.75

References

[1]

Yeh, I. C., & Lien, C. H. (2009). The comparisons of data mining techniques for the predictive accuracy of probability of default of credit card clients. Expert Systems with Applications, 36(2), 2473–2480.

[2]

Bahnsen, A. C., Aouada, D., & Ottersten, B. (2014). Example-dependent cost-sensitive logistic regression for credit scoring. In 2014 13th International Conference on Machine Learning and Applications (pp. 263-269).

Examples

import numpy as np
import pandas as pd
from empulse.datasets import fetch_default_credit_card_clients
from empulse.metrics import Metric, Cost

dataset = fetch_default_credit_card_clients(backend=pd)

# replace with your own model's predicted probabilities
y_score = np.random.default_rng(0).uniform(size=len(dataset.target))

metric = Metric(dataset.cost_matrix, Cost())
score = metric(dataset.target, y_score, **dataset.instance_costs)