6.3.15. Default of Credit Card Clients#

6.3.15.1. Summary#

Credit card clients of a bank in Taiwan, collected by Yeh & Lien [1] and published in the UCI Machine Learning Repository. Each row is a client, described by their credit limit, demographics and six months of repayment history (April to September 2005). The task is to predict whether they default on their next payment.

The credit limit of every client is known, which makes it a natural credit line for the instance-dependent cost matrix of Bahnsen et al. [2]. Vanderschueren et al. [3] use the dataset this way, and Empulse follows them.

Classes

2

Defaulters

6636

Non-defaulters

23364

Samples

30000

Features

23

6.3.15.2. Using the Dataset#

The dataset is fetched through fetch_default_credit_card_clients. It is downloaded from OpenML on first use and cached under ~/empulse_data (override with $EMPULSE_DATA_HOME or the data_home argument), so later calls work offline.

It returns a Dataset object with the following attributes:

  • data: the feature matrix

  • target: the target vector

  • cost_matrix: a CostMatrix with default values pre-filled

  • instance_costs: a dict of per-instance cost drivers ('cl', 'fp_cost')

  • feature_names: the feature names

  • target_names: the target names

  • DESCR: the full description of the dataset

import pandas as pd
from empulse.datasets import fetch_default_credit_card_clients

dataset = fetch_default_credit_card_clients(backend=pd)
X, y = dataset.data, dataset.target

The backend argument selects the dataframe library used for data and target. Pass the module itself — backend=pd for pandas or backend=pl for polars.

All features are numeric codes or amounts. Pass the cost matrix to the model as a Metric loss, and hand it the instance costs at fit time:

from empulse.metrics import Metric, Cost
from empulse.models import CSLogitClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

pipeline = Pipeline([
    ('scaler', StandardScaler()),
    ('model', CSLogitClassifier(loss=Metric(dataset.cost_matrix, Cost()))),
])
pipeline.fit(
    X,
    y,
    model__cl=dataset.instance_costs['cl'],
    model__fp_cost=dataset.instance_costs['fp_cost'],
)

6.3.15.3. Cost Matrix#

Actual positive \(y_i = 1\)

Actual negative \(y_i = 0\)

Predicted positive \(\hat{y}_i = 1\)

tp_cost \(= 0\)

fp_cost \(= r_i + -\bar{r} \cdot \pi_0 + \bar{Cl} \cdot L_{gd} \cdot \pi_1\)

Predicted negative \(\hat{y}_i = 0\)

fn_cost \(= Cl_i \cdot L_{gd}\)

tn_cost \(= 0\)

with
  • \(Cl_i\) : the client’s credit limit (limit_bal)

  • \(r_i\) : the profit lost by refusing a client who would have paid

  • \(\bar{r}\) : the average profit lost by refusing such a client

  • \(\pi_0\) : the share of non-defaulters

  • \(\pi_1\) : the share of defaulters

  • \(\bar{Cl}\) : the average credit limit

  • \(L_{gd}\) : the fraction of the credit line lost when the client defaults

Refusing a good client costs the profit their credit would have made, less what lending the money to an average alternative client would have earned instead [2]. The profit is computed with an interest rate of 4.79% and a cost of funds of 2.94% (the European rates Bahnsen et al. [2] use) and the two-year term they fix for credit cards, as Vanderschueren et al. [3] do. It is baked into 'fp_cost'.

The loss given default stays symbolic, with the default \(L_{gd} = 0.75\). Override it by passing its alias loss_given_default when evaluating the metric:

y_score = pipeline.predict_proba(X)[:, 1]

cost = Metric(dataset.cost_matrix, Cost())
default_lgd = cost(y, y_score, **dataset.instance_costs)
higher_lgd = cost(y, y_score, loss_given_default=0.9, **dataset.instance_costs)

6.3.15.4. Data Description#

Amounts are in New Taiwan dollars. Descriptions follow Yeh & Lien [1].

Feature

Description

Type

limit_bal

Amount of credit given, including the family’s supplementary credit; also the credit line

numeric

sex

1 = male, 2 = female

categorical

education

1 = graduate school, 2 = university, 3 = high school, 4 = others (the data also contains the undocumented codes 0, 5 and 6)

categorical

marriage

1 = married, 2 = single, 3 = others (the data also contains the undocumented code 0)

categorical

age

Age in years

numeric

pay_0, pay_2 – pay_6

Repayment status from September back to April 2005: -1 = paid duly, 1–9 = months of delay (the data also contains -2 and 0)

numeric

bill_amt1 – bill_amt6

Bill statement amount from September back to April 2005

numeric

pay_amt1 – pay_amt6

Amount paid from September back to April 2005

numeric

default (target)

Whether the client defaulted on their next payment (1 = yes, 0 = no)

binary

6.3.15.5. References#