6.3.15. Default of Credit Card Clients#
6.3.15.1. Summary#
Credit card clients of a bank in Taiwan, collected by Yeh & Lien [1] and published in the UCI Machine Learning Repository. Each row is a client, described by their credit limit, demographics and six months of repayment history (April to September 2005). The task is to predict whether they default on their next payment.
The credit limit of every client is known, which makes it a natural credit line for the instance-dependent cost matrix of Bahnsen et al. [2]. Vanderschueren et al. [3] use the dataset this way, and Empulse follows them.
Classes |
2 |
Defaulters |
6636 |
Non-defaulters |
23364 |
Samples |
30000 |
Features |
23 |
6.3.15.2. Using the Dataset#
The dataset is fetched through fetch_default_credit_card_clients. It is
downloaded from OpenML on first use and cached under ~/empulse_data (override with
$EMPULSE_DATA_HOME or the data_home argument), so later calls work offline.
It returns a Dataset object with the following attributes:
data: the feature matrixtarget: the target vectorcost_matrix: aCostMatrixwith default values pre-filledinstance_costs: a dict of per-instance cost drivers ('cl','fp_cost')feature_names: the feature namestarget_names: the target namesDESCR: the full description of the dataset
import pandas as pd
from empulse.datasets import fetch_default_credit_card_clients
dataset = fetch_default_credit_card_clients(backend=pd)
X, y = dataset.data, dataset.target
The backend argument selects the dataframe library used for data and target.
Pass the module itself — backend=pd for pandas or backend=pl for polars.
All features are numeric codes or amounts. Pass the cost matrix to the model as a
Metric loss, and hand it the instance costs at fit time:
from empulse.metrics import Metric, Cost
from empulse.models import CSLogitClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
pipeline = Pipeline([
('scaler', StandardScaler()),
('model', CSLogitClassifier(loss=Metric(dataset.cost_matrix, Cost()))),
])
pipeline.fit(
X,
y,
model__cl=dataset.instance_costs['cl'],
model__fp_cost=dataset.instance_costs['fp_cost'],
)
6.3.15.3. Cost Matrix#
Actual positive \(y_i = 1\) |
Actual negative \(y_i = 0\) |
|
Predicted positive \(\hat{y}_i = 1\) |
|
|
Predicted negative \(\hat{y}_i = 0\) |
|
|
- with
\(Cl_i\) : the client’s credit limit (
limit_bal)\(r_i\) : the profit lost by refusing a client who would have paid
\(\bar{r}\) : the average profit lost by refusing such a client
\(\pi_0\) : the share of non-defaulters
\(\pi_1\) : the share of defaulters
\(\bar{Cl}\) : the average credit limit
\(L_{gd}\) : the fraction of the credit line lost when the client defaults
Refusing a good client costs the profit their credit would have made, less what lending the money
to an average alternative client would have earned instead [2]. The profit is computed with an
interest rate of 4.79% and a cost of funds of 2.94% (the European rates Bahnsen et al. [2] use)
and the two-year term they fix for credit cards, as Vanderschueren et al. [3] do. It is baked into
'fp_cost'.
The loss given default stays symbolic, with the default \(L_{gd} = 0.75\). Override it by
passing its alias loss_given_default when evaluating the metric:
y_score = pipeline.predict_proba(X)[:, 1]
cost = Metric(dataset.cost_matrix, Cost())
default_lgd = cost(y, y_score, **dataset.instance_costs)
higher_lgd = cost(y, y_score, loss_given_default=0.9, **dataset.instance_costs)
6.3.15.4. Data Description#
Amounts are in New Taiwan dollars. Descriptions follow Yeh & Lien [1].
Feature |
Description |
Type |
|---|---|---|
|
Amount of credit given, including the family’s supplementary credit; also the credit line |
numeric |
|
1 = male, 2 = female |
categorical |
|
1 = graduate school, 2 = university, 3 = high school, 4 = others (the data also contains the undocumented codes 0, 5 and 6) |
categorical |
|
1 = married, 2 = single, 3 = others (the data also contains the undocumented code 0) |
categorical |
|
Age in years |
numeric |
|
Repayment status from September back to April 2005: -1 = paid duly, 1–9 = months of delay (the data also contains -2 and 0) |
numeric |
|
Bill statement amount from September back to April 2005 |
numeric |
|
Amount paid from September back to April 2005 |
numeric |
default (target) |
Whether the client defaulted on their next payment (1 = yes, 0 = no) |
binary |