6.3.14. South German Credit#
6.3.14.1. Summary#
1,000 consumer loans granted by a southern German bank in 1973–1975, in the corrected version published by Grömping [1]. The widely used Statlog “German Credit” data comes from the same loans, but many of its variables are wrongly coded; this version fixes them. The task is to predict whether a loan is a bad credit risk.
Bad credits are heavily oversampled, so the 30% default rate is far above the bank’s actual rate. Ballegeer et al. [2] use this dataset with the cost matrix of Bahnsen et al. [3], taking the credit amount as the credit line. Empulse follows them.
Classes |
2 |
Bad credit |
300 |
Good credit |
700 |
Samples |
1000 |
Features |
20 |
6.3.14.2. Using the Dataset#
The dataset is fetched through fetch_south_german_credit. It is
downloaded from the UCI Machine Learning Repository on first use and cached under
~/empulse_data (override with $EMPULSE_DATA_HOME or the data_home argument), so later
calls work offline.
It returns a Dataset object with the following attributes:
data: the feature matrixtarget: the target vectorcost_matrix: aCostMatrixwith default values pre-filledinstance_costs: a dict of per-instance cost drivers ('cl','fp_cost')feature_names: the feature namestarget_names: the target namesDESCR: the full description of the dataset
import pandas as pd
from empulse.datasets import fetch_south_german_credit
dataset = fetch_south_german_credit(backend=pd)
X, y = dataset.data, dataset.target
The backend argument selects the dataframe library used for data and target.
Pass the module itself — backend=pd for pandas or backend=pl for polars.
Every feature except duration, amount and age is an integer category code, so one-hot
encode those before fitting a linear model:
from empulse.metrics import Metric, Cost
from empulse.models import CSLogitClassifier
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
numeric = ['duration', 'amount', 'age']
categorical = [c for c in X.columns if c not in numeric]
pipeline = Pipeline([
('preprocessor', ColumnTransformer([
('num', StandardScaler(), numeric),
('cat', OneHotEncoder(handle_unknown='ignore', sparse_output=False), categorical),
])),
('model', CSLogitClassifier(loss=Metric(dataset.cost_matrix, Cost()))),
])
pipeline.fit(
X,
y,
model__cl=dataset.instance_costs['cl'],
model__fp_cost=dataset.instance_costs['fp_cost'],
)
6.3.14.3. Cost Matrix#
Actual positive \(y_i = 1\) |
Actual negative \(y_i = 0\) |
|
Predicted positive \(\hat{y}_i = 1\) |
|
|
Predicted negative \(\hat{y}_i = 0\) |
|
|
- with
\(Cl_i\) : the credit amount (
amount, in Deutsche Mark)\(r_i\) : the profit lost by rejecting what would have been a good loan
\(\bar{r}\) : the average profit lost by rejecting a good loan
\(\pi_0\) : the share of good credits
\(\pi_1\) : the share of bad credits
\(\bar{Cl}\) : the average credit amount
\(L_{gd}\) : the fraction of the credit amount lost when the borrower defaults
Rejecting a good applicant costs the profit their loan would have made, less what lending the money
to an average alternative applicant would have earned instead [3]. The profit is computed with an
interest rate of 4.79%, a cost of funds of 2.94% and a term of 24 months, and is baked into
'fp_cost'.
Because \(\pi_1\) enters the false positive cost, the oversampled bad credits make that cost larger than it would be at the bank’s real default rate.
The loss given default stays symbolic, with the default \(L_{gd} = 0.75\) of Ballegeer et
al. [2]. Override it by passing its alias loss_given_default when evaluating the metric:
y_score = pipeline.predict_proba(X)[:, 1]
cost = Metric(dataset.cost_matrix, Cost())
default_lgd = cost(y, y_score, **dataset.instance_costs)
higher_lgd = cost(y, y_score, loss_given_default=0.9, **dataset.instance_costs)
6.3.14.4. Data Description#
Feature names are the English names of Grömping [1], whose code tables define every category code.
Feature |
Description |
Type |
|---|---|---|
|
Status of the checking account (1 = no account, 2 = negative balance, 3 = up to 200 DM, 4 = 200 DM or more) |
categorical |
|
Duration of the credit, in months |
numeric |
|
History of compliance with previous credit contracts (0 = delays in the past, …, 4 = all credits at this bank paid back duly) |
categorical |
|
Purpose of the credit (0 = others, 1 = new car, 2 = used car, …, 10 = business) |
categorical |
|
Credit amount in Deutsche Mark; also the credit line |
numeric |
|
Savings (1 = unknown or none, …, 5 = 1,000 DM or more) |
categorical |
|
Time with the current employer (1 = unemployed, …, 5 = 7 years or more) |
categorical |
|
Instalments as a share of disposable income (1 = 35% or more, …, 4 = under 20%) |
categorical |
|
Combined sex and marital status |
categorical |
|
Other debtors or guarantors (1 = none, 2 = co-applicant, 3 = guarantor) |
categorical |
|
Time at the current residence (1 = under a year, …, 4 = 7 years or more) |
categorical |
|
Most valuable property (1 = unknown or none, …, 4 = real estate) |
categorical |
|
Age in years |
numeric |
|
Instalment plans with other providers (1 = bank, 2 = stores, 3 = none) |
categorical |
|
Type of housing (1 = for free, 2 = rent, 3 = own) |
categorical |
|
Number of credits at this bank (1 = one, …, 4 = six or more) |
categorical |
|
Quality of the job (1 = unemployed or unskilled non-resident, …, 4 = manager, self-employed or highly qualified) |
categorical |
|
Number of people financially dependent on the debtor (1 = three or more, 2 = up to two) |
categorical |
|
Whether a telephone is registered in the customer’s name (1 = no, 2 = yes) |
categorical |
|
Whether the debtor is a foreign worker (1 = yes, 2 = no) |
categorical |
default (target) |
Whether the credit is a bad risk (1 = bad, 0 = good) |
binary |