6.3.14. South German Credit#

6.3.14.1. Summary#

1,000 consumer loans granted by a southern German bank in 1973–1975, in the corrected version published by Grömping [1]. The widely used Statlog “German Credit” data comes from the same loans, but many of its variables are wrongly coded; this version fixes them. The task is to predict whether a loan is a bad credit risk.

Bad credits are heavily oversampled, so the 30% default rate is far above the bank’s actual rate. Ballegeer et al. [2] use this dataset with the cost matrix of Bahnsen et al. [3], taking the credit amount as the credit line. Empulse follows them.

Classes

2

Bad credit

300

Good credit

700

Samples

1000

Features

20

6.3.14.2. Using the Dataset#

The dataset is fetched through fetch_south_german_credit. It is downloaded from the UCI Machine Learning Repository on first use and cached under ~/empulse_data (override with $EMPULSE_DATA_HOME or the data_home argument), so later calls work offline.

It returns a Dataset object with the following attributes:

  • data: the feature matrix

  • target: the target vector

  • cost_matrix: a CostMatrix with default values pre-filled

  • instance_costs: a dict of per-instance cost drivers ('cl', 'fp_cost')

  • feature_names: the feature names

  • target_names: the target names

  • DESCR: the full description of the dataset

import pandas as pd
from empulse.datasets import fetch_south_german_credit

dataset = fetch_south_german_credit(backend=pd)
X, y = dataset.data, dataset.target

The backend argument selects the dataframe library used for data and target. Pass the module itself — backend=pd for pandas or backend=pl for polars.

Every feature except duration, amount and age is an integer category code, so one-hot encode those before fitting a linear model:

from empulse.metrics import Metric, Cost
from empulse.models import CSLogitClassifier
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

numeric = ['duration', 'amount', 'age']
categorical = [c for c in X.columns if c not in numeric]

pipeline = Pipeline([
    ('preprocessor', ColumnTransformer([
        ('num', StandardScaler(), numeric),
        ('cat', OneHotEncoder(handle_unknown='ignore', sparse_output=False), categorical),
    ])),
    ('model', CSLogitClassifier(loss=Metric(dataset.cost_matrix, Cost()))),
])
pipeline.fit(
    X,
    y,
    model__cl=dataset.instance_costs['cl'],
    model__fp_cost=dataset.instance_costs['fp_cost'],
)

6.3.14.3. Cost Matrix#

Actual positive \(y_i = 1\)

Actual negative \(y_i = 0\)

Predicted positive \(\hat{y}_i = 1\)

tp_cost \(= 0\)

fp_cost \(= r_i + -\bar{r} \cdot \pi_0 + \bar{Cl} \cdot L_{gd} \cdot \pi_1\)

Predicted negative \(\hat{y}_i = 0\)

fn_cost \(= Cl_i \cdot L_{gd}\)

tn_cost \(= 0\)

with
  • \(Cl_i\) : the credit amount (amount, in Deutsche Mark)

  • \(r_i\) : the profit lost by rejecting what would have been a good loan

  • \(\bar{r}\) : the average profit lost by rejecting a good loan

  • \(\pi_0\) : the share of good credits

  • \(\pi_1\) : the share of bad credits

  • \(\bar{Cl}\) : the average credit amount

  • \(L_{gd}\) : the fraction of the credit amount lost when the borrower defaults

Rejecting a good applicant costs the profit their loan would have made, less what lending the money to an average alternative applicant would have earned instead [3]. The profit is computed with an interest rate of 4.79%, a cost of funds of 2.94% and a term of 24 months, and is baked into 'fp_cost'.

Because \(\pi_1\) enters the false positive cost, the oversampled bad credits make that cost larger than it would be at the bank’s real default rate.

The loss given default stays symbolic, with the default \(L_{gd} = 0.75\) of Ballegeer et al. [2]. Override it by passing its alias loss_given_default when evaluating the metric:

y_score = pipeline.predict_proba(X)[:, 1]

cost = Metric(dataset.cost_matrix, Cost())
default_lgd = cost(y, y_score, **dataset.instance_costs)
higher_lgd = cost(y, y_score, loss_given_default=0.9, **dataset.instance_costs)

6.3.14.4. Data Description#

Feature names are the English names of Grömping [1], whose code tables define every category code.

Feature

Description

Type

status

Status of the checking account (1 = no account, 2 = negative balance, 3 = up to 200 DM, 4 = 200 DM or more)

categorical

duration

Duration of the credit, in months

numeric

credit_history

History of compliance with previous credit contracts (0 = delays in the past, …, 4 = all credits at this bank paid back duly)

categorical

purpose

Purpose of the credit (0 = others, 1 = new car, 2 = used car, …, 10 = business)

categorical

amount

Credit amount in Deutsche Mark; also the credit line

numeric

savings

Savings (1 = unknown or none, …, 5 = 1,000 DM or more)

categorical

employment_duration

Time with the current employer (1 = unemployed, …, 5 = 7 years or more)

categorical

installment_rate

Instalments as a share of disposable income (1 = 35% or more, …, 4 = under 20%)

categorical

personal_status_sex

Combined sex and marital status

categorical

other_debtors

Other debtors or guarantors (1 = none, 2 = co-applicant, 3 = guarantor)

categorical

present_residence

Time at the current residence (1 = under a year, …, 4 = 7 years or more)

categorical

property

Most valuable property (1 = unknown or none, …, 4 = real estate)

categorical

age

Age in years

numeric

other_installment_plans

Instalment plans with other providers (1 = bank, 2 = stores, 3 = none)

categorical

housing

Type of housing (1 = for free, 2 = rent, 3 = own)

categorical

number_credits

Number of credits at this bank (1 = one, …, 4 = six or more)

categorical

job

Quality of the job (1 = unemployed or unskilled non-resident, …, 4 = manager, self-employed or highly qualified)

categorical

people_liable

Number of people financially dependent on the debtor (1 = three or more, 2 = up to two)

categorical

telephone

Whether a telephone is registered in the customer’s name (1 = no, 2 = yes)

categorical

foreign_worker

Whether the debtor is a foreign worker (1 = yes, 2 = no)

categorical

default (target)

Whether the credit is a bad risk (1 = bad, 0 = good)

binary

6.3.14.5. References#