fetch_south_german_credit#

empulse.datasets.fetch_south_german_credit(*, backend, data_home=None, download_if_missing=True)[source]#

Fetch the South German Credit dataset from the UCI ML Repository (binary classification).

The goal is to predict whether a loan is a bad credit risk. Target variable: 1 = bad credit, 0 = good credit.

This is the corrected version of the Statlog “German Credit” data published by Grömping (2019), as used by Ballegeer et al. (2025). Every feature except duration, amount and age is an integer category code; the code tables are in Grömping (2019).

For additional information about the dataset, consult the User Guide.

Classes

2

Bad credit

300

Good credit

700

Samples

1000

Features

20

Parameters:
backendmodule

Dataframe library to use for data and target. Pass the library module directly, e.g. backend=polars or backend=pandas.

data_homestr or Path, optional

Directory used for caching downloaded data. Defaults to ~/empulse_data (or $EMPULSE_DATA_HOME).

download_if_missingbool, default=True

If False, raise an OSError when the data is not cached locally.

Returns:
datasetDataset

instance_costs contains:

  • 'cl': credit amount per borrower (amount, in Deutsche Mark).

  • 'fp_cost': precomputed false positive cost per borrower.

Notes

Cost matrix (Ballegeer et al. 2025, following Bahnsen et al. 2014):

Actual positive \(y_i = 1\)

Actual negative \(y_i = 0\)

Predicted positive \(\hat{y}_i = 1\)

tp_cost \(= 0\)

fp_cost (precomputed per borrower)

Predicted negative \(\hat{y}_i = 0\)

fn_cost \(= Cl_i \cdot L_{gd}\)

tn_cost \(= 0\)

The cost matrix uses symbolic parameters with the following defaults:

  • loss_given_default (\(L_{gd}\)) = 0.75

The false positive cost is precomputed with an annual interest rate of 4.79%, an annual cost of funds of 2.94% and a 24-month term.

References

[1]

Grömping, U. (2019). South German Credit Data: Correcting a widely used data set. Reports in Mathematics, Physics and Chemistry, Report 4/2019, Department II, Beuth University of Applied Sciences Berlin.

[2]

Ballegeer, M., Bogaert, M., & Benoit, D. F. (2025). Evaluating the stability of model explanations in instance-dependent cost-sensitive credit scoring. European Journal of Operational Research, 326(2), 630–640.

[3]

Bahnsen, A. C., Aouada, D., & Ottersten, B. (2014). Example-dependent cost-sensitive logistic regression for credit scoring. In 2014 13th International Conference on Machine Learning and Applications (pp. 263–269).

Examples

import numpy as np
import pandas as pd
from empulse.datasets import fetch_south_german_credit
from empulse.metrics import Metric, Cost

dataset = fetch_south_german_credit(backend=pd)

# replace with your own model's predicted probabilities
y_score = np.random.default_rng(0).uniform(size=len(dataset.target))

metric = Metric(dataset.cost_matrix, Cost())
score = metric(dataset.target, y_score, **dataset.instance_costs)