fetch_south_german_credit#
- empulse.datasets.fetch_south_german_credit(*, backend, data_home=None, download_if_missing=True)[source]#
Fetch the South German Credit dataset from the UCI ML Repository (binary classification).
The goal is to predict whether a loan is a bad credit risk. Target variable: 1 = bad credit, 0 = good credit.
This is the corrected version of the Statlog “German Credit” data published by Grömping (2019), as used by Ballegeer et al. (2025). Every feature except
duration,amountandageis an integer category code; the code tables are in Grömping (2019).For additional information about the dataset, consult the User Guide.
Classes
2
Bad credit
300
Good credit
700
Samples
1000
Features
20
- Parameters:
- backendmodule
Dataframe library to use for
dataandtarget. Pass the library module directly, e.g.backend=polarsorbackend=pandas.- data_homestr or Path, optional
Directory used for caching downloaded data. Defaults to
~/empulse_data(or$EMPULSE_DATA_HOME).- download_if_missingbool, default=True
If False, raise an
OSErrorwhen the data is not cached locally.
- Returns:
- dataset
Dataset instance_costscontains:'cl': credit amount per borrower (amount, in Deutsche Mark).'fp_cost': precomputed false positive cost per borrower.
- dataset
Notes
Cost matrix (Ballegeer et al. 2025, following Bahnsen et al. 2014):
Actual positive \(y_i = 1\)
Actual negative \(y_i = 0\)
Predicted positive \(\hat{y}_i = 1\)
tp_cost\(= 0\)fp_cost(precomputed per borrower)Predicted negative \(\hat{y}_i = 0\)
fn_cost\(= Cl_i \cdot L_{gd}\)tn_cost\(= 0\)The cost matrix uses symbolic parameters with the following defaults:
loss_given_default(\(L_{gd}\)) = 0.75
The false positive cost is precomputed with an annual interest rate of 4.79%, an annual cost of funds of 2.94% and a 24-month term.
References
[1]Grömping, U. (2019). South German Credit Data: Correcting a widely used data set. Reports in Mathematics, Physics and Chemistry, Report 4/2019, Department II, Beuth University of Applied Sciences Berlin.
[2]Ballegeer, M., Bogaert, M., & Benoit, D. F. (2025). Evaluating the stability of model explanations in instance-dependent cost-sensitive credit scoring. European Journal of Operational Research, 326(2), 630–640.
[3]Bahnsen, A. C., Aouada, D., & Ottersten, B. (2014). Example-dependent cost-sensitive logistic regression for credit scoring. In 2014 13th International Conference on Machine Learning and Applications (pp. 263–269).
Examples
import numpy as np import pandas as pd from empulse.datasets import fetch_south_german_credit from empulse.metrics import Metric, Cost dataset = fetch_south_german_credit(backend=pd) # replace with your own model's predicted probabilities y_score = np.random.default_rng(0).uniform(size=len(dataset.target)) metric = Metric(dataset.cost_matrix, Cost()) score = metric(dataset.target, y_score, **dataset.instance_costs)