6.3.6. VUB Credit Scoring#

6.3.6.1. Summary#

This dataset holds loans granted by a Romanian non-banking financial institution, published in anonymised form by the VUB Data Analytics Laboratory alongside Petrides et al. [1]. The goal is to predict whether a borrower will fall 45 or more days behind on a payment.

It has since become one of the standard benchmarks for instance-dependent cost-sensitive credit scoring [2] [3]. The dataset is bundled with Empulse and works offline.

Classes

2

Defaulters

3206

Non-defaulters

15711

Samples

18917

Features

16

6.3.6.2. Using the Dataset#

The dataset is loaded through the load_vub_credit_scoring function. This returns a Dataset object with the following attributes:

  • data: the feature matrix

  • target: the target vector

  • cost_matrix: a CostMatrix with default values pre-filled

  • instance_costs: a dict of per-instance cost drivers ('cl', 'fp_cost')

  • feature_names: the feature names

  • target_names: the target names

  • DESCR: the full description of the dataset

import numpy as np
import pandas as pd
from empulse.datasets import load_vub_credit_scoring
from empulse.metrics import Metric, Cost

dataset = load_vub_credit_scoring(backend=pd)

# replace with your own model's predicted probabilities
y_score = np.random.default_rng(0).uniform(size=len(dataset.target))

cost = Metric(dataset.cost_matrix, Cost())
default_lgd = cost(dataset.target, y_score, **dataset.instance_costs)
higher_lgd = cost(dataset.target, y_score, loss_given_default=0.9, **dataset.instance_costs)

6.3.6.3. Cost Matrix#

Actual positive \(y_i = 1\)

Actual negative \(y_i = 0\)

Predicted positive \(\hat{y}_i = 1\)

tp_cost \(= 0\)

fp_cost \(= r_i + -\bar{r} \cdot \pi_0 + \bar{Cl} \cdot L_{gd} \cdot \pi_1\)

Predicted negative \(\hat{y}_i = 0\)

fn_cost \(= Cl_i \cdot L_{gd}\)

tn_cost \(= 0\)

with
  • \(r_i\) : loss in profit by rejecting what would have been a good loan

  • \(\bar{r}\) : average loss in profit by rejecting what would have been a good loan

  • \(\pi_0\) : percentage of non-defaulters

  • \(\pi_1\) : percentage of defaulters

  • \(Cl_i\) : credit line of the borrower

  • \(\bar{Cl}\) : average credit line

  • \(L_{gd}\) : the fraction of the loan amount which is lost if the borrower defaults

This is the credit scoring cost matrix of Bahnsen et al. [4], with an interest rate of 4.79%, a cost of funds of 2.94%, a term of 24 months and a loss given default of 75%. The loss given default stays symbolic; the other parameters are baked into 'fp_cost'.

Petrides et al. [1] derived their own costs from the institution’s average return on investment and loss given default per business channel. That is not possible from the published data: every monetary column, including the loan amount and the Expected_loss and Expected_profit columns, is standardised to zero mean and unit variance. Following Vanderschueren et al. [2], the loan amount is shifted to be strictly positive, \(Cl_i = Loan\_amount_i - \min_j Loan\_amount_j + 10^{-9}\), and used as the credit line. The resulting costs are in arbitrary units: compare them relative to each other, not as money.

6.3.6.4. Data Description#

Variable Name

Description

Type

v1 – v8

Anonymised application variables

categorical

has_fico

Whether the applicant has a FICO score

binary

business_channel

Business channel through which the loan was granted (1, 2 or 3)

categorical

fico_score

FICO score, standardised; 0 when missing

numeric

loan_amount

Loan amount, standardised

numeric

monthly_income

Monthly income, standardised

numeric

age

Age of the borrower, standardised

numeric

gearing_coefficient

Gearing coefficient, standardised

numeric

max_gearing_ratio

Maximum gearing ratio, standardised

numeric

default

Whether the borrower fell 45 or more days behind on a payment (1 = yes, 0 = no)

binary

6.3.6.5. References#