fetch_home_equity#
- empulse.datasets.fetch_home_equity(*, backend, data_home=None, download_if_missing=True)[source]#
Fetch the Home Equity (HMEQ) dataset from OpenML (binary classification).
The goal is to predict whether a home equity loan applicant will default or be seriously delinquent. Target variable: 1 = default, 0 = no default.
For additional information about the dataset, consult the User Guide.
Classes
2
Defaults
1189
Non-defaults
4771
Samples
5960
Features
12
- Parameters:
- backendmodule
Dataframe library to use for
dataandtarget. Pass the library module directly, e.g.backend=polarsorbackend=pandas.- data_homestr or Path, optional
Directory used for caching downloaded data. Defaults to
~/empulse_data(or$EMPULSE_DATA_HOME).- download_if_missingbool, default=True
If False, raise an
OSErrorwhen the data is not cached locally.
- Returns:
- dataset
Dataset instance_costscontains:'cl': loan amount per borrower (LOAN).'fp_cost': precomputed false positive cost per borrower.
- dataset
Notes
Cost matrix (Bahnsen et al. 2014, Ballegeer et al. 2025):
Actual positive \(y_i = 1\)
Actual negative \(y_i = 0\)
Predicted positive \(\hat{y}_i = 1\)
tp_cost\(= 0\)fp_cost(precomputed per borrower)Predicted negative \(\hat{y}_i = 0\)
fn_cost\(= Cl_i \cdot L_{gd}\)tn_cost\(= 0\)The cost matrix uses symbolic parameters with the following defaults:
loss_given_default(\(L_{gd}\)) = 0.75
References
[1]Baesens, B., Roesch, D., & Scheule, H. (2016). Credit risk analytics: Measurement techniques, applications, and examples in SAS. John Wiley & Sons.
[2]Ballegeer, M., Bogaert, M., & Benoit, D. F. (2025). Evaluating the stability of model explanations in instance-dependent cost-sensitive credit scoring. European Journal of Operational Research, 326(2), 630–640.
Examples
import numpy as np import pandas as pd from empulse.datasets import fetch_home_equity from empulse.metrics import Metric, Cost dataset = fetch_home_equity(backend=pd) # replace with your own model's predicted probabilities y_score = np.random.default_rng(0).uniform(size=len(dataset.target)) metric = Metric(dataset.cost_matrix, Cost()) score = metric(dataset.target, y_score, **dataset.instance_costs)