fetch_credit_card_fraud#
- empulse.datasets.fetch_credit_card_fraud(*, backend, data_home=None, download_if_missing=True)[source]#
Fetch the Kaggle Credit Card Fraud (ULB MLG) dataset from OpenML (binary classification).
The goal is to predict whether a credit card transaction is fraudulent. Target variable: 1 = fraud, 0 = legitimate.
Transactions with zero amount are filtered out following Höppner et al. (2022) and Vanderschueren et al. (2022), resulting in 282,982 samples.
For additional information about the dataset, consult the User Guide.
Classes
2
Frauds
465
Legitimate
282517
Samples
282982
Features
29
- Parameters:
- backendmodule
Dataframe library to use for
dataandtarget. Pass the library module directly, e.g.backend=polarsorbackend=pandas.- data_homestr or Path, optional
Directory used for caching downloaded data. Defaults to
~/empulse_data(or$EMPULSE_DATA_HOME).- download_if_missingbool, default=True
If False, raise an
OSErrorwhen the data is not cached locally.
- Returns:
- dataset
Dataset instance_costscontains{'amount': array}— the transaction amount representing the loss if a fraudulent transaction is missed (false negative).
- dataset
Notes
Cost matrix (Höppner et al. 2022, Vanderschueren et al. 2022):
Actual positive \(y_i = 1\)
Actual negative \(y_i = 0\)
Predicted positive \(\hat{y}_i = 1\)
tp_cost\(= c_f\)fp_cost\(= c_f\)Predicted negative \(\hat{y}_i = 0\)
fn_cost\(= amount_i\)tn_cost\(= 0\)The cost matrix uses symbolic parameters with the following defaults:
investigation_cost(\(c_f\)) = 10.0 (cost of investigating an alert)
References
[1]Dal Pozzolo, A., Caelen, O., Johnson, R. A., & Bontempi, G. (2015). Calibrating probability with undersampling for unbalanced classification. In 2015 IEEE Symposium Series on Computational Intelligence (pp. 159-166).
[2]Höppner, S., Baesens, B., Verbeke, W., & Verdonck, T. (2022). Instance-dependent cost-sensitive learning for detecting transfer fraud. European Journal of Operational Research, 297(1), 291–300.
[3]Vanderschueren, T., Verdonck, T., Baesens, B., & Verbeke, W. (2022). Predict-then-optimize or predict-and-optimize? An empirical evaluation of cost-sensitive learning strategies. Information Sciences, 594, 400–415.
[4]De Vos, S., Vanderschueren, T., Verdonck, T., & Verbeke, W. (2023). Robust instance-dependent cost-sensitive classification. Advances in Data Analysis and Classification, 1–23.
Examples
import numpy as np import pandas as pd from empulse.datasets import fetch_credit_card_fraud from empulse.metrics import Metric, Cost dataset = fetch_credit_card_fraud(backend=pd) # replace with your own model's predicted probabilities y_score = np.random.default_rng(0).uniform(size=len(dataset.target)) metric = Metric(dataset.cost_matrix, Cost()) score = metric(dataset.target, y_score, **dataset.instance_costs)