fetch_credit_card_fraud#

empulse.datasets.fetch_credit_card_fraud(*, backend, data_home=None, download_if_missing=True)[source]#

Fetch the Kaggle Credit Card Fraud (ULB MLG) dataset from OpenML (binary classification).

The goal is to predict whether a credit card transaction is fraudulent. Target variable: 1 = fraud, 0 = legitimate.

Transactions with zero amount are filtered out following Höppner et al. (2022) and Vanderschueren et al. (2022), resulting in 282,982 samples.

For additional information about the dataset, consult the User Guide.

Classes

2

Frauds

465

Legitimate

282517

Samples

282982

Features

29

Parameters:
backendmodule

Dataframe library to use for data and target. Pass the library module directly, e.g. backend=polars or backend=pandas.

data_homestr or Path, optional

Directory used for caching downloaded data. Defaults to ~/empulse_data (or $EMPULSE_DATA_HOME).

download_if_missingbool, default=True

If False, raise an OSError when the data is not cached locally.

Returns:
datasetDataset

instance_costs contains {'amount': array} — the transaction amount representing the loss if a fraudulent transaction is missed (false negative).

Notes

Cost matrix (Höppner et al. 2022, Vanderschueren et al. 2022):

Actual positive \(y_i = 1\)

Actual negative \(y_i = 0\)

Predicted positive \(\hat{y}_i = 1\)

tp_cost \(= c_f\)

fp_cost \(= c_f\)

Predicted negative \(\hat{y}_i = 0\)

fn_cost \(= amount_i\)

tn_cost \(= 0\)

The cost matrix uses symbolic parameters with the following defaults:

  • investigation_cost (\(c_f\)) = 10.0 (cost of investigating an alert)

References

[1]

Dal Pozzolo, A., Caelen, O., Johnson, R. A., & Bontempi, G. (2015). Calibrating probability with undersampling for unbalanced classification. In 2015 IEEE Symposium Series on Computational Intelligence (pp. 159-166).

[2]

Höppner, S., Baesens, B., Verbeke, W., & Verdonck, T. (2022). Instance-dependent cost-sensitive learning for detecting transfer fraud. European Journal of Operational Research, 297(1), 291–300.

[3]

Vanderschueren, T., Verdonck, T., Baesens, B., & Verbeke, W. (2022). Predict-then-optimize or predict-and-optimize? An empirical evaluation of cost-sensitive learning strategies. Information Sciences, 594, 400–415.

[4]

De Vos, S., Vanderschueren, T., Verdonck, T., & Verbeke, W. (2023). Robust instance-dependent cost-sensitive classification. Advances in Data Analysis and Classification, 1–23.

Examples

import numpy as np
import pandas as pd
from empulse.datasets import fetch_credit_card_fraud
from empulse.metrics import Metric, Cost

dataset = fetch_credit_card_fraud(backend=pd)

# replace with your own model's predicted probabilities
y_score = np.random.default_rng(0).uniform(size=len(dataset.target))

metric = Metric(dataset.cost_matrix, Cost())
score = metric(dataset.target, y_score, **dataset.instance_costs)