fetch_ieee_fraud_detection#

empulse.datasets.fetch_ieee_fraud_detection(*, backend, data_home=None, download_if_missing=True)[source]#

Fetch the IEEE-CIS Fraud Detection dataset from OpenML (binary classification).

The goal is to predict whether an e-commerce transaction is fraudulent. Target variable: 1 = fraud, 0 = not fraud.

For additional information about the dataset, consult the User Guide.

Classes

2

Frauds

20663

Legitimate

569877

Samples

590540

Features

431

Parameters:
backendmodule

Dataframe library to use for data and target. Pass the library module directly, e.g. backend=polars or backend=pandas.

data_homestr or Path, optional

Directory used for caching downloaded data. Defaults to ~/empulse_data (or $EMPULSE_DATA_HOME).

download_if_missingbool, default=True

If False, raise an OSError when the data is not cached locally.

Returns:
datasetDataset

instance_costs contains {'amount': array} — the transaction amount representing the loss if a fraudulent transaction is missed (false negative).

Notes

Cost matrix (Höppner et al. 2022, Vanderschueren et al. 2022):

Actual positive \(y_i = 1\)

Actual negative \(y_i = 0\)

Predicted positive \(\hat{y}_i = 1\)

tp_cost \(= c_f\)

fp_cost \(= c_f\)

Predicted negative \(\hat{y}_i = 0\)

fn_cost \(= amount_i\)

tn_cost \(= 0\)

The cost matrix uses symbolic parameters with the following defaults:

  • investigation_cost (\(c_f\)) = 10.0 (cost of investigating an alert)

References

[1]

Höppner, S., Baesens, B., Verbeke, W., & Verdonck, T. (2022). Instance-dependent cost-sensitive learning for detecting transfer fraud. European Journal of Operational Research, 297(1), 291–300.

[2]

Vanderschueren, T., Verdonck, T., Baesens, B., & Verbeke, W. (2022). Predict-then-optimize or predict-and-optimize? An empirical evaluation of cost-sensitive learning strategies. Information Sciences, 594, 400–415.

Examples

import numpy as np
import pandas as pd
from empulse.datasets import fetch_ieee_fraud_detection
from empulse.metrics import Metric, Cost

dataset = fetch_ieee_fraud_detection(backend=pd)

# replace with your own model's predicted probabilities
y_score = np.random.default_rng(0).uniform(size=len(dataset.target))

metric = Metric(dataset.cost_matrix, Cost())
score = metric(dataset.target, y_score, **dataset.instance_costs)