6.3.11. Credit Card Fraud Detection (ULB)#

6.3.11.1. Summary#

Card transactions made by European cardholders over two days in September 2013, collected during a research collaboration between Worldline and the Machine Learning Group of the Université Libre de Bruxelles [1]. The task is to flag fraudulent transactions. For confidentiality, all features except the transaction amount are principal components of the original variables.

Fraud is rare: the raw data holds 492 frauds among 284,807 transactions. Following Höppner et al. [2] and Vanderschueren et al. [3], the 1,825 transactions with a zero amount are removed, since missing them costs nothing. That leaves 282,982 transactions, 465 of them fraudulent.

Classes

2

Frauds

465

Legitimate

282517

Samples

282982

Features

29

6.3.11.2. Using the Dataset#

The dataset is fetched through fetch_credit_card_fraud. It is downloaded from OpenML on first use (about 150 MB) and cached under ~/empulse_data (override with $EMPULSE_DATA_HOME or the data_home argument), so later calls work offline.

It returns a Dataset object with the following attributes:

  • data: the feature matrix

  • target: the target vector

  • cost_matrix: a CostMatrix with default values pre-filled

  • instance_costs: a dict of per-instance cost drivers ('amount')

  • feature_names: the feature names

  • target_names: the target names

  • DESCR: the full description of the dataset

import pandas as pd
from empulse.datasets import fetch_credit_card_fraud

dataset = fetch_credit_card_fraud(backend=pd)
X, y = dataset.data, dataset.target

The backend argument selects the dataframe library used for data and target. Pass the module itself — backend=pd for pandas or backend=pl for polars.

All features are numeric. Pass the cost matrix to the model as a Metric loss, and hand it the transaction amounts at fit time:

from empulse.metrics import Metric, Cost
from empulse.models import CSLogitClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

pipeline = Pipeline([
    ('scaler', StandardScaler()),
    ('model', CSLogitClassifier(loss=Metric(dataset.cost_matrix, Cost()))),
])
pipeline.fit(X, y, model__amount=dataset.instance_costs['amount'])

6.3.11.3. Cost Matrix#

Flagging a transaction triggers an investigation with a fixed administrative cost \(c_f\), whether or not it turns out to be fraud. A fraud that is not flagged costs its full amount \(A_i\) [2] [3].

Actual fraud \(y_i = 1\)

Actual legitimate \(y_i = 0\)

Predicted fraud \(\hat{y}_i = 1\)

tp_cost \(= c_f\)

fp_cost \(= c_f\)

Predicted legitimate \(\hat{y}_i = 0\)

fn_cost \(= A_i\)

tn_cost \(= 0\)

The investigation cost is a symbolic parameter with the default \(c_f = 10\) used in both papers, exposed under the alias investigation_cost. Override it by passing the alias when evaluating the metric:

y_score = pipeline.predict_proba(X)[:, 1]

cost = Metric(dataset.cost_matrix, Cost())
default_cost = cost(y, y_score, **dataset.instance_costs)
expensive_investigation = cost(y, y_score, investigation_cost=50, **dataset.instance_costs)

6.3.11.4. Data Description#

Feature

Description

v1 – v28

Principal components of the original, confidential transaction variables

amount

Transaction amount in euros; also the false-negative cost

fraud (target)

Whether the transaction is fraudulent (1 = fraud, 0 = legitimate)

The Time column of the original data, the seconds elapsed since the first transaction, is dropped.

6.3.11.5. References#