6.3.11. Credit Card Fraud Detection (ULB)#
6.3.11.1. Summary#
Card transactions made by European cardholders over two days in September 2013, collected during a research collaboration between Worldline and the Machine Learning Group of the Université Libre de Bruxelles [1]. The task is to flag fraudulent transactions. For confidentiality, all features except the transaction amount are principal components of the original variables.
Fraud is rare: the raw data holds 492 frauds among 284,807 transactions. Following Höppner et al. [2] and Vanderschueren et al. [3], the 1,825 transactions with a zero amount are removed, since missing them costs nothing. That leaves 282,982 transactions, 465 of them fraudulent.
Classes |
2 |
Frauds |
465 |
Legitimate |
282517 |
Samples |
282982 |
Features |
29 |
6.3.11.2. Using the Dataset#
The dataset is fetched through fetch_credit_card_fraud. It is downloaded
from OpenML on first use (about 150 MB) and cached under ~/empulse_data (override with
$EMPULSE_DATA_HOME or the data_home argument), so later calls work offline.
It returns a Dataset object with the following attributes:
data: the feature matrixtarget: the target vectorcost_matrix: aCostMatrixwith default values pre-filledinstance_costs: a dict of per-instance cost drivers ('amount')feature_names: the feature namestarget_names: the target namesDESCR: the full description of the dataset
import pandas as pd
from empulse.datasets import fetch_credit_card_fraud
dataset = fetch_credit_card_fraud(backend=pd)
X, y = dataset.data, dataset.target
The backend argument selects the dataframe library used for data and target.
Pass the module itself — backend=pd for pandas or backend=pl for polars.
All features are numeric. Pass the cost matrix to the model as a Metric
loss, and hand it the transaction amounts at fit time:
from empulse.metrics import Metric, Cost
from empulse.models import CSLogitClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
pipeline = Pipeline([
('scaler', StandardScaler()),
('model', CSLogitClassifier(loss=Metric(dataset.cost_matrix, Cost()))),
])
pipeline.fit(X, y, model__amount=dataset.instance_costs['amount'])
6.3.11.3. Cost Matrix#
Flagging a transaction triggers an investigation with a fixed administrative cost \(c_f\), whether or not it turns out to be fraud. A fraud that is not flagged costs its full amount \(A_i\) [2] [3].
Actual fraud \(y_i = 1\) |
Actual legitimate \(y_i = 0\) |
|
Predicted fraud \(\hat{y}_i = 1\) |
|
|
Predicted legitimate \(\hat{y}_i = 0\) |
|
|
The investigation cost is a symbolic parameter with the default \(c_f = 10\) used in both
papers, exposed under the alias investigation_cost. Override it by passing the alias when
evaluating the metric:
y_score = pipeline.predict_proba(X)[:, 1]
cost = Metric(dataset.cost_matrix, Cost())
default_cost = cost(y, y_score, **dataset.instance_costs)
expensive_investigation = cost(y, y_score, investigation_cost=50, **dataset.instance_costs)
6.3.11.4. Data Description#
Feature |
Description |
|---|---|
|
Principal components of the original, confidential transaction variables |
|
Transaction amount in euros; also the false-negative cost |
fraud (target) |
Whether the transaction is fraudulent (1 = fraud, 0 = legitimate) |
The Time column of the original data, the seconds elapsed since the first transaction, is
dropped.