6.3.12. IEEE-CIS Fraud Detection#
6.3.12.1. Summary#
Real-world e-commerce transactions from the payment service provider Vesta, released for the IEEE Computational Intelligence Society’s fraud detection competition on Kaggle [1]. Each row is an online transaction, described by its amount, the card, addresses, e-mail domains, many engineered features and, for part of the transactions, the device and network used. The task is to flag fraudulent transactions.
The data is the competition’s training set, the transaction table joined to the identity table. As in Vanderschueren et al. [2], the transaction identifier and the timestamp offset are dropped.
Classes |
2 |
Frauds |
20663 |
Legitimate |
569877 |
Samples |
590540 |
Features |
431 |
6.3.12.2. Using the Dataset#
The dataset is fetched through fetch_ieee_fraud_detection. It is
downloaded from OpenML on first use, which is a large download, and cached under
~/empulse_data (override with $EMPULSE_DATA_HOME or the data_home argument), so later
calls work offline.
It returns a Dataset object with the following attributes:
data: the feature matrixtarget: the target vectorcost_matrix: aCostMatrixwith default values pre-filledinstance_costs: a dict of per-instance cost drivers ('amount')feature_names: the feature namestarget_names: the target namesDESCR: the full description of the dataset
import pandas as pd
from empulse.datasets import fetch_ieee_fraud_detection
dataset = fetch_ieee_fraud_detection(backend=pd)
X, y = dataset.data, dataset.target
The backend argument selects the dataframe library used for data and target.
Pass the module itself — backend=pd for pandas or backend=pl for polars.
Most features have missing values, and fitting a model on all 431 of them and 590,540 rows is
slow. The example below imputes the numeric features of a sample of the transactions and trains
CSBoostClassifier on them:
from empulse.metrics import Metric, Cost
from empulse.models import CSBoostClassifier
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline
sample = X.sample(n=50_000, random_state=0).index
X_sample = X.loc[sample].select_dtypes(include=['number'])
y_sample = y.loc[sample]
amount_sample = dataset.instance_costs['amount'][sample]
model = Pipeline([
('imputer', SimpleImputer(strategy='median', keep_empty_features=True)),
('model', CSBoostClassifier(loss=Metric(dataset.cost_matrix, Cost()))),
])
model.fit(X_sample, y_sample, model__amount=amount_sample)
6.3.12.3. Cost Matrix#
Flagging a transaction triggers an investigation with a fixed administrative cost \(c_f\), whether or not it turns out to be fraud. A fraud that is not flagged costs its full amount \(A_i\). Vanderschueren et al. [2] use this cost matrix, from Höppner et al. [3], for this dataset.
Actual fraud \(y_i = 1\) |
Actual legitimate \(y_i = 0\) |
|
Predicted fraud \(\hat{y}_i = 1\) |
|
|
Predicted legitimate \(\hat{y}_i = 0\) |
|
|
The investigation cost is a symbolic parameter with the default \(c_f = 10\) of both papers,
exposed under the alias investigation_cost. Override it by passing the alias when evaluating
the metric:
y_score = model.predict_proba(X_sample)[:, 1]
cost = Metric(dataset.cost_matrix, Cost())
default_cost = cost(y_sample, y_score, amount=amount_sample)
expensive_investigation = cost(y_sample, y_score, investigation_cost=50, amount=amount_sample)
6.3.12.4. Data Description#
Most features are masked by Vesta. The groups below follow the description the competition host published [1].
Features |
Description |
Type |
|---|---|---|
|
Transaction amount in US dollars; also the false-negative cost |
numeric |
|
Product code of the transaction |
categorical |
|
Payment card information, such as card type, category, issuing bank and country |
categorical |
|
Billing address region and country |
categorical |
|
Distances, for example between addresses |
numeric |
|
E-mail domains of the purchaser and the recipient |
categorical |
|
Counts, such as how many addresses are associated with the card |
numeric |
|
Time deltas, such as days since the previous transaction |
numeric |
|
Matches, such as between the names on the card and the address |
categorical |
|
Features engineered by Vesta, including rankings, counts and entity relations |
numeric |
|
Identity information: network connection and digital signature |
numeric |
|
Identity information: network connection, browser, operating system and device |
categorical |
fraud (target) |
Whether the transaction is fraudulent (1 = fraud, 0 = legitimate) |
binary |
Identity information is only available for about a quarter of the transactions; for the others, those features are missing.