6. Validating the Whole Thing#
Everything so far used a single train/test split, which is fine for exposition and not enough to trust. This page cross-validates the pipeline and tunes it — with one complication worth understanding.
6.1. The complication: costs must follow the folds#
clv has one value per customer. When cross-validation splits the data, each fold needs its
own slice of those values — both to fit the model and to score it. Passing the full array would
silently misalign it with the fold’s rows.
scikit-learn solves this with metadata routing: estimators and scorers declare which extra parameters they want, and scikit-learn slices and delivers them per fold. It is opt-in.
from sklearn import set_config
set_config(enable_metadata_routing=True)
6.2. Cross-validating on savings#
Declare the requests with set_fit_request on the model and set_score_request on the scorer,
then hand the full array to cross_val_score via params.
import numpy as np
import pandas as pd
from empulse.datasets import fetch_iranian_churn
from empulse.metrics import Cost, Metric, Savings
from empulse.models import CSBoostClassifier
from sklearn.metrics import make_scorer
from sklearn.model_selection import cross_val_score
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
dataset = fetch_iranian_churn(backend=pd)
X, y = dataset.data, dataset.target
clv = dataset.instance_costs['clv']
expected_cost = Metric(dataset.cost_matrix, Cost())
savings = Metric(dataset.cost_matrix, Savings())
scorer = make_scorer(
savings, response_method='predict_proba', greater_is_better=True
).set_score_request(clv=True)
pipeline = Pipeline([
('scaler', StandardScaler()),
('model', CSBoostClassifier(loss=expected_cost).set_fit_request(clv=True)),
])
scores = cross_val_score(pipeline, X, y, cv=5, scoring=scorer, params={'clv': clv})
print(f'savings per fold: {np.round(scores, 3)}')
print(f'mean savings: {scores.mean():.4f} (+/- {scores.std():.4f})')
Mean savings of about 1.12, tight across folds — the single-split result holds up.
Note that clv is passed once, at the top level, as params={'clv': clv}. Routing takes it
from there: the pipeline forwards it to the model’s fit and the scorer’s __call__, each
time sliced to the right rows.
6.3. Tuning hyperparameters on business value#
The same scorer drives GridSearchCV, so model selection optimises
profit rather than accuracy.
from sklearn.model_selection import GridSearchCV
from xgboost import XGBClassifier
tuned_pipeline = Pipeline([
('scaler', StandardScaler()),
('model', CSBoostClassifier(
estimator=XGBClassifier(), loss=expected_cost
).set_fit_request(clv=True)),
])
search = GridSearchCV(
tuned_pipeline,
param_grid={'model__estimator__max_depth': [2, 4]},
cv=3,
scoring=scorer,
)
search.fit(X, y, clv=clv)
print(f'best params: {search.best_params_}')
print(f'best savings: {search.best_score_:.4f}')
Note
CSBoostClassifier is given an explicit
estimator=XGBClassifier() here. Its default is None, and
GridSearchCV cannot call set_params on None when
addressing nested parameters like model__estimator__max_depth. Pass the backend explicitly
whenever you tune its hyperparameters.
6.4. Tuning the business parameters too#
Because the cost matrix keeps its parameters symbolic, they are ordinary metric arguments — so you can ask what the campaign should look like, not just what the model should look like:
for accept_rate in (0.2, 0.3, 0.5):
score = savings(y, pipeline.fit(X, y, clv=clv).predict_proba(X)[:, 1],
clv=clv, accept_rate=accept_rate)
print(f'accept_rate={accept_rate}: savings {score:.4f}')
Treat these as sensitivity analysis rather than something to maximise: accept_rate is a fact
about the world, not a knob you control. What the sweep tells you is how much your conclusion
depends on an estimate you are not certain of.
6.5. Turning routing back off#
Metadata routing is a global setting. If other code in your session does not expect it, restore the default when finished:
set_config(enable_metadata_routing=False)
6.6. What we built#
Starting from a logistic regression with a 0.926 AUC that lost 7.20 per customer, we:
Wrote the campaign’s economics as a cost matrix.
Measured the model in money and found it unprofitable.
Trained
CSBoostClassifieron that cost matrix, reaching a profit of 2.78 per customer.Chose the operating point — contact the top 26% — from the profit curve instead of defaulting to 0.5.
Cross-validated the whole pipeline on savings, with per-customer costs routed through each fold.
The same features, the same algorithms; only the objective changed.
6.7. Where next#
Cross-Validation with Instance-dependent Costs — metadata routing in more depth.
Threshold Tuning — including
TunedThresholdClassifierCVwith Empulse metrics.Robust Cost-Sensitive Classification (RobustCS) — when your cost estimates contain outliers.
Define your own cost-sensitive or value metric — writing a cost matrix for your own domain.
User Guide — the full user guide.