Metric#
- class empulse.metrics.Metric(cost_matrix, strategy)[source]#
Class to create a custom value/cost-sensitive metric.
The metric is defined by a cost matrix and a strategy for computing the metric. The cost matrix defines the costs and benefits associated with each type of prediction outcome (true positive, true negative, false positive, false negative). The strategy defines how to compute the metric based on the cost matrix.
Read more in the User Guide.
- Parameters:
- cost_matrixCostMatrix
The cost matrix defining the costs and benefits associated with each type of prediction outcome.
- strategyMetricStrategy
The strategy to use for computing the metric. Several strategies come as a sign-flipped pair – one phrased as a profit to maximize, its sibling as a cost to minimize – that hand an estimator identical values and differ only in what they report.
Cost/Profitcompute the expected cost (or its negation, the expected profit) of a classifier. They support instance-dependent costs passed as array-likes. Any stochastic variable is reduced to its mean before use.MaxProfit/MinCostcompute the profit (or cost) at the profit-maximizing threshold, which the metric locates itself. They support stochastic variables, but reduce per-instance costs to their class means.EmpiricalMaxProfit/EmpiricalMinCostdo the same from the empirical score distribution rather than a parametric one, and keep per-instance costs.Savingscomputes the cost relative to a baseline classifier, scaled to1for a perfect model and0for the baseline. It has no sibling.AUEPCis the area under the empirical profit curve.
- Attributes:
- tp_benefitsympy.Expr
The benefit of a true positive. See
add_tp_benefitfor more details.- tn_benefitsympy.Expr
The benefit of a true negative. See
add_tn_benefitfor more details.- fp_benefitsympy.Expr
The benefit of a false positive. See
add_fp_benefitfor more details.- fn_benefitsympy.Expr
The benefit of a false negative. See
add_fn_benefitfor more details.- tp_costsympy.Expr
The cost of a true positive. See
add_tp_costfor more details.- tn_costsympy.Expr
The cost of a true negative. See
add_tn_costfor more details.- fp_costsympy.Expr
The cost of a false positive. See
add_fp_costfor more details.- fn_costsympy.Expr
The cost of a false negative. See
add_fn_costfor more details.directionDirectionWhether the metric is to be maximized or minimized.
Examples
Reimplementing
empc_scoreusing theMetricclass.import sympy as sp from empulse.metrics import Metric, MaxProfit, CostMatrix clv, d, f, alpha, beta = sp.symbols( 'clv d f alpha beta' ) # define deterministic variables gamma = sp.stats.Beta('gamma', alpha, beta) # define gamma to follow a Beta distribution cost_matrix = ( CostMatrix() .add_tp_benefit(gamma * (clv - d - f)) # when churner accepts offer .add_tp_benefit((1 - gamma) * -f) # when churner does not accept offer .add_fp_cost(d + f) # when you send an offer to a non-churner .alias({'incentive_cost': 'd', 'contact_cost': 'f'}) ) empc_score = Metric(cost_matrix, MaxProfit()) y_true = [1, 0, 1, 0, 1] y_proba = [0.9, 0.1, 0.8, 0.2, 0.7] empc_score(y_true, y_proba, clv=100, incentive_cost=10, contact_cost=1, alpha=6, beta=14)
Reimplementing
expected_cost_loss_churnusing theMetricclass.import sympy as sp from empulse.metrics import Metric, Cost, CostMatrix clv, delta, f, gamma = sp.symbols('clv delta f gamma') cost_matrix = ( CostMatrix() .add_tp_benefit(gamma * (clv - delta * clv - f)) # when churner accepts offer .add_tp_benefit((1 - gamma) * -f) # when churner does not accept offer .add_fp_cost(delta * clv + f) # when you send an offer to a non-churner .alias({'incentive_fraction': 'delta', 'contact_cost': 'f', 'accept_rate': 'gamma'}) ) cost_loss = Metric(cost_matrix, Cost()) y_true = [1, 0, 1, 0, 1] y_proba = [0.9, 0.1, 0.8, 0.2, 0.7] cost_loss( y_true, y_proba, clv=100, incentive_fraction=0.05, contact_cost=1, accept_rate=0.3 )
- __call__(y_true, y_score, *, validate=True, **parameters)[source]#
Compute the metric score or loss.
- Parameters:
- y_truearray-like of shape (n_samples,)
The ground truth labels.
- y_scorearray-like of shape (n_samples,)
The predicted labels, probabilities, or decision scores (based on the chosen metric).
For the ranking strategies (
MaxProfit,MinCost,EmpiricalMaxProfit,EmpiricalMinCost,AUEPC),y_scoreis used only to rank the samples, so any decision score works.For the probability strategies (
Cost,Profit,Savings,LogCost),y_scoremust be a calibrated probability.
- validatebool, default=True
Whether to check the labels, and the parameter values against the cost matrix’s domain. Pass
Falseonly when re-entering the metric on a training loop’s per-iteration path, where the labels and values have already been validated once at fit time. The scores are then only checked to be finite.- **parametersfloat or array-like of shape (n_samples,)
The parameter values for the costs and benefits defined in the metric. If any parameter is a stochastic variable, you should pass values for their distribution parameters. You can set the parameter values for either the symbol names or their aliases.
If
float, the same value is used for all samples (class-dependent).If
array-like, the values are used for each sample (instance-dependent).
- Returns:
- scorefloat
The computed metric score or loss.
- property capabilities#
The set of
Capabilitymembers this metric supports.Forwards to
strategy’s owncapabilities.MixtureMetricoverrides this to the intersection of its components’ capabilities, since a composite metric can only do what every component can.Use this instead of
isinstance(metric.strategy, SomeConcreteStrategy)to check whether a metric supports what a model needs, e.g.Capability.CLASS_COSTS in loss.capabilities.
- property direction#
Whether the metric is to be maximized or minimized.
- optimal_rate(y_true, y_score, *, validate=True, **parameters)[source]#
Compute the optimal predicted positive rate.
i.e., the fraction of observations that should be classified as positive to optimize the metric.
- Parameters:
- y_truearray-like of shape (n_samples,)
The ground truth labels.
- y_scorearray-like of shape (n_samples,)
The predicted labels, probabilities, or decision scores (based on the chosen metric).
For the ranking strategies (
MaxProfit,MinCost,EmpiricalMaxProfit,EmpiricalMinCost,AUEPC),y_scoreis used only to rank the samples, so any decision score works.For the probability strategies (
Cost,Profit,Savings,LogCost),y_scoremust be a calibrated probability.
- validatebool, default=True
Whether to check the labels, and the parameter values against the cost matrix’s domain. Pass
Falseonly when re-entering the metric on a training loop’s per-iteration path, where the labels and values have already been validated once at fit time. The scores are then only checked to be finite.- **parametersfloat or array-like of shape (n_samples,)
The parameter values for the costs and benefits defined in the metric. If any parameter is a stochastic variable, you should pass values for their distribution parameters. You can set the parameter values for either the symbol names or their aliases.
If
float, the same value is used for all samples (class-dependent).If
array-like, the values are used for each sample (instance-dependent).
- Returns:
- optimal_ratefloat
The optimal predicted positive rate.
- optimal_threshold(y_true, y_score, *, validate=True, **parameters)[source]#
Compute the optimal classification threshold(s).
i.e., the score threshold at which an observation should be classified as positive to optimize the metric. For instance-dependent costs and benefits, this will return an array of thresholds, one for each sample. For class-dependent costs and benefits, this will return a single threshold value.
- Parameters:
- y_truearray-like of shape (n_samples,)
The ground truth labels.
- y_scorearray-like of shape (n_samples,)
The predicted labels, probabilities, or decision scores (based on the chosen metric).
For the ranking strategies (
MaxProfit,MinCost,EmpiricalMaxProfit,EmpiricalMinCost,AUEPC),y_scoreis used only to rank the samples, so any decision score works.For the probability strategies (
Cost,Profit,Savings,LogCost),y_scoremust be a calibrated probability.
- validatebool, default=True
Whether to check the labels, and the parameter values against the cost matrix’s domain. Pass
Falseonly when re-entering the metric on a training loop’s per-iteration path, where the labels and values have already been validated once at fit time. The scores are then only checked to be finite.- **parametersfloat or array-like of shape (n_samples,)
The parameter values for the costs and benefits defined in the metric. If any parameter is a stochastic variable, you should pass values for their distribution parameters. You can set the parameter values for either the symbol names or their aliases.
If
float, the same value is used for all samples (class-dependent).If
array-like, the values are used for each sample (instance-dependent).
- Returns:
- optimal_thresholdfloat or NDArray of shape (n_samples,)
The optimal classification threshold(s).
- property parameter_names#
The set of all parameter names this metric accepts, including stochastic variables.
A public alias for
_all_symbols, for a caller that wants to know what to pass to this metric without reaching into a private name.
- score(y_true, y_score, *, validate=True, **parameters)#
Compute the metric score or loss (see
__call__).
- property strategy#
The strategy used to compute the metric.