Changelog#

Unreleased#

  • Feature Added EmpiricalMaxProfit and AUEPC strategies for building custom metrics that compute the empirical (convex-hull-based) maximum profit and the area under the empirical profit curve, respectively.

  • API Change empc, mpc, empa, mpa, empcs, mpcs, empb, make_objective_churn, make_objective_acquisition, and their supporting AECObjectiveChurn, AECMetricChurn, AECObjectiveAcquisition, and AECMetricAcquisition classes have been removed. empc_score, mpc_score, empa_score, mpa_score, empcs_score, mpcs_score, empb_score, and auepc_score are now prebuilt Metric/ MixtureMetric instances instead of hand-written functions. The decision threshold (previously returned alongside the score by the removed empc/mpc/empa/mpa/empcs/mpcs/empb functions) can still be obtained with the .optimal_rate(...) method on the corresponding *_score instance. Training a boosting model directly on any of these metrics (previously done through make_objective_churn/make_objective_acquisition) is now done by passing the metric as the loss argument to CSBoostClassifier.

  • API Change expected_cost_loss_churn and expected_cost_loss_acquisition now always return the mean cost per instance (previously they returned the summed cost by default, with an optional normalize=True argument to switch to the mean).

  • API Change empa_score’s beta parameter now represents the scale of the Gamma-distributed contribution (mean = alpha * beta), instead of the rate (mean = alpha / beta) used by the previous empa/empa_score functions. The default value has been adjusted accordingly, so calls that rely on the default are unaffected; only explicit non-default beta= overrides behave differently.

  • API Change The generic (domain-agnostic) native functions max_profit, make_objective_aec, and the supporting AECObjective and AECMetric classes have been removed. max_profit_score, expected_cost_loss, expected_log_cost_loss, and expected_savings_score are now prebuilt Metric instances instead of hand-written functions (cost_loss and savings_score, the hard-label variants that auto-threshold continuous scores, are unaffected and remain plain functions, since there is no Metric/MetricStrategy equivalent for that behavior). The decision threshold (previously returned alongside the score by the removed max_profit function) can still be obtained with the .optimal_rate(...) method on max_profit_score. Training a boosting or logistic-regression model directly on any of these metrics (previously done through make_objective_aec) is now done by passing the metric as the loss argument to a cost-sensitive model, e.g. CSBoostClassifier or CSLogitClassifier.

  • API Change max_profit_score now takes tp_cost/tn_cost parameters (costs) instead of the previous native max_profit/max_profit_score functions’ tp_benefit/tn_benefit parameters (benefits): tp_cost = -tp_benefit and tn_cost = -tn_benefit. fp_cost/fn_cost are unchanged. This matches the parameter naming already used throughout the rest of the package (e.g. the cost-sensitive models’ default loss).

  • API Change expected_cost_loss and expected_log_cost_loss now always return the mean cost per instance (previously they returned the summed cost by default, with an optional normalize=True argument to switch to the mean). This also affects the default (unweighted-loss) out-of-bag weighted-voting behavior of CSForestClassifier and CSBaggingClassifier, which use expected_cost_loss as their fallback per-estimator weight.

  • Fix Cost-sensitive models can now be trained with a loss metric whose cost matrix names one of its symbols (or aliases) tp_cost, tn_cost, fp_cost or fn_cost. Those names collide with the dedicated fit/predict parameters of the same name, so the value bound to the parameter and was silently discarded instead of reaching the metric, and training failed with TypeError: _lambdifygenerated() missing 1 required positional argument. This affected the cost matrices shipped with load_churn_tv_subscriptions, load_credit_scoring_pakdd and fetch_give_me_some_credit. Such values are now routed to the metric, and a cost argument that the metric does not use raises a warning instead of being dropped silently. Note that __init__-time costs are still not forwarded to a metric loss, since they default to 0.0 and would silently zero out a cost matrix term of the same name.

  • Fix Fix RobustCSClassifier not properly handling outlier sensitive costs when passing a custom loss function from Metric.

  • Fix Fix empb_score and auepc_score not always incurring the contact cost for customers who are not contacted.

  • Fix Fix expected_savings_score’s (and the Savings strategy’s) baseline='prior' option: it previously computed the baseline cost by hard-thresholding the constant prior probability (via the same logic as cost_loss), which usually made it silently degenerate to the same result as baseline='zero_one'/'one'. It now correctly evaluates the expected cost of predicting the prior probability of the majority or minority class, whichever is cheaper, as documented.

  • Fix Metric no longer shares mutable strategy state when the same MetricStrategy instance (e.g. a configured MaxProfit) is passed to more than one Metric. Previously, building a second Metric from an already-used strategy instance silently rebuilt that strategy in place, so the first Metric would start returning values computed from the second metric’s cost matrix instead of its own. Each Metric now builds and owns an independent copy of the strategy it is given.

  • Fix Metric no longer stays linked to the mutable CostMatrix it was constructed from. Previously, modifying the CostMatrix object after building a Metric from it silently desynchronized the metric’s advertised cost expressions (e.g. metric.fn_cost) from what it actually computed, so a parameter the metric claimed to depend on could be silently ignored when scoring. Metric now takes an independent copy of its cost matrix at construction time.

  • Fix Metric now raises a ValueError at construction time if the cost matrix uses a symbol name or alias reserved for internal use (y, s, F_0, F_1, pi_0, pi_1, N, i). Previously, a colliding symbol name was either silently fused with the identically-named internal variable (e.g. a user symbol named F_0 in a MaxProfit metric), producing a wrong score with no warning, or raised a confusing internal TypeError only once the metric was called (e.g. a symbol named y or s).

  • Fix Metric now raises a ValueError when called with both a symbol’s own name and one of its aliases (or with two different aliases for the same symbol). Previously, whichever keyword argument was seen last silently won, so the computed score depended on the order the keyword arguments happened to be passed in.

  • Fix Metric now raises a ValueError at construction time if an alias’s target, or a set_default parameter name, does not match any symbol used in the cost matrix. This also catches the ordering footgun documented on set_default, where a default keyed by an alias name that was set before the alias was registered used to be silently stored under that raw (untranslated) name and then silently ignored, instead of ever being applied.

  • Fix Metric now validates y_true/y_score.

  • Fix Metric now raises a ValueError naming the offending parameter when an instance-dependent (array-like) cost_matrix parameter’s length doesn’t match y_true/y_score and isn’t a single value broadcastable to every sample. Previously a length mismatch surfaced as an unrelated numpy broadcasting error, e.g. operands could not be broadcast together with shapes (3,) (5,), that didn’t say which parameter was wrong.

  • Fix Metric.optimal_threshold() and optimal_rate now raise a ValueError when the cost matrix is degenerate for the given parameters (fp_cost + tn_benefit + fn_cost + tp_benefit evaluates to 0, making the optimal threshold undefined).

  • Fix CSThresholdClassifier and CSRateClassifier no longer silently skip learning a cost-sensitive threshold/rate when fitted with an aliased Metric with every required parameter supplied through its alias.

  • Fix CSBaggingClassifier with combination='weighted_voting' no longer raises a shape error when a sub-estimator draws a feature subset (max_features < 1.0 or bootstrap_features=True).

  • Fix CSForestClassifier with combination='weighted_voting' no longer produces NaN out-of-bag estimator weights when max_samples is an int. numbers.Integral is a subclass of numbers.Real.

  • Fix Out-of-bag weighted voting (CSForestClassifier and CSBaggingClassifier with combination='weighted_voting') no longer gives more weight to worse estimators, and no longer raises when instance-dependent (array-like) costs are used together with weighted voting.

  • Fix CSTreeClassifier and CSForestClassifier no longer raises when fitted with a metric strategy containing a stochastic (sympy.stats) variable.

  • Fix CSBoostClassifier no longer mutates a caller-supplied fit_params dict in place.

  • Fix CSTreeClassifier no longer mutates a user-supplied custom criterion instance in place.

  • Fix CSBoostClassifier’s LightGBM and CatBoost backends now start from the intended probability. The internal base-score nudge was applied as a raw (log-odds) score for LightGBM’s init_score/CatBoost’s baseline but as a probability for XGBoost’s base_score, and neither LightGBM nor CatBoost persist that offset into the saved model, so predict_proba now adds it back manually before converting to a probability.

  • Fix BiasRelabelingClassifier, BiasResamplingClassifier, and BiasReweighingClassifier now raise a clear ValueError when sensitive_feature does not have the same length as y, instead of silently proceeding.

  • Fix CSThresholdClassifier and CSRateClassifier no longer mutate a user-supplied calibrator estimator instance in place; a clone is configured and fitted instead, so the constructor argument is safe to reuse or refit.

  • Fix CSLogitClassifier and ProfLogitClassifier documented soft_threshold as defaulting to False; the actual default is True.

0.11.1 (08-05-2026)#

  • Fix Fix empulse not running properly on scikit-learn versions lower than 1.6.

0.11.0 (08-05-2026)#

  • Feature Added ProfTreeClassifier to optimize a cost-sensitive metric using evolutionary trees (similar to ProfLogitClassifier).

  • Feature Added CSRateClassifier to optimize for a specific predicted positive rate.

  • Feature Allow CSThresholdClassifier to optimize the decision threshold at training time. Users can now either specify the decision threshold to use at prediction time or let the model choose the optimal threshold during training.

  • Feature (Experimental) When using the MaxProfit strategy, only one stochastic variable is present, and the cost/profit function is a polynomials of the stochastic variable: the metric will now be computed exactly instead of using numerical integration. This is currently supported for stochastic variables following the Normal, Log Normal, Uniform, Beta, Gamma, Chi Squared, Exponential, Weibull, Pareto, and Triangular distributions.

  • Feature (Experimental) Added support for the MaxProfit strategy metrics to be optimized through gradient descent methods in CSLogitClassifier and CSBoostClassifier. This is currently an experimental feature and is not recommended for use in production.

  • API Change Updated the ProfLogitClassifier interface to be more consistent with other models in the package. By default optimizes the maximum profit metric.

  • API Change CSLogitClassifier no longer takes a string argument for the loss function to be more consistent with other models in the package. Default value for the loss is None.

  • API Change MaxProfit now takes a numpy Generator instead of a RandomState instance.

  • Enhancement Metrics built with the MaxProfit strategy can now handle instance-dependent costs. They will automatically be averaged over the instances. Mathematically this is equivalent to recomputing the EMP score for each instance and then averaging the scores.

  • Enhancement Metrics built with the Cost and Savings strategies can now handle stochastic cost parameters. If the distribution allows it, the mean cost will be computed.

  • Fix Fix CSTreeClassifier and CSForestClassifier not properly training when costs were negative.

  • Fix Fix integration bounds inconsistently being calculated when the MaxProfit strategy was chosen.

  • Fix Fix CSBoostClassifier throwing errors when one or two of the Boosting libraries were not installed (XGBoost, LGBM & Catboost).

  • Fix Add __name__ attribute to Metric class to fix issues with scikit-learn compatibility.

  • Fix Fix metadata routing not working for scikit-learn>=1.8.0

  • Fix Fix MaxProfit strategy not calculating Log Normal distributed variables correctly when using quasi monte carlo.

  • Fix Fix some models not properly being able to be pickled when using a custom metric as the loss function.

  • Fix Fix some distributions not correctly computing the expected maximum profit score when using the MaxProfit strategy when using monte carlo or quasi monte carlo method.

0.10.4 (20-09-2025)#

0.9.0 (15-06-2025)#

  • Feature Added optimal_threshold and optimal_rate methods to calculate the optimal threshold(s) and optimal predicted positive rate for a given metric. This is useful for determining the best decision threshold and predicted positive rate for a cost-sensitive or value-driven model.

  • Feature CSTreeClassifier, CSForestClassifier, and CSBaggingClassifier can now take a Metric instance as their criterion to optimize.

  • Feature CSThresholdClassifier can now take a Metric instance to choose the optimal decision threshold.

  • Feature RobustCSClassifier can now take estimators with a Metric instance as the loss function or criterion. RobustCSClassifier will treat any cost marked as outlier sensitive. This can be done by using the mark_outlier_sensitive method.

  • Feature Allow savings metrics to be used in CSBoostClassifier and CSLogitClassifier as the objective function. Internally, the expected cost loss is used to train the model, since the expected savings score is just a transformation of the expected cost loss.

  • API Change kind argument to Metric has been replaced by strategy. The Metric class now takes a MetricStrategy instance. This change allows for more flexibility in defining the metric strategy. The currently available strategies are:

    • MaxProfit for the expected maximum profit score

    • Cost for the expected cost loss

    • Savings for the expected savings score

  • Fix Fix error when importing Empulse without any optional dependencies installed.

  • Fix Fix CSLogitClassifier not properly using the gradient when using a custom loss function from Metric.

  • Fix Fix models throwing errors when differently shaped costs are passed to the fit or predict method.

  • Fix Fix sympy distribution parameters not being properly translated to scipy distribution parameters when using the MaxProfit strategy (formerly kind=’max profit’) with the quasi monte-carlo integration method.

0.8.0 (01-06-2025)#

  • Feature CSBoostClassifier, CSLogitClassifier, and ProfLogitClassifier can now take a Metric instance as their loss function. Internally, the metric instance is converted to the appropriate loss function for the model. For more information, read the User Guide.

  • Feature Type hints are now available for all functions and classes.

  • Enhancement Add support for more than one stochastic variable when building maximum profit metrics with Metric

  • Enhancement Allow Metric to be used as a context manager. This ensures the metric is always built after defining the cost-benefit elements.

  • Fix Fix datasets not properly being packaged together with the package

  • Fix Fix RobustCSClassifier when array-like parameters are passed to fit method.

  • Fix Fix boosting models being biased towards the positive class.

0.7.0 (05-02-2025)#

  • Major Feature Add CSTreeClassifier, CSForestClassifier, and CSBaggingClassifier to support cost-sensitive decision tree and ensemble models

  • Enhancement Add support for scikit-learn 1.5.2 (previously Empulse only supported scikit-learn 1.6.0 and above).

  • API Change Removed the emp_score and emp functions from the metrics module. Use the Metric class instead to define custom expected maximum profit measures. For more information, read the User Guide.

  • API Change Removed numba as a dependency for Empulse. This will reduce the installation time and the size of the package.

  • Fix Fix Metric when defining stochastic variable with fixed values.

  • Fix Fix Metric when stochastic variable has infinite bounds.

  • Fix Fix CSThresholdClassifier when costs of predicting positive and negative classes are equal.

  • Fix Fix documentation linking issues to sklearn

0.6.0 (28-01-2025)#

  • Major Feature Add Metric to easily build your own value-driven and cost-sensitive metrics

  • Feature Add support for LightGBM and Catboost models in CSBoostClassifier and B2BoostClassifier

  • API Change make_objective_churn and make_objective_acquisition now take a model argument to calculate the objective for either XGBoost, LightGBM or Catboost models.

  • API Change XGBoost is now an optional dependency together with LightGBM and Catboost. To install the package with XGBoost, LightGBM and Catboost support, use the following command: pip install empulse[optional]

  • API Change Renamed y_pred_baseline and y_proba_baseline to baseline in savings_score and expected_savings_score. It now accepts the following arguments:

    • If 'zero_one', the baseline model is a naive model that predicts all zeros or all ones depending on which is better.

    • If 'prior', the baseline model is a model that predicts the prior probability of the majority or minority class depending on which is better (not available for savings score).

    • If array-like, target probabilities of the baseline model.

  • Feature Add parameter validation for all models and samplers

  • API Change Make all arguments of dataset loaders keyword-only

  • Fix Update the descriptions attached to each dataset to match information found in the user guide

  • Fix Improve type hints for functions and classes

0.5.2 (12-01-2025)#

  • Feature Allow savings_score and expected_savings_score to calculate the savings score over the baseline model instead of a naive model, by setting the y_pred_baseline and y_proba_baseline parameters, respectively.

  • Enhancement Reworked the user guide documentation to better explain the usage of value-driven and cost-sensitive models, samplers and metrics

  • API Change CSLogitClassifier and ProfLogitClassifier by default do not perform soft-thresholding on the regression coefficients. This can be enabled by setting the soft_threshold parameter to True.

  • Fix Prevent division by zero errors in expected_cost_loss

0.5.1 (05-01-2025)#

  • Fix Fixed documentation build issue

0.5.0 (05-01-2025)#