metrics¶
The t_boost.metrics module calculates metrics separately from the training. It needs only
NumPy: no scikit-learn.
All functions take the targets y, the predictions and optional weights as one-dimensional
array-like data.
Deviances¶
The deviances are the standard GLM unit deviances, numerically identical to scikit-learn's
functions of the same names, and the natural goodness-of-fit for the matching
objective. Lower is better. A value outside the domain of the
distribution (for example, a negative prediction) raises ValueError.
For a model trained with an exposure, pass the expected totals, predict(X) * exposure, as the
predictions.
mean_tweedie_deviance¶
The weighted mean Tweedie unit deviance with variance power power: 0 is the squared error, 1
the Poisson deviance, 2 the Gamma deviance, and a value between 1 and 2 the compound
Poisson-Gamma deviance. Powers between 0 and 1 are not supported.
Return value: float
mean_poisson_deviance¶
mean_tweedie_deviance with power=1: the goodness-of-fit of objective="poisson" frequency
models.
Return value: float
mean_gamma_deviance¶
mean_tweedie_deviance with power=2: the goodness-of-fit of objective="gamma" severity
models.
Return value: float
Ranking metrics¶
The ranking metrics measure how well the predictions order the objects. They are weight-aware, and degenerate or non-finite inputs give 0 rather than an error.
ordered_gini¶
The concentration Gini of y when the objects are ranked by pred, normalized by the Gini of
the perfect ranking (by y itself). 1 is a perfect ranking.
Return value: float
concentration_gini¶
The concentration (Lorenz) Gini of y when the objects are ranked by score, descending: the
weighted cumulative share of y against the weighted cumulative share of objects, as
\(2 \cdot area - 1\). y and the weights are clamped at 0.
Return value: float
lift_curve¶
The objects are ranked by pred, descending, and split into buckets groups with equal
numbers of objects.
Return value: a list with one dictionary per group: bucket (starting at 1), rows,
mean_y and mean_pred (weighted means) and lift (mean_y divided by the overall weighted
mean of y). An empty list for degenerate input.
top_bucket_lift¶
The lift of the first group of lift_curve (the highest predictions). 0 if the curve is empty.
Return value: float