actual_vs_expected¶
Calculate actual versus expected totals by rating-factor level, for every feature.
This shows where observed and predicted totals differ. For each feature, the objects are aggregated over the cells of the grid the deployed tables use (cell 0 holds missing values).
Note
The model prediction results will be correct only if the X parameter with feature values
contains all the features used in the model. For a polars DataFrame or LazyFrame, the
features are matched by name: extra columns are ignored and the order of the columns does
not matter (a LazyFrame collects only the columns the model needs). For other types, and
for a model trained without feature names, the features must be in the same number and
order as the columns provided during the training.
Method call format¶
Parameters¶
X¶
Description¶
Feature values data.
Possible types
- polars.DataFrame
- polars.LazyFrame
- numpy.ndarray of shape
(object_count, feature_count) - other array-like data of the same shape
Default value
Required parameter
y¶
Description¶
The target values of the objects (for a classifier, the class labels). A string names a column
of a polars X.
Possible types
- numpy.ndarray of shape
(object_count,) - polars.Series
- list
- string
Default value
Required parameter
sample_weight, exposure¶
Description¶
The weight and the exposure of each object. A string names a column of a polars X.
If the model was trained with a sample weight or an exposure, the corresponding value must be passed explicitly, even on the training data: the training values are never reused, because a matching number of objects does not prove the objects are the same. Pass a vector of ones for an unweighted or unit-exposure evaluation.
Possible types
- numpy.ndarray of shape
(object_count,) - polars.Series
- list
- string
Default value
None
Return value¶
A list with one dictionary per feature, holding one value per cell:
feature,raw— The feature name and its zero-based index.actual— \(\sum w \cdot t\).expected— \(\sum w \cdot prediction\), where the prediction includes the exposure for a model with a log link.mass— \(\sum w \cdot exposure\).rows— The number of objects.ae—actual / expected(nanwhereexpectedis 0).
Exact balance by factor level is not a general property of boosted models, so the ratios show where the model and the data disagree. For a classifier, only binary classification is supported.