Overview¶
These parameters are for the Python package classes TBoostRegressor and TBoostClassifier.
Both classes accept the same parameters; only the default value of objective differs.
n_trees, learning_rate and lambda_ can be passed by position. Every other parameter is
keyword-only.
Common parameters¶
objective¶
The metric to use in training. The specified value also determines the machine learning problem to solve.
tweedie_rho¶
The power parameter of the Tweedie distribution: the variance is proportional to
\(\mu^{\rho}\). Used only when objective="tweedie".
n_trees¶
The maximum number of trees that can be built when solving machine learning problems.
learning_rate¶
The learning rate.
Used for reducing the gradient step: the values in the leaves of every tree are multiplied by it before the tree is added to the model.
learning_rate_decay¶
The decay of the learning rate over the iterations.
seed¶
The random seed used for training.
It seeds every source of randomness in the fit: bagging, object and feature sampling, the folds
of the categorical encoding, DART and random_strength. The same seed and
the same input data reproduce a bit-identical model.
lambda_¶
Coefficient at the L2 regularization term of the cost function.
l1_leaf¶
Coefficient at the L1 regularization term of the leaf values. The sum of the gradients of a leaf is soft-thresholded by this value before the leaf value is calculated.
max_depth¶
Depth of the trees.
Every tree is symmetric (oblivious): each level applies one split, shared by all the nodes of
the level. Possible values are integers from 3 to 8, and the value must be at least
max_interaction_order.
max_interaction_order¶
The maximum number of distinct features a tree may use. This is the highest interaction order in the model, and so the maximum number of axes of a rating table.
min_data_in_leaf¶
The minimum number of training samples in a leaf. A split that would leave fewer samples on either side is rejected.
min_sum_hessian_in_leaf¶
The minimum sum of the Hessians of the objects in a leaf. A split that would leave less on either side is rejected.
min_weight_sum_in_leaf¶
The minimum sum of the sample weights of the objects in a leaf. A split that would leave less on either side is rejected.
min_split_gain¶
The minimum score a split must exceed to be selected. A level whose best split does not exceed it is not split, and its nodes stay leaves.
max_delta_step¶
The maximum absolute value of a leaf's Newton step, applied before
learning_rate. It keeps leaf values finite on sparse or zero-heavy targets.
max_delta_step_gated¶
A guard against the predicted rate of a small group of objects collapsing toward zero during the training of a log-link model.
path_smooth¶
Path smoothing. The value of a leaf is shrunk toward the value of its parent node with the credibility weight \(Z = \frac{n}{n + path\_smooth}\), so leaves with little data stay close to the path above them. 0 turns it off.
colsample_bytree¶
The fraction of features randomly sampled for each tree.
The value must be in the range \((0; 1]\).
subsample¶
Sample rate of the objects used to choose the structure of each tree.
mvs_min_rows¶
The minimum number of objects in the sample when subsample turns on minimal
variance sampling.
random_strength¶
The amount of randomness to use for scoring splits when the tree structure is selected. Use this parameter to avoid overfitting the model.
monotone_constraints¶
Impose monotonic constraints on numerical features.
leaf_refine_steps¶
t-boost might calculate leaf values using several Newton steps instead of a single one.
leaf_refine_backtracks¶
The maximum number of times a leaf refinement step is halved when it does not reduce the loss
(see leaf_refine_steps). It has no effect on multiclassification.
reanchor¶
Re-solve the intercept of the model exactly on the training data after the training. This removes the overall bias that shrinkage leaves in the predictions.
reanchor_slope¶
Recalibrate the raw score \(a\) to \(b_0 + b_1 \cdot a\) on the validation objects of early stopping after the training. This corrects a uniform compression of the score scale that shrinkage and early stopping can leave, without changing the ranking of the objects or the decomposition into rating tables.
Early stopping settings¶
validation_fraction¶
The fraction of the training objects set aside as the validation dataset of early stopping. Each bag sets aside its own validation objects from its own sample.
early_stopping_rounds¶
Stops the training after the specified number of iterations since the iteration with the optimal metric value. The metric is the mean deviance of the objective on the validation objects.
early_stopping_adaptive¶
Makes the number of iterations to wait grow with the iteration of the best result.
early_stopping_min_delta¶
The minimum relative improvement of the metric for an iteration to become the new best. The validation deviance must fall below \(best \cdot (1 - early\_stopping\_min\_delta)\); smaller improvements do not reset the count of iterations to wait and do not move the iteration the model is truncated at.
early_stopping¶
A single setting for the patience of early stopping. An int sets
early_stopping_rounds, a float sets
early_stopping_adaptive.
Quantization settings¶
max_bin¶
The maximum number of bins for numerical features, not counting the bin for missing values. Allowed values are integers from 2 to 254 inclusively.
Interaction settings¶
interaction_gain_hurdle¶
The interaction hurdle. While a tree grows, a split on a feature the tree does not use yet raises the tree's interaction order. Such a split is selected only if its gain is large enough relative to the gain of the tree's first split, and if it beats the best split on a feature the tree already uses. Otherwise the tree keeps refining the features it already has. This is soft heredity: interactions are admitted only on real evidence. A split that would raise the tree's order above 3 faces a doubled hurdle.
interaction_gain_hurdle_mode¶
How interaction_gain_hurdle is applied.
table_budget_cells¶
The cell budget of the table-size prior, which steers the trees toward smaller tables.
table_budget_order_shrink¶
How much table_budget_cells shrinks for each interaction order above 3.
A \(k\)-way table with \(k > 3\) is measured against
\(\frac{budget}{table\_budget\_order\_shrink^{k - 3}}\) cells, because a table with more axes is
harder to read at the same number of cells. 1 turns the shrinking off.
Categorical features settings¶
categorical_features¶
The features to treat as categorical.
For a polars DataFrame or LazyFrame, String, Categorical and Enum columns are
categorical automatically, and this parameter is only needed to treat a numeric column as
categorical. For other input types, every feature is numerical unless it is listed here.
unknown_category¶
How a categorical value that is absent from the training data is scored.
cat_smooth¶
The shrinkage strength \(m\) of the target statistic of a level toward the mean target of the whole dataset.
cat_target¶
The transformation of the target before the target statistic is calculated.
cat_leakage¶
The method used to keep an object's own target out of the target statistic it is trained on. At prediction time, the statistics calculated on the whole training dataset are always used.
cat_k¶
The number of folds of the kfold method of cat_leakage.
cat_n_perms¶
The number of random orders of the ordered method of cat_leakage.
cat_min_data_per_group¶
The minimum total weight (sample weight times exposure) of a level. Levels below it are
collapsed into one shared "<rare>" level before the encoding.
cat_direct_max_levels¶
The maximum number of levels of a low-cardinality feature. A feature with between 3 and this many levels (after rare levels are pooled) skips the cross-fitting and the shrinkage: each level keeps its target statistic calculated on the whole training dataset and gets its own bin. Binary features always use the regular path. 0 turns this off.
cat_channels¶
The numerical features (channels) built from each categorical feature.
cat_count_min_levels¶
The minimum number of levels (after rare levels are pooled) a feature must have to get the
count channel of cat_channels. Features with fewer levels behave as if the
channel was not requested. 0 gives the channel to every categorical feature.
cat_class_freq_min_levels¶
The minimum number of levels (after rare levels are pooled) a feature must have to get the
class_freq channels of cat_channels. Features with fewer levels keep the
target statistic. Used only for multiclassification.
Bagging settings¶
n_bags¶
The number of bags. Each bag is trained on its own sample of the objects, with its own early
stopping, and the bags are averaged into one model. Training costs about n_bags
times as much as training a single model.
bag_subsample¶
The fraction of the objects sampled for each bag. Only used when n_bags is greater
than 1.
cell_refit_base¶
The base penalty of the out-of-bag cell refit. After bagging, every cell of the averaged tables
is refit toward the residuals of the bags' out-of-bag objects under a ridge penalty shaped by
cell_refit_gamma, and the tables are purified again. The refit is kept
only if it improves the held-out deviance.
cell_refit_gamma¶
The adaptive exponent of the penalty of the out-of-bag cell refit. 0 applies the same ridge
penalty to every cell; larger values penalize cells with a strong signal less. Only used when
cell_refit_base is set.
Pruning settings¶
prune¶
Select the tables to deploy after the training, and deploy the smaller set.
prune_main_effects¶
Allow pruning to drop main effects too.
False deploys every main effect the training built and prunes interactions only. True makes
the main effects candidates as well, under hierarchy: a main effect is dropped only when it does
not earn its place and no kept interaction contains it, so a kept interaction always keeps its
main effects. A feature whose main effect is dropped, and which no kept interaction uses, no
longer affects predictions.
prune_selector¶
The method used to select the tables.
prune_path_fraction¶
The fraction of the out-of-bag improvement the deployed tables must capture. The improvement is
measured from the model with main effects only (the intercept-only model with
prune_main_effects=True) to the prefix with the lowest out-of-bag deviance.
prune_path_tolerance¶
The maximum relative excess of the deployed prefix's out-of-bag deviance over the best prefix's.
The deployed prefix is the larger of the smallest prefix that captures
prune_path_fraction of the improvement and the smallest prefix within
this tolerance.
prune_path_steps¶
The number of prefixes scored on the ranked path. The prefix sizes are spaced geometrically
between one table and all the candidate tables. Used only by the ranked_path selector.
prune_guard_min_rows¶
The minimum number of objects with honest evidence needed to judge the tables on.
prune_rebalance¶
After tables are dropped, re-solve the values of the remaining tables for the smaller structure
(one ridge IRLS step, followed by purification). The new values are kept only if they lower the
held-out deviance. Models with multi-channel categorical features (see
cat_channels) are not rebalanced.
prune_table_budget¶
The maximum number of deployed tables of order
prune_table_min_arity or higher.
prune_table_min_arity¶
The lowest interaction order counted by prune_table_budget and by
prune_lambda_tables. The default counts three-way
and higher tables, so main effects and pairs are never limited. Allowed values are integers
from 1 to 8.
prune_box_budget¶
The maximum total number of boxes in the deployed model.
An effect whose dense table would be too large is stored in factored form, as a sum of rank-one
boxes (regions of the feature space), and each box is one row of its exported rating table. A
deeper tree contributes more boxes, so max_depth mostly multiplies the number of boxes rather
than the number of tables.
Fold vote pruning settings¶
prune_n_folds¶
The number of cross-validation folds.
prune_fold_min_rows¶
The minimum number of objects in each fold (see prune_n_folds). Raising it
makes the number of folds adapt down earlier.
prune_fold_es_patience¶
The number of iterations early stopping of the fold models waits after the iteration with the
optimal metric value. The fold models only vote on which tables to keep, so they use a
shorter patience than the deployed model, which always uses
early_stopping_rounds.
prune_min_stability¶
The fraction of the folds in which a table must show signal to be kept. A table is kept when its
mean gain exceeds prune_min_mean_gain and either it was kept by at least
this fraction of the folds' own selections, or its gain was positive in at least this fraction of
the folds.
prune_min_mean_gain¶
The minimum mean gain a table must exceed to be kept, in units of deviance per unit of weight. 0 keeps every table with a positive mean gain.
prune_drop_z¶
The evidence bar for dropping a table. A table the vote would drop is kept instead unless its
gain is significantly negative: its mean over the folds must be below \(-z \cdot SE\), where \(z\) is
the value of this parameter and \(SE\) the standard error of the mean. Tables kept this way are
ranked by mean gain and limited by prune_keep_budget. Tables scored by
fewer than two folds keep the vote's verdict.
prune_keep_budget¶
The maximum number of tables after the evidence gate (prune_drop_z) adds its
tables: the limit is the larger of this value and the number of tables the vote kept. A set that
the vote already made larger than the budget is left as the vote chose it.
prune_lambda_tables¶
Alias: prune_size_penalty
The price of one kept table of order
prune_table_min_arity or higher, in units of held-out
deviance. The selection minimizes the held-out deviance plus this price times the number of
such tables. 0 turns it off.
prune_size_penalty¶
An alias of prune_lambda_tables. Setting both to different values
raises an error.
prune_lambda_boxes¶
The price of one deployed box (see prune_box_budget), in units
of held-out deviance per unit of weight. The selection minimizes the held-out deviance plus this
price times the number of boxes, so a large set of tables has to earn its size. Dense tables cost
no boxes. 0 turns it off.
prune_fold_fidelity¶
Score every candidate table in every fold. A fold model searches its own structure, so it may not build some of the tables the deployed model has, and those tables then get no evidence from that fold. With this parameter, the missing tables are given values in each fold by a ridge fit on the fold's training objects, so every candidate gets a held-out gain.
prune_guard¶
Check the selected set of tables as a whole. The selection judges tables one at a time, so it can drop a group of correlated tables that only matter together. The guard compares the deviance of the selected set with that of the full set on honest objects (the out-of-bag objects, or a shared holdout for grouped data). If the relative gap exceeds the tolerance, dropped tables are added back, best evidence first, until it does not.
prune_guard_tol¶
The tolerance of the no-harm guard: the largest acceptable relative increase of the deviance of the selected set over the full set.
prune_guard_z¶
The multiplier of the standard error that can raise the guard's tolerance above
prune_guard_tol. 0 uses the fixed tolerance.
prune_guard_z_dn¶
The multiplier of the standard error that can lower the guard's tolerance below
prune_guard_tol when the evidence is precise. It can only make the guard
act more often, never less. 0 turns it off.
prune_guard_tol_floor¶
The lowest tolerance prune_guard_z_dn can lower the guard's tolerance to,
so that very precise evidence cannot drive it to zero. Ignored when prune_guard_z_dn is 0.
prune_slope_eps¶
The dead band of the post-pruning slope correction. For a poisson, gamma or tweedie model
trained without an exposure, the scale \(b\) of the pruned model's score is estimated on the
out-of-bag objects after the guard. The score is rescaled only if \(|b - 1|\) exceeds this value
and \(b\) differs from 1 by at least prune_slope_min_z standard errors.
The outcome is recorded in pruning_report_["slope"].
prune_slope_min_z¶
The number of standard errors by which the scale \(b\) must differ from 1 for the post-pruning
slope correction to be applied (see prune_slope_eps).
Banding settings¶
band_tolerance¶
The tolerance of banding, as a multiple of the noise between the bags. The mean squared change
of the predictions caused by banding is held within \((band\_tolerance \cdot \sigma)^2\), and
within band_deviance_cap. Larger values give coarser bands.
band_deviance_cap¶
The maximum cost of banding, as a fraction of the model's training deviance. The noise tolerance alone could let a very noisy model move far; this cap bounds what that can cost.
Graduation settings¶
graduate¶
Smooth the deployed tables with Whittaker-Henderson graduation.
graduation_alpha¶
A fixed smoothing strength for every table, instead of the strength each table picks by
generalized cross-validation. 0 turns off the smoothing of the dense tables, which is useful
together with graduation_high_order_alpha.
graduation_high_order_alpha¶
The strength, in the range \([0; 1]\), of an additional neighbour smoothing step for effects of order 3 to 8 that are stored in factored form. Unlike the smoothing of the dense tables, its strength is fixed rather than selected by cross-validation.
Purification settings¶
ref_measure¶
The reference measure the tables are purified against.
measure_floor¶
The total mass of the floor of the exposure reference measure, as a fraction of the
training data's mass. It is spread evenly over the cells of each axis so that every weight is
strictly positive. The value must be finite and positive. Used only with
ref_measure="exposure".
Multiclassification settings¶
multiclass_prune_cv¶
The selection regime for multiclassification.
multiclass_prune_guard¶
Check the selected set of tables as a whole against the full set on honest objects, and add
dropped tables back while the gap exceeds the tolerance (see
multiclass_prune_guard_floor).
multiclass_prune_guard_floor¶
The floor of the guard's tolerance, as a fraction of the improvement of the full model's deviance over the class prior.
multiclass_prune_sel_bags¶
The number of bags of the models the selection is trained on (the fold models, or the selection model of the single-split regime). Must be at least 1. Any other value than 1 raises an error for regression and binary classification.
prune_validation_fraction¶
The fraction of the objects used to select the tables in the single-split regime
(multiclass_prune_cv=False). Any other value than the default raises an error in the other
regimes, and for regression and binary classification.
prune_se_rule¶
The selection rule, in standard errors of the held-out deviance estimate. 0 selects the set with the lowest held-out deviance; larger values (for example, the classic one-standard-error rule, 1.0) prune more aggressively and trade deviance for a smaller model.
Performance settings¶
n_jobs¶
The number of threads to use during the training and prediction.
hist_precision¶
The precision of the gradient histograms used to search for splits.
refine_closed_form_tier2¶
Calculate the leaf refinement steps (see
leaf_refine_steps) of a poisson model in closed form instead
of with an exact line search. This is faster, and the leaf values differ from the exact path by
about \(10^{-7}\). False uses the exact path.
incremental_mu¶
Update the predicted mean \(\mu = e^{a}\) of a poisson model incrementally from one iteration to
the next instead of recalculating it. This is faster, and the predictions differ from the exact
path by about \(10^{-9}\) to \(10^{-6}\).
Advanced settings¶
lambda_scale_invariant¶
Rescale lambda_ by the mean Hessian per object of each iteration instead
of using it as it is.
ridge_refit_l2¶
The L2 penalty of a fully corrective refit. After the trees are built, all their leaf values are re-solved jointly by regularized IRLS, with the tree structures fixed.
ridge_refit_max_iter¶
The maximum number of IRLS iterations of the fully corrective refit. Only used when
ridge_refit_l2 is set.
dart_drop_rate¶
Turns on DART (Dropout Additive Regression Trees): at each iteration, the trees built so far are dropped with this probability before the new tree is built, and the tree weights are normalized as in DART. The value must be in the range \([0; 1)\).
nesterov¶
Reserved for Nesterov-accelerated boosting, which is not implemented yet. Setting it to True
raises an error.
prune_refit_full¶
Deprecated, and without effect: the deployed model is always trained on all objects. Setting it
to True raises an error. It will be removed in a future release.