Skip to content

Parameter tuning

t-boost provides a flexible interface for parameter tuning and can be configured to suit different tasks. The default parameters are already a complete recipe: early stopping, bagging, pruning, banding and graduation are all turned on, so a model usually performs well without tuning.

This section contains some tips on the possible parameter settings.

Number of trees

It is recommended to check that there is no obvious underfitting or overfitting before tuning any other parameters. In order to do this it is necessary to analyze the metric value on the validation dataset and the number of trees each bag kept.

By default the number of trees (n_trees) is a large cap and early stopping decides when to stop. After the fit, check stopping_reason_per_bag_: bags that stopped with "max_trees" reached the cap before early stopping triggered, so raise n_trees or learning_rate. Pass an eval_set to fit to monitor the deviance of every iteration in evals_result_.

Learning rate

This setting is used for reducing the gradient step. It affects the overall time of training: the smaller the value, the more iterations are required for training. Choose the value based on the performance expectations.

Possible ways of adjusting the learning rate depending on the overfitting results:

  • There is no overfitting on the last iterations of training (the training does not converge) — increase the learning rate.
  • Overfitting is detected — decrease the learning rate.

Tree depth and interaction order

max_interaction_order decides the largest tables the model can have: 1 gives main effects only, 2 adds pairs, and the default 3 adds three-way tables. Higher orders are supported but quickly stop being readable.

max_depth (3 by default) controls the resolution of the trees, not their interaction order. Keep the default unless a held-out comparison shows a gain, and keep max_depth equal to max_interaction_order at order 4 and above.

L2 regularization

Try different values for the regularizer (lambda_) to find the best possible. With large sample weights or exposures, set lambda_scale_invariant=True so that the value keeps an effect.

Bagging

n_bags (8 by default) trades training time for accuracy: the training costs about n_bags times a single model. For the cheapest single-model baseline, pass n_bags=1, validation_fraction=None, prune=False.

Smaller rating structures

The following parameters give fewer or smaller tables:

Models for review or filing

  • Set min_data_in_leaf explicitly, so that no cell rests on a handful of objects.
  • Call check_bindings after the fit to make sure the parameters you set took effect.
  • Save the pricing report beside the serialized model, and record your context in the metadata attribute.

Methods for hyperparameter search

With scikit-learn installed, the estimators work with its model selection tools, such as GridSearchCV and RandomizedSearchCV:

from sklearn.model_selection import GridSearchCV
from t_boost import TBoostRegressor

grid = {"learning_rate": [0.05, 0.1], "lambda_": [1.0, 10.0]}
search = GridSearchCV(TBoostRegressor(), grid, cv=3).fit(X, y)
print(search.best_params_)

The default score is R2 for TBoostRegressor and accuracy for TBoostClassifier. For a model with an exposure, pass a scorer based on the deviances of t_boost.metrics. For panel data, use a group-aware splitter such as GroupKFold.