How training is performed¶
The goal of training is to select the model \(y\), depending on a set of features \(x_{i}\), that best solves the given problem (regression, classification, or multiclassification) for any input object. This model is found by using a training dataset, which is a set of objects with known features and label values. Accuracy is checked on the validation dataset, which has data in the same format as in the training dataset, but it is only used for evaluating the quality of training (it is not used for training).
t-boost is based on gradient boosted decision trees. During training, a set of decision trees is built consecutively. Each successive tree is built with reduced loss compared to the previous trees. The trees are constrained so that the trained model can be rewritten exactly as a set of rating tables.
The number of trees is controlled by the starting parameters. To prevent overfitting, use early stopping. When it is triggered, trees stop being built.
Building stages for a single tree:
- Preliminary calculation of splits: every numerical feature is quantized into bins once, before the training.
- (Optional) Transforming categorical features to numerical features.
- Choosing the tree structure.
- Calculating values in leaves: each leaf gets the Newton step
\(w = -\frac{G}{H + \lambda}\), where \(G\) and \(H\) are the sums of the gradients and the
Hessians of the objects in the leaf, optionally refined by further Newton steps (see
leaf_refine_steps). The values are multiplied by the learning rate before the tree is added to the model.
By default several models (bags) are trained this way, each on its own sample of the objects (see Bagging). Then the model is turned into the deployed set of rating tables:
- Decomposing the trees into rating tables: the averaged trees are rewritten exactly as an intercept plus one table per main effect and per interaction.
- Pruning, banding and graduation: the tables that contribute little are dropped, the interaction tables are condensed into bands, and the tables are smoothed.