Random forest
A random forest combines predictions from many randomized decision trees. In the standard design, trees train on bootstrap samples of rows and consider random subsets of features at each split, then aggregate their predictions.
A churn classifier could train many decision trees on different resampled customer rows. At each split, each tree considers only a random selection of features. These choices encourage differences among trees so aggregation can reduce the variance of relying on one tree.
Aggregation depends on the implementation. Breiman’s original classifier uses votes; scikit-learn averages class probabilities. For example, three trees giving churn probabilities 0.8, 0.6 and 0.1 produce an average of 0.5. A chosen threshold of 0.6 would label that customer non-churn. These numbers illustrate arithmetic, not measured or necessarily calibrated probabilities.
Trees can train independently, unlike gradient boosting, which adds learners in stages. A forest can capture nonlinear relationships, but it still needs held-out evaluation. Feature importance is a model-dependent diagnostic, not proof that a feature causes churn.
Sources
- scikit-learn: Forests of randomized trees — Explains bootstrap rows, random feature subsets, variance reduction and probability averaging versus the original vote rule.
Go deeper
- scikit-learn: Feature importances with a forest of trees docs
Compare impurity-based importance with held-out permutation importance and their limitations.