Generalization
In machine learning, generalization is a model's ability to perform well on examples beyond those used to fit it. That ability is measured for a particular task and data distribution, not for every possible future input.
AI Evaluation & ReliabilityModel Training
A support classifier should categorize tickets from customers whose messages were not used in training. If the same customer’s near-duplicate tickets appear in both splits, a high test score may reflect familiarity rather than performance on new customers. Hold out customers when that is the question the evaluation must answer.
Training fits the model; validation guides development choices; a separate test checks the result after those choices. Cross-validation repeats training and validation across different partitions. Choose the partitioning strategy to match the intended use, including time ordering or related records where relevant.
The distinction matters because a successful held-out test has a boundary. It does not establish reliability after a policy change or on a language missing from the data. A training-to-test gap can reveal overfitting, while a change in the incoming distribution can hurt a previously useful model. State what examples a result represents and keep checking outcomes after deployment.
Sources
- Google: Generalization — Frames generalization as prediction on new examples and motivates separate data partitions.
- scikit-learn: Cross-validation — Explains held-out evaluation, tuning separation and split strategies for grouped or ordered data.
Go deeper
- scikit-learn: Common pitfalls docs
See how preprocessing leakage and inconsistent transformations undermine unseen-data evaluation.