Training Data
Training data is the collection of examples used to fit a model's parameters. Depending on the training objective, examples can supply explicit labels or targets derived from the data itself.
A support classifier might learn from tickets paired with categories such as billing, access and outages. If most examples come from one product, it may have little evidence about another product’s vocabulary. More rows do not automatically fix missing coverage or incorrect labels.
Document where the examples came from, how they were selected and labeled, and which uses their collection permits. Keep a record of the version used for each training run so changes in behavior can be investigated. Synthetic examples also need review: they can reproduce the generator’s errors or omit difficult cases.
Keep validation and test examples separate from parameter fitting. For instance, hold out tickets from whole customers when the evaluation asks whether the classifier works for new customers. Near-duplicates across splits can make results look better than they are. This distinction matters because training measures what the model fitted, while evaluation asks whether that learning transfers to work it has not seen.
Sources
- Gebru et al.: Datasheets for Datasets — Proposes documenting dataset motivation, composition, collection, intended uses and maintenance.
Go deeper
- Google: Dividing the original dataset course
Separate training, validation and test examples and check for duplicate leakage.