Dataset
A dataset is a collection of data assembled for a purpose, such as analysis, model training or evaluation. It can contain text, images, measurements or other records, with or without labels.
Also known as: data set
Model TrainingAI Evaluation & Reliability
For a ticket classifier, the dataset might contain each customer’s message and a human-assigned category. Training examples fit the model; validation examples help choose settings; a separate test set estimates performance after those choices. These splits can belong to one documented collection.
The records need not be a tidy spreadsheet. Audio, images and documents are datasets too. Documentation should explain where the data came from, what it covers, how labels were assigned and what uses or permissions apply. The Datasheets for Datasets paper proposes a structured way to record this information.
More records do not automatically mean better evidence. Duplicate customers appearing in both training and test data can exaggerate performance, while missing a customer group can hide failures. Choose splits that reflect the intended use: testing on later tickets answers a different question from a random split across the same period.
Sources
- Gebru et al.: Datasheets for Datasets — Proposes documenting dataset motivation, composition, collection, uses and maintenance.
- Google: Dividing the original dataset — Explains training, validation and test partitions, representative evaluation and duplicate leakage.
- Hugging Face: Dataset features — Shows how a dataset represents typed columns, labels and image or audio records.
Go deeper
- Hugging Face Datasets: Quickstart docs
Load, inspect and transform a dataset, including its splits.