AI Glossary

Regularization

Regularization changes model training to discourage solutions that fit training details too closely and fail on new data. It includes penalties and constraints, as well as techniques such as dropout and early stopping.

· Updated · Chain of Thought

Model Training

A price-prediction model can use an L2 penalty that adds a cost for large coefficients to its training objective. The model must balance fitting the observed prices against that penalty. This favors smaller coefficients, rather than chasing every training fluctuation.

Regularization is broader than weight penalties. Dropout randomly omits selected activations during training. Early stopping ends training when a chosen validation measure stops improving and can retain the earlier model. These methods intervene differently; they are not interchangeable settings.

The benefit must be measured on held-out examples. Too little constraint can allow overfitting; too much can prevent the model from learning useful patterns. The regularization strength is a hyperparameter to select for the task. Neither a penalty nor early stopping repairs incorrect labels, leaked test data or an unrepresentative evaluation.

Fit the data, with a penalty on large weightsAn L2 training objective adds training loss to lambda times the sum of squared weights. Lambda controls the strength of the penalty. Too little constraint can allow overfitting; too much can underfit. Validation helps choose the balance, and penalties do not fix bad data. Fit the data, with a penalty on large weights L2 regularization changes what training tries to minimizeMINIMIZETraining loss+λ × sum of squared weightsλ (lambda) sets the penalty’s strength.Too little constraintCan overfit training detailsChoose using validationMeasure on held-out dataToo much constraintCan miss useful patternsA penalty cannot repair incorrect labels or leaked test data.
The L2 objective follows Google’s regularization course. Validation helps choose its strength; a penalty does not repair bad labels or leaked test data. Download the image

Sources

Go deeper

  • Keras: Dropout layer docs

    Inspect a regularizer that masks activations during training rather than penalizing weight size.