AI Glossary

Dropout

Dropout is a regularization technique that randomly sets selected neural-network activations to zero during training. In the common inverted-dropout implementation, retained activations are scaled up during training and ordinary inference uses them without masking.

· Updated · Chain of Thought

Dropout makes training less dependent on a particular combination of activations. The mask changes between training steps according to a chosen dropout rate; it does not permanently delete neurons or learned weights.

For example, use a drop probability of 0.5 on activations [2, 4, 6, 8]. One possible mask keeps the first and third values. Inverted dropout divides retained values by 1 − 0.5, giving [4, 0, 12, 0]. Each activation is preserved in expectation across masks, not necessarily in any one sample. Ordinary inference applies no mask and no extra doubling.

Dropout can help a classifier generalize, but its rate must be chosen using held-out data. Too much masking can prevent useful learning. It does not repair incorrect labels or an unrepresentative test set. Deliberately keeping dropout active for uncertainty estimation is a separate use; it is not the ordinary inference mode shown here.

Mask during training, use the full layer at inferenceTraining starts with activations 2, 4, 6, 8. A mask retains the first and third values, then scales them by two to produce 4, 0, 12, 0. Ordinary inference leaves the original activations 2, 4, 6, 8 unchanged. Scaling preserves each activation in expectation, not the total of this particular sample. Mask during training, use the full layer at inferenceInverted dropout with drop probability p = 0.5; one possible mask Training input: [2, 4, 6, 8]Inference input: [2, 4, 6, 8]Mask: [1, 0, 1, 0]Retained values × 1 / (1 − p)Output: [4, 0, 12, 0]No maskNo extra scalingOutput: [2, 4, 6, 8]Expected value is preserved across masks, not in every sample.Masks change during training; no neurons or weights are permanently removed.
A worked inverted-dropout example following the Keras layer contract. The inference column shows ordinary evaluation, not Monte Carlo dropout. Download the image

Sources

  • Keras: Dropout layer — Documents random training-time masking, inverted scaling and unmasked inference.

Go deeper