Dropout
Dropout is a regularization technique that randomly sets selected neural-network activations to zero during training. In the common inverted-dropout implementation, retained activations are scaled up during training and ordinary inference uses them without masking.
Dropout makes training less dependent on a particular combination of activations. The mask changes between training steps according to a chosen dropout rate; it does not permanently delete neurons or learned weights.
For example, use a drop probability of 0.5 on activations [2, 4, 6, 8]. One possible mask keeps the first and third values. Inverted dropout divides retained values by 1 − 0.5, giving [4, 0, 12, 0]. Each activation is preserved in expectation across masks, not necessarily in any one sample. Ordinary inference applies no mask and no extra doubling.
Dropout can help a classifier generalize, but its rate must be chosen using held-out data. Too much masking can prevent useful learning. It does not repair incorrect labels or an unrepresentative test set. Deliberately keeping dropout active for uncertainty estimation is a separate use; it is not the ordinary inference mode shown here.
Sources
- Keras: Dropout layer — Documents random training-time masking, inverted scaling and unmasked inference.
Go deeper
- Hinton et al.: Improving neural networks by preventing co-adaptation paper
Read the original dropout proposal and its explanation of co-adaptation.