Recurrent neural network
A recurrent neural network (RNN) processes a sequence by updating a hidden state from the current input and the previous state. Reusing the same transition weights across steps lets it process sequences of different lengths.
Also known as: RNN
An RNN could read “not very good” one word at a time. Each step combines the word’s input vector with the previous hidden state. A classifier reads the final state to predict sentiment. The transition at each position uses the same learned weights; the three steps are not three independently trained networks.
In a basic RNN, the state update applies a nonlinear function to weighted input and state values. Earlier words can influence later states, but the state is a learned summary rather than a verbatim transcript. LSTM variants add gates to control information retention.
Recurrence makes the state updates depend on preceding steps. Gradients through long chains can shrink or grow excessively, complicating training. Choose using the sequence length, memory needs and task results. Transformers instead connect positions through attention; neither architecture name establishes that all earlier information is used reliably.
Sources
- PyTorch: RNN — Gives the hidden-state recurrence and an equivalent loop with shared input and recurrent parameters.
- Hochreiter and Schmidhuber: Long Short-Term Memory — Describes gating and the vanishing or exploding error signals that complicate long-range recurrent learning.
Go deeper
- PyTorch: Classifying names with a character-level RNN docs
Build a sequence classifier and follow hidden-state outputs into a classification head.