← Back to concepts
7 min read

Recurrent neural networks (RNNs)

A recurrent neural network, or RNN, is a neural network designed to process a sequence one step at a time. At each step, it reads the current input and combines it with a hidden state that carries information from earlier steps.

This made RNNs useful for language, speech, time series, and other ordered data before transformers became dominant.

How recurrence works

Imagine processing a sentence from left to right. The RNN updates its hidden state for each token:

token 1 + state 0 -> state 1
token 2 + state 1 -> state 2
token 3 + state 2 -> state 3

The current state is a compact summary of what the network has seen so far. It is a fixed-size learned vector, not stored text. The starting state h0 is often all zeros or a learned initial vector. The model can use the current state to classify the sequence or predict the next item.

The same weights are reused at every step. This lets one network handle sequences of different lengths without creating a separate set of weights for every position.

A simple RNN update is often written like this:

h_t = tanh(W_x x_t + W_h h_(t-1) + b)

In plain words: the new hidden state h_t comes from the current input vector x_t, the previous hidden state h_(t-1), learned weights W_x and W_h, and a bias b. The tanh function squashes the result into a stable range. The same weights are reused at every time step.

An RNN unrolled across a sequence The same RNN cell is reused at four time steps. State h0 and token The produce h1. Then h1 and token server produce h2. Then h2 and token failed produce h3. Then h3 and token today produce h4, which can be used for a prediction. The hidden state arrow carries information from each step to the next. One RNN cell reused over time h0 RNN cell same weights token: The RNN cell same weights token: server RNN cell same weights token: failed RNN cell same weights token: today h1 h2 h3 h4 for prediction
Unrolling shows the key dependency: each step waits for the hidden state produced by the previous step.

The memory problem

The hidden state has limited capacity. As a sequence becomes longer, important information from the beginning can be diluted by later updates. Because the state vector has a fixed size, it must compress the past into a lossy summary.

During training, RNNs use backpropagation through time. The loop is unrolled into a long chain, then trained like a deep network. For long sequences, teams often use truncated BPTT, which backpropagates through only a fixed number of recent steps to save memory and reduce instability.

The model’s error signal must travel backward through many time steps. Repeated multiplication can make that signal become extremely small, a problem called the vanishing gradient. If the signal becomes too small, early steps receive almost no useful update. The 0.6 multiplier below is illustrative; real behavior depends on the learned weights and activation functions.

Very large repeated multiplications can create exploding gradients, where updates become unstable. A common defense is gradient clipping, which caps the update size before it damages training.

A long backward path can shrink the training signal A forward chain runs from h0 to h4. During backpropagation through time, an illustrative gradient starts at 1.00 near h4 and is multiplied by 0.6 at each step as it moves backward, becoming 0.60, 0.36, 0.22, and 0.13 by h0. The early token receives a much weaker update. Vanishing gradient, simplified Illustrative multiplier: each backward step keeps 60% of the signal. h0 h1 h2 h3 h4 loss 1.00 -> 0.60 0.36 0.22 0.13 The exact numbers vary by model, but repeated multiplication can make early-step updates tiny.
Backpropagation through many recurrent steps can weaken the signal that teaches early tokens what to remember.

Variants such as LSTM and GRU add gates that help decide what to keep, update, or forget. They can remember longer relationships than a simple RNN, but they still process a sequence step by step.

An LSTM, or long short-term memory network, keeps a separate cell state alongside the hidden state. Gates decide what to forget, what new information to write, and what part of the cell state to expose as the next hidden state. A GRU, or gated recurrent unit, is a simpler gated design that combines some of those decisions into two main gates.

How gates help an LSTM carry information A simplified LSTM step. The previous cell state c_(t-1) runs along a top path. The forget gate scales how much of that old memory is kept, and the input gate adds new information, producing the new cell state c_t, which carries on to the next step. All three gates read the same inputs: the previous hidden state h_(t-1) and the current input x_t. The output gate decides how much of the updated memory, passed through tanh, is exposed as the new hidden state h_t. One LSTM step, simplified c_(t-1) cell state path x + c_t to t+1 Forget gate scale old memory Input gate add new info Output gate reveal memory h_t hidden output Inputs h_(t-1), x_t Every gate reads h_(t-1) and x_t. GRUs use two gates (update and reset) for a similar goal.
The forget and input gates edit the cell state path; the output gate decides how much of the updated memory becomes the hidden state.

RNNs versus transformers

RNNs can also be bidirectional: one RNN reads left to right, another reads right to left, and their states are combined. This helps understanding tasks where the whole input is available, but it cannot be used the same way for left-to-right generation.

RNNs and transformers both process sequences, but they organize computation differently:

Model Main approach Practical effect
RNN Carry a hidden state from one step to the next Natural sequence order, but limited parallelism
Transformer Use attention between token positions Better parallel training and direct long-range connections

Because an RNN must usually finish one step before starting the next, it is harder to train efficiently on large datasets. A transformer can process many positions in parallel during training using attention. Transformer generation is still one token at a time during inference, but training is much more parallel.

Sequential recurrence versus parallel attention The top lane shows an RNN processing token positions one after another, where step four directly receives only step three's state, which carries a summary of earlier steps. The bottom lane shows a transformer training pass where all four token positions are available together and attention can connect distant positions directly. Different ways to organize sequence computation RNN must carry state step by step step 1 step 2 step 3 step 4 Transformer positions together pos 1 pos 2 pos 3 pos 4 direct attention links
RNNs make sequence order explicit through a state chain; transformers trade that chain for attention between positions.

RNNs can still be a good fit for small streaming systems, compact time-series models, or devices with strict resource limits. The best architecture depends on the task and constraints.

Historically, many sequence-to-sequence translation systems used RNN encoders and decoders. Attention was first added to help those decoders look back at the most relevant source states, a pattern that later became central to encoder-decoder models. Some modern recurrent-style and state-space ideas are also returning for long sequences, usually to reduce the memory cost of full attention.

The key idea

An RNN processes ordered data by carrying a hidden state through a sequence. This idea provides useful memory, but long-range learning and sequential computation are difficult. LSTMs and GRUs improve the design with gates, while transformers use attention to connect positions more directly.