A type of neural network designed to process sequential data by maintaining an internal “memory” of previous inputs, allowing it to handle tasks where context and order matter, such as language, time series, and speech.
Imagine you’re reading a book aloud to a friend. As you read each word, you don’t just think about that word in isolation — you remember all the words that came before it. That’s why you can understand pronouns like “he” or “she,” and why you can follow a story that unfolds over many pages.
An RNN works similarly. When it processes information, it doesn’t just look at the current input — it also remembers what it saw before. It has a kind of “memory” that carries forward from one step to the next.
This is really useful for things that happen in sequence, like sentences in a sentence, notes in a song, or stock prices over time. The RNN can use what it learned earlier to help understand what’s happening now.
But there’s a catch: just like you might forget the beginning of a very long story, RNNs can struggle to remember things from far back in a sequence. That’s why newer versions like LSTM and GRU were invented — they have better “long-term memory.”
RNNs process sequential data by maintaining a hidden state that captures information about previous time steps. At each step, the network takes both the current input and the previous hidden state as inputs, producing a new hidden state and an output.
How it works:
The hidden state acts as the network’s memory, allowing information to persist across time steps.
Types of RNNs:
Common applications:
Limitations:
While Transformers have largely replaced RNNs for many NLP tasks, RNNs and their variants (LSTM, GRU) remain relevant in specific enterprise scenarios:
Where RNNs still excel:
Business considerations:
When to consider RNNs vs. Transformers:
Following a recipe while cooking. You don’t just look at the current step — you remember what you did before. If step 5 says “add the mixture from step 3,” you need to remember what you did in step 3. Your memory of previous steps helps you understand and execute the current step correctly.
# Simple RNN for sequence classification using PyTorch
import torch
import torch.nn as nn
class SimpleRNN(nn.Module):
def __init__(self, input_size, hidden_size, num_layers, num_classes):
super(SimpleRNN, self).__init__()
self.hidden_size = hidden_size
self.num_layers = num_layers
# RNN layer
self.rnn = nn.RNN(
input_size=input_size,
hidden_size=hidden_size,
num_layers=num_layers,
batch_first=True
)
# Output layer
self.fc = nn.Linear(hidden_size, num_classes)
def forward(self, x):
# Initialize hidden state with zeros
batch_size = x.size(0)
h0 = torch.zeros(self.num_layers, batch_size, self.hidden_size)
# Forward pass through RNN
out, _ = self.rnn(x, h0)
# Get output from last time step
out = self.fc(out[:, -1, :])
return out
# Create model with sample parameters
model = SimpleRNN(
input_size=10, # Number of input features
hidden_size=64, # Number of hidden units
num_layers=2, # Number of RNN layers
num_classes=5 # Number of output classes
)
# Test with sample input
batch_size = 32
sequence_length = 20
input_features = 10
sample_input = torch.randn(batch_size, sequence_length, input_features)
output = model(sample_input)
print("Input shape:", sample_input.shape)
print("Output shape:", output.shape)
print("Output:", output[0]) # First sample's predictions
Reality: For most NLP tasks, Transformers have surpassed RNNs in performance and efficiency. RNNs are now mainly used for specific use cases like real-time processing or edge deployment.
Reality: Vanilla RNNs struggle with long-range dependencies due to vanishing gradients. LSTM and GRU help but still have limits.