🎉 75% of content is free forever — Unlock Premium from $10/mo →
CW
Search courses…
💼 Servicesℹ️ About✉️ ContactView Pricing Plansfrom $10

RNN Deep Dive — Vanilla RNN, BPTT and Vanishing/Exploding Gradients

Sequence ModelsRNNs🟢 Free Lesson

Advertisement

Sequence Models

RNNs Deep Dive — Understanding Recurrent Neural Networks

Recurrent Neural Networks process sequential data by maintaining a hidden state that carries information across time steps. This tutorial covers vanilla RNNs and their fundamental limitations.

  • BPTT Unrolls Through Time — Backpropagation through time treats the RNN as a deep feedforward network
  • Vanishing Gradients Limit Memory — The product of Jacobians shrinks exponentially, limiting to ~10-20 steps
  • Orthogonal Initialization Helps — Setting recurrent weights to orthogonal matrices preserves gradient magnitude

RNN Deep Dive — Vanilla RNN, BPTT and Vanishing/Exploding Gradients

Recurrent Neural Networks process sequential data by maintaining a hidden state that carries information across time steps. This tutorial covers vanilla RNNs and their fundamental limitations.


Why Recurrence?


Vanilla RNN

RNN: Unrolled Through Timet=1h₁RNN Cellt=2h₂Same cellt=3h₃Same cellt=4h₄Same cellh₁h₂h₃x₁x₂x₃x₄y₁y₂y₃y₄Same parameters W_hh, W_xh shared across all time steps

Backpropagation Through Time (BPTT)

BPTT: Gradient Flow Through TimeForward Pass (left to right)h₁h₂h₃h₄W_hhW_hhW_hhBackward Pass (right to left)∂L₄∂L₃∂L₂∂L₁W_hhᵀW_hhᵀW_hhᵀGradient throughT steps:(W_hhᵀ)ᵀ × ...

Vanishing and Exploding Gradients

Vanishing vs Exploding GradientsTime Steps (T)Gradient MagnitudeGood (‖W‖=1)Vanishing‖W‖=0.9Exploding‖W‖=1.1~20 steps: vanishing gradient makes learning impossiblePractical limit

Solutions to Vanishing Gradients


Applications


Practical Considerations


Summary

  • Vanilla RNN maintains a hidden state that evolves over time through recurrence
  • BPTT unrolls the RNN and applies backpropagation through the unrolled graph
  • Vanishing gradients limit vanilla RNNs to ~10-20 time steps
  • Exploding gradients can be mitigated with gradient clipping
  • Solutions: LSTM/GRU (gating), orthogonal initialization, skip connections
  • Transformers have largely replaced RNNs for most sequence tasks due to parallelization

Next: LSTM Networks

Need Expert Deep Learning Help?

Get personalized tutoring, project support, or professional consulting.

Advertisement