This paper introduces a new technique for mapping Deep Recurrent Neural Networks (RNN) efficiently onto GPUs by caching recurrent weights in on-chip registers, achieving a 30x speedup over the previous state of the art at low mini-batch sizes and enabling larger, deeper RNNs to be trained efficiently.