Embodiments of end-to-end deep learning systems for speech recognition (Deep Speech 2) that handle a diverse variety of speech including noisy environments, accents and different languages, using recurrent/convolutional neural networks trained with CTC.