Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗Lexical Matching of Queries and Ads Bid Terms in Sponsored Search
Lecture notes in computer science · DOI 10.1007/978-3-319-46049-9_22 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates, Guoqiang Wang · 3 authors totalRefactoring Software Development Process Terminology Through the Use of Ontology
EuroSPI · DOI 10.1007/978-3-319-44817-6_4 · Source: dblp+adapt-autodesk-authorityAlex O'Connor, Paul M. Clarke, Antoni-Lluís Mesquida Calafat, Damjan Ekert, Joseph J. Ekstrom, Tatjana Gornostaja, Milos Jovanovic 0003, Jørn Johansen · 16 authors totalModeling and Processing of Time Interval Data for Data-Driven Decision Support
DOI 10.1007/978-3-319-42620-4_69 · 5 citations · Source: openalex+career-authorityPhilipp Meisen, Tobias Meisen, Marco Recchioni, Daniel Schilberg, Sabina Jeschke · 5 authors totalBitmap-Based On-Line Analytical Processing of Time Interval Data
DOI 10.1007/978-3-319-42620-4_68 · 0 citations · Source: openalex+career-authorityPhilipp Meisen, Tobias Meisen, Diane Keng, Marco Recchioni, Sabina Jeschke · 5 authors totalImplementing a Volunteer Notification System into a Scalable, Analytical Realtime Data Processing Environment
DOI 10.1007/978-3-319-42620-4_64 · 1 citations · Source: openalex+career-authorityPhilipp Meisen, J. Elsner, Tomas Sivicki, Tobias Meisen, Sabina Jeschke · 5 authors totalUsing Off-the-Shelf Medical Devices for Biomedical Signal Monitoring in a Telemedicine System for Emergency Medical Services
DOI 10.1007/978-3-319-42620-4_61 · 9 citations · Source: openalex+career-authorityPhilipp Meisen, Sebastian Thelén, Michael Czaplik, Daniel Schilberg, Sabina Jeschke · 5 authors totalAUDIME: Augmented Disaster Medicine
DOI 10.1007/978-3-319-42620-4_47 · 5 citations · Source: openalex+career-authorityPhilipp Meisen, Alexander Paulus, Michael Czaplik, F. Hirsch, Tobias Meisen, Sabina Jeschke · 6 authors totalTIDAQL: A Query Language Enabling On-line Analytical Processing of Time Interval Data
DOI 10.1007/978-3-319-42620-4_42 · 5 citations · Source: openalex+career-authorityPhilipp Meisen, Diane Keng, Tobias Meisen, Marco Recchioni, Sabina Jeschke · 5 authors totalTowards Evaluating the Impact of Anaphora Resolution on Text Summarisation from a Human Perspective
NLDB · DOI 10.1007/978-3-319-41754-7_16 · Source: dblp+adapt-autodesk-authorityAlex O'Connor, Mostafa Bayomi, Killian Levacher, M. Rami Ghorab, Peter Lavin, Alexander O'Connor, Séamus Lawless · 7 authors totalAn Investigation of Software Development Process Terminology
SPICE · DOI 10.1007/978-3-319-38980-6_25 · Source: dblp+adapt-autodesk-authorityAlex O'Connor, Paul M. Clarke, Antoni-Lluís Mesquida, Damjan Ekert, Joseph J. Ekstrom, Tatjana Gornostaja, Milos Jovanovic 0003, Jørn Johansen · 16 authors totalThe Essence of Dependent Object Types
Lecture Notes in Computer Science · DOI 10.1007/978-3-319-30936-1_14 · 71 citations · Source: openalexMartin Odersky, Nada Amin, Samuel Grütter, Tiark Rompf, Sandro Stucki · 5 authors totalToward Reproducible Baselines: The Open-Source IR Reproducibility Challenge
European Conference on Information Retrieval · DOI 10.1007/978-3-319-30671-1_30 · 111 citations · Source: semantic-scholar+dblp+lucidworksGrant Ingersoll, Jimmy J. Lin, Matt Crane, A. Trotman, Jamie Callan, Ishan Chattopadhyaya, John Foley, Craig Macdonald · 9 authors totalMotion Planning Under Uncertainty Using Differential Dynamic Programming in Belief Space
ISRR 2013 (Springer Tracts in Advanced Robotics 114) · DOI 10.1007/978-3-319-29363-9_27 · 50 citations · Source: openalexJur van den Berg, Sachin Patil, Ron Alterovitz · 3 authors totalSafe Motion Planning for Imprecise Robotic Manipulators by Minimizing Probability of Collision
ISRR 2013 (Springer Tracts in Advanced Robotics 114) · DOI 10.1007/978-3-319-28872-7_39 · 40 citations · Source: openalexJur van den Berg, Wen Sun, Luis G. Torres, Ron Alterovitz · 4 authors totalEmergency Evacuation Plan Maintenance
Encyclopedia of GIS · DOI 10.1007/978-3-319-23519-6_344-2 · 0 citations · Source: openalex+first-party-career-authorityMark Samuel Tuttle, Cheng Liu, Mark S. Tuttle · 3 authors totalConcluding Remarks
Synthesis lectures on information concepts, retrieval, and services · DOI 10.1007/978-3-031-02298-2_5 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, B. Barla Cambazoğlu, Ricardo Baeza‐Yates · 3 authors totalThe Query Processing System
Synthesis lectures on information concepts, retrieval, and services · DOI 10.1007/978-3-031-02298-2_4 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, B. Barla Cambazoğlu, Ricardo Baeza‐Yates · 3 authors totalThe Indexing System
Synthesis lectures on information concepts, retrieval, and services · DOI 10.1007/978-3-031-02298-2_3 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, B. Barla Cambazoğlu, Ricardo Baeza‐Yates · 3 authors totalThe Web Crawling System
Synthesis lectures on information concepts, retrieval, and services · DOI 10.1007/978-3-031-02298-2_2 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, B. Barla Cambazoğlu, Ricardo Baeza‐Yates · 3 authors totalIntroduction
Synthesis lectures on information concepts, retrieval, and services · DOI 10.1007/978-3-031-02298-2_1 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, B. Barla Cambazoğlu, Ricardo Baeza‐Yates · 3 authors totalScalability Challenges in Web Search Engines
Synthesis lectures on information concepts, retrieval, and services · DOI 10.1007/978-3-031-02298-2 · 77 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, B. Barla Cambazoğlu, Ricardo Baeza‐Yates · 3 authors totalMulti-datacenter Consistency Properties
Encyclopedia of Database Systems · DOI 10.1007/978-1-4899-7993-3_80643-1 · 0 citations · Source: openalex+authoritative-profilePeter Bailis · 1 author totalStructured Document Retrieval
Encyclopedia of Database Systems · DOI 10.1007/978-1-4899-7993-3_378-2 · 10 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Mounia Lalmas, Ricardo Baeza‐Yates · 3 authors totalA Strategic Approach for Managing Oil and Gas Assets
Oil, Gas & Energy Quarterly · DOI 10.1002/GAS.21930 · 0 citations · Source: semantic-scholarBrad Powley, Sang-Won Kim, L. Bertrand · 3 authors totalStory‐focused reading in online news and its potential for user engagement
Journal of the Association for Information Science and Technology · DOI 10.1002/asi.23707 · 23 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Janette Lehmann, Carlos Castillo, Mounia Lalmas, Ricardo Baeza‐Yates · 5 authors totalSpark: Big Data Cluster Computing in Production
Wiley (book) · DOI 10.1002/9781119254805 · 12 citations · Source: semantic-scholarIlya Ganelin, Ema Orhian, Kai Sasaki, Brennon York · 4 authors totalActive Learning for Speech Recognition: the Power of Gradients
NIPS 2016 Workshop · arXiv 1612.03226 · 69 citations · Source: arxiv+semantic-scholarIn training speech recognition systems, labeling audio clips can be expensive, and not all data is equally valuable. Active learning aims to label only the most informative samples to reduce cost. For speech recognition, confidence scores and other likelihood-based active learning methods have been shown to be effective. Gradient-based active learning methods, however, are still not well-understood. This work investigates the Expected Gradient Length (EGL) approach in active learning for end-to-end speech recognition. We justify EGL from a variance reduction perspective, and observe that EGL's measure of informativeness picks novel samples uncorrelated with confidence scores. Experimentally, we show that EGL can reduce word errors by 11\%, or alternatively, reduce the number of samples to label by 50\%, when compared to random sampling.
Sanjeev Satheesh, Jiaji Huang, Rewon Child, Vinay Rao, Hairong Liu, Adam Coates · 6 authors totalProbabilistic Neural Programs
CoRR · arXiv 1612.00712 · Source: dblp+author-first-party+semantic-machines-career-authorityJayant Krishnamurthy, Kenton W. Murray · 2 authors totalDynamic Coattention Networks For Question Answering
International Conference on Learning Representations · arXiv 1611.01604 · 699 citations · Source: semantic-scholarSeveral deep learning models have been proposed for question answering. However, due to their single-pass nature, they have no way to recover from local maxima corresponding to incorrect answers. To address this problem, we introduce the Dynamic Coattention Network (DCN) for question answering. The DCN first fuses co-dependent representations of the question and the document in order to focus on relevant parts of both. Then a dynamic pointing decoder iterates over potential answer spans. This iterative procedure enables the model to recover from initial local maxima corresponding to incorrect answers. On the Stanford question answering dataset, a single DCN model improves the previous state of the art from 71.0% F1 to 75.9%, while a DCN ensemble obtains 80.4% F1.
Richard Socher, Caiming Xiong, Victor Zhong, R. Socher · 4 authors totalQuasi-Recurrent Neural Networks
International Conference on Learning Representations · arXiv 1611.01576 · 474 citations · Source: semantic-scholarRecurrent neural networks are a powerful tool for modeling sequential data, but the dependence of each timestep's computation on the previous timestep's output limits parallelism and makes RNNs unwieldy for very long sequences. We introduce quasi-recurrent neural networks (QRNNs), an approach to neural sequence modeling that alternates convolutional layers, which apply in parallel across timesteps, and a minimalist recurrent pooling function that applies in parallel across channels. Despite lacking trainable recurrent layers, stacked QRNNs have better predictive accuracy than stacked LSTMs of the same hidden size. Due to their increased parallelism, they are up to 16 times faster at train and test time. Experiments on language modeling, sentiment classification, and character-level neural machine translation demonstrate these advantages and underline the viability of QRNNs as a basic building block for a variety of sequence tasks.
Stephen Merity, James Bradbury, Caiming Xiong, R. Socher · 4 authors totalTensorLy: Tensor Learning in Python
Journal of machine learning research · arXiv 1610.09555 · 430 citations · Source: semantic-scholarTensors are higher-order extensions of matrices. While matrix methods form the cornerstone of traditional machine learning and data analysis, tensor methods have been gaining increasing traction. However, software support for tensor operations is not on the same footing. In order to bridge this gap, we have developed TensorLy, a Python library that provides a high-level API for tensor methods and deep tensorized neural networks. TensorLy aims to follow the same standards adopted by the main projects of the Python scientific community, and to seamlessly integrate with them. Its BSD license makes it suitable for both academic and commercial applications. TensorLy's backend system allows users to perform computations with several libraries such as NumPy or PyTorch to name but a few. They can be scaled on multiple CPU or GPU machines. In addition, using the deep-learning frameworks as backend allows to easily design and train deep tensorized neural networks. TensorLy is available at https://github.com/tensorly/tensorly
Jean Kossaifi, Yannis Panagakis, M. Pantic · 3 authors totalTowards a continuous modeling of natural language domains
arXiv (Cornell University) · DOI 10.48550/arxiv.1610.09158 · 2 citations · Source: openalex+career-authorityParsa Ghaffari, Sebastian Ruder, John G. Breslin · 3 authors totalTransfer from Simulation to Real World through Learning Deep Inverse Dynamics Model
arXiv · DOI 10.48550/arXiv.1610.03518 · arXiv 1610.03518 · 272 citations · Source: openalex+arxivDeveloping control policies in simulation is often more practical and safer than directly running experiments in the real world. This applies to policies obtained from planning and optimization, and even more so to policies obtained from reinforcement learning, which is often very data demanding. However, a policy that succeeds in simulation often doesn't work when deployed on a real robot. Nevertheless, often the overall gist of what the policy does in simulation remains valid in the real world. In this paper we investigate such settings, where the sequence of states traversed in simulation remains reasonable for the real world, even if the details of the controls are not, as could be the case when the key differences lie in detailed friction, contact, mass and geometry properties. During execution, at each time step our approach computes what the simulation-based control policy would do, but then, rather than executing these controls on the real robot, our approach computes what the simulation expects the resulting next state(s) will be, and then relies on a learned deep inverse dynamics model to decide which real-world action is most suitable to achieve those next states. Deep models are only as good as their training data, and we also propose an approach for data collection to (incrementally) learn the deep inverse dynamics model. Our experiments shows our approach compares favorably with various baselines that have been developed for dealing with simulation to real world model discrepancy, including output error control and Gaussian dynamics adaptation.
Josh Tobin, Paul Christiano, Zain Shah, Igor Mordatch, Jonas Schneider, Trevor Blackwell, Joshua Tobin, Pieter Abbeel · 8 authors totalLogic as a distributive law
CoRR · arXiv 1610.02247 · Source: dblp+arxiv+career-authorityLucius Gregory Meredith, Mike Stay · 2 authors totalDiscriminative Information Retrieval for Knowledge Discovery
arXiv.org · arXiv 1610.01901 · 0 citations · Source: semantic-scholarWe propose a framework for discriminative Information Retrieval (IR) atop linguistic features, trained to improve the recall of tasks such as answer candidate passage retrieval, the initial step in text-based Question Answering (QA). We formalize this as an instance of linear feature-based IR (Metzler and Croft, 2007), illustrating how a variety of knowledge discovery tasks are captured under this approach, leading to a 44% improvement in recall for candidate triage for QA.
Tongfei Chen, Benjamin Van Durme · 2 authors totalTechnical Report on the CleverHans v2.1.0 Adversarial Examples Library
arXiv 1610.00768 · 546 citations · Source: semantic-scholarCleverHans is a software library that provides standardized reference implementations of adversarial example construction techniques and adversarial training. The library may be used to develop more robust machine learning models and to provide standardized benchmarks of models' performance in the adversarial setting. Benchmarks constructed without a standardized implementation of adversarial example construction are not comparable to each other, because a good result may indicate a robust model or it may merely indicate a weak implementation of the adversarial example construction procedure. This technical report is structured as follows. Section 1 provides an overview of adversarial examples in machine learning and of the CleverHans software. Section 2 presents the core functionalities of the library: namely the attacks based on adversarial examples and defenses to improve the robustness of machine learning models to these attacks. Section 3 describes how to report benchmark results using the library. Section 4 describes the versioning system.
Tom Brown, Nicolas Papernot, Fartash Faghri, Nicholas Carlini, I. Goodfellow, Reuben Feinman, Alexey Kurakin, Cihang Xie · 26 authors totalPointer Sentinel Mixture Models
International Conference on Learning Representations · arXiv 1609.07843 · 4,453 citations · Source: semantic-scholarRecent neural network sequence models with softmax classifiers have achieved their best language modeling performance only with very large hidden states and large vocabularies. Even then they struggle to predict rare or unseen words even if the context makes the prediction unambiguous. We introduce the pointer sentinel mixture architecture for neural sequence models which has the ability to either reproduce a word from the recent context or produce a word from a standard softmax classifier. Our pointer sentinel-LSTM model achieves state of the art language modeling performance on the Penn Treebank (70.9 perplexity) while using far fewer parameters than a standard softmax LSTM. In order to evaluate how well language models can exploit longer contexts and deal with more realistic vocabularies and larger corpora we also introduce the freely available WikiText corpus.
Richard Socher, Stephen Merity, Caiming Xiong, James Bradbury, R. Socher · 5 authors totalCharacter-level and Multi-channel Convolutional Neural Networks for Large-scale Authorship Attribution
arXiv (Cornell University) · DOI 10.48550/arxiv.1609.06686 · 88 citations · Source: openalex+career-authorityParsa Ghaffari, Sebastian Ruder, John G. Breslin · 3 authors totalINSIGHT-1 at SemEval-2016 Task 5: Deep Learning for Multilingual\n Aspect-based Sentiment Analysis
arXiv (Cornell University) · DOI 10.48550/arxiv.1609.02748 · 2 citations · Source: openalex+career-authorityParsa Ghaffari, Sebastian Ruder, John G. Breslin · 3 authors totalINSIGHT-1 at SemEval-2016 Task 4: Convolutional Neural Networks for\n Sentiment Classification and Quantification
arXiv (Cornell University) · DOI 10.48550/arxiv.1609.02746 · 0 citations · Source: openalex+career-authorityParsa Ghaffari, Sebastian Ruder, John G. Breslin · 3 authors totalLearning a Driving Simulator
arXiv (comma.ai technical report) · arXiv 1608.01230 · 232 citations · Source: semantic-scholarcomma.ai's approach to Artificial Intelligence for self-driving cars is based on an agent that learns to clone driver behaviors and plans maneuvers by simulating future events in the road. This paper illustrates one of our research approaches for driving simulation. One where we learn to simulate. Here we investigate variational autoencoders with classical and learned cost functions using generative adversarial networks for embedding road frames. Afterwards, we learn a transition model in the embedded space using action conditioned Recurrent Neural Networks. We show that our approach can keep predicting realistic looking video for several frames despite the transition model being optimized without a cost function in the pixel space.
George Hotz, Eder Santana · 2 authors totalInformation-theoretical label embeddings for large-scale image classification
arXiv · arXiv 1607.05691 · 17 citations · Source: semantic-scholarWe present a method for training multi-label, massively multi-class image classification models, that is faster and more accurate than supervision via a sigmoid cross-entropy loss (logistic regression). Our method consists in embedding high-dimensional sparse labels onto a lower-dimensional dense sphere of unit-normed vectors, and treating the classification problem as a cosine proximity regression problem on this sphere. We test our method on a dataset of 300 million high-resolution images with 17,000 labels, where it yields considerably faster convergence, as well as a 7% higher mean average precision compared to logistic regression.
Francois Chollet, François Chollet · 2 authors totalDSD: Dense-Sparse-Dense Training for Deep Neural Networks
ICLR · arXiv 1607.04381 · 219 citations · Source: arxiv+semantic-scholarModern deep neural networks have a large number of parameters, making them very hard to train. We propose DSD, a dense-sparse-dense training flow, for regularizing deep neural networks and achieving better optimization performance. In the first D (Dense) step, we train a dense network to learn connection weights and importance. In the S (Sparse) step, we regularize the network by pruning the unimportant connections with small weights and retraining the network given the sparsity constraint. In the final D (re-Dense) step, we increase the model capacity by removing the sparsity constraint, re-initialize the pruned parameters from zero and retrain the whole dense network. Experiments show that DSD training can improve the performance for a wide range of CNNs, RNNs and LSTMs on the tasks of image classification, caption generation and speech recognition. On ImageNet, DSD improved the Top1 accuracy of GoogLeNet by 1.1%, VGG-16 by 4.3%, ResNet-18 by 1.2% and ResNet-50 by 1.1%, respectively. On the WSJ'93 dataset, DSD improved DeepSpeech and DeepSpeech2 WER by 2.0% and 1.1%. On the Flickr-8K dataset, DSD improved the NeuralTalk BLEU score by over 1.7. DSD is easy to use in practice: at training time, DSD incurs only one extra hyper-parameter: the sparsity ratio in the S step. At testing time, DSD doesn't change the network architecture or incur any inference overhead. The consistent and significant performance gain of DSD experiments shows the inadequacy of the current training methods for finding the best local optimum, while DSD effectively achieves superior optimization performance for finding a better solution. DSD models are available to download at https://songhan.github.io/DSD.
Erich Elsen, Song Han, Jeff Pool, Sharan Narang, Huizi Mao, Enhao Gong, Shijian Tang, Peter Vajda · 12 authors totalActionable and Political Text Classification using Word Embeddings and LSTM
arXiv preprint · arXiv 1607.02501 · 67 citations · Source: semantic-scholar+arxivIn this work, we apply word embeddings and neural networks with Long Short-Term Memory (LSTM) to text classification problems, where the classification criteria are decided by the context of the application. We examine two applications in particular. The first is that of Actionability, where we build models to classify social media messages from customers of service providers as Actionable or Non-Actionable. We build models for over 30 different languages for actionability, and most of the models achieve accuracy around 85%, with some reaching over 90% accuracy. We also show that using LSTM neural networks with word embeddings vastly outperform traditional techniques. Second, we explore classification of messages with respect to political leaning, where social media messages are classified as Democratic or Republican. The model is able to classify messages with a high accuracy of 87.57%. As part of our experiments, we vary different hyperparameters of the neural networks, and report the effect of such variation on the accuracy. These actionability models have been deployed to production and help company agents provide customer support by prioritizing which messages to respond to. The model for political leaning has been opened and made available for wider use.
Adithya Rao, Nemanja Spasojevic · 2 authors totalModel-Agnostic Interpretability of Machine Learning
arXiv.org · arXiv 1606.05386 · 972 citations · Source: semantic-scholarUnderstanding why machine learning models behave the way they do empowers both system designers and end-users in many ways: in model selection, feature engineering, in order to trust and act upon the predictions, and in more intuitive user interfaces. Thus, interpretability has become a vital concern in machine learning, and work in the area of interpretable models has found renewed interest. In some applications, such models are as accurate as non-interpretable ones, and thus are preferred for their transparency. Even when they are not accurate, they may still be preferred when interpretability is of paramount importance. However, restricting machine learning to interpretable models is often a severe limitation. In this paper we argue for explaining machine learning predictions using model-agnostic approaches. By treating the machine learning models as black-box functions, these approaches provide crucial flexibility in the choice of models, explanations, and representations, improving debugging, comparison, and interfaces for a variety of users and models. We also outline the main challenges for such methods, and review a recently-introduced model-agnostic explanation approach (LIME) that addresses these challenges.
Carlos Guestrin, Sameer Singh, Marco Tulio Ribeiro · 3 authors totalDeepMath - Deep Sequence Models for Premise Selection
NeurIPS 2016 · arXiv 1606.04442 · 267 citations · Source: semantic-scholarWe study the effectiveness of neural sequence models for premise selection in automated theorem proving, one of the main bottlenecks in the formalization of mathematics. We propose a two stage approach for this task that yields good results for the premise selection task on the Mizar corpus while avoiding the hand-engineered features of existing state-of-the-art models. To our knowledge, this is the first time deep learning has been applied to theorem proving on a large scale.
Francois Chollet, G. Irving, Christian Szegedy, Alexander A. Alemi, N. Eén, François Chollet, J. Urban · 7 authors totalCooperative Inverse Reinforcement Learning
Neural Information Processing Systems · arXiv 1606.03137 · 795 citations · Source: semantic-scholarFor an autonomous system to be helpful to humans and to pose no unwarranted risks, it needs to align its values with those of the humans in its environment in such a way that its actions contribute to the maximization of value for the humans. We propose a formal definition of the value alignment problem as cooperative inverse reinforcement learning (CIRL). A CIRL problem is a cooperative, partial-information game with two agents, human and robot; both are rewarded according to the human's reward function, but the robot does not initially know what this is. In contrast to classical IRL, where the human is assumed to act optimally in isolation, optimal CIRL solutions produce behaviors such as active teaching, active learning, and communicative actions that are more effective in achieving value alignment. We show that computing optimal joint policies in CIRL games can be reduced to solving a POMDP, prove that optimality in isolation is suboptimal in CIRL, and derive an approximate CIRL algorithm.
Stuart Russell, Dylan Hadfield-Menell, Stuart J. Russell, P. Abbeel, A. Dragan · 5 authors totalMixing Dirichlet Topic Models and Word Embeddings to Make lda2vec
arXiv preprint · arXiv 1605.02019 · 211 citations · Source: arxiv+semantic-scholarDistributed dense word vectors have been shown to be effective at capturing token-level semantic and syntactic regularities in language, while topic models can form interpretable representations over documents. In this work, we describe lda2vec, a model that learns dense word vectors jointly with Dirichlet-distributed latent document-level mixtures of topic vectors. In contrast to continuous dense document representations, this formulation produces sparse, interpretable document mixtures through a non-negative simplex constraint. Our method is simple to incorporate into existing automatic differentiation frameworks and allows for unsupervised document representations geared for use by scientists while simultaneously learning word vectors and the linear relationships between them.
Christopher Erick Moody, Christopher E Moody · 2 authors totalTraining Deep Nets with Sublinear Memory Cost
arXiv.org · arXiv 1604.06174 · 1,535 citations · Source: semantic-scholarWe propose a systematic approach to reduce the memory consumption of deep neural network training. Specifically, we design an algorithm that costs O( √ n) memory to train a n layer network, with only the computational cost of an extra forward pass per mini-batch. As many of the state-of-the-art models hit the upper bound of the GPU memory, our algorithm allows deeper and more complex models to be explored, and helps advance the innovations in deep learning research. We focus on reducing the memory cost to store the intermediate feature maps and gradients during training. Computation graph analysis is used for automatic in-place operation and memory sharing optimizations. We show that it is possible to trade computation for memory giving a more memory efficient training algorithm with a little extra computation cost. In the extreme case, our analysis also shows that the memory consumption can be reduced to O(logn) with as little as O(n logn) extra cost for forward computation. Our experiments show that we can reduce the memory cost of a 1,000-layer deep residual network from 48G to 7G on ImageNet problems. Similarly, significant memory cost reduction is observed in training complex recurrent neural networks on very long sequences.
Carlos Guestrin, Tianqi Chen, Bing Xu, Chiyuan Zhang · 4 authors total