Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗Meta-Graph: Few Shot Link Prediction via Meta Learning
arXiv · DOI 10.48550/arxiv.1912.09867 · arXiv 1912.09867 · 30 citations · Source: openalexWe consider the task of few shot link prediction on graphs. The goal is to learn from a distribution over graphs so that a model is able to quickly infer missing edges in a new graph after a small amount of training. We show that current link prediction methods are generally ill-equipped to handle this task. They cannot effectively transfer learned knowledge from one graph to another and are unable to effectively learn from sparse samples of edges. To address this challenge, we introduce a new gradient-based meta learning framework, Meta-Graph. Our framework leverages higher-order gradients along with a learned graph signature function that conditionally generates a graph neural network initialization. Using a novel set of few shot link prediction benchmarks, we show that Meta-Graph can learn to quickly adapt to a new graph using only a small sample of true edges, enabling not only fast adaptation but also improved results at convergence.
Piero Molino, Avishek Joey Bose, Ankit Kumar Jain, William L. Hamilton · 4 authors totalStatistically Robust Neural Network Classification
UAI 2021 · arXiv 1912.04884 · 23 citations · Source: arxiv+semantic-scholarRecently there has been much interest in quantifying the robustness of neural network classifiers through adversarial risk metrics. However, for problems where test-time corruptions occur in a probabilistic manner, rather than being generated by an explicit adversary, adversarial metrics typically do not provide an accurate or reliable indicator of robustness. To address this, we introduce a statistically robust risk (SRR) framework which measures robustness in expectation over both network inputs and a corruption distribution. Unlike many adversarial risk metrics, which typically require separate applications on a point-by-point basis, the SRR can easily be directly estimated for an entire network and used as a training objective in a stochastic gradient scheme. Furthermore, we show both theoretically and empirically that it can scale to higher-dimensional networks by providing superior generalization performance compared with comparable adversarial risks.
Stefan Webb, Benjie Wang, Tom Rainforth · 3 authors totalPlug and Play Language Models: A Simple Approach to Controlled Text\n Generation
arXiv · DOI 10.48550/arxiv.1912.02164 · arXiv 1912.02164 · 407 citations · Source: openalexLarge transformer-based language models (LMs) trained on huge text corpora\nhave shown unparalleled generation capabilities. However, controlling\nattributes of the generated language (e.g. switching topic or sentiment) is\ndifficult without modifying the model architecture or fine-tuning on\nattribute-specific data and entailing the significant cost of retraining. We\npropose a simple alternative: the Plug and Play Language Model (PPLM) for\ncontrollable language generation, which combines a pretrained LM with one or\nmore simple attribute classifiers that guide text generation without any\nfurther training of the LM. In the canonical scenario we present, the attribute\nmodels are simple classifiers consisting of a user-specified bag of words or a\nsingle learned layer with 100,000 times fewer parameters than the LM. Sampling\nentails a forward and backward pass in which gradients from the attribute model\npush the LM's hidden activations and thus guide the generation. Model samples\ndemonstrate control over a range of topics and sentiment styles, and extensive\nautomated and human annotated evaluations show attribute alignment and fluency.\nPPLMs are flexible in that any combination of differentiable attribute models\nmay be used to steer text generation, which will allow for diverse and creative\napplications beyond the examples given in this paper.\n
Piero Molino, Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Jason Yosinski, Rosanne Liu · 8 authors totalPyTorch: An Imperative Style, High-Performance Deep Learning Library
NeurIPS · arXiv 1912.01703 · Source: neurips+arxiv+dblp+pytorch-first-partyGregory Chanan, Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Trevor Killeen, Zeming Lin · 21 authors totalJust Ask: An Interactive Learning Framework for Vision and Language Navigation
CoRR · arXiv 1912.00915 · Source: dblp+first-party-homepageMihail Eric, Ta-Chung Chi, Seokhwan Kim, Minmin Shen, Dilek Hakkani-Tür · 5 authors totalSingle Headed Attention RNN: Stop Thinking With Your Head
arXiv.org · arXiv 1911.11423 · 71 citations · Source: semantic-scholarThe leading approaches in language modeling are all obsessed with TV shows of my youth - namely Transformers and Sesame Street. Transformers this, Transformers that, and over here a bonfire worth of GPU-TPU-neuromorphic wafer scale silicon. We opt for the lazy path of old and proven techniques with a fancy crypto inspired acronym: the Single Headed Attention RNN (SHA-RNN). The author's lone goal is to show that the entire field might have evolved a different direction if we had instead been obsessed with a slightly different acronym and slightly different result. We take a previously strong language model based only on boring LSTMs and get it to within a stone's throw of a stone's throw of state-of-the-art byte level language model results on enwik8. This work has undergone no intensive hyperparameter optimization and lived entirely on a commodity desktop machine that made the author's small studio apartment far too warm in the midst of a San Franciscan summer. The final results are achievable in plus or minus 24 hours on a single GPU as the author is impatient. The attention mechanism is also readily extended to large contexts with minimal computation. Take that Sesame Street.
Stephen Merity · 1 author totalRigging the Lottery: Making All Tickets Winners
ICML · arXiv 1911.11134 · 771 citations · Source: arxiv+semantic-scholarMany applications require sparse neural networks due to space or inference time restrictions. There is a large body of work on training dense networks to yield sparse networks for inference, but this limits the size of the largest trainable sparse model to that of the largest trainable dense model. In this paper we introduce a method to train sparse neural networks with a fixed parameter count and a fixed computational cost throughout training, without sacrificing accuracy relative to existing dense-to-sparse training methods. Our method updates the topology of the sparse network during training by using parameter magnitudes and infrequent gradient calculations. We show that this approach requires fewer floating-point operations (FLOPs) to achieve a given level of accuracy compared to prior techniques. We demonstrate state-of-the-art sparse training results on a variety of networks and datasets, including ResNet-50, MobileNets on Imagenet-2012, and RNNs on WikiText-103. Finally, we provide some insights into why allowing the topology to change during the optimization can overcome local minima encountered when the topology remains static. Code used in our work can be found in github.com/google-research/rigl.
Erich Elsen, Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro · 5 authors totalSAVEHR: Self Attention Vector Representations for EHR based Personalized Chronic Disease Onset Prediction and Interpretability
arXiv.org · arXiv 1911.05370 · 2 citations · Source: semantic-scholarChronic disease progression is emerging as an important area of investment for healthcare providers. As the quantity and richness of available clinical data continue to increase along with advances in machine learning, there is great potential to advance our approaches to caring for patients. An ideal approach to this problem should generate good performance on at least three axes namely, a) perform across many clinical conditions without requiring deep clinical expertise or extensive data scientist effort, b) generalization across populations, and c) be explainable (model interpretability). We present SAVEHR, a self-attention based architecture on heterogeneous structured EHR data that achieves $>$ 0.51 AUC-PR and $>$ 0.87 AUC-ROC gains on predicting the onset of four clinical conditions (CHF, Kidney Failure, Diabetes and COPD) 15-months in advance, and transfers with high performance onto a new population. We demonstrate that SAVEHR model performs superior to ten baselines on all three axes stated formerly.
Sunil Mallya, S. Mallya, J. Overhage, S. Bodapati, Navneet Srivastava, Sahika Genc · 6 authors totalImproving Robustness of Task Oriented Dialog Systems
arXiv · arXiv 1911.05153 · 22 citations · Source: semantic-scholarTask oriented language understanding in dialog systems is often modeled using intents (task of a query) and slots (parameters for that task). Intent detection and slot tagging are, in turn, modeled using sentence classification and word tagging techniques respectively. Similar to adversarial attack problems with computer vision models discussed in existing literature, these intent-slot tagging models are often over-sensitive to small variations in input -- predicting different and often incorrect labels when small changes are made to a query, thus reducing their accuracy and reliability. However, evaluating a model's robustness to these changes is harder for language since words are discrete and an automated change (e.g. adding `noise') to a query sometimes changes the meaning and thus labels of a query. In this paper, we first describe how to create an adversarial test set to measure the robustness of these models. Furthermore, we introduce and adapt adversarial training methods as well as data augmentation using back-translation to mitigate these issues. Our experiments show that both techniques improve the robustness of the system substantially and can be combined to yield the best results.
Sonal Gupta, Arash Einolghozati, Mrinal Mohit, Rushin Shah · 4 authors totalGeometry-Aware Neural Rendering
NeurIPS · DOI 10.48550/arXiv.1911.04554 · arXiv 1911.04554 · 25 citations · Source: openalex+semantic-scholarUnderstanding the 3-dimensional structure of the world is a core challenge in computer vision and robotics. Neural rendering approaches learn an implicit 3D model by predicting what a camera would see from an arbitrary viewpoint. We extend existing neural rendering to more complex, higher dimensional scenes than previously possible. We propose Epipolar Cross Attention (ECA), an attention mechanism that leverages the geometry of the scene to perform efficient non-local operations, requiring only $O(n)$ comparisons per spatial dimension instead of $O(n^2)$. We introduce three new simulated datasets inspired by real-world robotics and demonstrate that ECA significantly improves the quantitative and qualitative performance of Generative Query Networks (GQN).
Josh Tobin, Joshua Tobin, OpenAI Robotics, Pieter Abbeel · 4 authors totalDeepRacer: Educational Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning
arXiv.org · arXiv 1911.01562 · 71 citations · Source: semantic-scholarDeepRacer is a platform for end-to-end experimentation with RL and can be used to systematically investigate the key challenges in developing intelligent control systems. Using the platform, we demonstrate how a 1/18th scale car can learn to drive autonomously using RL with a monocular camera. It is trained in simulation with no additional tuning in physical world and demonstrates: 1) formulation and solution of a robust reinforcement learning algorithm, 2) narrowing the reality gap through joint perception and dynamics, 3) distributed on-demand compute architecture for training optimal policies, and 4) a robust evaluation method to identify when to stop training. It is the first successful large-scale deployment of deep reinforcement learning on a robotic control agent that uses only raw camera images as observations and a model-free learning method to perform robust path planning. We open source our code and video demo on GitHub: this https URL.
Sunil Mallya, Bharathan Balaji, S. Mallya, Sahika Genc, Saurabh Gupta, Leo Dirac, Vineet Khare, Gourav Roy · 13 authors totalOn the Measure of Intelligence
arXiv · arXiv 1911.01547 · Source: arxivTo make deliberate progress towards more intelligent and more human-like artificial systems, we need to be following an appropriate feedback signal: we need to be able to define and evaluate intelligence in a way that enables comparisons between two systems, as well as comparisons with humans. ... we articulate a new formal definition of intelligence based on Algorithmic Information Theory, describing intelligence as skill-acquisition efficiency ... we introduce the ARC benchmark.
Francois Chollet · 1 author totalFast Structured Decoding for Sequence Models
NeurIPS 2019 · arXiv 1910.11555 · 136 citations · Source: arxiv+semantic-scholarAutoregressive sequence models achieve state-of-the-art performance in domains like machine translation. However, due to the autoregressive factorization nature, these models suffer from heavy latency during inference. Recently, non-autoregressive sequence models were proposed to reduce the inference time. However, these models assume that the decoding process of each token is conditionally independent of others. Such a generation process sometimes makes the output sentence inconsistent, and thus the learned non-autoregressive models could only achieve inferior accuracy compared to their autoregressive counterparts. To improve then decoding consistency and reduce the inference cost at the same time, we propose to incorporate a structured inference module into the non-autoregressive models. Specifically, we design an efficient approximation for Conditional Random Fields (CRF) for non-autoregressive sequence models, and further propose a dynamic transition technique to model positional contexts in the CRF. Experiments in machine translation show that while increasing little latency (8~14ms), our model could achieve significantly better translation performance than previous non-autoregressive models on different translation datasets. In particular, for the WMT14 En-De dataset, our model obtains a BLEU score of 26.80, which largely outperforms the previous non-autoregressive baselines and is only 0.61 lower in BLEU than purely autoregressive models.
Zhuohan Li, Zhiqing Sun, Haoqing Wang, Zi Lin, Di He, Zhihong Deng · 6 authors totalQuantum Hamiltonian-Based Models and the Variational Quantum Thermalizer Algorithm
arXiv · arXiv 1910.02071 · Source: arxiv+dblp+stanford-career-authorityJacob Marks, Guillaume Verdon, Sasha Nanda, Stefan Leichenauer, Jack Hidary · 5 authors totalMLPerf Training Benchmark
Conference on Machine Learning and Systems · DOI 10.48550/arxiv.1910.01500 · arXiv 1910.01500 · 398 citations · Source: semantic-scholar+openalexMachine learning (ML) needs industry-standard performance benchmarks to support design and competitive evaluation of the many emerging software and hardware solutions for ML. But ML training presents three unique benchmarking challenges absent from other domains: optimizations that improve training throughput can increase the time to solution, training is stochastic and time to solution exhibits high variance, and software and hardware systems are so diverse that fair benchmarking with the same binary, code, and even hyperparameters is difficult. We therefore present MLPerf, an ML benchmark that overcomes these challenges. Our analysis quantitatively evaluates MLPerf's efficacy at driving performance and scalability improvements across two rounds of results from multiple vendors.
Matei Zaharia, Peter Bailis, Peter Mattson, Christine Cheng, Cody A. Coleman, G. Diamos, P. Micikevicius, David A. Patterson · 34 authors totalGradient Descent: The Ultimate Optimizer
arXiv (Cornell University) · DOI 10.48550/arxiv.1909.13371 · arXiv 1909.13371 · 20 citations · Source: openalex+semantic-scholarWorking with any gradient-based machine learning algorithm involves the tedious task of tuning the optimizer's hyperparameters, such as its step size. Recent work has shown how the step size can itself be optimized alongside the model parameters by manually deriving expressions for "hypergradients" ahead of time. We show how to automatically compute hypergradients with a simple and elegant modification to backpropagation. This allows us to easily apply the method to other optimizers and hyperparameters (e.g. momentum coefficients). We can even recursively apply the method to its own hyper-hyperparameters, and so on ad infinitum. As these towers of optimizers grow taller, they become less sensitive to the initial choice of hyperparameters. We present experiments validating this for MLPs, CNNs, and RNNs. Finally, we provide a simple PyTorch implementation of this algorithm (see people.csail.mit.edu/kach/gradient-descent-the-ultimate-optimizer).
Erik Meijer, Kartik Chandra, Xie, Audrey, Ragan-Kelley, Jonathan · 4 authors totalHigh Fidelity Speech Synthesis with Adversarial Networks
ICLR · arXiv 1909.11646 · 264 citations · Source: arxiv+semantic-scholarGenerative adversarial networks have seen rapid development in recent years and have led to remarkable improvements in generative modelling of images. However, their application in the audio domain has received limited attention, and autoregressive models, such as WaveNet, remain the state of the art in generative modelling of audio signals such as human speech. To address this paucity, we introduce GAN-TTS, a Generative Adversarial Network for Text-to-Speech. Our architecture is composed of a conditional feed-forward generator producing raw speech audio, and an ensemble of discriminators which operate on random windows of different sizes. The discriminators analyse the audio both in terms of general realism, as well as how well the audio corresponds to the utterance that should be pronounced. To measure the performance of GAN-TTS, we employ both subjective human evaluation (MOS - Mean Opinion Score), as well as novel quantitative metrics (Fréchet DeepSpeech Distance and Kernel DeepSpeech Distance), which we find to be well correlated with MOS. We show that GAN-TTS is capable of generating high-fidelity speech with naturalness comparable to the state-of-the-art models, and unlike autoregressive models, it is highly parallelisable thanks to an efficient feed-forward generator. Listen to GAN-TTS reading this abstract at https://storage.googleapis.com/deepmind-media/research/abstract.wav.
Erich Elsen, Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Norman Casagrande, Luis C. Cobo, Karen Simonyan · 8 authors totalFine-Tuning Language Models from Human Preferences
arXiv.org · arXiv 1909.08593 · 2,691 citations · Source: semantic-scholarReward learning enables the application of reinforcement learning (RL) to tasks where reward is defined by human judgment, building a model of reward by asking humans questions. Most work on reward learning has used simulated environments, but complex information about values is often expressed in natural language, and we believe reward learning for language is a key to making RL practical and safe for real-world tasks. In this paper, we build on advances in generative pretraining of language models to apply reward learning to four natural language tasks: continuing text with positive sentiment or physically descriptive language, and summarization tasks on the TL;DR and CNN/Daily Mail datasets. For stylistic continuation we achieve good results with only 5,000 comparisons evaluated by humans. For summarization, models trained with 60,000 comparisons copy whole sentences from the input but skip irrelevant preamble; this leads to reasonable ROUGE scores and very good performance according to our human labelers, but may be exploiting the fact that labelers rely on simple heuristics.
Tom Brown, Daniel M. Ziegler, Nisan Stiennon, Jeff Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano · 8 authors totalLudwig: a type-based declarative deep learning toolbox
arXiv · DOI 10.48550/arxiv.1909.07930 · arXiv 1909.07930 · 28 citations · Source: openalexIn this work we present Ludwig, a flexible, extensible and easy to use toolbox which allows users to train deep learning models and use them for obtaining predictions without writing code. Ludwig implements a novel approach to deep learning model building based on two main abstractions: data types and declarative configuration files. The data type abstraction allows for easier code and sub-model reuse, and the standardized interfaces imposed by this abstraction allow for encapsulation and make the code easy to extend. Declarative model definition configuration files enable inexperienced users to obtain effective models and increase the productivity of expert users. Alongside these two innovations, Ludwig introduces a general modularized deep learning architecture called Encoder-Combiner-Decoder that can be instantiated to perform a vast amount of machine learning tasks. These innovations make it possible for engineers, scientists from other fields and, in general, a much broader audience to adopt deep learning models for their tasks, concretely helping in its democratization.
Piero Molino, Y. O. Dudin, Sai Sumanth Miryala · 3 authors totalEfficient 2.5D Hand Pose Estimation via Auxiliary Multi-Task Training for Embedded Devices
arXiv preprint (Magic Leap) · arXiv 1909.05897 · 3 citations · Source: arxiv+dblp2D Key-point estimation is an important precursor to 3D pose estimation problems for human body and hands. In this work, we discuss the data, architecture, and training procedure necessary to deploy extremely efficient 2.5D hand pose estimation on embedded devices with highly constrained memory and compute envelope, such as AR/VR wearables. Our 2.5D hand pose estimation consists of 2D key-point estimation of joint positions on an egocentric image, captured by a depth sensor, and lifted to 2.5D using the corresponding depth values. Our contributions are two fold: (a) We discuss data labeling and augmentation strategies, the modules in the network architecture that collectively lead to $3\%$ the flop count and $2\%$ the number of parameters when compared to the state of the art MobileNetV2 architecture. (b) We propose an auxiliary multi-task training strategy needed to compensate for the small capacity of the network while achieving comparable performance to MobileNetV2. Our 32-bit trained model has a memory footprint of less than 300 Kilobytes, operates at more than 50 Hz with less than 35 MFLOPs.
Adithya Rao, Prajwal Chidananda, Ayan Sinha, Douglas Lee, A. Rabinovich · 5 authors totalCTRL: A Conditional Transformer Language Model for Controllable Generation
arXiv.org · arXiv 1909.05858 · 1,466 citations · Source: semantic-scholarLarge-scale language models show promising text generation capabilities, but users cannot easily control particular aspects of the generated text. We release CTRL, a 1.63 billion-parameter conditional transformer language model, trained to condition on control codes that govern style, content, and task-specific behavior. Control codes were derived from structure that naturally co-occurs with raw text, preserving the advantages of unsupervised learning while providing more explicit control over text generation. These codes also allow CTRL to predict which parts of the training data are most likely given a sequence. This provides a potential method for analyzing large amounts of data via model-based source attribution. We have released multiple full-sized, pretrained versions of CTRL at this https URL.
Richard Socher, N. Keskar, Bryan McCann, L. Varshney, Caiming Xiong, R. Socher · 6 authors totalAvaya Conversational Intelligence: A Real-Time System for Spoken Language Understanding in Human-Human Call Center Conversations
Interspeech 2019 (Show & Tell) · arXiv 1909.02851 · 2 citations · Source: semantic-scholarAvaya Conversational Intelligence(ACI) is an end-to-end, cloud-based solution for real-time Spoken Language Understanding for call centers. It combines large vocabulary, real-time speech recognition, transcript refinement, and entity and intent recognition in order to convert live audio into a rich, actionable stream of structured events. These events can be further leveraged with a business rules engine, thus serving as a foundation for real-time supervision and assistance applications. After the ingestion, calls are enriched with unsupervised keyword extraction, abstractive summarization, and business-defined attributes, enabling offline use cases, such as business intelligence, topic mining, full-text search, quality assurance, and agent training. ACI comes with a pretrained, configurable library of hundreds of intents and a robust intent training environment that allows for efficient, cost-effective creation and customization of customer-specific intents.
Yishay Carmiel, Jan Mizgajski, Adrian Szymczak, R. Glowski, Piotr Szymański, Piotr Żelasko, Lukasz Augustyniak, Mikolaj Morzy · 17 authors totalCloudy with high chance of DBMS: A 10-year prediction for Enterprise-Grade ML
CIDR 2020 · arXiv 1909.00084 · 44 citations · Source: arxivAshvin Agrawal, Rony Chatterjee, Carlo Curino, Avrilia Floratou, Neha Godwal, Matteo Interlandi, Alekh Jindal, Konstantinos Karanasos · 22 authors totalOpenSpiel: A Framework for Reinforcement Learning in Games
arXiv preprint · DOI 10.48550/arXiv.1908.09453 · arXiv 1908.09453 · 332 citations · Source: arxiv+dblpOpenSpiel is a collection of environments and algorithms for research in general reinforcement learning and search/planning in games. OpenSpiel supports n-player (single- and multi- agent) zero-sum, cooperative and general-sum, one-shot and sequential, strictly turn-taking and simultaneous-move, perfect and imperfect information games, as well as traditional multiagent environments such as (partially- and fully- observable) grid worlds and social dilemmas. OpenSpiel also includes tools to analyze learning dynamics and other common evaluation metrics. This document serves both as an overview of the code base and an introduction to the terminology, core concepts, and algorithms across the fields of reinforcement learning, computational game theory, and search.
Brennan Saeta, Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, Vinicius Zambaldi, Satyaki Upadhyay, Julien Perolat, Sriram Srinivasan · 12 authors totalRelease Strategies and the Social Impacts of Language Models
arXiv · arXiv 1908.09203 · Source: arxiv+dblp+openai-career-authorityIrene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Jasmine Wang · 8 authors totalOn the importance of system-view centric validation for the design and operation of a crypto-based digital economy
arXiv · DOI 10.48550/arXiv.1908.08675 · arXiv 1908.08675 · 5 citations · Source: arxiv+openalex+bosch-career-authorityNik Scharmann, Alexander Poddey · 2 authors totalTesting Robustness Against Unforeseen Adversaries
arXiv.org · arXiv 1908.08016 · 145 citations · Source: semantic-scholarTom Brown, Daniel Kang, Yi Sun, Dan Hendrycks, Tom B. Brown, J. Steinhardt · 6 authors totalTowards Better Understanding of Spontaneous Conversations: Overcoming Automatic Speech Recognition Errors With Intent Recognition
arXiv (Interspeech 2019 track) · arXiv 1908.07888 · 0 citations · Source: semantic-scholarIn this paper, we present a method for correcting automatic speech recognition (ASR) errors using a finite state transducer (FST) intent recognition framework. Intent recognition is a powerful technique for dialog flow management in turn-oriented, human-machine dialogs. This technique can also be very useful in the context of human-human dialogs, though it serves a different purpose of key insight extraction from conversations. We argue that currently available intent recognition techniques are not applicable to human-human dialogs due to the complex structure of turn-taking and various disfluencies encountered in spontaneous conversations, exacerbated by speech recognition errors and scarcity of domain-specific labeled data. Without efficient key insight extraction techniques, raw human-human dialog transcripts remain significantly unexploited. Our contribution consists of a novel FST for intent indexing and an algorithm for fuzzy intent search over the lattice - a compact graph encoding of ASR's hypotheses. We also develop a pruning strategy to constrain the fuzziness of the FST index search. Extracted intents represent linguistic domain knowledge and help us improve (rescore) the original transcript. We compare our method with a baseline, which uses only the most likely transcript hypothesis (best path), and find an increase in the total number of recognized intents by 25%.
Yishay Carmiel, Piotr Żelasko, Jan Mizgajski, Mikolaj Morzy, Adrian Szymczak, Piotr Szymański, Lukasz Augustyniak · 7 authors totalCovering up bias in CelebA-like datasets with Markov blankets: A post-hoc cure for attribute prior avoidance
arXiv (Cornell University) · DOI 10.48550/arxiv.1907.12917 · 3 citations · Source: openalex+orcid+dblp-identityJohn Whaley, Vinay Uday Prabhu, Dian Ang Yap, Alexander Wang · 4 authors totalFoundational Patterns for Efficient Quantum Computing
arXiv · arXiv 1907.11513 · Source: arxiv+jpmorgan-career-authorityConstantin Gonciulea, Austin Gilliam, Charlene Venci, Sreraman Muralidharan, Vitaliy Dorum, Eric May, Rajesh Narasimhan · 7 authors totalUnderstanding Adversarial Robustness Through Loss Landscape Geometries
arXiv (Cornell University) · DOI 10.48550/arxiv.1907.09061 · 11 citations · Source: openalex+orcid+dblp-identityJohn Whaley, Vinay Uday Prabhu, Dian Ang Yap, Joyce Xu · 4 authors totalCollaborative Multi-Agent Dialogue Model Training Via Reinforcement\n Learning
arXiv · DOI 10.48550/arxiv.1907.05507 · arXiv 1907.05507 · 0 citations · Source: openalexWe present the first complete attempt at concurrently training conversational\nagents that communicate only via self-generated language. Using DSTC2 as seed\ndata, we trained natural language understanding (NLU) and generation (NLG)\nnetworks for each agent and let the agents interact online. We model the\ninteraction as a stochastic collaborative game where each agent (player) has a\nrole ("assistant", "tourist", "eater", etc.) and their own objectives, and can\nonly interact via natural language they generate. Each agent, therefore, needs\nto learn to operate optimally in an environment with multiple sources of\nuncertainty (its own NLU and NLG, the other agent's NLU, Policy, and NLG). In\nour evaluation, we show that the stochastic-game agents outperform deep\nlearning based supervised baselines.\n
Piero Molino, Alexandros Papangelis, Yi‐Chia Wang, Gökhan Tür · 4 authors totalApplying a Pre-trained Language Model to Spanish Twitter Humor Prediction
IberLEF@SEPLN 2019 · arXiv 1907.03187 · 7 citations · Source: semantic-scholarOur entry into the HAHA 2019 Challenge placed $3^{rd}$ in the classification task and $2^{nd}$ in the regression task. We describe our system and innovations, as well as comparing our results to a Naive Bayes baseline. A large Twitter based corpus allowed us to train a language model from scratch focused on Spanish and transfer that knowledge to our competition model. To overcome the inherent errors in some labels we reduce our class confidence with label smoothing in the loss function. All the code for our project is included in a GitHub repository for easy reference and to enable replication by others.
Jeremy Howard, Bobak Farzin, Piotr Czapla · 3 authors totalMultiWOZ 2.1: Multi-Domain Dialogue State Corrections and State Tracking Baselines
CoRR · arXiv 1907.01669 · Source: dblp+first-party-homepageMihail Eric, Rahul Goel, Shachi Paul, Abhishek Sethi, Sanchit Agarwal, Shuyang Gao, Dilek Hakkani-Tür · 7 authors totalSelection via Proxy: Efficient Data Selection for Deep Learning
arXiv (Cornell University) · DOI 10.48550/arxiv.1906.11829 · 76 citations · Source: openalex+authoritative-profilePeter Bailis, Cody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman, Percy Liang, Jure Leskovec, Matei Zaharia · 8 authors totalThe Difficulty of Training Sparse Neural Networks
arXiv · arXiv 1906.10732 · 110 citations · Source: arxiv+semantic-scholarWe investigate the difficulties of training sparse neural networks and make new observations about optimization dynamics and the energy landscape within the sparse regime. Recent work of \citep{Gale2019, Liu2018} has shown that sparse ResNet-50 architectures trained on ImageNet-2012 dataset converge to solutions that are significantly worse than those found by pruning. We show that, despite the failure of optimizers, there is a linear path with a monotonically decreasing objective from the initialization to the "good" solution. Additionally, our attempts to find a decreasing objective path from "bad" solutions to the "good" ones in the sparse subspace fail. However, if we allow the path to traverse the dense subspace, then we consistently find a path between two solutions. These findings suggest traversing extra dimensions may be needed to escape stationary points found in the sparse subspace.
Erich Elsen, Utku Evci, Fabian Pedregosa, Aidan Gomez · 4 authors totalNon-Differentiable Supervised Learning with Evolution Strategies and Hybrid Methods
arXiv · arXiv 1906.03139 · 22 citations · Source: arxiv+semantic-scholarIn this work we show that Evolution Strategies (ES) are a viable method for learning non-differentiable parameters of large supervised models. ES are black-box optimization algorithms that estimate distributions of model parameters; however they have only been used for relatively small problems so far. We show that it is possible to scale ES to more complex tasks and models with millions of parameters. While using ES for differentiable parameters is computationally impractical (although possible), we show that a hybrid approach is practically feasible in the case where the model has both differentiable and non-differentiable parameters. In this approach we use standard gradient-based methods for learning differentiable weights, while using ES for learning non-differentiable parameters - in our case sparsity masks of the weights. This proposed method is surprisingly competitive, and when parallelized over multiple devices has only negligible training time overhead compared to training with gradient descent. Additionally, this method allows to train sparse models from the first training step, so they can be much larger than when using methods that require training dense models first. We present results and analysis of supervised feed-forward models (such as MNIST and CIFAR-10 classification), as well as recurrent models, such as SparseWaveRNN for text-to-speech.
Erich Elsen, Karel Lenc, Tom Schaul, Karen Simonyan · 4 authors totalUnderstanding and Improving Transformer From a Multi-Particle Dynamic System Point of View
arXiv.org · arXiv 1906.02762 · 237 citations · Source: arxiv+semantic-scholarThe Transformer architecture is widely used in natural language processing. Despite its success, the design principle of the Transformer remains elusive. In this paper, we provide a novel perspective towards understanding the architecture: we show that the Transformer can be mathematically interpreted as a numerical Ordinary Differential Equation (ODE) solver for a convection-diffusion equation in a multi-particle dynamic system. In particular, how words in a sentence are abstracted into contexts by passing through the layers of the Transformer can be interpreted as approximating multiple particles' movement in the space using the Lie-Trotter splitting scheme and the Euler's method. Given this ODE's perspective, the rich literature of numerical analysis can be brought to guide us in designing effective structures beyond the Transformer. As an example, we propose to replace the Lie-Trotter splitting scheme by the Strang-Marchuk splitting scheme, a scheme that is more commonly used and with much lower local truncation errors. The Strang-Marchuk splitting scheme suggests that the self-attention and position-wise feed-forward network (FFN) sub-layers should not be treated equally. Instead, in each layer, two position-wise FFN sub-layers should be used, and the self-attention sub-layer is placed in between. This leads to a brand new architecture. Such an FFN-attention-FFN layer is "Macaron-like", and thus we call the network with this new architecture the Macaron Net. Through extensive experiments, we show that the Macaron Net is superior to the Transformer on both supervised and unsupervised learning tasks. The reproducible codes and pretrained models can be found at this https URL
Zhuohan Li, Yiping Lu, Di He, Zhiqing Sun, Bin Dong, Tao Qin, Liwei Wang, Tie-Yan Liu · 8 authors totalWillump: A Statistically-Aware End-to-end Optimizer for Machine Learning Inference
arXiv (Cornell University) · DOI 10.48550/arxiv.1906.01974 · 8 citations · Source: openalex+authoritative-profilePeter Bailis, Peter Kraft, Daniel Kang, Deepak Narayanan, Shoumik Palkar, Matei Zaharia · 6 authors totalAdversarial Policies: Attacking Deep Reinforcement Learning
International Conference on Learning Representations · arXiv 1905.10615 · 441 citations · Source: semantic-scholarDeep reinforcement learning (RL) policies are known to be vulnerable to adversarial perturbations to their observations, similar to adversarial examples for classifiers. However, an attacker is not usually able to directly modify another agent's observations. This might lead one to wonder: is it possible to attack an RL agent simply by choosing an adversarial policy acting in a multi-agent environment so as to create natural observations that are adversarial? We demonstrate the existence of adversarial policies in zero-sum games between simulated humanoid robots with proprioceptive observations, against state-of-the-art victims trained via self-play to be robust to opponents. The adversarial policies reliably win against the victims but generate seemingly random and uncoordinated behavior. We find that these policies are more successful in high-dimensional environments, and induce substantially different activations in the victim policy network than when the victim plays against a normal opponent. Videos are available at this https URL.
Stuart Russell, Adam Gleave, Michael Dennis, Neel Kant, Cody Wild, S. Levine, Stuart J. Russell · 7 authors totalDesigning a Symbolic Intermediate Representation for Neural Surface Realization
CoRR · arXiv 1905.10486 · Source: dblp+adapt-autodesk-authorityAlex O'Connor, Henry Elder, Jennifer Foster, James Barry, Alexander O'Connor · 5 authors totalFonts-2-Handwriting: A Seed-Augment-Train framework for universal digit classification
arXiv (Cornell University) · DOI 10.48550/arxiv.1905.08633 · 5 citations · Source: openalex+orcid+dblp-identityJohn Whaley, Vinay Uday Prabhu, Sanghyun Han, Dian Ang Yap, Mihail Douhaniaris, Preethi Seshadri · 6 authors totalCrossTrainer: Practical Domain Adaptation with Loss Reweighting
arXiv (Cornell University) · DOI 10.48550/arxiv.1905.02304 · 1 citations · Source: openalex+authoritative-profilePeter Bailis, Justin Chen, Edward Gan, Kexin Rong, Sahaana Suri · 5 authors totalTransfer of Adversarial Robustness Between Perturbation Types
arXiv.org · arXiv 1905.01034 · 54 citations · Source: semantic-scholarWe study the transfer of adversarial robustness of deep neural networks between different perturbation types. While most work on adversarial examples has focused on $L_\infty$ and $L_2$-bounded perturbations, these do not capture all types of perturbations available to an adversary. The present work evaluates 32 attacks of 5 different types against models adversarially trained on a 100-class subset of ImageNet. Our empirical results suggest that evaluating on a wide range of perturbation sizes is necessary to understand whether adversarial robustness transfers between perturbation types. We further demonstrate that robustness against one perturbation type may not always imply and may sometimes hurt robustness against other perturbation types. In light of these results, we recommend evaluation of adversarial defenses take place on a diverse range of perturbation types and sizes.
Tom Brown, Daniel Kang, Yi Sun, Tom B. Brown, Dan Hendrycks, J. Steinhardt · 6 authors totalRouting Networks and the Challenges of Modular and Compositional Computation
arXiv.org · arXiv 1904.12774 · 94 citations · Source: semantic-scholarCompositionality is a key strategy for addressing combinatorial complexity and the curse of dimensionality. Recent work has shown that compositional solutions can be learned and offer substantial gains across a variety of domains, including multi-task learning, language modeling, visual question answering, machine comprehension, and others. However, such models present unique challenges during training when both the module parameters and their composition must be learned jointly. In this paper, we identify several of these issues and analyze their underlying causes. Our discussion focuses on routing networks, a general approach to this problem, and examines empirically the interplay of these challenges and a variety of design decisions. In particular, we consider the effect of how the algorithm decides on module composition, how the algorithm updates the modules, and if the algorithm uses regularization.
Ignacio Cases, C. Rosenbaum, M. Riemer, Tim Klinger · 4 authors totalRelay: A High-Level Compiler for Deep Learning
arXiv · arXiv 1904.08368 · 42 citations · Source: semantic-scholar+dblpFrameworks for writing, compiling, and optimizing deep learning (DL) models have recently enabled progress in areas like computer vision and natural language processing. Extending these frameworks to accommodate the rapidly diversifying landscape of DL models and hardware platforms presents challenging tradeoffs between expressiveness, composability, and portability. We present Relay, a new intermediate representation (IR) and compiler framework for DL models. The functional, statically-typed Relay IR unifies and generalizes existing DL IRs and can express state-of-the-art models. Relay's expressive IR required careful design of the type system, automatic differentiation, and optimizations. Relay's extensible compiler can eliminate abstraction overhead and target new hardware platforms. The design insights from Relay can be applied to existing frameworks to develop IRs that support extension without compromising on expressivity, composibility, and portability. Our evaluation demonstrates that the Relay prototype can already provide competitive performance for a broad class of models running on CPUs, GPUs, and FPGAs.
Jared Roesch, Steven Lyubomirsky, Marisa Kirisame, Josh Pollock, Logan Weber, Ziheng Jiang, Tianqi Chen, T. Moreau · 9 authors totalImproved training of binary networks for human pose estimation and image recognition
arXiv.org · arXiv 1904.05868 · 49 citations · Source: semantic-scholarBig neural networks trained on large datasets have advanced the state-of-the-art for a large variety of challenging problems, improving performance by a large margin. However, under low memory and limited computational power constraints, the accuracy on the same problems drops considerable. In this paper, we propose a series of techniques that significantly improve the accuracy of binarized neural networks (i.e networks where both the features and the weights are binary). We evaluate the proposed improvements on two diverse tasks: fine-grained recognition (human pose estimation) and large-scale image recognition (ImageNet classification). Specifically, we introduce a series of novel methodological changes including: (a) more appropriate activation functions, (b) reverse-order initialization, (c) progressive quantization, and (d) network stacking and show that these additions improve existing state-of-the-art network binarization techniques, significantly. Additionally, for the first time, we also investigate the extent to which network binarization and knowledge distillation can be combined. When tested on the challenging MPII dataset, our method shows a performance improvement of more than 4% in absolute terms. Finally, we further validate our findings by applying the proposed techniques for large-scale object recognition on the Imagenet dataset, on which we report a reduction of error rate by 4%.
Jean Kossaifi, Adrian Bulat, Georgios Tzimiropoulos, M. Pantic · 4 authors totalMLSys: The New Frontier of Machine Learning Systems
arXiv (Cornell University) · DOI 10.48550/arxiv.1904.03257 · arXiv 1904.03257 · 21 citations · Source: openalex+semantic-scholarMachine learning (ML) techniques are enjoying rapidly increasing adoption. However, designing and implementing the systems that support ML models in real-world deployments remains a significant obstacle, in large part due to the radically different development and deployment profile of modern ML methods, and the range of practical concerns that come with broader adoption. We propose to foster a new systems machine learning research community at the intersection of the traditional systems and ML communities, focused on topics such as hardware systems for ML, software systems for ML, and ML optimized for metrics beyond predictive accuracy. To do this, we describe a new conference, MLSys, that explicitly targets research at the intersection of systems and machine learning with a program committee split evenly between experts in systems and ML, and an explicit focus on topics at the intersection of the two.
Erik Meijer, Evan R. Sparks, Peter Bailis, Alexander Ratner, Dan Alistarh, Gustavo Alonso, David G. Andersen, Sarah Bird · 69 authors total