Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗Bean Machine: A Declarative Probabilistic Programming Language For Efficient Programmable Inference.
Probabilistic Graphical Models · 9 citations · Source: openalex+semantic-scholarErik Meijer, Nazanin Khosravani Tehrani, Nimar S. Arora, Yucen Lily Li, Kinjal Divesh Shah, David Noursi, Michael Tingley, Narjes Torabi · 10 authors totalBash Command Line and Shell Scripts Pocket Primer
Mercury Learning & Information · Source: open-library+publisher-catalogOswald Campesato · 1 author totalArtificial Intelligence, Machine Learning, and Deep Learning
Mercury Learning & Information · Source: open-library+publisher-catalogOswald Campesato · 1 author totalApache Mahout: Machine Learning on Distributed Dataflow Systems
Journal of Machine Learning Research 21(127):1-6 · Source: jmlrApache Mahout is a library for scalable machine learning (ML) on distributed dataflow systems, offering various implementations of classification, clustering, dimensionality reduction and recommendation algorithms. Mahout was a pioneer in large-scale machine learning in 2008, when it started and targeted MapReduce, which was the predominant abstraction for scalable computing in industry at that time. Mahout has been widely used by leading web companies and is part of several commercial cloud offerings. In recent years, Mahout migrated to a general framework enabling a mix of dataflow programming and linear algebraic computations on backends such as Apache Spark and Apache Flink. This design allows users to execute data preprocessing and model training in a single, unified dataflow system, instead of requiring a complex integration of several specialized systems. Mahout is maintained as a community-driven open source project at the Apache Software Foundation, and is available under https://mahout.apache.org.
Trevor Grant, Robin Anil, Gokhan Capan, Isabel Drost-Fromm, Ted Dunning, Ellen Friedman, Shannon Quinn, Paritosh Ranjan · 10 authors totalAngular and Machine Learning Pocket Primer
Mercury Learning & Information · Source: open-library+publisher-catalogOswald Campesato · 1 author totalAngular and Deep Learning Pocket Primer
Mercury Learning & Information · Source: open-library+publisher-catalogOswald Campesato · 1 author totalActive Online Domain Adaptation.
CoRR · Source: dblp+stanford-authorityTengyu Ma, Yining Chen, Haipeng Luo, Tengyu Ma 0001, Chicheng Zhang · 5 authors totalA Progress Report on Activities at the Linguistic Data Consortium Benefitting the LREC Community
Language Resources and Evaluation · 2 citations · Source: openalex+first-party-career-authorityMark Liberman, Christopher Cieri, James Fiumara, Stephanie Strassel, Jonathan Wright, Denise DiPersio · 6 authors totalFrom Copernicus Big Data to Extreme Earth Analytics
EDBT · DOI 10.5441/002/EDBT.2019.88 · Source: dblp+first-party-career-authorityJim Dowling, Manolis Koubarakis, Konstantina Bereta, Dimitris Bilidas, Konstantinos Giannousis, Theofilos Ioannidis, Despina-Athanasia Pantazi, George Stamoulis 0001 · 26 authors totalAdapting Linear Hashing for Flash Memory Resource-constrained Embedded Devices.
ICEIS (1) · DOI 10.5220/0007709301760181 · Source: dblp+ubc-authorityRamon Lawrence, Andrew Feltham, Spencer MacBeth, Scott Fazackerley · 4 authors totalComprehensive Overview of Neural Networks and Its Applications in Autonomous Vehicles
Advances in Computational Intelligence and Robotics (IGI Global book chapter) · DOI 10.4018/978-1-5225-7955-7.ch007 · 4 citations · Source: crossref+openalex+semanticscholarDeep learning and Artificial intelligence (AI) have been trending these days due to the capability and state-of-the-art results that they provide. They have replaced some highly skilled professionals with neural network-powered AI, also known as deep learning algorithms. Deep learning majorly works on neural networks. This chapter discusses about the working of a neuron, which is a unit component of neural network. There are numerous techniques that can be incorporated while designing a neural network, such as activation functions, training, etc. to improve its features, which will be explained in detail. It has some challenges such as overfitting, which are difficult to neglect but can be overcome using proper techniques and steps that have been discussed. The chapter will help the academician, researchers, and practitioners to further investigate the associated area of deep learning and its applications in the autonomous vehicle industry.
Jay Rodge, Swati Jaiswal · 2 authors totalMeeting Compliance Requirements While Using Cloud Services
Cloud Security · DOI 10.4018/978-1-4666-5788-5.CH007 · 1 citations · Source: openalex+semantic-scholarCompliance with government and industry regulations is an essential part of conducting business in several sectors. Many of the requirements revolve around financial, privacy, or security aspects. Most of the requirements are due to federal regulations in USA while some are industry requirements that are applicable globally. Even some of the federal regulations in USA apply to service providers abroad when they are providing service to entities in USA. In that sense, all of the compliance requirements discussed here apply to a global audience. In this chapter, the authors discuss in detail the scope of the Health Insurance Portability and Accountability Act, Sarbanes-Oxley Act, Federal Information Security Management Act, Gramm-Leach-Bliley Act, Payment Card Industry Requirements, and the Statement on Auditing Standards 70. These compliance requirements concern protecting the customer data stored in the cloud with respect to confidentiality and integrity. Several of these requirements have significant enforcement powers associated with them, and businesses need to take these requirements seriously and comply. The compliance aspect involves gathering and reporting appropriate information on a regular basis. The authors present details on all these aspects in this chapter.
S. Srinivasan · 1 author totalCómo funciona la web
Universidad de Chile · DOI 10.34720/a4ya-2k31 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Claudio Gutiérrez, Ricardo Baeza‐Yates, José Miguel Piquer Gardner, Gonzalo Navarro, Mauricio Marı́n, Marcelo Arenas, Andrea M. Rodríguez Tastets · 10 authors totalThe Practice of Crowdsourcing
DOI 10.2200/S00904ED1V01Y201903ICR066 · 19 citations · Source: dblp+semantic-scholarAbstract Many data-intensive applications that use machine learning or artificial intelligence techniques depend on humans providing the initial dataset, enabling algorithms to process the rest or ...
Omar Alonso · 1 author totalAutomatic Detection of Prosodic Focus in American English
DOI 10.21437/interspeech.2019-1668 · 2 citations · Source: openalex+first-party-career-authorityMark Liberman, Sunghye Cho, Yong-cheol Lee · 3 authors totalAutomatic Detection of Autism Spectrum Disorder in Children Using Acoustic and Text Features from Brief Natural Conversations
DOI 10.21437/interspeech.2019-1452 · 47 citations · Source: openalex+first-party-career-authorityMark Liberman, Sunghye Cho, Neville Ryant, Meredith Cola, Robert T. Schultz, Julia Parish‐Morris · 6 authors totalThe Second DIHARD Diarization Challenge: Dataset, Task, and Baselines
DOI 10.21437/interspeech.2019-1268 · 27 citations · Source: openalex+first-party-career-authorityMark Liberman, Neville Ryant, Kenneth Church, Christopher Cieri, Alejandrina Cristià, Jun Du, Sriram Ganapathy · 7 authors totalIn Pursuit of Good Governance for the Energy Industry Blockchain
Journal of Energy Markets · DOI 10.21314/JEM.2019.192 · Source: doi+fefa-authorityAna Trbovich, Ana S. Trbovich · 2 authors totalFlexibly-Structured Model for Task-Oriented Dialogues
DOI 10.18653/v1/w19-5922 · arXiv 1908.02402 · 2 citations · Source: openalexThis paper proposes a novel end-to-end architecture for task-oriented dialogue systems. It is based on a simple and practical yet very effective sequence-to-sequence approach, where language understanding and state tracking tasks are modeled jointly with a structured copy-augmented sequential decoder and a multi-label decoder for each slot. The policy engine and language generation tasks are modeled jointly following that. The copyaugmented sequential decoder deals with new or unknown values in the conversation, while the multi-label decoder combined with the sequential decoder ensures the explicit assignment of values to slots. On the generation part, slot binary classifiers are used to improve performance. This architecture is scalable to real-world scenarios and is shown through an empirical evaluation to achieve state-of-the-art performance on both the Cambridge Restaurant dataset and the Stanford in-car assistant dataset 1 .
Piero Molino, Lei Shu, Mahdi Namazifar, Hu Xu, Bing Liu, Huaixiu Zheng, Gökhan Tür · 7 authors totalCollaborative Multi-Agent Dialogue Model Training Via Reinforcement Learning
DOI 10.18653/v1/w19-5912 · arXiv 1907.05507 · 6 citations · Source: openalexWe present the first complete attempt at concurrently training conversational agents that communicate only via self-generated language. Using DSTC2 as seed data, we trained natural language understanding (NLU) and generation (NLG) networks for each agent and let the agents interact online. We model the interaction as a stochastic collaborative game where each agent (player) has a role ("assistant", "tourist", "eater", etc.) and their own objectives, and can only interact via natural language they generate. Each agent, therefore, needs to learn to operate optimally in an environment with multiple sources of uncertainty (its own NLU and NLG, the other agent's NLU, Policy, and NLG). In our evaluation, we show that the stochastic-game agents outperform deep learning based supervised baselines.
Piero Molino, Alexandros Papangelis, Yi‐Chia Wang, Gökhan Tür · 4 authors totalImproving Long Distance Slot Carryover in Spoken Dialogue Systems
Proceedings of the First Workshop on NLP for Conversational AI · DOI 10.18653/v1/W19-4111 · arXiv 1906.01149 · 11 citations · Source: semantic-scholarTracking the state of the conversation is a central component in task-oriented spoken dialogue systems. One such approach for tracking the dialogue state is slot carryover, where a model makes a binary decision if a slot from the context is relevant to the current turn. Previous work on the slot carryover task used models that made independent decisions for each slot. A close analysis of the results show that this approach results in poor performance over longer context dialogues. In this paper, we propose to jointly model the slots. We propose two neural network architectures, one based on pointer networks that incorporate slot ordering information, and the other based on transformer networks that uses self attention mechanism to model the slot interdependencies. Our experiments on an internal dialogue benchmark dataset and on the public DSTC2 dataset demonstrate that our proposed models are able to resolve longer distance slot references and are able to achieve competitive performance.
Tongfei Chen, Chetan Naik, Hua He, Pushpendre Rastogi, Lambert Mathias · 5 authors totalParallax: Visualizing and Understanding the Semantics of Embedding Spaces via Algebraic Formulae
DOI 10.18653/v1/p19-3028 · 12 citations · Source: openalexEmbeddings are a fundamental component of many modern machine learning and natural language processing models. Understanding them and visualizing them is essential for gathering insights about the information they capture and the behavior of the models. In this paper, we introduce Parallax, a tool explicitly designed for this task. Parallax allows the user to use both state-of-the-art embedding analysis methods (PCA and t-SNE) and a simple yet effective task-oriented approach where users can explicitly define the axes of the projection through algebraic formulae. %consists in projecting them in two-dimensional planes without any interpretable semantics associated to the axes of the projection, which makes detailed analyses and comparison among multiple sets of embeddings challenging. In this approach, embeddings are projected into a semantically meaningful subspace, which enhances interpretability and allows for more fine-grained analysis. We demonstrate the power of the tool and the proposed methodology through a series of case studies and a user study.
Piero Molino, Yan Wang, Jiawei Zhang · 3 authors totalA Corpus for Reasoning about Natural Language Grounded in Photographs
ACL · DOI 10.18653/v1/P19-1644 · arXiv 1811.00491 · 740 citations · Source: openalex+semanticscholarWe introduce a new dataset for joint reasoning about natural language and images, with a focus on semantic diversity, compositionality, and visual reasoning challenges. The data contains 107,292 examples of English sentences paired with web photographs. The task is to determine whether a natural language caption is true about a pair of photographs. We crowdsource the data using sets of visually rich images and a compare-and-contrast task to elicit linguistically diverse language. Qualitative analysis shows the data requires compositional joint reasoning, including about quantities, comparisons, and relations. Evaluation using state-of-the-art visual reasoning methods shows the data presents a strong challenge.
Iris Zhang, Alane Suhr, Stephanie Zhou, Ally Zhang, Huajun Bai, Yoav Artzi · 6 authors totalExplain Yourself! Leveraging Language Models for Commonsense Reasoning
Annual Meeting of the Association for Computational Linguistics · DOI 10.18653/v1/P19-1487 · arXiv 1906.02361 · 653 citations · Source: semantic-scholarDeep learning models perform poorly on tasks that require commonsense reasoning, which often necessitates some form of world-knowledge or reasoning over information not immediately present in the input. We collect human explanations for commonsense reasoning in the form of natural language sequences and highlighted annotations in a new dataset called Common Sense Explanations (CoS-E). We use CoS-E to train language models to automatically generate explanations that can be used during training and inference in a novel Commonsense Auto-Generated Explanation (CAGE) framework. CAGE improves the state-of-the-art by 10% on the challenging CommonsenseQA task. We further study commonsense reasoning in DNNs using both human and auto-generated explanations including transfer to out-of-domain tasks. Empirical results indicate that we can effectively leverage language models for commonsense reasoning.
Richard Socher, Nazneen Rajani, Bryan McCann, Caiming Xiong, R. Socher · 5 authors totalLearning to Rank for Plausible Plausibility
Annual Meeting of the Association for Computational Linguistics · DOI 10.18653/v1/P19-1475 · arXiv 1906.02079 · 25 citations · Source: semantic-scholarResearchers illustrate improvements in contextual encoding strategies via resultant performance on a battery of shared Natural Language Understanding (NLU) tasks. Many of these tasks are of a categorical prediction variety: given a conditioning context (e.g., an NLI premise), provide a label based on an associated prompt (e.g., an NLI hypothesis). The categorical nature of these tasks has led to common use of a cross entropy log-loss objective during training. We suggest this loss is intuitively wrong when applied to plausibility tasks, where the prompt by design is neither categorically entailed nor contradictory given the context. Log-loss naturally drives models to assign scores near 0.0 or 1.0, in contrast to our proposed use of a margin-based loss. Following a discussion of our intuition, we describe a confirmation study based on an extreme, synthetically curated task derived from MultiNLI. We find that a margin-based loss leads to a more plausible model of plausibility. Finally, we illustrate improvements on the Choice Of Plausible Alternative (COPA) task through this change in loss.
Tongfei Chen, Zhongyang Li, Benjamin Van Durme · 3 authors totalCompositional Questions Do Not Necessitate Multi-hop Reasoning
Annual Meeting of the Association for Computational Linguistics · DOI 10.18653/v1/P19-1416 · arXiv 1906.02900 · 175 citations · Source: semantic-scholarMulti-hop reading comprehension (RC) questions are challenging because they require reading and reasoning over multiple paragraphs. We argue that it can be difficult to construct large multi-hop RC datasets. For example, even highly compositional questions can be answered with a single hop if they target specific entity types, or the facts needed to answer them are redundant. Our analysis is centered on HotpotQA, where we show that single-hop reasoning can solve much more of the dataset than previously thought. We introduce a single-hop BERT-based RC model that achieves 67 F1—comparable to state-of-the-art multi-hop models. We also design an evaluation setting where humans are not shown all of the necessary paragraphs for the intended multi-hop reasoning but can still answer over 80% of questions. Together with detailed error analysis, these results suggest there should be an increasing focus on the role of evidence in multi-hop reasoning and possibly even a shift towards information retrieval style evaluations with large and diverse evidence collections.
Sameer Singh, Sewon Min, Eric Wallace, Matt Gardner, Hannaneh Hajishirzi, Luke Zettlemoyer · 6 authors totalScaling Multi-Domain Dialogue State Tracking via Query Reformulation
North American Chapter of the Association for Computational Linguistics · DOI 10.18653/v1/N19-2013 · arXiv 1903.05164 · 47 citations · Source: semantic-scholarWe present a novel approach to dialogue state tracking and referring expression resolution tasks. Successful contextual understanding of multi-turn spoken dialogues requires resolving referring expressions across turns and tracking the entities relevant to the conversation across turns. Tracking conversational state is particularly challenging in a multi-domain scenario when there exist multiple spoken language understanding (SLU) sub-systems, and each SLU sub-system operates on its domain-specific meaning representation. While previous approaches have addressed the disparate schema issue by learning candidate transformations of the meaning representation, in this paper, we instead model the reference resolution as a dialogue context-aware user query reformulation task – the dialog state is serialized to a sequence of natural language tokens representing the conversation. We develop our model for query reformulation using a pointer-generator network and a novel multi-task learning setup. In our experiments, we show a significant improvement in absolute F1 on an internal as well as a, soon to be released, public benchmark respectively.
Tongfei Chen, Pushpendre Rastogi, Arpit Gupta, Lambert Mathias · 4 authors totalRecursive Routing Networks: Learning to Compose Modules for Language Understanding
North American Chapter of the Association for Computational Linguistics · DOI 10.18653/v1/N19-1365 · 30 citations · Source: semantic-scholarWe introduce Recursive Routing Networks (RRNs), which are modular, adaptable models that learn effectively in diverse environments. RRNs consist of a set of functions, typically organized into a grid, and a meta-learner decision-making component called the router. The model jointly optimizes the parameters of the functions and the meta-learner’s policy for routing inputs through those functions. RRNs can be incorporated into existing architectures in a number of ways; we explore adding them to word representation layers, recurrent network hidden layers, and classifier layers. Our evaluation task is natural language inference (NLI). Using the MultiNLI corpus, we show that an RRN’s routing decisions reflect the high-level genre structure of that corpus. To show that RRNs can learn to specialize to more fine-grained semantic distinctions, we introduce a new corpus of NLI examples involving implicative predicates, and show that the model components become fine-tuned to the inferential signatures that are characteristic of these predicates.
Ignacio Cases, C. Rosenbaum, M. Riemer, Atticus Geiger, Tim Klinger, Alex Tamkin, Olivia Li, S. Agarwal · 12 authors totalDROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs
North American Chapter of the Association for Computational Linguistics · DOI 10.18653/v1/N19-1246 · arXiv 1903.00161 · 1,343 citations · Source: semantic-scholarReading comprehension has recently seen rapid progress, with systems matching humans on the most popular datasets for the task. However, a large body of work has highlighted the brittleness of these systems, showing that there is much work left to be done. We introduce a new reading comprehension benchmark, DROP, which requires Discrete Reasoning Over the content of Paragraphs. In this crowdsourced, adversarially-created, 55k-question benchmark, a system must resolve references in a question, perhaps to multiple input positions, and perform discrete operations over them (such as addition, counting, or sorting). These operations require a much more comprehensive understanding of the content of paragraphs, as they remove the paraphrase-and-entity-typing shortcuts available in prior datasets. We apply state-of-the-art methods from both the reading comprehension and semantic parsing literatures on this dataset and show that the best systems only achieve 38.4% F1 on our generalized accuracy metric, while expert human performance is 96%. We additionally present a new model that combines reading comprehension methods with simple numerical reasoning to achieve 51% F1.
Sameer Singh, Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Matt Gardner · 6 authors totalImproved Lexically Constrained Decoding for Translation and Monolingual Rewriting
North American Chapter of the Association for Computational Linguistics · DOI 10.18653/v1/N19-1090 · 156 citations · Source: semantic-scholarLexically-constrained sequence decoding allows for explicit positive or negative phrase-based constraints to be placed on target output strings in generation tasks such as machine translation or monolingual text rewriting. We describe vectorized dynamic beam allocation, which extends work in lexically-constrained decoding to work with batching, leading to a five-fold improvement in throughput when working with positive constraints. Faster decoding enables faster exploration of constraint strategies: we illustrate this via data augmentation experiments with a monolingual rewriter applied to the tasks of natural language inference, question answering and machine translation, showing improvements in all three.
Tongfei Chen, J. Hu, Huda Khayrallah, Ryan Culkin, Patrick Xia, Matt Post, Benjamin Van Durme · 7 authors totalEvaluating Question Answering Evaluation
Conference on Empirical Methods in Natural Language Processing · DOI 10.18653/v1/D19-5817 · 147 citations · Source: semantic-scholarAs the complexity of question answering (QA) datasets evolve, moving away from restricted formats like span extraction and multiple-choice (MC) to free-form answer generation, it is imperative to understand how well current metrics perform in evaluating QA. This is especially important as existing metrics (BLEU, ROUGE, METEOR, and F1) are computed using n-gram similarity and have a number of well-known drawbacks. In this work, we study the suitability of existing metrics in QA. For generative QA, we show that while current metrics do well on existing datasets, converting multiple-choice datasets into free-response datasets is challenging for current metrics. We also look at span-based QA, where F1 is a reasonable metric. We show that F1 may not be suitable for all extractive QA tasks depending on the answer types. Our study suggests that while current metrics may be suitable for existing QA datasets, they limit the complexity of QA datasets that can be created. This is especially true in the context of free-form QA, where we would like our models to be able to generate more complex and abstractive answers, thus necessitating new metrics that go beyond n-gram based matching. As a step towards a better QA metric, we explore using BERTScore, a recently proposed metric for evaluating translation, for QA. We find that although it fails to provide stronger correlation with human judgements, future work focused on tailoring a BERT-based metric to QA evaluation may prove fruitful.
Sameer Singh, Anthony Chen, Gabriel Stanovsky, Matt Gardner · 4 authors totalAllenNLP Interpret: A Framework for Explaining Predictions of NLP Models
Conference on Empirical Methods in Natural Language Processing · DOI 10.18653/v1/D19-3002 · arXiv 1909.09251 · 146 citations · Source: semantic-scholarNeural NLP models are increasingly accurate but are imperfect and opaque—they break in counterintuitive ways and leave end users puzzled at their behavior. Model interpretation methods ameliorate this opacity by providing explanations for specific model predictions. Unfortunately, existing interpretation codebases make it difficult to apply these methods to new models and tasks, which hinders adoption for practitioners and burdens interpretability researchers. We introduce AllenNLP Interpret, a flexible framework for interpreting NLP models. The toolkit provides interpretation primitives (e.g., input gradients) for any AllenNLP model and task, a suite of built-in interpretation methods, and a library of front-end visualization components. We demonstrate the toolkit’s flexibility and utility by implementing live demos for five interpretation methods (e.g., saliency maps and adversarial attacks) on a variety of models and tasks (e.g., masked language modeling using BERT and reading comprehension using BiDAF). These demos, alongside our code and tutorials, are available at https://allennlp.org/interpret.
Sameer Singh, Eric Wallace, Jens Tuyls, Junlin Wang, Sanjay Subramanian, Matt Gardner · 6 authors totalHint-Based Training for Non-Autoregressive Machine Translation
EMNLP 2019 · DOI 10.18653/v1/D19-1573 · arXiv 1909.06708 · 77 citations · Source: arxiv+semantic-scholarDue to the unparallelizable nature of the autoregressive factorization, AutoRegressive Translation (ART) models have to generate tokens sequentially during decoding and thus suffer from high inference latency. Non-AutoRegressive Translation (NART) models were proposed to reduce the inference time, but could only achieve inferior translation accuracy. In this paper, we proposed a novel approach to leveraging the hints from hidden states and word alignments to help the training of NART models. The results achieve significant improvement over previous NART models for the WMT14 En-De and De-En datasets and are even comparable to a strong LSTM-based ART baseline but one order of magnitude faster in inference.
Zhuohan Li, Zi Lin, Di He, Fei Tian, Tao Qin, Liwei Wang, Tie-Yan Liu · 7 authors totalMultiFiT: Efficient Multi-lingual Language Model Fine-tuning
EMNLP 2019 · DOI 10.18653/v1/D19-1572 · arXiv 1909.04761 · 107 citations · Source: semantic-scholarPretrained language models are promising particularly for low-resource languages as they only require unlabelled data. However, training existing models requires huge amounts of compute, while pretrained cross-lingual models often underperform on low-resource languages. We propose Multi-lingual language model Fine-Tuning (MultiFiT) to enable practitioners to train and fine-tune language models efficiently in their own language. In addition, we propose a zero-shot method using an existing pretrained cross-lingual model. We evaluate our methods on two widely used cross-lingual classification datasets where they outperform models pretrained on orders of magnitude more data and compute. We release all models and code.
Jeremy Howard, Julian Martin Eisenschlos, Sebastian Ruder, Piotr Czapla, Marcin Kardas, Sylvain Gugger · 6 authors totalDo NLP Models Know Numbers? Probing Numeracy in Embeddings
Conference on Empirical Methods in Natural Language Processing · DOI 10.18653/v1/D19-1534 · arXiv 1909.07940 · 314 citations · Source: semantic-scholarThe ability to understand and work with numbers (numeracy) is critical for many complex reasoning tasks. Currently, most NLP models treat numbers in text in the same way as other tokens—they embed them as distributed vectors. Is this enough to capture numeracy? We begin by investigating the numerical reasoning capabilities of a state-of-the-art question answering model on the DROP dataset. We find this model excels on questions that require numerical reasoning, i.e., it already captures numeracy. To understand how this capability emerges, we probe token embedding methods (e.g., BERT, GloVe) on synthetic list maximum, number decoding, and addition tasks. A surprising degree of numeracy is naturally present in standard embeddings. For example, GloVe and word2vec accurately encode magnitude for numbers up to 1,000. Furthermore, character-level embeddings are even more precise—ELMo captures numeracy the best for all pre-trained methods—but BERT, which uses sub-word units, is less exact.
Sameer Singh, Eric Wallace, Yizhong Wang, Sujian Li, Matt Gardner · 5 authors totalPosing Fair Generalization Tasks for Natural Language Inference
Conference on Empirical Methods in Natural Language Processing · DOI 10.18653/v1/D19-1456 · arXiv 1911.00811 · 50 citations · Source: semantic-scholarDeep learning models for semantics are generally evaluated using naturalistic corpora. Adversarial testing methods, in which models are evaluated on new examples with known semantic properties, have begun to reveal that good performance at these naturalistic tasks can hide serious shortcomings. However, we should insist that these evaluations be fair – that the models are given data sufficient to support the requisite kinds of generalization. In this paper, we define and motivate a formal notion of fairness in this sense. We then apply these ideas to natural language inference by constructing very challenging but provably fair artificial datasets and showing that standard neural models fail to generalize in the required ways; only task-specific models that jointly compose the premise and hypothesis are able to achieve high performance, and even these models do not solve the task perfectly.
Ignacio Cases, Atticus Geiger, L. Karttunen, Christopher Potts · 4 authors totalUniversal Adversarial Triggers for Attacking and Analyzing NLP
Conference on Empirical Methods in Natural Language Processing · DOI 10.18653/v1/D19-1221 · arXiv 1908.07125 · 1,124 citations · Source: semantic-scholarAdversarial examples highlight model vulnerabilities and are useful for evaluation and interpretation. We define universal adversarial triggers: input-agnostic sequences of tokens that trigger a model to produce a specific prediction when concatenated to any input from a dataset. We propose a gradient-guided search over tokens which finds short trigger sequences (e.g., one word for classification and four words for language modeling) that successfully trigger the target prediction. For example, triggers cause SNLI entailment accuracy to drop from 89.94% to 0.55%, 72% of “why” questions in SQuAD to be answered “to kill american people”, and the GPT-2 language model to spew racist output even when conditioned on non-racial contexts. Furthermore, although the triggers are optimized using white-box access to a specific model, they transfer to other models for all tasks we consider. Finally, since triggers are input-agnostic, they provide an analysis of global model behavior. For instance, they confirm that SNLI models exploit dataset biases and help to diagnose heuristics learned by reading comprehension models.
Sameer Singh, Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner · 5 authors totalExecuting Instructions in Situated Collaborative Interactions
EMNLP-IJCNLP · DOI 10.18653/v1/D19-1218 · arXiv 1910.03655 · 100 citations · Source: openalex+semanticscholarIris Zhang, Alane Suhr, Claudia Yan, Jack Schluger, Stanley Yu, Hadi Khader, Marwa Mouallem, Yoav Artzi · 8 authors totalSpan-based Hierarchical Semantic Parsing for Task-Oriented Dialog
EMNLP 2019 · DOI 10.18653/v1/d19-1163 · 21 citations · Source: semantic-scholarWe propose a semantic parser for parsing compositional utterances into Task Oriented Parse (TOP), a tree representation that has intents and slots as labels of nesting tree nodes. Our parser is span-based: it scores labels of the tree nodes covering each token span independently, but then decodes a valid tree globally. In contrast to previous sequence decoding approaches and other span-based parsers, we (1) improve the training speed by removing the need to run the decoder at training time; and (2) introduce edge scores, which model relations between parent and child labels, to mitigate the independence assumption between node labels and improve accuracy. Our best parser outperforms previous methods on the TOP dataset of mixed-domain task-oriented utterances in both accuracy and training speed.
Sonal Gupta, Panupong Pasupat, S. Gupta, Karishma Mandyam, Rushin Shah, Michael Lewis, Luke Zettlemoyer · 7 authors totalModeling Multi-Action Policy for Task-Oriented Dialogues
DOI 10.18653/v1/d19-1130 · 9 citations · Source: openalexLei Shu, Hu Xu, Bing Liu, Piero Molino. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Piero Molino, Lei Shu, Hu Xu, Bing Liu · 4 authors totalKnowledge Enhanced Contextual Word Representations
Conference on Empirical Methods in Natural Language Processing · DOI 10.18653/v1/D19-1005 · arXiv 1909.04164 · 735 citations · Source: semantic-scholarContextual word representations, typically trained on unstructured, unlabeled text, do not contain any explicit grounding to real world entities and are often unable to remember facts about those entities. We propose a general method to embed multiple knowledge bases (KBs) into large scale models, and thereby enhance their representations with structured, human-curated knowledge. For each KB, we first use an integrated entity linker to retrieve relevant entity embeddings, then update contextual word representations via a form of word-to-entity attention. In contrast to previous approaches, the entity linkers and self-supervised language modeling objective are jointly trained end-to-end in a multitask setting that combines a small amount of entity linking supervision with a large amount of raw text. After integrating WordNet and a subset of Wikipedia into BERT, the knowledge enhanced BERT (KnowBert) demonstrates improved perplexity, ability to recall facts as measured in a probing task and downstream performance on relationship extraction, entity typing, and word sense disambiguation. KnowBert’s runtime is comparable to BERT’s and it scales to large KBs.
Sameer Singh, Matthew E. Peters, Mark Neumann, IV RobertL.Logan, Roy Schwartz, Vidur Joshi, Noah A. Smith · 7 authors totalReading the Manual: Event Extraction as Definition Comprehension
SPNLP · DOI 10.18653/v1/2020.spnlp-1.9 · arXiv 1912.01586 · 68 citations · Source: semantic-scholarWe ask whether text understanding has progressed to where we may extract event information through incremental refinement of bleached statements derived from annotation manuals. Such a capability would allow for the trivial construction and extension of an extraction framework by intended end-users through declarations such as, “Some person was born in some location at some time.” We introduce an example of a model that employs such statements, with experiments illustrating we can extract events under closed ontologies and generalize to unseen event types simply by reading new definitions.
Tongfei Chen, Yunmo Chen, Seth Ebner, Benjamin Van Durme · 4 authors totalEvaluating the Factual Consistency of Abstractive Text Summarization
Conference on Empirical Methods in Natural Language Processing · DOI 10.18653/v1/2020.emnlp-main.750 · arXiv 1910.12840 · 1,028 citations · Source: semantic-scholarCurrently used metrics for assessing summarization algorithms do not account for whether summaries are factually consistent with source documents. We propose a weakly-supervised, model-based approach for verifying factual consistency and identifying conflicts between source documents and a generated summary. Training data is generated by applying a series of rule-based transformations to the sentences of source documents. The factual consistency model is then trained jointly for three tasks: 1) identify whether sentences remain factually consistent after transformation, 2) extract a span in the source documents to support the consistency prediction, 3) extract a span in the summary sentence that is inconsistent if one exists. Transferring this model to summaries generated by several state-of-the art models reveals that this highly scalable approach substantially outperforms previous models, including those trained with strong supervision using standard datasets for natural language inference and fact checking. Additionally, human evaluation shows that the auxiliary span extraction tasks provide useful assistance in the process of verifying factual consistency.
Richard Socher, Wojciech Kryscinski, Bryan McCann, Caiming Xiong, R. Socher · 5 authors totalUncertain Natural Language Inference
Annual Meeting of the Association for Computational Linguistics · DOI 10.18653/v1/2020.acl-main.774 · arXiv 1909.03042 · 62 citations · Source: semantic-scholarWe introduce Uncertain Natural Language Inference (UNLI), a refinement of Natural Language Inference (NLI) that shifts away from categorical labels, targeting instead the direct prediction of subjective probability assessments. We demonstrate the feasibility of collecting annotations for UNLI by relabeling a portion of the SNLI dataset under a probabilistic scale, where items even with the same categorical label differ in how likely people judge them to be true given a premise. We describe a direct scalar regression modeling approach, and find that existing categorically-labeled NLI data can be used in pre-training. Our best models correlate well with humans, demonstrating models are capable of more subtle inferences than the categorical bin assignment employed in current NLI tasks.
Tongfei Chen, Zhengping Jiang, Keisuke Sakaguchi, Benjamin Van Durme · 4 authors totalERASER: A Benchmark to Evaluate Rationalized NLP Models
Annual Meeting of the Association for Computational Linguistics · DOI 10.18653/v1/2020.acl-main.408 · arXiv 1911.03429 · 880 citations · Source: semantic-scholarState-of-the-art models in NLP are now predominantly based on deep neural networks that are opaque in terms of how they come to make predictions. This limitation has increased interest in designing more interpretable deep models for NLP that reveal the ‘reasoning’ behind model outputs. But work in this direction has been conducted on different datasets and tasks with correspondingly unique aims and metrics; this makes it difficult to track progress. We propose the Evaluating Rationales And Simple English Reasoning (ERASER a benchmark to advance research on interpretable models in NLP. This benchmark comprises multiple datasets and tasks for which human annotations of “rationales” (supporting evidence) have been collected. We propose several metrics that aim to capture how well the rationales provided by models align with human rationales, and also how faithful these rationales are (i.e., the degree to which provided rationales influenced the corresponding predictions). Our hope is that releasing this benchmark facilitates progress on designing more interpretable NLP systems. The benchmark, code, and documentation are available at https://www.eraserbenchmark.com/
Richard Socher, Jay DeYoung, Sarthak Jain, Nazneen Rajani, Eric P. Lehman, Caiming Xiong, R. Socher, Byron C. Wallace · 8 authors totalThe TechQA Dataset
Annual Meeting of the Association for Computational Linguistics · DOI 10.18653/v1/2020.acl-main.117 · arXiv 1911.02984 · 59 citations · Source: arxiv+semantic-scholarWe introduce TechQA, a domain-adaptation question answering dataset for the technical support domain. The TechQA corpus highlights two real-world issues from the automated customer support domain. First, it contains actual questions posed by users on a technical forum, rather than questions generated specifically for a competition or a task. Second, it has a real-world size -- 600 training, 310 dev, and 490 evaluation question/answer pairs -- thus reflecting the cost of creating large labeled datasets with actual data. Consequently, TechQA is meant to stimulate research in domain adaptation rather than being a resource to build QA systems from scratch. The dataset was obtained by crawling the IBM Developer and IBM DeveloperWorks forums for questions with accepted answers that appear in a published IBM Technote---a technical document that addresses a specific technical issue. We also release a collection of the 801,998 publicly available Technotes as of April 4, 2019 as a companion resource that might be used for pretraining, to learn representations of the IT domain language.
Rosario Uceda-Sosa, Vittorio Castelli, Rishav Chakravarti, Saswati Dana, Anthony Ferritto, Radu Florian, Martin Franz, Dinesh Garg · 21 authors totalUnsupervised Large‐Scale Search for Similar Earthquake Signals
Bulletin of the Seismological Society of America · DOI 10.1785/0120190006 · 10 citations · Source: openalex+authoritative-profilePeter Bailis, Clara E. Yoon, Karianne J. Bergen, Kexin Rong, Hashem Elezabi, William L. Ellsworth, Gregory C. Beroza, Philip Levis · 8 authors totalIncremental multi-domain learning with network latent tensor factorization
AAAI Conference on Artificial Intelligence · DOI 10.1609/AAAI.V34I07.6617 · arXiv 1904.06345 · 34 citations · Source: semantic-scholarThe prominence of deep learning, large amount of annotated data and increasingly powerful hardware made it possible to reach remarkable performance for supervised classification tasks, in many cases saturating the training sets. However the resulting models are specialized to a single very specific task and domain. Adapting the learned classification to new domains is a hard problem due to at least three reasons: (1) the new domains and the tasks might be drastically different; (2) there might be very limited amount of annotated data on the new domain and (3) full training of a new model for each new task is prohibitive in terms of computation and memory, due to the sheer number of parameters of deep CNNs. In this paper, we present a method to learn new-domains and tasks incrementally, building on prior knowledge from already learned tasks and without catastrophic forgetting. We do so by jointly parametrizing weights across layers using low-rank Tucker structure. The core is task agnostic while a set of task specific factors are learnt on each new domain. We show that leveraging tensor structure enables better performance than simply using matrix operations. Joint tensor modelling also naturally leverages correlations across different layers. Compared with previous methods which have focused on adapting each layer separately, our approach results in more compact representations for each new task/domain. We apply the proposed method to the 10 datasets of the Visual Decathlon C
Jean Kossaifi, Adrian Bulat, Georgios Tzimiropoulos, M. Pantic · 4 authors totalLikelihood Ratios and Generative Classifiers for Unsupervised Out-of-Domain Detection In Task Oriented Dialog
AAAI 2020 · DOI 10.1609/aaai.v34i05.6280 · arXiv 1912.12800 · 59 citations · Source: semantic-scholarThe task of identifying out-of-domain (OOD) input examples directly at test-time has seen renewed interest recently due to increased real world deployment of models. In this work, we focus on OOD detection for natural language sentence inputs to task-based dialog systems. Our findings are three-fold:First, we curate and release ROSTD (Real Out-of-Domain Sentences From Task-oriented Dialog) - a dataset of 4K OOD examples for the publicly available dataset from (Schuster et al. 2019). In contrast to existing settings which synthesize OOD examples by holding out a subset of classes, our examples were authored by annotators with apriori instructions to be out-of-domain with respect to the sentences in an existing dataset.Second, we explore likelihood ratio based approaches as an alternative to currently prevalent paradigms. Specifically, we reformulate and apply these approaches to natural language inputs. We find that they match or outperform the latter on all datasets, with larger improvements on non-artificial OOD benchmarks such as our dataset. Our ablations validate that specifically using likelihood ratios rather than plain likelihood is necessary to discriminate well between OOD and in-domain data.Third, we propose learning a generative classifier and computing a marginal likelihood (ratio) for OOD detection. This allows us to use a principled likelihood while at the same time exploiting training-time labels. We find that this approach outperforms both simple likelihood (ratio) based and other prior approaches. We are hitherto the first to investigate the use of generative classifiers for OOD detection at test-time.
Sonal Gupta, Varun Prashant Gangal, Abhinav Arora, Arash Einolghozati, S. Gupta · 5 authors totalOn the Role of Weight Sharing During Deep Option Learning
AAAI Conference on Artificial Intelligence · DOI 10.1609/AAAI.V34I04.6003 · arXiv 1912.13408 · 23 citations · Source: semantic-scholarThe options framework is a popular approach for building temporally extended actions in reinforcement learning. In particular, the option-critic architecture provides general purpose policy gradient theorems for learning actions from scratch that are extended in time. However, past work makes the key assumption that each of the components of option-critic has independent parameters. In this work we note that while this key assumption of the policy gradient theorems of option-critic holds in the tabular case, it is always violated in practice for the deep function approximation setting. We thus reconsider this assumption and consider more general extensions of option-critic and hierarchical option-critic training that optimize for the full architecture with each update. It turns out that not assuming parameter independence challenges a belief in prior work that training the policy over options can be disentangled from the dynamics of the underlying options. In fact, learning can be sped up by focusing the policy over options on states where options are actually likely to terminate. We put our new algorithms to the test in application to sample efficient learning of Atari games, and demonstrate significantly improved stability and faster convergence when learning long options. 1
Ignacio Cases, M. Riemer, C. Rosenbaum, Miao Liu, G. Tesauro · 5 authors total