Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗Towards Assessing Changes in Degree of Depression through Facebook
DOI 10.3115/v1/w14-3214 · 261 citations · Source: openalexH. Andrew Schwartz, Johannes Eichstaedt, Margaret L. Kern, Gregory Park, Maarten Sap, David Stillwell, Michal Kosinski, Lyle Ungar. Proceedings of the Workshop on Computational Linguistics and Clinical Psychology: From Linguistic Signal to Clinical Reality. 2014.
Lyle Ungar, H. Andrew Schwartz, Johannes C. Eichstaedt, Margaret L. Kern, Gregory Park, Maarten Sap, David Stillwell, Michał Kosiński · 8 authors totalComputing Affect in Metaphors
DOI 10.3115/v1/W14-2306 · 20 citations · Source: semantic-scholarThis article describes a novel approach to automated determination of affect associated with metaphorical language. Affect in language is understood to mean the attitude toward a topic that a writer attempts to convey to the reader by using a particular metaphor. This affect, which we will classify as positive, negative or neutral with various degrees of intensity, may arise from the target of the metaphor, from the choice of words used to describe it, or from other elements in its immediate linguistic context. We attempt to capture all these contributing elements in an Affect Calculus and demonstrate experimentally that the resulting method can accurately approximate human judgment. The work reported here is part of a larger effort to develop a highly accurate system for identifying, classifying, and comparing metaphors occurring in large volumes of text across four different languages: English, Spanish, Russian, and Farsi.
Ignacio Cases, T. Strzalkowski, Samira Shaikh, Kit Cho, G. Broadwell, L. Feldman, Sarah M. Taylor, B. Yamrom · 11 authors totalImproved Pattern Learning for Bootstrapped Entity Extraction
CoNLL 2014 · DOI 10.3115/v1/w14-1611 · 142 citations · Source: semantic-scholarBootstrapped pattern learning for entity extraction usually starts with seed entities and iteratively learns patterns and entities from unlabeled text. Patterns are scored by their ability to extract more positive entities and less negative entities. A problem is that due to the lack of labeled data, unlabeled entities are either assumed to be negative or are ignored by the existing pattern scoring measures. In this paper, we improve pattern scoring by predicting the labels of unlabeled entities. We use various unsupervised features based on contrasting domain-specific and general text, and exploiting distributional similarity and edit distances to learned entities. Our system outperforms existing pattern scoring algorithms for extracting drug-andtreatment entities from four medical forums.
Sonal Gupta, S. Gupta, Christopher D. Manning · 3 authors totalKeyword Highlighting Improves Comprehension for People with Dyslexia
DOI 10.3115/v1/w14-1204 · 23 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Luz Rello, Horacio Saggion, Ricardo Baeza‐Yates · 4 authors totalMeerkat Mafia: Multilingual and Cross-Level Semantic Textual Similarity Systems
International Workshop on Semantic Evaluation · DOI 10.3115/v1/S14-2072 · 28 citations · Source: semantic-scholarWe describe UMBC’s systems developed for the SemEval 2014 tasks on Multilingual Semantic Textual Similarity (Task 10) and Cross-Level Semantic Similarity (Task 3). Our best submission in the Multilingual task ranked second in both English and Spanish subtasks using an unsupervised approach. Our best systems for Cross-Level task ranked second in Paragraph-Sentence and first in both Sentence-Phrase and Word-Sense subtask. The system ranked first for the PhraseWord subtask but was not included in the official results due to a late submission.
Abhay Kashyap, Abhay L. Kashyap, Lushan Han, Roberto Yus, Jennifer Sleeman, Taneeya Satyapanich, Sunil Gandhi, Tim Finin · 8 authors totalThe Stanford CoreNLP Natural Language Processing Toolkit
ACL · DOI 10.3115/V1/P14-5010 · Source: dblp+stanford-nlp+mixpanel-career-authorityJenny Finkel, Christopher D. Manning, Mihai Surdeanu, John Bauer, Jenny Rose Finkel, Steven Bethard, David McClosky · 7 authors totalParser Evaluation Using Derivation Trees: A Complement to evalb
DOI 10.3115/v1/p14-2109 · 2 citations · Source: openalex+first-party-career-authorityMark Liberman, Seth Kulick, Ann Bies, Justin L. Mott, Anthony Kroch, Beatrice Santorini · 6 authors totalJoint Syntactic and Semantic Parsing with Combinatory Categorial Grammar
ACL · DOI 10.3115/V1/P14-1112 · Source: dblp+author-first-party+semantic-machines-career-authorityJayant Krishnamurthy, Tom M. Mitchell · 2 authors totalLess Grammar, More Features
Annual Meeting of the Association for Computational Linguistics · DOI 10.3115/v1/P14-1022 · 59 citations · Source: semantic-scholarWe present a parser that relies primarily on extracting information directly from surface spans rather than on propagating information through enriched grammar structure. For example, instead of creating separate grammar symbols to mark the definiteness of an NP, our parser might instead capture the same information from the first word of the NP. Moving context out of the grammar and onto surface features can greatly simplify the structural component of the parser: because so many deep syntactic cues have surface reflexes, our system can still parse accurately with context-free backbones as minimal as Xbar grammars. Keeping the structural backbone simple and moving features to the surface also allows easy adaptation to new languages and even to new tasks. On the SPMRL 2013 multilingual constituency parsing shared task (Seddah et al., 2013), our system outperforms the top single parser system of Bjorkelund et al. (2013) on a range of languages. In addition, despite being designed for syntactic analysis, our system also achieves stateof-the-art numbers on the structural sentiment task of Socher et al. (2013). Finally, we show that, in both syntactic parsing and sentiment analysis, many broad linguistic trends can be captured via surface features.
David Hall, David Leo Wright Hall, Greg Durrett, D. Klein · 4 authors totalSparser, Better, Faster GPU Parsing
Annual Meeting of the Association for Computational Linguistics · DOI 10.3115/v1/P14-1020 · 21 citations · Source: semantic-scholarDue to their origin in computer graphics, graphics processing units (GPUs) are highly optimized for dense problems, where the exact same operation is applied repeatedly to all data points. Natural language processing algorithms, on the other hand, are traditionally constructed in ways that exploit structural sparsity. Recently, Canny et al. (2013) presented an approach to GPU parsing that sacrifices traditional sparsity in exchange for raw computational power, obtaining a system that can compute Viterbi parses for a high-quality grammar at about 164 sentences per second on a mid-range GPU. In this work, we reintroduce sparsity to GPU parsing by adapting a coarse-to-fine pruning approach to the constraints of a GPU. The resulting system is capable of computing over 404 Viterbi parses per second—more than a 2x speedup—on the same hardware. Moreover, our approach allows us to efficiently implement less GPU-friendly minimum Bayes risk inference, improving throughput for this more accurate algorithm from only 32 sentences per second unpruned to over 190 sentences per second using pruning—nearly a 6x speedup.
David Hall, David Leo Wright Hall, Taylor Berg-Kirkpatrick, D. Klein · 4 authors totalGloVe: Global Vectors for Word Representation
Conference on Empirical Methods in Natural Language Processing · DOI 10.3115/v1/D14-1162 · 34,948 citations · Source: semantic-scholarRecent methods for learning vector space representations of words have succeeded in capturing fine-grained semantic and syntactic regularities using vector arithmetic, but the origin of these regularities has remained opaque. We analyze and make explicit the model properties needed for such regularities to emerge in word vectors. The result is a new global logbilinear regression model that combines the advantages of the two major model families in the literature: global matrix factorization and local context window methods. Our model efficiently leverages statistical information by training only on the nonzero elements in a word-word cooccurrence matrix, rather than on the entire sparse matrix or on individual context windows in a large corpus. The model produces a vector space with meaningful substructure, as evidenced by its performance of 75% on a recent word analogy task. It also outperforms related models on similarity tasks and named entity recognition.
Richard Socher, Jeffrey Pennington, R. Socher, Christopher D. Manning · 4 authors totalDeveloping Age and Gender Predictive Lexica over Social Media
DOI 10.3115/v1/d14-1121 · 241 citations · Source: openalexMaarten Sap, Gregory Park, Johannes Eichstaedt, Margaret Kern, David Stillwell, Michal Kosinski, Lyle Ungar, Hansen Andrew Schwartz. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014.
Lyle Ungar, Maarten Sap, Gregory Park, Johannes C. Eichstaedt, Margaret L. Kern, David Stillwell, Michał Kosiński, Hansen Andrew Schwartz · 8 authors totalIncorporating Vector Space Similarity in Random Walk Inference over Knowledge Bases
EMNLP · DOI 10.3115/V1/D14-1044 · Source: dblp+author-first-party+semantic-machines-career-authorityJayant Krishnamurthy, Matt Gardner 0001, Partha Pratim Talukdar, Tom M. Mitchell · 4 authors totalBuilding Trust in Cloud Computing: Challenges in the Midst of Outages
DOI 10.28945/2018 · 10 citations · Source: openalex+semantic-scholarS. Srinivasan, Jesse H. Jones · 2 authors totalDemand Uncertainty and Cost Behavior
DOI 10.2308/ACCR-50661 · 189 citations · Source: semantic-scholarJose Plehn, R. Banker, Dmitri Byzalov, J. Plehn-Dujowich · 4 authors totalHighly Accurate Mandarin Tone Classification In The Absence of Pitch Information
DOI 10.21437/speechprosody.2014-123 · 33 citations · Source: openalex+first-party-career-authorityMark Liberman, Neville Ryant, Malcolm Slaney, Elizabeth Shriberg, Jiahong Yuan · 5 authors totalThe Changing Face of International Dispute Resolution: An Analysis of Factors Driving Trends in Investor-State Dispute Resolution and the WTO Dispute Resolution Mechanism
SSRN working paper · DOI 10.2139/ssrn.2438955 · Source: semantic-scholar+crossrefThis paper provides a landscape analysis of international arbitration trends, and places special attention on the proliferation of investor-state dispute resolution (hereinafter “ISDR”) in free trade agreements (hereinafter “FTAs”) and, more widely, in international investment agreements (hereinafter “IIA”). The dispute settlement provisions of the proposed and pending Trans Pacific Partnership Agreement (hereinafter “TPP”) help provide context for these developments. This paper sources opinions and data from a number works, with a special attention to those capturing statistics on how ISDR changes the dynamic of international trade disputes and FTA negotiation. Over the past decade, FTA’s have increasingly segregated dispute resolution responsibility away from centralized World Trade Organization (hereinafter “WTO”) processes. Passage of an ISDR provision in the TPP would substantially further this trend by binding eleven more countries to ISDR requirements similar to those in sister agreements, including NAFTA, the Energy Charter Treaty and the Argentina-United States Bilateral Investment Treaty (hereinafter “BIT”). However, recent national sovereignty and public interest concerns over ISDR may mean a tapering off of this trend and a return to the WTO Dispute Settlement Mechanism (hereinafter “DSM”) and preference for judicial review by national courts. An analysis of whether ISDR infringes on national sovereignty and the public interest is provided to help frame and guide this discussion.
Nicole Shanahan · 1 author totalAADL Fault Modeling and Analysis Within an ARP4761 Safety Assessment
US Dept of the Air Force · DOI 10.21236/ada610294 · 39 citations · Source: semantic-scholar+openalexAbstract : SAE Standard Aerospace Recommended Practice (ARP) 4761, Guidelines and Methods for Conducting the Safety Assessment Process on Civil Airborne Systems and Equipment, provides general guidance on evaluating the safety aspects of a design and identifies processes, methods, and tools to support the evaluation. The Architecture Analysis and Design Language (AADL) Error Model Annex defines features to enable specification of risk mitigation methods in an architecture and assessments of system properties such as safety and reliability. This report describes how the AADL Error Model Annex supports the safety assessment processes and techniques presented in SAE Standard ARP4761. It provides a mapping between constructs of the AADL Error Model Annex and the assessment techniques identified in ARP4761 and presents examples of using the Error Model Annex with those techniques. The processes and techniques of the ARP4761 standard that this report addresses are the Functional Hazard Assessment, Preliminary System Safety Assessment, System Safety Assessment, Fault Tree Analysis, Failure Modes and Effects Analysis, Markov Analysis, and Dependence Diagrams, also referred to as Reliability Block Diagrams.
Julien Delange, Peter H. Feiler, David P. Gluch, John Hudak · 4 authors totalEfficient Informative Sensing using Multiple Robots
Journal of Artificial Intelligence Research · DOI 10.1613/jair.2674 · arXiv 1401.3462 · 385 citations · Source: semantic-scholarThe need for efficient monitoring of spatio-temporal dynamics in large environmental applications, such as the water quality monitoring in rivers and lakes, motivates the use of robotic sensors in order to achieve sufficient spatial coverage. Typically, these robots have bounded resources, such as limited battery or limited amounts of time to obtain measurements. Thus, careful coordination of their paths is required in order to maximize the amount of information collected, while respecting the resource constraints. In this paper, we present an efficient approach for near-optimally solving the NP-hard optimization problem of planning such informative paths. In particular, we first develop eSIP (efficient Single-robot Informative Path planning), an approximation algorithm for optimizing the path of a single robot. Hereby, we use a Gaussian Process to model the underlying phenomenon, and use the mutual information between the visited locations and remainder of the space to quantify the amount of information collected. We prove that the mutual information collected using paths obtained by using eSIP is close to the information obtained by an optimal solution. We then provide a general technique, sequential allocation, which can be used to extend any single robot planning algorithm, such as eSIP, for the multi-robot problem. This procedure approximately generalizes any guarantees for the single-robot problem to the multi-robot case. We extensively evaluate the effectiveness of our approach on several experiments performed infield for two important environmental sensing applications, lake and river monitoring, and simulation experiments performed using several real world sensor network data sets.
Carlos Guestrin, Amarjeet Singh, Andreas Krause, W. Kaiser · 4 authors totalUsing Worker Quality Scores to Improve Stopping Rules
HCOMP · DOI 10.1609/hcomp.v2i1.13201 · 25 citations · Source: dblp+semantic-scholarWe consider the crowdsourcing task of learning the answer to simple multiple-choice microtasks. In order to provide statistically significant results, one often needs to ask multiple workers to answer the same microtask. A stopping rule is an algorithm that for a given microtask decides for any given set of worker answers if the system should stop and output an answer or iterate and ask one more worker. A quality score for a worker is a score that reflects the historic performance of that worker. In this paper we investigate how to devise better stopping rules given such quality scores. We conduct a data analysis on a large-scale industrial crowdsourcing platform, and use the observations from this analysis to design new stopping rules that use the workers’ quality scores in a non-trivial manner. We then conduct a simulation based on a real-world workload, showing that our algorithm performs better than the more naive approaches.
Omar Alonso, Ittai Abraham, Vasileios Kandylas, Rajesh Patel, Steven Shelford, Aleksandrs Slivkins · 6 authors totalCan Oral Nutritional Supplements Improve Medicare Patient Outcomes in the Hospital?
Forum for Health Economics & Policy · DOI 10.1515/fhep-2014-0011 · Source: publisher+career-authorityDaniella Perlroth, Darius N. Lakdawalla, Julia T. Snider, Daniella J. Perlroth, Chris LaVallee, Mark T. Linthicum, Tomas J. Philipson, Jamie S. Partridge · 8 authors totalCoordination avoidance in database systems
Proceedings of the VLDB Endowment · DOI 10.14778/2735508.2735509 · 191 citations · Source: openalex+authoritative-profilePeter Bailis, Alan Fekete, Michael J. Franklin, Ali Ghodsi, Joseph M. Hellerstein, Ion Stoica · 6 authors totalA Partitioning Framework for Aggressive Data Skipping
Proceedings of the VLDB Endowment · DOI 10.14778/2733004.2733044 · 11 citations · Source: semantic-scholarReynold Xin, Liwen Sun, S. Krishnan, M. Franklin · 4 authors totalSummingbird: A Framework for Integrating Batch and Online MapReduce Computations
Proc. VLDB Endow. · DOI 10.14778/2733004.2733016 · 117 citations · Source: semantic-scholar+dblpOscar Boykin, P. Oscar Boykin, Sam Ritchie, Ian O'Connell, Jimmy Lin · 5 authors totalReal-Time Twitter Recommendation: Online Motif Detection in Large Dynamic Graphs
Proceedings of the VLDB Endowment (VLDB 2014) · DOI 10.14778/2733004.2733010 · 69 citations · Source: dblp+semantic-scholarAjeet Grewal, Pankaj Gupta, Venu Satuluri, Siva Gurumurthy, Volodymyr Zhabiuk, Quannan Li, Jimmy Lin · 7 authors totalComparative Genome Analyses Reveal Distinct Structure in the Saltwater Crocodile MHC
Plos One · DOI 10.1371/JOURNAL.PONE.0114631 · Source: orcidJohn St. John, Jaratlerdsiri, Weerachai, Deakin, Janine, Godinez, Ricardo M., Shan, Xueyan, Peterson, Daniel G., Marthey, Sylvain, Lyons, Eric · 15 authors totalDissecting noncoding and pathogen RNA–protein interactomes
RNA · DOI 10.1261/rna.047803.114 · 82 citations · Source: openalex+stanford-first-party+career-authorityLance Martin, Ryan A. Flynn, Robert C. Spitale, T. Brian, Selena M. Sagan, Brian Zarnegar, Kun Qu, Paul A. Khavari · 11 authors totalClinical validation of a comprehensive cancer genomics analysis for lung cancer patients.
Journal of Clinical Oncology · DOI 10.1200/jco.2014.32.15_suppl.e22122 · 1 citations · Source: openalex+authoritative-profilePetros Giannikopoulos, Catherine K. Foo, John A. St. John, Oscar Westesson, Nicholas Hahner, Aleah F. Caulin, Mitchell E. Skinner, Jeffrey Catalano · 17 authors totalIntegrated genomic analysis for revealing broad remodeling of EGFR-targeted therapy resistant lung cancers.
Journal of Clinical Oncology · DOI 10.1200/jco.2014.32.15_suppl.8083 · 0 citations · Source: openalex+authoritative-profilePetros Giannikopoulos, John A. St. John, Joel S. Parker, Oscar Westesson, Nicholas Hahner, Niki Karachaliou, Carlota Costa, Aleah F. Caulin · 20 authors totalPsychological Strategies for Winning a Geopolitical Forecasting Tournament
Psychological Science · DOI 10.1177/0956797614524255 · 327 citations · Source: openalexFive university-based research groups competed to recruit forecasters, elicit their predictions, and aggregate those predictions to assign the most accurate probabilities to events in a 2-year geopolitical forecasting tournament. Our group tested and found support for three psychological drivers of accuracy: training, teaming, and tracking. Probability training corrected cognitive biases, encouraged forecasters to use reference classes, and provided forecasters with heuristics, such as averaging when multiple estimates were available. Teaming allowed forecasters to share information and discuss the rationales behind their beliefs. Tracking placed the highest performers (top 2% from Year 1) in elite teams that worked together. Results showed that probability training, team collaboration, and tracking improved both calibration and resolution. Forecasting is often viewed as a statistical problem, but forecasts can be improved with behavioral interventions. Training, teaming, and tracking are psychological interventions that dramatically increased the accuracy of forecasts. Statistical algorithms (reported elsewhere) improved the accuracy of the aggregation. Putting both statistics and psychology to work produced the best forecasts 2 years in a row.
Lyle Ungar, Barbara A. Mellers, Jonathan Baron, Jaime Ramos, Burcu Gürçay, Katrina Fincher, Sydney Scott, Don A. Moore · 13 authors totalGrounded Compositional Semantics for Finding and Describing Images with Sentences
Transactions of the Association for Computational Linguistics · DOI 10.1162/tacl_a_00177 · 904 citations · Source: semantic-scholarPrevious work on Recursive Neural Networks (RNNs) shows that these models can produce compositional feature vectors for accurately representing and classifying sentences or images. However, the sentence vectors of previous models cannot accurately represent visually grounded meaning. We introduce the DT-RNN model which uses dependency trees to embed sentences into a vector space in order to retrieve images that are described by those sentences. Unlike previous RNN-based models which use constituency trees, DT-RNNs naturally focus on the action and agents in a sentence. They are better able to abstract from the details of word order and syntactic expression. DT-RNNs outperform other recursive and recurrent neural networks, kernelized CCA and a bag-of-words baseline on the tasks of finding an image that fits a sentence description and vice versa. They also give more similar representations to sentences that describe the same image.
Richard Socher, R. Socher, A. Karpathy, Quoc V. Le, Christopher D. Manning, A. Ng · 6 authors totalIntegrated genomic analysis by whole exome and transcriptome sequencing of tumor samples from EGFR-mutant non-small-cell lung cancer patients with acquired resistance to erlotinib
Cancer Research · DOI 10.1158/1538-7445.AM2014-954 · Source: orcidJohn St. John, Petros Giannikopoulos, Giannikopoulos, Petros, St John, John, Hahner, Nicholas, Parker, Joel S., Karachaliou, Niki, Costa, Carlota · 28 authors totalComprehensive integrated genomic analysis
Cancer Research · DOI 10.1158/1538-7445.AM2014-4707 · Source: orcidJohn St. John, Petros Giannikopoulos, Foo, Catherine K., John, John St., Hahner, Nicholas, Westesson, Oscar, Skinner, Mitchell E., Parikh, Urvish · 19 authors totalThe Impact of EGFR T790M Mutations and BIM mRNA Expression on Outcome in Patients with EGFR -Mutant NSCLC Treated with Erlotinib or Chemotherapy in the Randomized Phase III EURTAC Trial
Clinical Cancer Research · DOI 10.1158/1078-0432.ccr-13-2233 · 228 citations · Source: openalex+authoritative-profilePetros Giannikopoulos, Carlota Costa, Miguel Angel Molina, Ana Drozdowskyj, Ana Giménez‐Capitán, Jordi Bertrán-Alamillo, Niki Karachaliou, Radj Gervais · 20 authors totalSoftware defined unified monitoring and management of clouds
IBM Journal of Research and Development · DOI 10.1147/jrd.2014.2305313 · 3 citations · Source: crossref+semantic-scholarRuchi Mahindru, R. Mahindru, S. Sarkar, M. Viswanathan · 4 authors totalSession details: Session 6c: users vs. models
DOI 10.1145/3255814 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates · 2 authors totalSession details: Temporal web analytics workshop (TempWeb'14)
DOI 10.1145/3254754 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Marc Spaniol, Julien Masanés, Ricardo Baeza‐Yates · 4 authors totalLate data layout
Conference on Object-Oriented Programming Systems, Languages, and Applications · DOI 10.1145/2714064.2660197 · 7 citations · Source: semantic-scholar+epfl-infoscienceEugene Burmako, Vlad Ureche, E. Burmako, Martin Odersky · 4 authors totalWhat is Tumblr: a statistical overview and comparison
SIGKDD Explor. · DOI 10.1145/2674026.2674030 · Source: dblp+asu-first-party+career-authorityLei Tang, Yi Chang, Yoshiyuki Inagaki, Yan Liu · 4 authors totalTachyon: Reliable, Memory Speed Storage for Cluster Computing Frameworks
ACM Symposium on Cloud Computing · DOI 10.1145/2670979.2670985 · 368 citations · Source: semantic-scholarTachyon is a distributed file system enabling reliable data sharing at memory speed across cluster computing frameworks. While caching today improves read workloads, writes are either network or disk bound, as replication is used for fault-tolerance. Tachyon eliminates this bottleneck by pushing lineage, a well-known technique, into the storage layer. The key challenge in making a long-running lineage-based storage system is timely data recovery in case of failures. Tachyon addresses this issue by introducing a checkpointing algorithm that guarantees bounded recovery cost and resource allocation strategies for recomputation under commonly used resource schedulers. Our evaluation shows that Tachyon outperforms in-memory HDFS by 110x for writes. It also improves the end-to-end latency of a realistic workflow by 4x. Tachyon is open source and is deployed at multiple companies.
Haoyuan Li, Matei Zaharia, A. Ghodsi, M. Zaharia, S. Shenker, Ion Stoica · 6 authors total