Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗Simplicitly: foundations and applications of implicit function types
Proceedings of the ACM on Programming Languages · DOI 10.1145/3158130 · 32 citations · Source: openalexUnderstanding a program entails understanding its context; dependencies, configurations and even implementations are all forms of contexts. Modern programming languages and theorem provers offer an array of constructs to define contexts, implicitly. Scala offers implicit parameters which are used pervasively, but which cannot be abstracted over. This paper describes a generalization of implicit parameters to implicit function types , a powerful way to abstract over the context in which some piece of code is run. We provide a formalization based on bidirectional type-checking that closely follows the semantics implemented by the Scala compiler. To demonstrate their range of abstraction capabilities, we present several applications that make use of implicit function types. We show how to encode the builder pattern, tagless interpreters, reader and free monads and we assess the performance of the monadic structures presented.
Heather, Martin Odersky, Olivier Blanvillain, Fengyun Liu, Aggelos Biboudis, Heather Miller, Sandro Stucki · 7 authors totalLow Latency Stream Processing: Apache Heron with Infiniband & Intel Omni-Path
UCC · DOI 10.1145/3147213.3147232 · 15 citations · Source: semantic-scholar+dblpKarthik Ramasamy, Supun Kamburugamuve, Martin Swany, Geoffrey C. Fox · 4 authors totalSpark and Scala (keynote)
SCALA@SPLASH · DOI 10.1145/3136000.3148042 · 1 citations · Source: semantic-scholarReynold Xin · 1 author totalThe limitations of type classes as subtyped implicits (short paper)
SCALA@SPLASH · DOI 10.1145/3136000.3136006 · 1 citations · Source: dblp+semantic-scholarType classes in Scala are encoded with implicit parameters and subtyping. This short paper describes limitations that arise from encoding type classes as subtyped implicits, in particular around coherence and ambiguity.
Adelbert Chang · 1 author totalInteractive Development Using the Dotty Compiler
Scala Symposium · DOI 10.1145/3136000.3136005 · Source: acm+dblp+personal-first-partyGuillaume Martres · 1 author totalTypesafe abstractions for tensor operations (short paper)
SCALA@SPLASH · DOI 10.1145/3136000.3136001 · arXiv 1710.06892 · 14 citations · Source: semantic-scholarWe propose a typesafe abstraction to tensors (i.e. multidimensional arrays) exploiting the type-level programming capabilities of Scala through heterogeneous lists (HList), and showcase typesafe abstractions of common tensor operations and various neural layers such as convolution or recurrent neural networks. This abstraction could lay the foundation of future typesafe deep learning frameworks that runs on Scala/JVM.
Tongfei Chen · 1 author totalXML and JSON Are Like Cardboard
ACM Queue · DOI 10.1145/3134434.3143320 · 4 citations · Source: semantic-scholarCardboard is what you wrap things in for shipping; likewise semi-structured formats are packaging for data in motion.
Pat Helland · 1 author totalFA*IR
DOI 10.1145/3132847.3132938 · 490 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Meike Zehlike, Francesco Bonchi, Carlos Castillo, Sara Hajian, M. Megahed, Ricardo Baeza‐Yates · 7 authors totalResearch for practice
Communications of the ACM · DOI 10.1145/3132257 · 1 citations · Source: openalex+authoritative-profilePeter Bailis, John Regehr · 2 authors totalSelf-Regulating Streaming Systems: Challenges and Opportunities
BIRTE @ VLDB · DOI 10.1145/3129292.3129295 · 2 citations · Source: semantic-scholarAshvin Agrawal, Avrilia Floratou · 2 authors totalMitigating Poisoning Attacks on Machine Learning Models: A Data Provenance Based Approach
AISec@CCS · DOI 10.1145/3128572.3140450 · 127 citations · Source: semantic-scholarNathalie Baracaldo, Bryant Chen, Heiko Ludwig, Jaehoon Amir Safavi · 4 authors totalSLAQ: quality-driven scheduling for distributed machine learning
ACM Symposium on Cloud Computing · DOI 10.1145/3127479.3127490 · arXiv 1802.04819 · 156 citations · Source: semantic-scholarTraining machine learning (ML) models with large datasets can incur significant resource contention on shared clusters. This training typically involves many iterations that continually improve the quality of the model. Yet in exploratory settings, better models can be obtained faster by directing resources to jobs with the most potential for improvement. We describe SLAQ, a cluster scheduling system for approximate ML training jobs that aims to maximize the overall job quality. When allocating cluster resources, SLAQ explores the quality-runtime trade-offs across multiple jobs to maximize system-wide quality improvement. To do so, SLAQ leverages the iterative nature of ML training algorithms, by collecting quality and resource usage information from concurrent jobs, and then generating highly-tailored quality-improvement predictions for future iterations. Experiments show that SLAQ achieves an average quality improvement of up to 73% and an average delay reduction of up to 44% on a large set of ML training jobs, compared to resource fairness schedulers.
Andrew Or, Haoyu Zhang, Logan Stafman, M. Freedman · 4 authors totalHaskell a language for modern times
XRDS · DOI 10.1145/3123764 · 0 citations · Source: semantic-scholarMihai Maruseac · 1 author totalA metaprogramming framework for formal verification
Proc. ACM Program. Lang. (ICFP) · DOI 10.1145/3110278 · 101 citations · Source: semantic-scholar+dblpJared Roesch, Gabriel Ebner, Sebastian Ullrich, J. Avigad, L. D. Moura · 5 authors totalDéjà Vu: The Importance of Time and Causality in Recommender Systems
ACM Conference on Recommender Systems · DOI 10.1145/3109859.3109922 · 14 citations · Source: semantic-scholar+openalexTime plays a key role in recommendation. Handling it properly is especially critical when using recommender systems in real-world applications, which may not be as clear when doing research with historical data. In this talk, we will discuss some of the important challenges of handling time in recommendation algorithms at Netflix. We will focus on challenges related to how our users, items, and systems all change over time. We will then discuss some strategies for tackling these challenges, which revolves around proper treatment of causality in our systems.
Justin Basilico, Yves Raimond · 2 authors totalBenchmarks and Process Management in Data Science: Will We Ever Get Over the Mess?
KDD 2017 (Panel) · DOI 10.1145/3097983.3120998 · 4 citations · Source: semantic-scholarArno Candel, Eduardo Ariño de la Rubia, Usama Fayyad, Eduardo Arino de la Rubia, Szilard Pafka, Anthony Chong, Jeong-Yoon Lee · 7 authors totalMore than the Sum of its Parts: Building Domino Data Lab
KDD 2017 (Invited talk abstract) · DOI 10.1145/3097983.3106682 · 0 citations · Source: semantic-scholarEduardo Ariño de la Rubia, Eduardo Arino de la Rubia · 2 authors totalAutomated Fault Tree Analysis from AADL Models
ACM SIGAda Ada Letters · DOI 10.1145/3092893.3092900 · 26 citations · Source: semantic-scholar+openalexCyber-physical systems, used in domains such as avionics or medical devices, perform critical functions where a fault might have catastrophic consequences (mission failure, severe injuries, etc.). Their development is guided by rigorous practice standards that prescribe safety analysis methods in order to verify that failure have been correctly evaluated and/or mitigated. This laborintensive practice typically focuses system safety analysis on system engineering activities. As reliance on software for system operation grows, embedded software systems have become a major source of hazard contributors. Studies show that late discovery of errors in embedded software system have resulted in costly rework, making up as much as 50% of the total software system cost. Automation of the safety analysis process is key to extending safety analysis to the software system and to accommodate system evolution. In this paper we discuss three elements that are key to safety analysis automation in the context of fault tree analysis (FTA). First, generation of fault trees from annotated architecture models consistently reflects architecture changes in safety analysis results. Second, use of a taxonomy of failure effects ensures coverage of potential hazard contributors is achieved. Third, common cause failures are identified based on architecture information and reflected appropriately in probabilistic fault tree analysis. The approach utilizes the SAE Architecture Analysis & Design Language (AADL) standard and the recently published revised Error Model Annex V2 (EMV2) standard to represent annotated architecture models of systems and embedded software systems. The approach takes into account error sources specified with an EMV2 error propagation type taxonomy and occurrence probabilities as well as direct and indirect propagation paths between system components identified in the architecture model to generate a fault graph and apply transformations into a fault tree representatio
Julien Delange, Peter H. Feiler · 2 authors totalSide Effects, Front and Center!
ACM Queue · DOI 10.1145/3084693.3099561 · 2 citations · Source: semantic-scholarPat Helland · 1 author totalResearch for practice
Communications of the ACM · DOI 10.1145/3080188 · 2 citations · Source: openalex+authoritative-profilePeter Bailis, Tawanna R. Dillahunt, Stefanie Mueller, Patrick Baudisch · 4 authors totalDetection of Trending Topic Communities
DOI 10.1145/3078714.3078735 · 6 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Lorena Recalde, David Nettleton, Ricardo Baeza‐Yates, Ludovico Boratto · 5 authors totalDiagnosing Machine Learning Pipelines with Fine-grained Lineage
ACM International Symposium on High-Performance Parallel and Distributed Computing · DOI 10.1145/3078597.3078603 · Source: acm+dblp+berkeley-career-authorityEvan R. Sparks, Zhao Zhang, Evan Randall Sparks, Michael J. Franklin · 4 authors totalSemantic Query Understanding
DOI 10.1145/3077136.3096472 · 28 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates · 2 authors totalThe Lucene for Information Access and Retrieval Research (LIARR) Workshop at SIGIR 2017
Annual International ACM SIGIR Conference on Research and Development in Information Retrieval · DOI 10.1145/3077136.3084374 · 19 citations · Source: semantic-scholar+dblp+lucidworksGrant Ingersoll, L. Azzopardi, Matt Crane, Hui Fang, Jimmy J. Lin, Yashar Moshfeghi, Harrisen Scells, Peilin Yang · 9 authors totalTowards the Prediction of Dyslexia by a Web-based Game with Musical Elements
DOI 10.1145/3058555.3058565 · 18 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Maria Rauschenberger, Luz Rello, Ricardo Baeza‐Yates, Emília Gómez, Jeffrey P. Bigham · 6 authors totalProvable learning of noisy-OR networks.
STOC · DOI 10.1145/3055399.3055482 · Source: dblp+stanford-authorityTengyu Ma, Sanjeev Arora, Rong Ge 0001, Tengyu Ma 0001, Andrej Risteski · 5 authors totalFinding approximate local minima faster than gradient descent.
STOC · DOI 10.1145/3055399.3055464 · Source: dblp+stanford-authorityTengyu Ma, Naman Agarwal, Zeyuan Allen Zhu, Brian Bullins, Elad Hazan, Tengyu Ma 0001 · 6 authors totalToo Big NOT to Fail
ACM Queue · DOI 10.1145/3055301.3077383 · 6 citations · Source: semantic-scholarEmbracing failure as normal: how hyperscale services are engineered so that constant partial failure is survivable.
Pat Helland, Simon Weaver, Ed Harris · 3 authors totalResearch for practice
Communications of the ACM · DOI 10.1145/3052942 · 11 citations · Source: openalex+authoritative-profilePeter Bailis, Peter Alvaro, Sumit Gulwani · 3 authors totalGlobal Entity Ranking Across Multiple Languages
WWW Companion · DOI 10.1145/3041021.3054213 · arXiv 1703.06108 · 5 citations · Source: semantic-scholar+arxivWe present work on building a global long-tailed ranking of entities across multiple languages using Wikipedia and Freebase knowledge bases. We identify multiple features and build a model to rank entities using a ground-truth dataset of more than 10 thousand labels. The final system ranks 27 million entities with 75% precision and 48% F1 score. We provide performance evaluation and empirical evidence of the quality of ranking across languages, and open the final ranked lists for future research.
Nemanja Spasojevic, Prantik Bhattacharyya · 2 authors totalDAWT: Densely Annotated Wikipedia Texts Across Multiple Languages
WWW Companion · DOI 10.1145/3041021.3053367 · arXiv 1703.00948 · 9 citations · Source: semantic-scholar+arxivIn this work, we open up the DAWT dataset - Densely Annotated Wikipedia Texts across multiple languages. The annotations include labeled text mentions mapping to entities (represented by their Freebase machine ids) as well as the type of the entity. The data set contains total of 13.6M articles, 5.0B tokens, 13.8M mention entity co-occurrences. DAWT contains 4.8 times more anchor text to entity links than originally present in the Wikipedia markup. Moreover, it spans several languages including English, Spanish, Italian, German, French and Arabic. We also present the methodology used to generate the dataset which enriches Wikipedia markup in order to increase number of links. In addition to the main dataset, we open up several derived datasets including mention entity co-occurrence counts and entity embeddings, as well as mappings between Freebase ids and Wikidata item ids. We also discuss two applications of these datasets and hope that opening them up would prove useful for the Natural Language Processing and Information Retrieval communities, as well as facilitate multi-lingual research.
Nemanja Spasojevic, Preeti Bhargava, Guoning Hu · 3 authors totalExploring Query Auto-Completion and Click Logs for Contextual-Aware Web Search and Query Suggestion
DOI 10.1145/3038912.3052593 · 49 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Liangda Li, Hongbo Deng, Anlei Dong, Yi Chang, Ricardo Baeza‐Yates, Hongyuan Zha · 7 authors totalAn Architecture Supporting Formal and Compositional Binary Analysis
ASPLOS · DOI 10.1145/3037697.3037733 · 6 citations · Source: semantic-scholar+dblpJared Roesch, Joseph McMahan, Michael Christensen, L. Nichols, Sung-Yee Guo, Ben Hardekopf, T. Sherwood · 7 authors totalACIDRain
DOI 10.1145/3035918.3064037 · 50 citations · Source: openalex+authoritative-profilePeter Bailis, Todd Warszawski · 2 authors totalScalable Kernel Density Classification via Threshold-Based Pruning
DOI 10.1145/3035918.3064035 · 26 citations · Source: openalex+authoritative-profilePeter Bailis, Edward Gan · 2 authors totalZipG: A Memory-efficient Graph Store for Interactive Queries
SIGMOD Conference · DOI 10.1145/3035918.3064012 · 52 citations · Source: semantic-scholarAnurag Khandelwal, Zongheng Yang, Evan Ye, Rachit Agarwal, Ion Stoica · 5 authors totalDemonstration
DOI 10.1145/3035918.3056446 · 3 citations · Source: openalex+authoritative-profilePeter Bailis, Edward Gan, Kexin Rong, Sahaana Suri · 4 authors totalMacroBase
DOI 10.1145/3035918.3035928 · 102 citations · Source: openalex+authoritative-profilePeter Bailis, Edward Gan, Samuel Madden, Deepak Narayanan, Kexin Rong, Sahaana Suri · 6 authors totalDifferentially-Private Big Data Analytics for High-Speed Research Network Traffic Measurement
Conference on Data and Application Security and Privacy · DOI 10.1145/3029806.3029841 · 2 citations · Source: semantic-scholarMihai Maruseac, Oana Niculaescu, Gabriel Ghinita · 3 authors totalSubcontracting Microwork
CHI · DOI 10.1145/3025453.3025687 · Source: dblp+corestory-authorityAnand Kulkarni, Meredith Ringel Morris, Jeffrey P. Bigham, Robin Brewer, Jonathan Bragg, Jessie Li, Saiph Savage · 7 authors totalResearch for practice
Communications of the ACM · DOI 10.1145/3024928 · 39 citations · Source: openalex+authoritative-profilePeter Bailis, Arvind Narayanan, Andrew Miller, Song Han · 4 authors totalNext generation JDBC database drivers for performance, transparent caching, load balancing, and scale-out.
SAC · DOI 10.1145/3019612.3019870 · Source: dblp+ubc-authorityRamon Lawrence, Roland Lee, Erik Brandsberg · 3 authors totalTen Years of Wisdom
DOI 10.1145/3018661.3022744 · 2 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates · 2 authors totalResearch for practice
Communications of the ACM · DOI 10.1145/3009832 · 4 citations · Source: openalex+authoritative-profilePeter Bailis, Irene Zhang, Fadel Adib · 3 authors totalWhat happens when f0 movements and prosodic units were randomly aligned?
The Journal of the Acoustical Society of America · DOI 10.1121/1.5014199 · 0 citations · Source: openalex+first-party-career-authorityMark Liberman, Wei Lai, Nari Rhee · 3 authors totalThe Interdisciplinarity of Collaborations in Cognitive Science
Cognitive Science · DOI 10.1111/cogs.12352 · 32 citations · Source: semantic-scholarTill Bergmann, Rick Dale, Negin Sattari, Evan Heit, Harish S. Bhat · 5 authors totalLow-current Spin Transfer Torque MRAM
International Symposium on VLSI Technology, Systems, and Applications · DOI 10.1109/VLSI-TSA.2017.7942445 · 3 citations · Source: semantic-scholarAnthony Annunziata, G. Hu, Janusz J. Nowak, G. Lauer, J. Lee, J. Sun, J. Harms, A. Annunziata · 20 authors totalLow-current Spin Transfer Torque MRAM
International Symposium on VLSI Design, Automation and Test · DOI 10.1109/VLSI-DAT.2017.7939701 · 9 citations · Source: semantic-scholarAnthony Annunziata, G. Hu, J. Nowak, G. Lauer, J. Lee, J. Sun, J. Harms, A. Annunziata · 20 authors totalDomain randomization for transferring deep neural networks from simulation to the real world
IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) · DOI 10.1109/IROS.2017.8202133 · arXiv 1703.06907 · 3,995 citations · Source: openalex+semantic-scholarBridging the ‘reality gap’ that separates simulated robotics from experiments on hardware could accelerate robotic research through improved data availability. This paper explores domain randomization, a simple technique for training models on simulated images that transfer to real images by randomizing rendering in the simulator. With enough variability in the simulator, the real world may appear to the model as just another variation. We focus on the task of object localization, which is a stepping stone to general robotic manipulation skills. We find that it is possible to train a real-world object detector that is accurate to 1.5 cm and robust to distractors and partial occlusions using only data from a simulator with non-realistic random textures. To demonstrate the capabilities of our detectors, we show they can be used to perform grasping in a cluttered environment. To our knowledge, this is the first successful transfer of a deep neural network trained only on simulated RGB images (without pre-training on real images) to the real world for the purpose of robotic control.
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, Pieter Abbeel · 6 authors totalFinding Bottlenecks: Predicting Student Attrition with Unsupervised Classifier
IntelliSys 2017 (Intelligent Systems Conference) · DOI 10.1109/INTELLISYS.2017.8324279 · arXiv 1705.02687 · 3 citations · Source: arxiv+semantic-scholarWith pressure to increase graduation rates and reduce time to degree in higher education, it is important to identify at-risk students early. Automated early warning systems are therefore highly desirable. In this paper, we use unsupervised clustering techniques to predict the graduation status of declared majors in five departments at California State University Northridge (CSUN), based on a minimal number of lower division courses in each major. In addition, we use the detected clusters to identify hidden bottleneck courses.
Chris McKinlay, Seyed Sajjadi, Bruce Shapiro, Christopher McKinlay, Allen Sarkisyan, Carol Shubin, Efunwande Osoba · 7 authors total