Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗Robust estimation of mutation burden
Cancer Research · DOI 10.1158/1538-7445.AM2015-2173 · Source: orcidJohn St. John, Petros Giannikopoulos, Westesson, Oscar, Nielsen, Rasmus, St John, John, Caulin, Aleah, Hahner, Nicholas, Stewart, Stewart · 15 authors totalCritical BEOL Aspects of the Fabrication of a Thermally-Assisted MRAM Device
DOI 10.1149/06903.0127ECST · 1 citations · Source: semantic-scholarAnthony Annunziata, E. O'Sullivan, D. Edelstein, N. Marchack, M. Lofaro, M. Gaidis, E. Joseph, A. Annunziata · 14 authors totalMorris Halle: An Appreciation
Annual Review of Linguistics · DOI 10.1146/annurev-linguistics-060515-105131 · 1 citations · Source: openalex+first-party-career-authorityMark Liberman · 1 author totalSession details: Tutorials
DOI 10.1145/3261081 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates, Meeyoung Cha · 3 authors totalSession details: TempWeb 2015
DOI 10.1145/3261075 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Marc Spaniol, Ricardo Baeza‐Yates, Julien Masanés · 4 authors totalSession details: Keynote
DOI 10.1145/3255933 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates · 2 authors totalImmutability Changes Everything
Conference on Innovative Data Systems Research (CIDR) · DOI 10.1145/2857274.2884038 · 64 citations · Source: cidrdbArgues that falling storage costs have made append-only, immutable data the organizing principle of modern systems, from SSDs and log-structured storage to distributed data lakes and versioned datasets.
Pat Helland · 1 author totalData-centric metaprogramming in object-oriented languages
ICOOOLPS@ECOOP · DOI 10.1145/2843915.2843916 · 0 citations · Source: semantic-scholarVlad Ureche · 1 author totalDisseminating Location Privacy Threats and Protection Techniques through Capture-the-Flag Games (Vision Paper)
GeoPrivacy@SIGSPATIAL · DOI 10.1145/2830834.2830836 · 1 citations · Source: semantic-scholarRecent years witnessed a tremendous growth in the area of mobile computing. Users with mobile devices are able to access services customized to their geographical coordinates, and to engage in complex interactions with other users in their proximity. However, in addition to its many benefits, sharing location with service providers and other users also introduces serious privacy threats. If not properly addressed, the loss of location privacy can bring significant harm to mobile users. Currently, there is a low level of awareness among mobile users, and even among computer professionals, with respect to the contingent threats on location privacy, and to the approaches available to mitigate such threats. We envision GeoCTF, an educational capture-the-flag (CTF) - style game designed to raise the level of awareness about the dangers of uncontrolled sharing of location data, and to illustrate prominent location protection techniques. The game-based approach of GeoCTF makes it an effective and engaging educational tool, suitable for high-school and University students, as well as computer-literate general population mobile users.
Mihai Maruseac, Gabriel Ghinita · 2 authors totalClassification Models for New Language Communities
DOI 10.1145/2830629.2830633 · 0 citations · Source: openalex+personal-publication-listRob Munro, Robert Munro, Jessica Long, Nicholas Gaylord · 4 authors totalVuvuzela: scalable private messaging resistant to traffic analysis
Symposium on Operating Systems Principles · DOI 10.1145/2815400.2815417 · 335 citations · Source: semantic-scholarPrivate messaging over the Internet has proven challenging to implement, because even if message data is encrypted, it is difficult to hide metadata about who is communicating in the face of traffic analysis. Systems that offer strong privacy guarantees, such as Dissent [36], scale to only several thousand clients, because they use techniques with superlinear cost in the number of clients (e.g., each client broadcasts their message to all other clients). On the other hand, scalable systems, such as Tor, do not protect against traffic analysis, making them ineffective in an era of pervasive network monitoring. Vuvuzela is a new scalable messaging system that offers strong privacy guarantees, hiding both message data and metadata. Vuvuzela is secure against adversaries that observe and tamper with all network traffic, and that control all nodes except for one server. Vuvuzela's key insight is to minimize the number of variables observable by an attacker, and to use differential privacy techniques to add noise to all observable variables in a way that provably hides information about which users are communicating. Vuvuzela has a linear cost in the number of clients, and experiments show that it can achieve a throughput of 68,000 messages per second for 1 million users with a 37-second end-to-end latency on commodity servers.
Matei Zaharia, Jelle van den Hooff, David Lazar, M. Zaharia, N. Zeldovich · 5 authors totalAutomating ad hoc data representation transformations
Conference on Object-Oriented Programming Systems, Languages, and Applications · DOI 10.1145/2814270.2814271 · 15 citations · Source: semantic-scholarVlad Ureche, Aggelos Biboudis, Y. Smaragdakis, Martin Odersky · 4 authors totalBetter Malware Ground Truth: Techniques for Weighting Anti-Virus Vendor Labels
AISec@CCS · DOI 10.1145/2808769.2808780 · 119 citations · Source: semantic-scholarWe examine the problem of aggregating the results of multiple anti-virus (AV) vendors' detectors into a single authoritative ground-truth label for every binary. To do so, we adapt a well-known generative Bayesian model that postulates the existence of a hidden ground truth upon which the AV labels depend. We use training based on Expectation Maximization for this fully unsupervised technique. We evaluate our method using 279,327 distinct binaries from VirusTotal, each of which appeared for the first time between January 2012 and June 2014. Our evaluation shows that our statistical model is consistently more accurate at predicting the future-derived ground truth than all unweighted rules of the form "k out of n" AV detections. In addition, we evaluate the scenario where partial ground truth is available for model building. We train a logistic regression predictor on the partial label information. Our results show that as few as a 100 randomly selected training instances with ground truth are enough to achieve 80% true positive rate for 0.1% false positive rate. In comparison, the best unweighted threshold rule provides only 60% true positive rate at the same false positive rate.
Vaishaal Shankar, Alex Kantchelian, Michael Carl Tschantz, U. Berkeley, Brad Miller, Rekha Bachwani Netflix, UC AnthonyD.Joseph, Berkeley J. D. Tygar · 8 authors totalImproving the Interoperation between Generics Translations
Principles and Practice of Programming in Java · DOI 10.1145/2807426.2807436 · 6 citations · Source: semantic-scholarVlad Ureche, Miloš Stojanović, Romain Béguet, Nicolas Stucki, Martin Odersky · 5 authors totalAutomating Model Search for Large Scale Machine Learning
ACM Symposium on Cloud Computing · DOI 10.1145/2806777.2806945 · Source: acm+dblp+berkeley-career-authorityEvan R. Sparks, Ameet Talwalkar, Daniel Haas, Michael J. Franklin, Michael I. Jordan, Tim Kraska · 6 authors totalScalable Recommender Systems: Where Machine Learning Meets Search
RecSys · DOI 10.1145/2792838.2799672 · Source: publisher+dblp+first-party-career-authorityJoaquin Delgado · 1 author totalScalable Recommender Systems: Where Machine Learning Meets Search!
9th ACM Conference on Recommender Systems · DOI 10.1145/2792838.2792842 · Source: acm-crossref+recsys-program+speaker-career-authorityDiana Hu, Si Ying Diana Hu, Joaquin Delgado · 3 authors totalRRB vector: a practical general purpose immutable sequence
ACM SIGPLAN International Conference on Functional Programming · DOI 10.1145/2784731.2784739 · 27 citations · Source: semantic-scholarState-of-the-art immutable collections have wildly differing performance characteristics across their operations, often forcing programmers to choose different collection implementations for each task. Thus, changes to the program can invalidate the choice of collections, making code evolution costly. It would be desirable to have a collection that performs well for a broad range of operations. To this end, we present the RRB-Vector, an immutable sequence collection that offers good performance across a large number of sequential and parallel operations. The underlying innovations are: (1) the Relaxed-Radix-Balanced (RRB) tree structure, which allows efficient structural reorganization, and (2) an optimization that exploits spatio-temporal locality on the RRB data structure in order to offset the cost of traversing the tree. In our benchmarks, the RRB-Vector speedup for parallel operations is lower bounded by 7x when executing on 4 CPUs of 8 cores each. The performance for discrete operations, such as appending on either end, or updating and removing elements, is consistently good and compares favorably to the most important immutable sequence collections in the literature and in use today. The memory footprint of RRB-Vector is on par with arrays and an order of magnitude less than competing collections.
Vlad Ureche, Nicolas Stucki, Tiark Rompf, P. Bagwell · 4 authors totalWhen-To-Post on Social Networks
KDD · DOI 10.1145/2783258.2788584 · arXiv 1506.02089 · 74 citations · Source: semantic-scholar+arxivFor many users on social networks, one of the goals when broadcasting content is to reach a large audience. The probability of receiving reactions to a message differs for each user and depends on various factors, such as location, daily and weekly behavior patterns and the visibility of the message. While previous work has focused on overall network dynamics and message flow cascades, the problem of recommending personalized posting times has remained an under-explored topic of research. In this study, we formulate a when-to-post problem, where the objective is to find the best times for a user to post on social networks in order to maximize the probability of audience responses. To understand the complexity of the problem, we examine user behavior in terms of post-to-reaction times, and compare cross-network and cross-city weekly reaction behavior for users in different cities, on both Twitter and Facebook. We perform this analysis on over a billion posted messages and observed reactions, and propose multiple approaches for generating personalized posting schedules. We empirically assess these schedules on a sampled user set of 0.5 million active users and more than 25 million messages observed over a 56 day period. We show that users see a reaction gain of up to 17% on Facebook and 4% on Twitter when the recommended posting times are used. We open the dataset used in this study, which includes timestamps for over 144 million posts and over 1.1 billion reactions. The personalized schedules derived here are used in a fully deployed production system to recommend posting times for millions of users every day.
Adithya Rao, Nemanja Spasojevic, Zhisheng Li, Prantik Bhattacharyya · 4 authors totalTransfer Learning for Bilingual Content Classification
Knowledge Discovery and Data Mining · DOI 10.1145/2783258.2788575 · 16 citations · Source: semantic-scholarVita Markman, Qian Sun, Mohammad Shafkat Amin, B. Yan, C. Martell, V. Markman, Anmol Bhasin, Jieping Ye · 8 authors totalReferential integrity with Scala types
DOI 10.1145/2774975.2774979 · 0 citations · Source: openalex+career-authorityPatrick Prémont · 1 author totalIncremental Sampling of Query Logs
DOI 10.1145/2766462.2776780 · 29 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates · 2 authors totalPractical Lessons for Gathering Quality Labels at Scale
SIGIR · DOI 10.1145/2766462.2776778 · 14 citations · Source: dblp+semantic-scholarOmar Alonso · 1 author totalAnalyzing User's Sequential Behavior in Query Auto-Completion via Markov Processes
DOI 10.1145/2766462.2767723 · 51 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Liangda Li, Hongbo Deng, Anlei Dong, Yi Chang, Hongyuan Zha, Ricardo Baeza‐Yates · 7 authors totalBias and Reciprocity in Online Reviews: Evidence From Field Experiments on Airbnb
16th ACM Conference on Economics and Computation · DOI 10.1145/2764468.2764528 · Source: acm+dblp+author-first-partyDave Holtz, Andrey Fradkin, Elena Grewal, Matthew Pearson · 4 authors totalDebugging a Crowdsourced Task with Low Inter-Rater Agreement
JCDL · DOI 10.1145/2756406.2757741 · 21 citations · Source: dblp+semantic-scholarOmar Alonso, C. Marshall, Marc Najork · 3 authors totalUser Identification Across Social Media
ACM Trans. Knowl. Discov. Data · DOI 10.1145/2747880 · Source: dblp+asu-first-party+career-authorityLei Tang, Reza Zafarani, Huan Liu · 3 authors totalA plug-in to aid online reading in Spanish
DOI 10.1145/2745555.2746661 · 13 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Luz Rello, Roberto Carlini, Ricardo Baeza‐Yates, Jeffrey P. Bigham · 5 authors totalLarge-scale Contextual Query-to-Ad Matching and Retrieval System for Sponsored Search
DOI 10.1145/2740908.2744112 · 0 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates, Nemanja Djuric, Mihajlo Grbovic, Vladan Radosavljević, Fabrizio Silvestri · 6 authors totalTroubleshooting blackbox SDN control software with minimal causal sequences
Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication · DOI 10.1145/2740070.2626304 · 19 citations · Source: semantic-scholarAndrew Or, Colin Scott, Andreas Wundsam, B. Raghavan, Aurojit Panda, J. Lai, Eugene Huang, Zhi Liu · 13 authors totalEssential Web Pages Are Easy to Find
DOI 10.1145/2736277.2741100 · 24 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates, Paolo Boldi, Flavio Chierichetti · 4 authors totalChallenges with Label Quality for Supervised Learning
ACM J. Data Inf. Qual. · DOI 10.1145/2724721 · 32 citations · Source: dblp+semantic-scholarOmar Alonso · 1 author totalA Modern Student Experience inSystems Programming
ACM Conference on Learning @ Scale · DOI 10.1145/2724660.2728665 · 1 citations · Source: semantic-scholarVaishaal Shankar, D. Culler · 2 authors totalResource Elasticity for Large-Scale Machine Learning
SIGMOD · DOI 10.1145/2723372.2749432 · 57 citations · Source: semantic-scholar+dblpDeclarative large-scale machine learning (ML) aims at flexible specification of ML algorithms and automatic generation of hybrid runtime plans ranging from single node, in-memory computations to distributed computations on MapReduce (MR) or similar frameworks. State-of-the-art compilers in this context are very sensitive to memory constraints of the master process and MR cluster configuration. Different memory configurations can lead to significant performance differences. Interestingly, resource negotiation frameworks like YARN allow us to explicitly request preferred resources including memory. This capability enables automatic resource elasticity, which is not just important for performance but also removes the need for a static cluster configuration, which is always a compromise in multi-tenancy environments. In this paper, we introduce a simple and robust approach to automatic resource elasticity for large-scale ML. This includes (1) a resource optimizer to find near-optimal memory configurations for a given ML program, and (2) dynamic plan migration to adapt memory configurations during runtime. These techniques adapt resources according to data, program, and cluster characteristics. Our experiments demonstrate significant improvements up to 21x without unnecessary over-provisioning and low optimization overhead.
Frederick Reiss, Botong Huang, Matthias Boehm, Yuanyuan Tian, B. Reinwald, S. Tatikonda · 6 authors totalSpark SQL: Relational Data Processing in Spark
SIGMOD Conference · DOI 10.1145/2723372.2742797 · 1,494 citations · Source: semantic-scholarJoseph Bradley, Matei Zaharia, Reynold Xin, Michael Armbrust, Cheng Lian, Yin Huai, Davies Liu, Joseph K. Bradley · 11 authors totalTwitter Heron: Stream Processing at Scale
SIGMOD Conference · DOI 10.1145/2723372.2742788 · 631 citations · Source: semantic-scholar+dblpStorm has long served as the main platform for real-time analytics at Twitter. However, as the scale of data being processed in real-time at Twitter has increased, along with an increase in the diversity and the number of use cases, many limitations of Storm have become apparent. We need a system that scales better, has better debug-ability, has better performance, and is easier to manage -- all while working in a shared cluster infrastructure. We considered various alternatives to meet these needs, and in the end concluded that we needed to build a new real-time stream data processing system. This paper presents the design and implementation of this new system, called Heron. Heron is now the de facto stream data processing engine inside Twitter, and in this paper we also share our experiences from running Heron in production. In this paper, we also provide empirical evidence demonstrating the efficiency and scalability of Heron.
Karthik Ramasamy, Sanjeev Kulkarni, Nikunj Bhagat, Maosong Fu, Vikas Kedigehalli, Christopher Kellogg, Sailesh Mittal, Jignesh M. Patel · 9 authors totalRethinking Data-Intensive Science Using Scalable Analytics Systems
SIGMOD 2015 · DOI 10.1145/2723372.2742787 · 101 citations · Source: openalex"Next generation" data acquisition technologies are allowing scientists to collect exponentially more data at a lower cost. These trends are broadly impacting many scientific fields, including genomics, astronomy, and neuroscience. We can attack the problem caused by exponential data growth by applying horizontally scalable techniques from current analytics systems to accelerate scientific processing pipelines. In this paper, we describe ADAM, an example genomics pipeline that leverages the open-source Apache Spark and Parquet systems to achieve a 28x speedup over current genomics pipelines, while reducing cost by 63%. From building this system, we were able to distill a set of techniques for implementing scientific analyses efficiently using commodity "big data" systems. To demonstrate the generality of our architecture, we then implement a scalable astronomy image processing system which achieves a 2.8--8.9x improvement over the state-of-the-art MPI-based system.
Frank Austin Nothaft, Matt D. Massie, Timothy Danford, Zhao Zhang, Uri Laserson, Carl Yeksigian, Jey Kottalam, Arun Ahuja · 13 authors totalFeral Concurrency Control
DOI 10.1145/2723372.2737784 · 72 citations · Source: openalex+authoritative-profilePeter Bailis, Alan Fekete, Michael J. Franklin, Ali Ghodsi, Joseph M. Hellerstein, Ion Stoica · 6 authors totalSocial Networks Meet Distributed Systems
DOI 10.1145/2714576.2714606 · 6 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Nitin Chiluka, Nazareno Andrade, Henk Sips · 4 authors totalWisdom of the Crowd or Wisdom of a Few?
DOI 10.1145/2700171.2791056 · 46 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates, Diego Sáez-Trumper · 3 authors totalDocument Spanners
JACM · DOI 10.1145/2699442 · 60 citations · Source: semantic-scholar+dblpFrederick Reiss, Ronald Fagin, B. Kimelfeld, Stijn Vansummeren · 4 authors totalDifferentially-Private Mining of Moderately-Frequent High-Confidence Association Rules
Conference on Data and Application Security and Privacy · DOI 10.1145/2699026.2699102 · 6 citations · Source: semantic-scholarMihai Maruseac, Gabriel Ghinita · 2 authors totalTrust-based collection of information in distributed reputation networks
DOI 10.1145/2695664.2695868 · 2 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Dimitra Gkorou, Dick Epema · 3 authors totalScalability and Efficiency Challenges in Large-Scale Web Search Engines
DOI 10.1145/2684822.2697039 · 13 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, B. Barla Cambazoğlu, Ricardo Baeza‐Yates · 3 authors totalPredicting The Next App That You Are Going To Use
DOI 10.1145/2684822.2685302 · 167 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates, Di Jiang, Fabrizio Silvestri, Beverly L. Harrison · 5 authors totalSteal This Courseware: FOSS, Github, Python, and OpenShift (Abstract Only)
SIGCSE '15: ACM Technical Symposium on Computer Science Education · DOI 10.1145/2676723.2678305 · 1 citations · Source: semantic-scholarRemy DeCausemaker, Stephen Jacobs · 2 authors totalForum77: An Analysis of an Online Health Forum Dedicated to Addiction Recovery
CSCW 2015 · DOI 10.1145/2675133.2675146 · 109 citations · Source: semantic-scholarSonal Gupta, Diana L. MacLean, S. Gupta, A. Lembke, Christopher D. Manning, Jeffrey Heer · 6 authors totalOrganic vs. Sponsored Content: From Ads to Native Ads
2015 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT), pp. 229-230 · DOI 10.1109/WI-IAT.2015.163 · 0 citations · Source: crossref+dblp+semantic-scholarShort paper studying the difference between organic and sponsored (paid) content in a web content-discovery setting and what it takes to make sponsored recommendations behave like native content rather than ads.
Ashok Venkatesan, Soumyava Das, Akshay Soni, Debora Donato · 4 authors total