Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗PyTorch Distributed: Experiences on Accelerating Data Parallel Training
PVLDB (VLDB 2020) · DOI 10.14778/3415478.3415530 · arXiv 2006.15704 · 296 citations · Source: arxivThis paper presents the design, implementation, and evaluation of the PyTorch distributed data parallel module. PyTorch is a widely-adopted scientific computing package used in deep learning research and applications. Recent advances in deep learning argue for the value of large datasets and large models, which necessitates the ability to scale out model training to more computational resources.
Omkar Salpekar, Shen Li, Yanli Zhao, Rohan Varma, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith · 11 authors totalA demonstration of willump
Proceedings of the VLDB Endowment · DOI 10.14778/3415478.3415487 · 7 citations · Source: openalex+authoritative-profilePeter Bailis, Peter Kraft, Daniel Kang, Deepak Narayanan, Shoumik Palkar, Matei Zaharia · 6 authors totalApproximate partition selection for big-data workloads using summary statistics
Proceedings of the VLDB Endowment · DOI 10.14778/3407790.3407848 · 0 citations · Source: openalex+authoritative-profilePeter Bailis, Kexin Rong, Yao Lu, Srikanth Kandula, Phil Levis · 5 authors totalCoopStore
Proceedings of the VLDB Endowment · DOI 10.14778/3407790.3407817 · 10 citations · Source: openalex+authoritative-profilePeter Bailis, Edward Gan, Moses Charikar · 3 authors totalApproximate selection with guarantees using proxies
Proceedings of the VLDB Endowment · DOI 10.14778/3407790.3407804 · 28 citations · Source: openalex+authoritative-profilePeter Bailis, Daniel Kang, Edward Gan, Tatsunori Hashimoto, Matei Zaharia · 5 authors totalPredicting risk of dyslexia with an online gamified test
PLoS ONE · DOI 10.1371/journal.pone.0241687 · 75 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Luz Rello, Ricardo Baeza‐Yates, Abdullah Ali, Jeffrey P. Bigham, Miquel Serra‐Burriel · 6 authors totalAcoustic Prosodic Measures in Natural Speech of Progressive Supranuclear Palsy and Corticobasal Spectrum Disorders (4421)
Neurology · DOI 10.1212/wnl.94.15_supplement.4421 · 0 citations · Source: openalex+first-party-career-authorityMark Liberman, Natalia Parjane, Sharon Ash, Sunghye Cho, Murray Grossman, Naomi Nevler · 6 authors totalAutomatic Analysis of Lexical Features in Speech of Patients with Primary Progressive Aphasia (1965)
Neurology · DOI 10.1212/wnl.94.15_supplement.1965 · 0 citations · Source: openalex+first-party-career-authorityMark Liberman, Sunghye Cho, Naomi Nevler, Sharon Ash, Murray Grossman · 5 authors totalAutomated Analysis of Natural Speech in Amyotrophic Lateral Sclerosis (1903)
Neurology · DOI 10.1212/wnl.94.15_supplement.1903 · 0 citations · Source: openalex+first-party-career-authorityMark Liberman, Naomi Nevler, Sharon Ash, Corey T. McMillan, Lauren Elman, Leo McCluskey, David J. Irwin, Sunghye Cho · 9 authors totalAutomated analysis of natural speech in amyotrophic lateral sclerosis spectrum disorders
Neurology · DOI 10.1212/wnl.0000000000010366 · 36 citations · Source: openalex+first-party-career-authorityMark Liberman, Naomi Nevler, Sharon Ash, Corey T. McMillan, Lauren Elman, Leo McCluskey, David J. Irwin, Sunghye Cho · 9 authors totalUsing Fish Assemblages in a State Biological Assessment and Criteria Program: Essential Concepts and Considerations
DOI 10.1201/9781003068013-3 · 20 citations · Source: openalex+first-party-career-authorityMarc Smith, Chris O. Yoder, Marc A. Smith · 3 authors totalLearning dexterous in-hand manipulation
The International Journal of Robotics Research · DOI 10.1177/0278364919887447 · arXiv 1808.00177 · 2,260 citations · Source: openalex+crossrefWe use reinforcement learning (RL) to learn dexterous in-hand manipulation policies that can perform vision-based object reorientation on a physical Shadow Dexterous Hand. The training is performed in a simulated environment in which we randomize many of the physical properties of the system such as friction coefficients and an object’s appearance. Our policies transfer to the physical robot despite being trained entirely in simulation. Our method does not rely on any human demonstrations, but many behaviors found in human manipulation emerge naturally, including finger gaiting, multi-finger coordination, and the controlled use of gravity. Our results were obtained using the same distributed RL system that was used to train OpenAI Five. We also include a video of our results: https://youtu.be/jwSbzNHGflM.
Josh Tobin, Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron · 16 authors totalSelectivity and Robustness of Sparse Coding Networks
Journal of Vision · DOI 10.1167/jov.20.12.10 · Source: pubmed+berkeley+author-first-partyCharles Frye, Dylan M. Paiton, Charles G. Frye, Sheng Y. Lundquist, Joel D. Bowen, Ryan Zarcone, Bruno A. Olshausen · 7 authors totalSummEval: Re-evaluating Summarization Evaluation
Transactions of the Association for Computational Linguistics · DOI 10.1162/tacl_a_00373 · arXiv 2007.12626 · 1,057 citations · Source: semantic-scholarAbstract The scarcity of comprehensive up-to-date studies on evaluation metrics for text summarization and the lack of consensus regarding evaluation protocols continue to inhibit progress. We address the existing shortcomings of summarization evaluation methods along five dimensions: 1) we re-evaluate 14 automatic evaluation metrics in a comprehensive and consistent fashion using neural summarization model outputs along with expert and crowd-sourced human annotations; 2) we consistently benchmark 23 recent summarization models using the aforementioned automatic evaluation metrics; 3) we assemble the largest collection of summaries generated by models trained on the CNN/DailyMail news dataset and share it in a unified format; 4) we implement and share a toolkit that provides an extensible and unified API for evaluating summarization models across a broad range of automatic metrics; and 5) we assemble and share the largest and most diverse, in terms of model types, collection of human judgments of model-generated summaries on the CNN/Daily Mail dataset annotated by both expert judges and crowd-source workers. We hope that this work will help promote a more complete evaluation protocol for text summarization as well as advance research in developing evaluation metrics that better correlate with human judgments.
Richard Socher, A. R. Fabbri, Wojciech Kryscinski, Bryan McCann, R. Socher, Dragomir R. Radev · 6 authors totalTask-Oriented Dialogue as Dataflow Synthesis
Transactions of the Association for Computational Linguistics · DOI 10.1162/tacl_a_00333 · arXiv 2009.11423 · 177 citations · Source: semantic-scholarAbstract We describe an approach to task-oriented dialogue in which dialogue state is represented as a dataflow graph. A dialogue agent maps each user utterance to a program that extends this graph. Programs include metacomputation operators for reference and revision that reuse dataflow fragments from previous turns. Our graph-based state enables the expression and manipulation of complex user intents, and explicit metacomputation makes these intents easier for learned models to predict. We introduce a new dataset, SMCalFlow, featuring complex dialogues about events, weather, places, and people. Experiments show that dataflow graphs and metacomputation substantially improve representability and predictability in these natural dialogues. Additional experiments on the MultiWOZ dataset show that our dataflow representation enables an otherwise off-the-shelf sequence-to-sequence model to match the best existing task-specific state tracking model. The SMCalFlow dataset, code for replicating experiments, and a public leaderboard are available at https://www.microsoft.com/en-us/research/project/dataflow-based-dialogue-semantic-machines.
David Hall, Jayant Krishnamurthy, Jacob Andreas, J. Bufe, David Burkett, Charles C. Chen, Joshua Clausman, Jean Crawford · 45 authors totalMachine Learned Cellular Phenotypes in Cardiomyopathy Predict Sudden Death
Circulation Research · DOI 10.1161/circresaha.120.317345 · 54 citations · Source: openalex+authoritative-profilePeter Bailis, Albert J. Rogers, Anojan Selvalingam, Mahmood Alhusseini, David E. Krummen, Cesare Corrado, Firas Abuzaid, Tina Baykaner · 16 authors totalMachine Learning to Classify Intracardiac Electrical Patterns During Atrial Fibrillation
Circulation Arrhythmia and Electrophysiology · DOI 10.1161/circep.119.008160 · 72 citations · Source: openalex+authoritative-profilePeter Bailis, Mahmood Alhusseini, Firas Abuzaid, Albert J. Rogers, Junaid Zaman, Tina Baykaner, Paul Clopton, Matei Zaharia · 11 authors totalA Machine Learning Approach to Predict Air Quality in California
Complexity · DOI 10.1155/2020/8049504 · Source: publisher+nova-university+ydata-career-authorityFabiana Clemente, Mauro Castelli, Fabiana Martins Clemente, Ales Popovic, Sara Silva, Leonardo Vanneschi · 6 authors totalClinical Performance and Role of Expert Supervision of Deep Learning for Cardiac Ventricular Volumetry: A Validation Study.
Radiology: Artificial Intelligence · DOI 10.1148/ryai.2020190064 · 22 citations · Source: openalexPurpose To evaluate the performance of a deep learning (DL) algorithm for clinical measurement of right and left ventricular volume and function across cardiac MR images obtained for a range of clinical indications and pathologies. Materials and Methods A retrospective, Health Insurance Portability and Accountability Act-compliant study was conducted using the first 200 noncongenital clinical cardiac MRI examinations from June 2015 to June 2017 for which volumetry was available. Images were analyzed using commercially available software for automated DL-based and manual contouring of biventricular volumes. Fully automated measurements were compared using Pearson correlations, relative volume errors, and Bland-Altman analyses. Manual, automated, and expert revised contours for 50 MR images were examined by comparing regional Dice coefficients at the base, midventricle, and apex to further analyze the contour quality. Results Fully automated and manual left ventricular volumes were strongly correlated for end-systolic volume (ESV: Pearson r = 0.99, P < .001), end-diastolic volume (EDV: r = 0.97, P < .001), and ejection fraction (EF: r = 0.94, P < .001). Right ventricular measurements were also correlated for ESV (r = 0.93, P < .001), EDV (r = 0.92, P < .001), and EF (r = 0.73, P < .001). Visual inspection of segmentation quality showed most errors (73%) occurred at the cardiac base. Mean Dice coefficients between manual, automated, and expert revised contours ranged from 0.92 to 0.95, with greatest variance at the base and apex. Conclusion Fully automated ventricular segmentation by the tested algorithm provides contours and ventricular volumes that could be used to aid expert segmentation, but can benefit from expert supervision, particularly to resolve errors at the basal and apical slices. Supplemental material is available for this article. © RSNA, 2020.
Daniel Golden, T. Retson, Evan M. Masutani, A. Hsiao · 4 authors totalCall for Code: Developers tackle natural disasters with software
IBM Journal of Research and Development · DOI 10.1147/jrd.2019.2960241 · 2 citations · Source: semantic-scholar+arxivNatural disasters are increasing as highlighted in many reports including the Borgen Project. In 2018, David Clark Cause as creator and IBM as founding partner, in partnership with the United Nations Human Rights Office, the American Red Cross International Team, and The Linux Foundation, issued a “Call for Code” to developers to create robust projects that prepare communities for natural disasters and help them respond more quickly in their aftermath. This article covers the steps and tools used to engage with developers, the results from the first of five competitions to be run by the Call for Code Global Initiative over five years, and how the winners were selected. Insights from the mobilization of 100,000 developers toward this cause are described, as well as the lessons learned from running large-scale hackathons.
Susan Malaika, D. Krook, S. Malaika · 3 authors totalBirdsong Learning and Culture: Analogies with Human Spoken Language
Annual Review of Linguistics · DOI 10.1146/annurev-linguistics-090420-121034 · 51 citations · Source: openalex+first-party-career-authorityMark Liberman, Julia Hyland Bruno, Erich D. Jarvis, Ofer Tchernichovski · 4 authors totalHoplite: efficient and fault-tolerant collective communication for task-based distributed systems
SIGCOMM 2021 · DOI 10.1145/3452296.3472897 · arXiv 2002.05814 · 35 citations · Source: arxiv+semantic-scholarTask-based distributed frameworks (e.g., Ray, Dask, Hydro) have become increasingly popular for distributed applications that contain asynchronous and dynamic workloads, including asynchronous gradient descent, reinforcement learning, and model serving. As more data-intensive applications move to run on top of task-based systems, collective communication efficiency has become an important problem. Unfortunately, traditional collective communication libraries (e.g., MPI, Horovod, NCCL) are an ill fit, because they require the communication schedule to be known before runtime and they do not provide fault tolerance. We design and implement Hoplite, an efficient and fault-tolerant collective communication layer for task-based distributed systems. Our key technique is to compute data transfer schedules on the fly and execute the schedules efficiently through fine-grained pipelining. At the same time, when a task fails, the data transfer schedule adapts quickly to allow other tasks to keep making progress. We apply Hoplite to a popular task-based distributed framework, Ray. We show that Hoplite speeds up asynchronous stochastic gradient descent, reinforcement learning, and serving an ensemble of machine learning models that are difficult to execute efficiently with traditional collective communication by up to 7.8x, 3.9x, and 3.3x, respectively.
Zhuohan Li, Siyuan Zhuang, Danyang Zhuo, Stephanie Wang, Eric Liang, Robert Nishihara, Philipp Moritz, Ion Stoica · 8 authors totalBaleen Analytics
ACM Queue · DOI 10.1145/3442632.3446917 · 0 citations · Source: semantic-scholarPat Helland · 1 author totalHopsFS-S3: Extending Object Stores with POSIX-like Semantics and more (industry track)
Middleware Industry · DOI 10.1145/3429357.3430521 · Source: dblp+first-party-career-authorityJim Dowling, Mahmoud Ismail, Salman Niazi, Gautier Berthou, Mikael Ronström, Seif Haridi · 6 authors totalTowards the Science of Essential Decentralised Infrastructures
DOI 10.1145/3428662.3429744 · 3 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse · 1 author totalConTrib
DOI 10.1145/3428662.3428789 · 1 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Martijn de Vos · 2 authors totalMaggy: Scalable Asynchronous Parallel Hyperparameter Search
DistributedML@CoNEXT · DOI 10.1145/3426745.3431338 · Source: dblp+first-party-career-authorityJim Dowling, Moritz Meister, Sina Sheikholeslami, Amir Hossein Payberah, Vladimir Vlassov · 5 authors totalScalaPy: seamless Python interoperability for cross-platform Scala programs
SCALA@SPLASH · DOI 10.1145/3426426.3428485 · 3 citations · Source: semantic-scholarIn recent years, Python has become the language of choice for data scientists with its many high-quality scientific libraries and Scala has become the go-to language for big data systems. In this paper, we bridge these languages with ScalaPy, a system for interoperability between Scala and Python. With ScalaPy, developers can use Python libraries in Scala by treating Python values as Scala objects and exposing Scala values to Python. ScalaPy supports both Scala on the JVM and Scala Native, enabling its usage from data experiments in interactive notebook environments to performance-critical production systems. In this paper, we explore the challenges involved with mixing the semantics and implementations of these two disparate languages.
Shadaj Laddad, Koushik Sen · 2 authors totalFluid quotes: metaprogramming across abstraction boundaries with dependent types
International Conference on Generative Programming: Concepts and Experiences · DOI 10.1145/3425898.3426953 · 2 citations · Source: semantic-scholarObject-oriented programming, functional programming, and metaprogramming each offer a unique axis of abstraction that enables modular code. Macros, a common technique for metaprogramming, capture ASTs as quotes to let users manipulate them in the host language. However, macros are often at odds with other programming techniques since they can only process code written at the call-site and cannot analyze code behind abstraction boundaries such as variables and methods. Furthermore, the quotes generated for macro expansion exist only at compile-time and cannot be passed around in user code. Multi-stage programming treats quotes as runtime values to address this problem, but introduces the cost of running the compiler when splicing quotes. This forces developers to choose between low runtime overhead and modularity. What if we could have the best of both worlds? We introduce fluid quotes, a new technique that uses dependent types to let users pass quotes through abstraction boundaries in runtime code while splicing them ahead-of-time. This technique enables new metaprogramming capabilities by eliminating the traditional requirement of co-locating parameter expressions with call-sites. Fluid quotes capture not only source code but also associated runtime context to ensure correctness. In addition, they can be composed into larger expressions without any macro code. We demonstrate the capabilities of fluid quotes through two specific applications: optimizing data processing pipelines and making language integrated queries more flexible.
Shadaj Laddad, Koushik Sen · 2 authors totalMATCH
DOI 10.1145/3423211.3425678 · 3 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Martijn de Vos, Georgy Ishmaev · 3 authors totalServerless linear algebra
ACM Symposium on Cloud Computing · DOI 10.1145/3419111.3421287 · 119 citations · Source: semantic-scholarDatacenter disaggregation provides numerous benefits to both the datacenter operator and the application designer. However switching from the server-centric model to a disaggregated model requires developing new programming abstractions that can achieve high performance while benefiting from the greater elasticity. To explore the limits of datacenter disaggregation, we study an application area that near-maximally benefits from current server-centric datacenters: dense linear algebra. We build NumPyWren, a system for linear algebra built on a disaggregated serverless programming model, and LAmbdaPACK, a companion domain-specific language designed for serverless execution of highly parallel linear algebra algorithms. We show that, for a number of linear algebra algorithms such as matrix multiply, singular value decomposition, Cholesky decomposition, and QR decomposition, NumPyWren's performance (completion time) is within a factor of 2 of optimized server-centric MPI implementations, and has up to 15% greater compute efficiency (total CPU-hours), while providing fault tolerance.
Vaishaal Shankar, K. Krauth, Kailas Vodrahalli, Qifan Pu, B. Recht, Ion Stoica, Jonathan Ragan-Kelley, Eric Jonas · 9 authors totalDevelopments in MLflow: A System to Accelerate the Machine Learning Lifecycle
DEEM@SIGMOD 2020 · DOI 10.1145/3399579.3399867 · 146 citations · Source: semantic-scholarMLflow is a popular open source platform for managing ML development, including experiment tracking, reproducibility, and deployment. In this paper, we discuss user feedback collected since MLflow was launched in 2018, as well as three major features we have introduced in response to this feedback: a Model Registry for collaborative model management and review, tools for simplifying ML code instrumentation, and experiment analytics functions for extracting insights from millions of ML experiments.
Aaron Davidson, Andrew Chen, Andy Chow, Arjun DCunha, Ali Ghodsi, Sue Ann Hong, Andy Konwinski, Clemens Mewald · 18 authors totalColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Annual International ACM SIGIR Conference on Research and Development in Information Retrieval · DOI 10.1145/3397271.3401075 · arXiv 2004.12832 · 2,419 citations · Source: semantic-scholar+openalexRecent progress in Natural Language Understanding (NLU) is driving fast-paced advances in Information Retrieval (IR), largely owed to fine-tuning deep language models (LMs) for document ranking. While remarkably effective, the ranking models based on these LMs increase computational cost by orders of magnitude over prior approaches, particularly as they must feed each query-document pair through a massive neural network to compute a single relevance score. To tackle this, we present ColBERT, a novel ranking model that adapts deep LMs (in particular, BERT) for efficient retrieval. ColBERT introduces a late interaction architecture that independently encodes the query and the document using BERT and then employs a cheap yet powerful interaction step that models their fine-grained similarity. By delaying and yet retaining this fine-granular interaction, ColBERT can leverage the expressiveness of deep LMs while simultaneously gaining the ability to pre-compute document representations offline, considerably speeding up query processing. Crucially, ColBERT's pruning-friendly interaction mechanism enables leveraging vector-similarity indexes for end-to-end retrieval directly from millions of documents. We extensively evaluate ColBERT using two recent passage search datasets. Results show that ColBERT's effectiveness is competitive with existing BERT-based models (and outperforms every non-BERT baseline), while executing two orders-of-magnitude faster and requiring up to four orders-of-magnitude fewer FLOPs per query.
Matei Zaharia, O. Khattab, M. Zaharia · 3 authors totalScaffle: bug localization on millions of files
International Symposium on Software Testing and Analysis · DOI 10.1145/3395363.3397356 · 35 citations · Source: openalex+semantic-scholarDespite all efforts to avoid bugs, software sometimes crashes in the field, leaving crash traces as the only information to localize the problem. Prior approaches on localizing where to fix the root cause of a crash do not scale well to ultra-large scale, heterogeneous code bases that contain millions of code files written in multiple programming languages. This paper presents Scaffle, the first scalable bug localization technique, which is based on the key insight to divide the problem into two easier sub-problems. First, a trained machine learning model predicts which lines of a raw crash trace are most informative for localizing the bug. Then, these lines are fed to an information retrieval-based search engine to retrieve file paths in the code base, predicting which file to change to address the crash. The approach does not make any assumptions about the format of a crash trace or the language that produces it. We evaluate Scaffle with tens of thousands of crash traces produced by a large-scale industrial code base at Facebook that contains millions of possible bug locations and that powers tools used by billions of people. The results show that the approach correctly predicts the file to fix for 40% to 60% (50% to 70%) of all crash traces within the top-1 (top-5) predictions. Moreover, Scaffle improves over several baseline approaches, including an existing classification-based approach, a scalable variant of existing information retrieval-based approaches, and a set of hand-tuned, industrially deployed heuristics.
Erik Meijer, Michael Pradel, Vijayaraghavan Murali, Rebecca Qian, Mateusz Machalica, Satish Chandra · 6 authors totalSparkFuzz: searching correctness regressions in modern query engines
DBTest@SIGMOD · DOI 10.1145/3395032.3395327 · 19 citations · Source: semantic-scholarWith more than 1200 contributors, Apache Spark is one of the most actively developed open source projects. At this scale and pace of development, mistakes are bound to happen. In this paper we present SparkFuzz, a toolkit we developed at Databricks for uncovering correctness errors in the Spark SQL engine. To guard the system against correctness errors, SparkFuzz takes a fuzzing approach to testing by generating random data and queries. Spark-Fuzz executes the generated queries on a reference database system such as PostgreSQL which is then used as a test oracle to verify the results returned by Spark SQL. We explain the approach we take to data and query generation and we analyze the coverage of SparkFuzz. We show that SparkFuzz achieves its current maximum coverage relatively fast by generating a small number of queries.
Reynold Xin, Bogdan Ghit, Nicolás Poggi, Josh Rosen, P. Boncz · 5 authors totalVamsa: Automated Provenance Tracking in Data Science Scripts
KDD · DOI 10.1145/3394486.3403205 · arXiv 2001.01861 · 73 citations · Source: semantic-scholarAshvin Agrawal, Mohammad Hossein Namaki, Avrilia Floratou, Fotis Psallidas, Subru Krishnan, Yinghui Wu, Yiwen Zhu, Markus Weimer · 8 authors totalWES: agent-based user interaction simulation on real infrastructure
International Conference on Software Engineering · DOI 10.1145/3387940.3392089 · arXiv 2004.05363 · 25 citations · Source: openalex+semantic-scholarWe introduce the Web-Enabled Simulation (WES) research agenda, and describe FACEBOOK's WW system. We describe the application of WW to reliability, integrity and privacy at FACEBOOK1, where it is used to simulate social media interactions on an infrastructure consisting of hundreds of millions of lines of code. The WES agenda draws on research from many areas of study, including Search Based Software Engineering, Machine Learning, Programming Languages, Multi Agent Systems, Graph Theory, Game AI, and AI Assisted Game Play. We conclude with a set of open problems and research challenges to motivate wider investigation.
Erik Meijer, John Ahlgren, Maria Eugenia Berezin, Kinga Bojarczuk, Elena Dulskyte, Inna Dvortsova, Johann George, Natalija Gucevska · 12 authors totalPersonalization, Bias and Privacy
DOI 10.1145/3386392.3399994 · 25 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates · 2 authors totalThe Best Place to Build a Subway
ACM Queue · DOI 10.1145/3386269 · 1 citations · Source: semantic-scholarPat Helland · 1 author totalThe Seattle Report on Database Research
ACM SIGMOD Record · DOI 10.1145/3385658.3385668 · 71 citations · Source: openalex+authoritative-profilePeter Bailis, Daniel J. Abadi, Anastasia Ailamaki, David F. Andersen, Magdalena Bałazińska, Philip A. Bernstein, Peter Boncz, Surajit Chaudhuri · 33 authors totalBias in Search and Recommender Systems
DOI 10.1145/3383313.3418435 · 76 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates · 2 authors totalOptimal Data Placement for Heterogeneous Cache, Memory, and Storage Systems
Proceedings of the ACM on Measurement and Analysis of Computing Systems · DOI 10.1145/3379472 · Source: acm+sigmetricsIrfan Ahmad, Lei Zhang, Reza Karimi, Ymir Vigfusson · 4 authors totalBias on the web and beyond
DOI 10.1145/3371300.3385335 · 4 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates · 2 authors totalScreening risk of dyslexia through a web-game using language-independent content and machine learning
DOI 10.1145/3371300.3383342 · 32 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Maria Rauschenberger, Ricardo Baeza‐Yates, Luz Rello · 4 authors totalTiFL: A Tier-based Federated Learning System
IEEE International Symposium on High-Performance Parallel Distributed Computing · DOI 10.1145/3369583.3392686 · arXiv 2001.09249 · 376 citations · Source: semantic-scholar+arxivFederated Learning (FL) enables learning a shared model acrossmany clients without violating the privacy requirements. One of the key attributes in FL is the heterogeneity that exists in both resource and data due to the differences in computation and communication capacity, as well as the quantity and content of data among different clients. We conduct a case study to show that heterogeneity in resource and data has a significant impact on training time and model accuracy in conventional FL systems. To this end, we propose TiFL, a Tier-based Federated Learning System, which divides clients into tiers based on their training performance and selects clients from the same tier in each training round to mitigate the straggler problem caused by heterogeneity in resource anddata quantity. To further tame the heterogeneity caused by non-IID (Independent and Identical Distribution) data and resources, TiFL employs an adaptive tier selection approach to update the tiering on-the-fly based on the observed training performance and accuracy. We prototype TiFL in a FL testbed following Google's FL architecture and evaluate it using the state-of-the-art FL benchmarks. Experimental evaluation shows that TiFL outperforms the conventional FL in various heterogeneous conditions. With the proposed adaptive tier selection policy, we demonstrate that TiFL achieves much faster training performance while achieving the same or better test accuracy across the board.
Nathalie Baracaldo, Zheng Chai, Ahsan Ali, Syed Zawad, Stacey Truex, Ali Anwar, Yi Zhou, Heiko Ludwig · 10 authors totalBiases on Social Media Data
Companion Proceedings of the Web Conference 2020 · DOI 10.1145/3366424.3383564 · 2 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates · 2 authors totalRepresentativeness of Abortion Legislation Debate on Twitter: A Case Study in Argentina and Chile
Companion Proceedings of the Web Conference 2020 · DOI 10.1145/3366424.3383561 · 17 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Eduardo Graells-Garrido, Ricardo Baeza‐Yates, Mounia Lalmas · 4 authors totalSpur: Mitigating Slow Instances in Large-Scale Streaming Pipelines
SIGMOD · DOI 10.1145/3318464.3386142 · 5 citations · Source: semantic-scholarBing's monetization pipeline is one of the largest and most critical streaming workloads deployed in Microsoft's internal data lake. The pipeline runs 24/7 at a scale of 3500 YARN containers and is required to meet a Service Level Objective of low tail latency.
Ashvin Agrawal, Ke Wang, Avrilia Floratou, Daniel Musgrave · 4 authors totalLe Taureau: Deconstructing the Serverless Landscape & A Look Forward
SIGMOD Conference · DOI 10.1145/3318464.3383130 · 18 citations · Source: semantic-scholarAkin to the natural evolution of programming in assembly language to high-level languages, serverless computing represents the next frontier in the evolution of cloud computing: bare metal -> virtual machines -> containers -> serverless. The genesis of serverless computing can be traced back to the fundamental need of enabling a programmer to singularly focus on writing application code in a high-level language and isolating all facets of system management (for example, but not limited to, instance selection, scaling, deployment, logging, monitoring, fault tolerance and so on). This is particularly critical in light of today's, increasingly tightening, time-to-market constraints. Currently, serverless computing is supported by leading public cloud vendors, such as AWS Lambda, Google Cloud Functions, Azure Cloud Functions and others. While this is an important step in the right direction, there are many challenges going forward. For instance, but not limited to, how to enable support for dynamic optimization, how to extend support for stateful computation, how to efficiently bin-pack applications, how to support hardware heterogeneity (this will be key especially in light of the emergence of hardware accelerators for deep learning workloads). Inspired by Picasso's Le Taureau, in the tutorial proposed herein, we shall deconstruct evolution of serverless --- the overarching intent being to facilitate better understanding of the serverless landscape. This, we hope, would help push the innovation frontier on both fronts, the paradigm itself and the applications built atop of it.
Anurag Khandelwal, Karthik Ramasamy, Arun Kejariwal, Karthikeyan Ramasamy · 4 authors totalProposal Learning for Semi-Supervised Object Detection
IEEE Workshop/Winter Conference on Applications of Computer Vision · DOI 10.1109/WACV48630.2021.00234 · arXiv 2001.05086 · 111 citations · Source: arxiv+semantic-scholarIn this paper, we focus on semi-supervised object detection to boost performance of proposal-based object detectors (a.k.a. two-stage object detectors) by training on both labeled and unlabeled data. However, it is non-trivial to train object detectors on unlabeled data due to the un-availability of ground truth labels. To address this problem, we present a proposal learning approach to learn proposal features and predictions from both labeled and unlabeled data. The approach consists of a self-supervised proposal learning module and a consistency-based proposal learning module. In the self-supervised proposal learning module, we present a proposal location loss and a contrastive loss to learn context-aware and noise-robust proposal features respectively. In the consistency-based proposal learning module, we apply consistency losses to both bounding box classification and regression predictions of proposals to learn noise-robust proposal features and predictions. Our approach enjoys the following benefits: 1) encouraging more context information to be delivered in the proposals learning procedure; 2) noisy proposal features and enforcing consistency to allow noise-robust object detection; 3) building a general and high-performance semi-supervised object detection framework, which can be easily adapted to proposal-based object detectors with different backbone architectures. Experiments are conducted on the COCO dataset with all available labeled and unlabeled data. Results demonstrate that our approach consistently improves the performance of fully-supervised baselines. In particular, after combining with data distillation [39], our approach improves AP by about 2.0% and 0.9% on average compared to fully-supervised baselines and data distillation baselines respectively.
Ran Xu, Peng Tang, Chetan Ramaiah, Caiming Xiong · 4 authors total