Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗Managing data transfers in computer clusters with orchestra
Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication · DOI 10.1145/2018436.2018448 · 677 citations · Source: semantic-scholar+openalexCluster computing applications like MapReduce and Dryad transfer massive amounts of data between their computation stages. These transfers can have a significant impact on job performance, accounting for more than 50% of job completion times. Despite this impact, there has been relatively little work on optimizing the performance of these data transfers, with networking researchers traditionally focusing on per-flow traffic management. We address this limitation by proposing a global management architecture and a set of algorithms that (1) improve the transfer times of common communication patterns, such as broadcast and shuffle, and (2) allow scheduling policies at the transfer level, such as prioritizing a transfer over other transfers. Using a prototype implementation, we show that our solution improves broadcast completion times by up to 4.5X compared to the status quo in Hadoop. We also show that transfer-level scheduling can reduce the completion time of high-priority transfers by 1.7X.
Matei Zaharia, Mosharaf Chowdhury, M. Zaharia, Justin Ma, Michael I. Jordan, Ion Stoica · 6 authors totalBuilding a highly consumable semantic model for smarter cities
AIIP '11 · DOI 10.1145/2018316.2018319 · 37 citations · Source: crossref+semantic-scholarRosario Uceda-Sosa, Rosario A. Uceda-Sosa, Biplav Srivastava, Robert J. Schloss · 4 authors totalWeb retrieval
DOI 10.1145/2009916.2010172 · 46 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates, Yoelle Maarek · 3 authors totalCrowdsourcing for information retrieval: principles, methods, and applications
SIGIR · DOI 10.1145/2009916.2010170 · 32 citations · Source: dblp+semantic-scholarOmar Alonso, Matthew Lease · 2 authors totalScalable multi-dimensional user intent identification using tree structured distributions
DOI 10.1145/2009916.2009971 · 10 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Vinay Jethava, Liliana Calderón-Benavides, Ricardo Baeza‐Yates, Chiranjib Bhattacharyya, Devdatt Dubhashi · 6 authors totalThe world according to LINQ
Communications of the ACM · DOI 10.1145/2001269.2001285 · 71 citations · Source: openalex+semantic-scholarBig data is about more than size, and LINQ is more than up to the task.
Erik Meijer · 1 author totalInterpreting Relational Databases in the RDF Domain
K-CAP · DOI 10.1145/1999676.1999699 · Source: acm+dblp+w3c-authorityAlexandre Bertails, Eric Prud'hommeaux · 2 authors totalVMFlock: Virtual Machine Co-Migration for the Cloud
ACM International Symposium on High Performance Distributed Computing · DOI 10.1145/1996130.1996153 · Source: acm+dblp+ibm-career-authorityDinesh Subhraveti, Samer Al-Kiswany, Prasenjit Sarkar, Matei Ripeanu · 4 authors totalPersonalisation in the wild: providing personalisation across semantic, social and open-web resources
HT · DOI 10.1145/1995966.1995979 · Source: dblp+adapt-autodesk-authorityAlex O'Connor, Ben Steichen, Alexander O'Connor, Vincent Wade · 4 authors totalRecord and Transplay: Partial Checkpointing for Replay Debugging Across Heterogeneous Systems
ACM SIGMETRICS · DOI 10.1145/1993744.1993757 · Source: acm+dblp+columbia-career-authorityDinesh Subhraveti, Jason Nieh · 2 authors totalThe SystemT IDE: an integrated development environment for information extraction rules
SIGMOD · DOI 10.1145/1989323.1989479 · 19 citations · Source: semantic-scholar+dblpInformation Extraction (IE)-the problem of extracting structured information from unstructured text - has become the key enabler for many enterprise applications such as semantic search, business analytics and regulatory compliance. While rule-based IE systems are widely used in practice due to their well-known "explainability," developing high-quality information extraction rules is known to be a labor-intensive and time-consuming iterative process. Our demonstration showcases SystemT IDE, the integrated development environment for SystemT, a state-of-the-art rule-based IE system from IBMResearch that has been successfully embedded in multiple IBM enterprise products. SystemT IDE facilitates the development, test and analysis of high-quality IE rules by means of sophisticated techniques, ranging from data management to machine learning. We show how to build high-quality IE annotators using a suite of tools provided by SystemT IDE, including computing data provenance, learning basic features such as regular expressions and dictionaries, and automatically refining rules based on labeled examples.
Frederick Reiss, Laura Chiticariu, Vivian Chu, Sajib Dasgupta, Thilo W. Goetz, C. T. H. Ho, R. Krishnamurthy, Alexander Lang · 13 authors totalCrowdDB: answering queries with crowdsourcing
ACM SIGMOD Conference · DOI 10.1145/1989323.1989331 · 726 citations · Source: semantic-scholarSome queries cannot be answered by machines only. Processing such queries requires human input for providing information that is missing from the database, for performing computationally difficult functions, and for matching, ranking, or aggregating results based on fuzzy criteria. CrowdDB uses human input via crowdsourcing to process queries that neither database systems nor search engines can adequately answer. It uses SQL both as a language for posing complex queries and as a way to model data. While CrowdDB leverages many aspects of traditional database systems, there are also important differences. Conceptually, a major change is that the traditional closed-world assumption for query processing does not hold for human input. From an implementation perspective, human-oriented query operators are needed to solicit, integrate and cleanse crowdsourced data. Furthermore, performance and cost depend on a number of new factors including worker affinity, training, fatigue, motivation and location. We describe the design of CrowdDB, report on an initial set of experiments using Amazon Mechanical Turk, and outline important avenues for future work in the development of crowdsourced query processing systems.
Reynold Xin, M. Franklin, Donald Kossmann, Tim Kraska, Sukriti Ramesh · 5 authors totalEstimating dyslexia in the web
DOI 10.1145/1969289.1969300 · 26 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates, Luz Rello · 3 authors totalParallel symbolic execution for automated real-world software testing
European Conference on Computer Systems · DOI 10.1145/1966445.1966463 · 262 citations · Source: semantic-scholarVlad Ureche, Stefan Bucur, Cristian Zamfir, George Candea · 4 authors totalThe 1st temporal web analytics workshop (TWAW)
DOI 10.1145/1963192.1963325 · 16 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates, Julien Masanés, Marc Spaniol · 4 authors totalDistributed web retrieval
DOI 10.1145/1963192.1963310 · 1 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates · 2 authors totalIf You Have Too Much Data, then 'Good Enough' Is Good Enough
Communications of the ACM · DOI 10.1145/1953122.1953140 · 39 citations · Source: semantic-scholarArgues that at very large scale, classic single-truth relational semantics give way to approximate, loosely-coupled and eventually-consistent answers - and that this is usually acceptable.
Pat Helland · 1 author totalFinding social roles in Wikipedia
Proceedings of the 2011 iConference · DOI 10.1145/1940761.1940778 · 240 citations · Source: openalex+first-party-career-authorityMarc Smith, Howard T. Welser, Dan Cosley, Gueorgi Kossinets, Austin Lin, Fedor A. Dokshin, Geri Gay, Marc A. Smith · 8 authors totalEfficient online ad serving in a display advertising exchange
WSDM · DOI 10.1145/1935826.1935864 · Source: publisher+dblp+first-party-career-authorityJoaquin Delgado, Kevin J. Lang, Dongming Jiang, Bhaskar Ghosh, Shirshanka Das, Amita Gajewar, Swaroop Jagadish, Arathi Seshan · 12 authors totalBatch query processing for web search engines
DOI 10.1145/1935826.1935858 · 32 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Shuai Ding, Josh Attenberg, Ricardo Baeza‐Yates, Torsten Suel · 5 authors totalWeb retrieval
DOI 10.1145/1935826.1935835 · 4 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates, Yoelle Maarek · 3 authors totalCrowdsourcing 101: putting the WSDM of crowds to work for you
WSDM · DOI 10.1145/1935826.1935831 · 47 citations · Source: dblp+semantic-scholarCrowdsourcing has emerged in recent years as an exciting new avenue for leveraging the tremendous potential and resources of today's digitally-connected, diverse, distributed workforce. Generally speaking, crowdsourcing describes outsourcing of tasks to a large group of people instead of assigning such tasks to an in-house employee or contractor. Crowdsourcing platforms such as Amazon Mechanical Turk and CrowdFlower have gained particular attention as active online market places for reaching and tapping into this glut of a still largely under-utilized workforce. Crowdsourcing offers intriguing new opportunities for accomplishing different kinds of tasks or achieving broader participation than previously possible, as well as completing standard tasks more accurately in less time and at lower cost. Unlocking the potential of crowdsourcing in practice, however, requires a tri-partite understanding of principles, platforms, and best practices. This tutorial will introduce the opportunities and challenges of crowdsourcing while discussing the three issues above. This will provide attendees with a basic foundation to begin applying crowdsourcing in the context of their own particular tasks.
Omar Alonso, Matthew Lease · 2 authors totalA co-relational model of data for large shared data banks
Communications of the ACM · DOI 10.1145/1924421.1924436 · 66 citations · Source: openalex+semantic-scholarContrary to popular belief, SQL and noSQL are really just two sides of the same coin.
Erik Meijer, Gavin Bierman · 2 authors totalMassive multiplayer human computation for fun, money, and survival
ICWE Workshops · DOI 10.1145/1869086.1869093 · 26 citations · Source: semantic-scholarLukas Biewald · 1 author totalDiscovery and Preclinical Validation of Drug Indications Using Compendia of Public Gene Expression Data
Science Translational Medicine · DOI 10.1126/scitranslmed.3001318 · Source: nih+ucsfAtul Butte, Marina Sirota, Joel T. Dudley, Jeewon Kim, Annie P. Chiang, Alex A. Morgan, Alejandro Sweet-Cordero, Julien Sage · 8 authors totalUnequal channel error protection of multiple description codes for wireless media streaming
DOI 10.1109/vcip.2011.6115914 · 0 citations · Source: openalex+career-authorityNima Sarshar, Tanay Dey, Abdul Bais · 3 authors totalThe Hybrid Reciprocal Velocity Obstacle
IEEE Transactions on Robotics · DOI 10.1109/tro.2011.2120810 · 457 citations · Source: openalexWe present the hybrid reciprocal velocity obstacle for collision-free and oscillation-free navigation of multiple mobile robots or virtual agents. Each robot senses its surroundings and acts independently without central coordination or communication with other robots. Our approach uses both the current position and the velocity of other robots to compute their future trajectories in order to avoid collisions. Moreover, our approach is reciprocal and avoids oscillations by explicitly taking into account that the other robots sense their surroundings as well and change their trajectories accordingly. We apply hybrid reciprocal velocity obstacles to iRobot Create mobile robots and demonstrate direct, collision-free, and oscillation-free navigation.
Jur van den Berg, Jamie Snape, Stephen J. Guy, Dinesh Manocha · 4 authors totalModeling Unconnectable Peers in Private BitTorrent Communities
DOI 10.1109/pdp.2011.21 · 2 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Kornél Csernai, Márk Jelasity, Tamás Vinkó · 4 authors totalUsing Social Sensing to Understand the Links between Sleep, Mood, and Sociability
2011 IEEE Third International Conference on Privacy, Security, Risk and Trust and 2011 IEEE Third International Conference on Social Computing · DOI 10.1109/PASSAT/SocialCom.2011.200 · 79 citations · Source: semantic-scholar+dblp+career-authoritySai Moturu, S. Moturu, Inas S. Khayal, Nadav Aharony, Wei Pan, A. Pentland · 6 authors totalLitter: A Lightweight Peer-to-Peer Microblogging Service
SocialCom/PASSAT · DOI 10.1109/PASSAT/SocialCom.2011.192 · 9 citations · Source: semantic-scholar+dblpOscar Boykin, Pierre St. Juste, David Wolinsky, P. Oscar Boykin, Renato J. O. Figueiredo · 5 authors totalGroup-in-a-Box Layout for Multi-faceted Analysis of Communities
DOI 10.1109/passat/socialcom.2011.139 · 59 citations · Source: openalex+first-party-career-authorityMarc Smith, Eduarda Mendes Rodrigues, Nataša Milić-Frayling, Marc A. Smith, Ben Shneiderman, Derek L. Hansen · 6 authors totalRegulating Locality vs. Parallelism Tradeoffs in Multiple Memory Controller Environments
DOI 10.1109/pact.2011.33 · 8 citations · Source: openalexThe presence of multiple MCs and their integration into the on-chip network fabric creates a highly concurrent system that can support significant levels of memory level parallelism (MLP) across cores. This work exposes the trade-off between DRAM parameters, bank level parallelism (BLP), and row buffer hit rate that exposes the amount of effective BLP that is necessary to approximate a 100% hit rate. We further study how this trade-off can be controlled and propose a class of global (system) and local (within an MC) address mappings that can be tuned to optimize the performance across a set of multiprogrammed benchmarks.
Dhruv Choudhary, Syed Minhaj Hassan, Mitchelle Rasquinha, Sudhakar Yalamanchili · 4 authors totalInter-swarm resource allocation in BitTorrent communities
DOI 10.1109/p2p.2011.6038748 · 12 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Mihai Capotă, Nazareno Andrade, Tamás Vinkó, Flávio Roberto Santos, Dick Epema · 6 authors totalFast download but eternal seeding: The reward and punishment of Sharing Ratio Enforcement
DOI 10.1109/p2p.2011.6038746 · 27 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Adele Lu Jia, Rameez Rahman, Tamás Vinkó, Dick Epema · 5 authors totalIdentifying, analyzing, and modeling flashcrowds in BitTorrent
DOI 10.1109/p2p.2011.6038742 · 31 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Boxun Zhang, Alexandru Iosup, Dick Epema · 4 authors totalTribler: Search and stream
DOI 10.1109/p2p.2011.6038729 · 4 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Niels Zeilemaker, Mihai Capotă, Arno Bakker · 4 authors totalTemporal classification of events in cricket videos
National Conference on Communications · DOI 10.1109/NCC.2011.5734784 · 29 citations · Source: crossref+semantic-scholarSanjeev Satheesh, N. Harikrishna, S. Satheesh, D. Sriram, K. Easwarakumar · 5 authors totalImplementing Domain-Specific Languages for Heterogeneous Parallel Computing
IEEE Micro · DOI 10.1109/mm.2011.68 · 86 citations · Source: openalexDomain-specific languages offer a solution to the performance and the productivity issues in heterogeneous computing systems. The Delite compiler framework simplifies the process of building embedded parallel DSLs. DSL developers can implement domain-specific operations by extending the DSL framework, which provides static optimizations and code generation for heterogeneous hardware. The Delite runtime automatically schedules and executes DSL operations on heterogeneous hardware.
Martin Odersky, HyoukJoong Lee, Kevin Brown, Arvind K. Sujeeth, Hassan Chafi, Tiark Rompf, Kunle Olukotun · 7 authors totalScala Web Frameworks: Looking Beyond Lift
IEEE Internet Computing · DOI 10.1109/MIC.2011.104 · 4 citations · Source: semantic-scholarDean Wampler, D. Wampler · 2 authors totalGLive: The Gradient Overlay as a Market Maker for Mesh-Based P2P Live Streaming
ISPDC · DOI 10.1109/ISPDC.2011.31 · Source: dblp+first-party-career-authorityJim Dowling, Amir Hossein Payberah, Seif Haridi · 3 authors totalDeftpack: A Robust Piece-Picking Algorithm for Scalable Video Coding in P2P Systems
DOI 10.1109/ism.2011.52 · 6 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Riccardo Petrocco, Michael Eberhard, Dick Epema · 4 authors totalEnhanced Visual Scene Understanding through Human-Robot Dialog
IEEE/RSJ International Conference on Intelligent Robots and Systems · DOI 10.1109/IROS.2011.6094596 · Source: kth+mpi+crossrefBabak Rasolzadeh, Matthew Johnson-Roberson, Jeannette Bohg, Gabriel Skantze, Joakim Gustafson, Rolf Carlson, Danica Kragic · 7 authors totalBetweenness Centrality Approximations for an Internet Deployed P2P Reputation System
DOI 10.1109/ipdps.2011.317 · 8 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Dimitra Gkorou, Dick Epema · 3 authors totalParallel Processing Framework on a P2P System Using Map and Reduce Primitives
IPDPS Workshops · DOI 10.1109/IPDPS.2011.315 · 17 citations · Source: semantic-scholar+dblpOscar Boykin, Kyungyong Lee 0001, Tae Woong Choi, Arijit Ganguly, David Wolinsky, P. Oscar Boykin, Renato J. O. Figueiredo · 7 authors totalNetwork warehouses: Efficient information distribution to mobile users
IEEE INFOCOM · DOI 10.1109/INFCOM.2011.5935015 · 11 citations · Source: openalex+semanticscholarWe consider the problem of distributing time-sensitive information from a collection of sources to mobile users traversing a wireless mesh network. Our strategy is to distributively select a set of well-placed nodes (warehouses) to act as intermediaries between the information sources and clusters of users. Warehouses are selected via the distributed construction of Hierarchical Well-Separated Trees (HSTs), which are sparse structures that induce a natural spatial clustering of the network. Unlike many traditional multicast protocols, our approach is not data driven. Rather, it is agnostic to the number and position of sources as well as to the mobility patterns of users. Whereas source-rooted tree multicast algorithms construct a separate routing infrastructure to support each source, our sparse and flexible infrastructure is precomputed and efficiently reused by sources and users, its cost amortized over time. Moreover, the route acquisition delay inherent in on-demand wireless ad hoc network protocols is avoided by exploiting the HST addressing scheme. Our algorithm ensures with high probability a guaranteed stretch bound for the information delivery path, and is robust to lossy links and node failure by providing alternative HST-induced routes. Nearby users are clustered and their requests aggregated, further reducing communication overhead.
Ian Downes, Arik Motskin, Branislav Kusy, Omprakash Gnawali, Leonidas Guibas · 5 authors totalSleep, mood and sociability in a healthy population
Annual International Conference of the IEEE Engineering in Medicine and Biology Society · DOI 10.1109/IEMBS.2011.6091303 · 33 citations · Source: semantic-scholar+dblp+career-authoritySai Moturu, S. Moturu, Inas S. Khayal, Nadav Aharony, Wei Pan, A. Pentland · 6 authors totalRacetrack memory cell array with integrated magnetic tunnel junction readout
International Electron Devices Meeting · DOI 10.1109/IEDM.2011.6131604 · 90 citations · Source: semantic-scholarAnthony Annunziata, A. Annunziata, M. Gaidis, Luc Thomas, C. Chien, C. Hung, P. Chevalier, E. O'Sullivan · 15 authors totalReciprocal collision avoidance with acceleration-velocity obstacles
ICRA 2011 · DOI 10.1109/icra.2011.5980408 · 288 citations · Source: openalexWe present an approach for collision avoidance for mobile robots that takes into account acceleration constraints. We discuss both the case of navigating a single robot among moving obstacles, and the case of multiple robots reciprocally avoiding collisions with each other while navigating a common workspace. Inspired by the concept of velocity obstacles [3], we introduce the acceleration-velocity obstacle (AVO) to let a robot avoid collisions with moving obstacles while obeying acceleration constraints. AVO characterizes the set of new velocities the robot can safely reach and adopt using proportional control of the acceleration. We extend this concept to reciprocal collision avoidance for multi-robot settings, by letting each robot take half of the responsibility of avoiding pairwise collisions. Our formulation guarantees collision-free navigation even as the robots act independently and simultaneously, without coordination. Our approach is designed for holonomic robots, but can also be applied to kinematically constrained non-holonomic robots such as cars. We have implemented our approach, and we show simulation results in challenging environments with large numbers of robots and obstacles.
Jur van den Berg, Jamie Snape, Stephen J. Guy, Dinesh Manocha · 4 authors totalCOMET: A Recipe for Learning and Using Large Ensembles on Massive Data
2011 IEEE 11th International Conference on Data Mining · DOI 10.1109/ICDM.2011.39 · arXiv 1103.2068 · 38 citations · Source: semantic-scholar+openalexCOMET is a single-pass MapReduce algorithm for learning on large-scale data. It builds multiple random forest ensembles on distributed blocks of data and merges them into a mega-ensemble. This approach is appropriate when learning from massive-scale data that is too large to fit on a single machine. To get the best accuracy, IVoting should be used instead of bagging to generate the training subset for each decision tree in the random forest. Experiments with two large datasets (5GB and 50GB compressed) show that COMET compares favorably (in both accuracy and training time) to learning on a sub sample of data using a serial algorithm. Finally, we propose a new Gaussian approach for lazy ensemble evaluation which dynamically decides how many ensemble members to evaluate per data point, this can reduce evaluation cost by 100X or more.
Justin Basilico, Justin D. Basilico, M. Arthur Munson, Tamara G. Kolda, Kevin R. Dixon, W. Philip Kegelmeyer · 6 authors total