Papers.
Research connected to its authors, projects, companies, talks, events, and the rest of the graph.
Add a paper ↗SparkBench - A Spark Performance Testing Suite
TPCTC · DOI 10.1007/978-3-319-31409-9_3 · 26 citations · Source: semantic-scholar+dblpFrederick Reiss, D. Agrawal, A. Butt, K. Doshi, J. Larriba-Pey, Min Li, F. Raab, Berni Schiefer · 10 authors totalQuerying Time Interval Data
Lecture notes in business information processing · DOI 10.1007/978-3-319-29133-8_3 · 1 citations · Source: openalex+career-authorityPhilipp Meisen, Diane Keng, Tobias Meisen, Marco Recchioni, Sabina Jeschke · 5 authors totalFeasibility of Word Difficulty Prediction
Lecture notes in computer science · DOI 10.1007/978-3-319-23826-5_34 · 3 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates, Martí Mayo-Casademont, Luz Rello · 4 authors totalPrivacy-Preserving Detection of Anomalous Phenomena in Crowdsourced Environmental Sensing
International Symposium on Spatial and Temporal Databases · DOI 10.1007/978-3-319-22363-6_17 · 6 citations · Source: semantic-scholarMihai Maruseac, Gabriel Ghinita, Besim Avci, Goce Trajcevski, P. Scheuermann · 5 authors totalLeader Election Using NewSQL Database Systems
DAIS · DOI 10.1007/978-3-319-19129-4_13 · Source: dblp+first-party-career-authorityJim Dowling, Salman Niazi, Mahmoud Ismail, Gautier Berthou · 4 authors totalThe Structures of Twitter Crowds and Conversations
DOI 10.1007/978-3-319-18552-1_5 · 25 citations · Source: openalex+first-party-career-authorityMarc Smith, Marc A. Smith, Itai Himelboim, Lee Rainie, Ben Shneiderman · 5 authors totalDiscussion of BigBench: A Proposed Industry Standard Performance Benchmark for Big Data
Lecture notes in computer science · DOI 10.1007/978-3-319-15350-6_4 · 35 citations · Source: openalex+career-authorityMilind Bhandarkar, Chaitanya Baru, Carlo Curino, Manuel Danisch, Michael Frank, Bhaskar Gowda, Hans‐Arno Jacobsen, Huang Jie · 18 authors totalRecommender Systems in Industry: A Netflix Case Study
Recommender Systems Handbook · DOI 10.1007/978-1-4899-7637-6_11 · 117 citations · Source: semantic-scholar+openalexJustin Basilico, Xavier Amatriain · 2 authors totalPractical Graph Analytics with Apache Giraph
Apress (book) · DOI 10.1007/978-1-4842-1251-6 · 62 citations · Source: semantic-scholarRoman Shaposhnik, Claudio Martella, Dionysios Logothetis · 3 authors totalBig Data Analytics with Spark: A Practitioner's Guide to Using Spark for Large Scale Data Analysis
Apress · DOI 10.1007/978-1-4842-0964-6 · Source: apress+conference-biographyMohammed Guller · 1 author totalEstimating user interaction strength in distributed online networks
Concurrency and Computation Practice and Experience · DOI 10.1002/cpe.3575 · 1 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Adele Lu Jia, Boudewijn Schoon, Dick Epema · 4 authors totalDistributed Bayesian Learning with Stochastic Natural Gradient Expectation Propagation and the Posterior Server
Journal of Machine Learning Research (2017) · arXiv 1512.09327 · 73 citations · Source: arxiv+semantic-scholarThis paper makes two contributions to Bayesian machine learning algorithms. Firstly, we propose stochastic natural gradient expectation propagation (SNEP), a novel alternative to expectation propagation (EP), a popular variational inference algorithm. SNEP is a black box variational algorithm, in that it does not require any simplifying assumptions on the distribution of interest, beyond the existence of some Monte Carlo sampler for estimating the moments of the EP tilted distributions. Further, as opposed to EP which has no guarantee of convergence, SNEP can be shown to be convergent, even when using Monte Carlo moment estimates. Secondly, we propose a novel architecture for distributed Bayesian learning which we call the posterior server. The posterior server allows scalable and robust Bayesian learning in cases where a data set is stored in a distributed manner across a cluster, with each compute node containing a disjoint subset of data. An independent Monte Carlo sampler is run on each compute node, with direct access only to the local data subset, but which targets an approximation to the global posterior distribution given all data across the whole cluster. This is achieved by using a distributed asynchronous implementation of SNEP to pass messages across the cluster. We demonstrate SNEP and the posterior server on distributed Bayesian learning of logistic regression and neural networks. Keywords: Distributed Learning, Large Scale Learning, Deep Learning, Bayesian Learn- ing, Variational Inference, Expectation Propagation, Stochastic Approximation, Natural Gradient, Markov chain Monte Carlo, Parameter Server, Posterior Server.
Stefan Webb, Leonard Hasenclever, Thibaut Lienart, S. Vollmer, Balaji Lakshminarayanan, C. Blundell, Y. Teh · 7 authors totalDeep Speech 2: End-to-End Speech Recognition in English and Mandarin
ICML · arXiv 1512.02595 · 3,186 citations · Source: arxiv+semantic-scholarWe show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech--two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, resulting in a 7x speedup over our previous system. Because of this efficiency, experiments that previously took weeks now run in days. This enables us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale.
Erich Elsen, Sanjeev Satheesh, Dario Amodei, Rishita Anubhai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro · 34 authors totalSurvey of robust and resilient social media tools on Android
arXiv (Cornell University) · DOI 10.48550/arxiv.1512.00071 · 0 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, P. W. G. Brussee · 2 authors totalEmbarrassingly Parallel Time Series Analysis for Large Scale Weak Memory Systems
arXiv · arXiv 1511.06493 · Source: arxiv+dblp+berkeley-career-authorityEvan R. Sparks, Francois Belletti, Evan Randall Sparks, Michael J. Franklin, Alexandre M. Bayen · 5 authors totalAutonomous smartphone apps: self-compilation, mutation, and viral spreading
arXiv (Cornell University) · DOI 10.48550/arxiv.1511.00444 · 2 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, P. W. G. Brussee · 2 authors totalAsynchronous Complex Analytics in a Distributed Dataflow Architecture
arXiv (Cornell University) · DOI 10.48550/arxiv.1510.07092 · 10 citations · Source: openalex+authoritative-profilePeter Bailis, Joseph E. Gonzalez, Michael I. Jordan, Michael J. Franklin, Joseph M. Hellerstein, Ali Ghodsi, Ion Stoica · 7 authors totalSPRIGHT: A Fast and Robust Framework for Sparse Walsh-Hadamard Transform
arXiv.org · arXiv 1508.06336 · 17 citations · Source: semantic-scholar+openalexWe consider the problem of computing the Walsh-Hadamard Transform (WHT) of some $N$-length input vector in the presence of noise, where the $N$-point Walsh spectrum is $K$-sparse with $K = {O}(N^{\delta})$ scaling sub-linearly in the input dimension $N$ for some $0<\delta<1$. Over the past decade, there has been a resurgence in research related to the computation of Discrete Fourier Transform (DFT) for some length-$N$ input signal that has a $K$-sparse Fourier spectrum. In particular, through a sparse-graph code design, our earlier work on the Fast Fourier Aliasing-based Sparse Transform (FFAST) algorithm computes the $K$-sparse DFT in time ${O}(K\log K)$ by taking ${O}(K)$ noiseless samples. Inspired by the coding-theoretic design framework, Scheibler et al. proposed the Sparse Fast Hadamard Transform (SparseFHT) algorithm that elegantly computes the $K$-sparse WHT in the absence of noise using ${O}(K\log N)$ samples in time ${O}(K\log^2 N)$. However, the SparseFHT algorithm explicitly exploits the noiseless nature of the problem, and is not equipped to deal with scenarios where the observations are corrupted by noise. Therefore, a question of critical interest is whether this coding-theoretic framework can be made robust to noise. Further, if the answer is yes, what is the extra price that needs to be paid for being robust to noise? In this paper, we show, quite interestingly, that there is {\it no extra price} that needs to be paid for being robust to noise other than a constant factor. In other words, we can maintain the same sample complexity ${O}(K\log N)$ and the computational complexity ${O}(K\log^2 N)$ as those of the noiseless case, using our SParse Robust Iterative Graph-based Hadamard Transform (SPRIGHT) algorithm.
Joseph Bradley, Xiao Li, Joseph K. Bradley, Sameer Pawar, Kannan Ramchandran · 5 authors totalA survey of P2P multidimensional indexing structures
arXiv (Cornell University) · DOI 10.48550/arxiv.1507.05501 · 2 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Ewout Bongers · 2 authors totalPerformance analysis of a Tor-like onion routing implementation
arXiv (Cornell University) · DOI 10.48550/arxiv.1507.00245 · 2 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Quinten Stokkink, Harmjan Treep · 3 authors totalAsk Me Anything: Dynamic Memory Networks for Natural Language Processing
International Conference on Machine Learning · arXiv 1506.07285 · 1,221 citations · Source: semantic-scholarMost tasks in natural language processing can be cast into question answering (QA) problems over language input. We introduce the dynamic memory network (DMN), a neural network architecture which processes input sequences and questions, forms episodic memories, and generates relevant answers. Questions trigger an iterative attention process which allows the model to condition its attention on the inputs and the result of previous iterations. These results are then reasoned over in a hierarchical recurrent sequence model to generate answers. The DMN can be trained end-to-end and obtains state-of-the-art results on several types of tasks and datasets: question answering (Facebook's bAbI dataset), text classification for sentiment analysis (Stanford Sentiment Treebank) and sequence modeling for part-of-speech tagging (WSJ-PTB). The training for these different tasks relies exclusively on trained word vector representations and input-question-answer triplets.
Richard Socher, A. Kumar, Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury, Ishaan Gulrajani, Victor Zhong · 9 authors totalLinear Types Can Change the Blockchain
CoRR · arXiv 1506.01001 · Source: dblp+arxiv+career-authorityLucius Gregory Meredith · 1 author totalFinding Intermediary Topics Between People of Opposing Views: A Case Study
arXiv (Cornell University) · DOI 10.48550/arxiv.1506.00963 · 7 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Eduardo Graells-Garrido, Mounia Lalmas, Ricardo Baeza‐Yates · 4 authors totalAnonymous online purchases with exhaustive operational security
arXiv (Cornell University) · DOI 10.48550/arxiv.1505.07370 · 3 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Vincent Van Mieghem · 2 authors totalMLlib: Machine Learning in Apache Spark
Journal of machine learning research · DOI 10.48550/arxiv.1505.06807 · arXiv 1505.06807 · 1,868 citations · Source: semantic-scholar+openalexApache Spark is a popular open-source platform for large-scale data processing that is well-suited for iterative machine learning tasks. In this paper we present MLlib, Spark's open-source distributed machine learning library. MLlib provides efficient functionality for a wide range of learning settings and includes several underlying statistical, optimization, and linear algebra primitives. Shipped with Spark, MLlib supports several languages and provides a high-level API that leverages Spark's rich ecosystem to simplify the development of end-to-end machine learning pipelines. MLlib has experienced a rapid growth due to its vibrant open-source community of over 140 contributors, and includes extensive documentation to support further growth and to let users quickly get up to speed.
Evan R. Sparks, Joseph Bradley, Matei Zaharia, Reynold Xin, Xiangrui Meng, Joseph K. Bradley, B. Yavuz, S. Venkataraman · 16 authors totalEstimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence
Journal of machine learning research · arXiv 1505.01462 · 192 citations · Source: semantic-scholar+openalexData in the form of pairwise comparisons arises in many domains, including preference elicitation, sporting competitions, and peer grading among others. We consider parametric ordinal models for such pairwise comparison data involving a latent vector w* e Rd that represents the "qualities" of the d items being compared; this class of models includes the two most widely used parametric models|the Bradley-Terry-Luce (BTL) and the Thurstone models. Working within a standard minimax framework, we provide tight upper and lower bounds on the optimal error in estimating the quality score vector w* under this class of models. The bounds depend on the topology of the comparison graph induced by the subset of pairs being compared, via the spectrum of the Laplacian of the comparison graph. Thus, in settings where the subset of pairs may be chosen, our results provide principled guidelines for making this choice. Finally, we compare these error rates to those under cardinal measurement models and show that the error rates in the ordinal and cardinal settings have identical scalings apart from constant pre-factors.
Joseph Bradley, Nihar B. Shah, Sivaraman Balakrishnan, Joseph K. Bradley, Abhay Parekh, Kannan Ramchandran, Martin J. Wainwright · 7 authors totalHigher category models of the pi-calculus
CoRR · arXiv 1504.04311 · Source: dblp+arxiv+career-authorityLucius Gregory Meredith, Mike Stay · 2 authors totalA Self-Compiling Android Data Obfuscation Tool
arXiv (Cornell University) · DOI 10.48550/arxiv.1502.01625 · 2 citations · Source: openalex+orcid+dblp-identityJohan Pouwelse, Olivier Hokke, Alex Kolpa, Joris van den Oever, Alex Walterbos · 5 authors totalWorkforce Issues: Profiles of Specialty Crop Farms in New York State
0 citations · Source: openalex+first-party-career-authorityMarc Smith, Thomas R. Maloney, Marc A. Smith, Rachel Saputo, Bradley J. Rickard, Charles H. Dyson · 6 authors totalWisdom of Crowds or Wisdom of a Few
2 citations · Source: openalex+authoritative-profileRicardo Baeza-Yates, Ricardo Baeza‐Yates · 2 authors totalWhy are deep nets reversible: A simple theory, with implications for training.
CoRR · Source: dblp+stanford-authorityTengyu Ma, Sanjeev Arora, Yingyu Liang, Tengyu Ma 0001 · 4 authors totalWatching Fictive Motion in Action: Discourse Data from the TV News Archive
Proceedings of the 37th Annual Conference of the Cognitive Science Society · 1 citations · Source: escholarshipTill Bergmann, Teenie Matlock · 2 authors totalViolence Metaphors in Presidential Debates
Proceedings of the 37th Annual Conference of the Cognitive Science Society · 0 citations · Source: escholarshipTill Bergmann, Chelsea Coe, Teenie Matlock · 3 authors totalTraversal Query Language For Scala.Meta
1 citations · Source: semantic-scholar+epfl-infoscienceEugene Burmako, Eric Béguet, E. Burmako · 3 authors totalThe Missing Piece in Complex Analytics: Low Latency, Scalable Model Management and Serving with Velox
CIDR · Source: haoyuanli-personal+dblpHaoyuan Li, Dan Crankshaw, Peter Bailis, Joseph Gonzalez, Zhao Zhang, Michael J. Franklin, Ali Ghodsi, Michael I. Jordan · 8 authors totalThe Lake-dwellings of Europe
Bulletin of Miscellaneous Information (Royal Gardens Kew) · 2 citations · Source: openalex+personal-publication-listRob Munro, Robert Munro · 2 authors totalThe Case for Invariant-Based Concurrency Control.
Conference on Innovative Data Systems Research · 4 citations · Source: openalex+authoritative-profilePeter Bailis · 1 author totalSum-of-Squares Lower Bounds for Sparse PCA.
NIPS · Source: dblp+stanford-authorityTengyu Ma, Tengyu Ma 0001, Avi Wigderson · 3 authors totalSuccinct: Enabling Queries on Compressed Data
NSDI · 63 citations · Source: openalexAnurag Khandelwal, Rachit Agarwal, Ion Stoica · 3 authors totalStyle Checking With Scala.Meta
0 citations · Source: semantic-scholar+epfl-infoscienceEugene Burmako, M. Demarne, E. Burmako · 3 authors totalStreaming@Twitter
IEEE Data Eng. Bull. · Source: dblpKarthik Ramasamy, Maosong Fu, Sailesh Mittal, Vikas Kedigehalli, Michael Barry, Andrew Jorgensen, Christopher Kellogg, Neng Lu · 10 authors totalSimple, Efficient, and Neural Algorithms for Sparse Coding.
COLT · Source: dblp+stanford-authorityTengyu Ma, Sanjeev Arora, Rong Ge 0001, Tengyu Ma 0001, Ankur Moitra · 5 authors totalSEEDB: Efficient Data-Driven Visualization Recommendations to Support Visual Analytics
PVLDB · Source: dblpManasi Vartak, Sajjadur Rahman, Samuel Madden, Aditya G. Parameswaran, Neoklis Polyzotis · 5 authors totalScalable Transactions for Scalable Distributed Database Systems
University of California, Berkeley PhD dissertation · Source: uc-berkeley+dblp+alluxio-career-authorityGene Pang · 1 author totalsbt in Action
Manning · Source: publisher+first-partyJosh Suereth, Joshua Suereth, Matthew Farwell · 3 authors totalRobust Sparse Walsh-Hadamard Transform : The SPRIGHT Algorithm
0 citations · Source: semantic-scholar+openalexJoseph Bradley, Xiao Li, Joseph K. Bradley, S. Pawar, K. Ramchandran · 5 authors totalRemote patient monitoring to improve health: challenges and opportunities
Conference of the Centre for Advanced Studies on Collaborative Research · 1 citations · Source: semantic-scholarRandy Giffen, M. Fung, M. Rotaru · 3 authors totalRegion-based off-heap memory for Scala
3 citations · Source: openalexThis work explores support for region-based memory in Scala through macro-based domain-specific encoding that exploits a number of extensible language features such as implicits and value classes.
Denys Shabalin, Martin Odersky · 2 authors total