Matei Zaharia — retrospective
Community-maintained public research about Matei Zaharia; it has no identity claim or authority to speak for them. How community accounts work ↗
Public graph
Centered on Matei Zaharia
On 12/12/2011, Alexy Khrabrov hosted the very first talk by Matei Zaharia at the Scala for Startups meetup at Klout. Alexy founded Scala for Startups and later merged it with SF Scala to build the world's largest Scala meetup. After hosting Matei, Alexy proposed establishing a new meetup, the Bay Area Spark Meetup, which they co-founded.
This playlist has been found and uploaded 15 years later. It includes talks from the original Scala for Startups and then the first Spark Users meetup. Matei presents Spark design, does the famous Berkeley vs Stanford demo, and answers questions. We also have talks from the early adopters Klout (where Alexy pioneered it and hosted the meetup) and Quantifind.\n\n## Original event descriptions
The first Spark User Meetup
On January 31st, we’re hosting the first Spark User Meetup at Klout in San Francisco. Spark is a cluster programming framework that provides in-memory computing for iterative and interactive analytics and a high-level programming interface in the Scala language. This meetup will be a chance to learn Spark from the developers, hear about other peoples’ experiences with Spark, network, and hear about future development plans.
At the first meeting, we’re planning to have a Spark tutorial by Matei Zaharia followed by a presentation from Karthik Thiyagarajan of Quantifind on their experience using Spark to replace Pig for predictive analytics.
If you want to follow along, we strongly recommend you sign up for Amazon EC2 and make sure you can launch instances. This way you’ll be able to bring up your own Mesos/Spark cluster on EC2 and play with a Wikipedia dataset in real time!
Time: Tuesday January 31. Pizza and beer at 6:30, talks at 7 PM.
Location: Klout, 77 Stillman St, San Francisco, CA 94107
About Quantifind
Quantifind is a small start up into predictive analytics. We are building a platform that automatically identifies relevant signals in both structured and unstructured content, contextualizes them based on what matters most to the customer, and derives insights about past and future events in a way that is directly relevant and actionable.
The talk highlights our experience using Spark in various ways ranging from a distributed batch processing framework powering our analytics pipeline to an interactive computing infrastructure serving some of our internal exploratory tools.
Over the course of moving our analytics pipeline from Pig to Spark, we realized that Spark's inherent characteristics of low latency and interactivity can be leveraged as an agile way to create new computing services. So we experimented with creating web services on top of Spark which answered queries in real time by performing operations on a cached RDD. Today, Spark acts as a sharded in memory infrastructure for many of the services we use internally helping us explore our data and prototype algorithms in an agile manner.
Scala Spark, or What’s Next in Big Data
On Monday, December 12, Matei Zaharia of UC Berkeley will talk about Spark, a new cluster computing framework for big data. Matei is the lead committer on Spark as well as an author of Mesos, the open-source cluster management system that Spark runs on top of. Spark uses in-memory caching and the Scala interpreter to effectively give you a prompt to explore terabytes of data interactively.
We'll possibly have other members of the Spark group talk about their uses of Spark for Machine Learning, Data Mining, and production use. The meetup will also be a starting point of the Spark Users Group. Pizza and beer will be served!
Future meetups: The meetup will rotate among locations in San Francisco, Silicon Valley, and Berkeley.
Talks & videos
12 connected recordings
Faster and Cheaper Training for Large Models
Two broad lines of research to make large-scale ML accessible: (1) pipeline and hybrid parallelism as in PipeDream and FlexFlow for cheaper training of existing DNN models; (2) retrieval-based NLP models like ColBERT that search through a corpus of documents at inference time rather than memorizing all knowledge in parameters.
Scale By The Bay 2020:Matei Zaharia,Scaling Databricks to Run Data & AI Workloads on Millions of VMs
ai.bythebay.io Nov 2025, Oakland, full-stack AI conference Title: Scaling Databricks to Run Data and AI Workloads on Millions of VMs Cloud service developers need to handle massive scale workloads from thousands of customers with no downtime or regressions. In this talk, I’ll present our experience building a very large-scale cloud service at Databricks, which provides a data and ML platform service used by many of the largest enterprises in the world. Databricks manages millions of cloud VMs that process exabytes of data per day for interactive, streaming and batch production applications. This means that our control plane has to handle a wide range of workload patterns and cloud issues such as outages. We will describe how we built our control plane for Databricks using Scala services and open source infrastructure such as Kubernetes, Envoy, and Prometheus, and various design patterns and engineering processes that we learned along the way. In addition, I’ll describe how we have adapted data analytics systems themselves to improve reliability and manageability in the cloud, such as creating an ACID storage system that is as reliable as the underlying cloud object store (Delta Lake) and adding autoscaling and auto-shutdown features for Apache Spark. Matei Zaharia Databricks Chief Technologist Matei Zaharia is an Assistant Professor of Computer Science at Stanford and Co-founder and Chief Technologist at Databricks. He started the Apache Spark project during his PhD at UC Berkeley, and has worked on other widely used open source data analytics and AI software including M…
Photos
0 connected galleries