1st Spark: Alexy Khrabrov hosts the first Apache Spark meetup talk by Matei at Scala for Startups
Uh, welcome everybody. I >> think we'll start this show in a minute. [clears throat] So this is u scholar startups meetup. It's a fairly recent meetup. We started in September and we started with about maybe 20 25 people in the first uh meeting and then uh second we held jointly with Mike Sllin's uh scholar for startups uh San Francisco scholar meetup uh which is an older group which he revived recently. So there are two groups in uh Soma uh and there is another group uh called uh Bay Area Skull Enthusiasts which gathers uh alternatively uh here uh >> not anymore just south >> and salt and now it's basically been kept salt. So I'll let Mike say a few words >> about that about that. Uh >> okay so I organized SF Scala Chase the area enthusiast SF Scala shares the with this group and that shares the city
It's a big city. There's room for both. Uh and base is all going south. Um and base used to alternate between southern state peninsula location and San Francisco. No more. Now it's strictly south. It will be every month. And SSA is also going to be every month up here
So, uh there's links between the two. We're friendly. Uh we're cooperative. And so if you go to any of the sites you can find all the other sites. >> Yep. Thanks Mike. Uh so the focus of this meetup is Scola and startups and the combination uh is very fitting to this area. We had uh people from all kinds of startups and this not limited to startups right there are of course people from large companies like LinkedIn and Twitter
I mean there were startups at some point were interested in the process by which you grow from a startup right and scale. Scala has scaling in this nature. Uh so that's I think one of the interesting uh intersections of these two things is how do you leverage all these new technologies? How do you leverage the language, its infrastructure, its libraries, uh it's all its hookups into Java world to build something scalable quickly and do something useful and interesting, right? So so we're fairly pragmatic. Uh we're interested in making things work. Uh and uh also interested in stories. So anybody with a story how you use Scala in a startup you know one person startup five people you know 20 to actually make something useful uh it's very interesting to to us so please uh contact me uh there will be a contact uh email uh or just contact me through the uh scholar for startups meetup page uh if you would like to present uh in the future meetups about how scholar made your startup work or what plans do you have what questions do you have and for We're very fortunate to have here in today's talk actually folks from an existing startup Canviva who will uh tell us about uh the usage of skull and and the uh spark technologies. Uh so uh just a quick summary um right so that's just to summarize it's not necessarily about scola itself we're interested in functional programming languages languages on JVM so if you have an interesting story how you used Ruby how you switch to Scala you know things like this closure hasll or camel whatever I think there is a sort of set of issues which are relevant uh among all functional programming languages Right. Bascal is very pragmatic
So it's very kind of in in a good uh in a good spot in this area. But we're definitely not limited just to this. So I'd welcome all kind of functional programming uh experiences uh to to this uh meetup. Uh so today the specific topic is uh spark and big data. Uh so I'll let mate tell us more about uh about that. But uh we at cloud are consumers of uh big data, right? We uh uh so cloud is hosting this uh uh basically it's the uh standard for uh uh social influence uh it takes all the data from existing uh social networks and computes the influence score and this is a typical application for big data essentially you need uh uh Hadoop you need HDFS you need all these uh things and and uh once I started cloud uh I realized that uh Hadoop in some way is forrun of today right basically you launch your application and you go get a cup of coffee and it cranks and cranks and cranks uh and uh as as a data scientist we want to play with something interactively. So spark is uh an amazing example uh which lets you do this right and Scala has this interactive prompt. So I think it's it's great that these things came together
I started looking for something in that area and found uh that spark is here. It's developed at UC Berkeley. Uh there is a group of people which is growing uh starting to use it. So uh this uh also will be a start of the spark user group. Uh so we'll uh uh actually uh uh ask people to to specifically join this uh spark user group if there is a subset right but we'll announce it through the meetup. I wanted to ask folks who use Spark uh already here some sizable number and uh who would like to have a Spark user group uh separately going as a special meetup. All right. And who would like to have this user group going together with the Scala for startups meet up? So you think it's it deserves its own kind of right its own notice
Okay. So there is already a Google group uh for for Spark users and I think we'll announce that again through uh this meetup. So you'll get basically the coordinates and then we can synchronize uh uh about specific meetups for Spark going forward. Uh so that's my contact email. Uh if you have interest in uh the in presenting and uh interest in cloud uh uh Spark user group, please contact me. Uh and uh now I'll let uh Derek uh tell us more about uh uh Scala at cloud and uh after that we'll proceed with uh mate's talk on spark and uh folks from conviva about actual usage of spark. So, let me >> upgrade the uh computer. >> I think it's only it's for the camera, not for um >> Yeah
>> Yeah. Well, well, if if you can't hear us, we'll just yell louder. It's the uh the current plan. >> Is this right? >> Oh. Um I'm just gonna yell louder. Cool. Let's see. All right
Cool. So um basically I just was um uh giving a little presentation and um I don't know if you saw this but actually just on uh this past Friday we had a blog post about how we're using um basically elastic search play and Scala to power our search on the clout.com website and we just thought we'd kind of a take a little bit of a look into how that works and you know um how that plays into everything and so um Lexi cool. So um first let's just talk a little bit about clout. So of course our theory of cloud is everyone has clout and as Alexi alluded to our main project here is right is to measure sort of your influence on social networks. So that's basically about you know when you you know when you post your content on the social network you know how does that drive other content how does that drive people's opinions and you know how how does that affect how does that affect the world and um if we and if we move on you know we've actually already overcored over 100 million users on cloud and that number is um that number is actually from September so we've actually scored a bit more users than that already and um just to uh sort of to give people a little bit of a flavor of how that's calculated um we actually use this um people rank algorithm and again if you take a look right the whole idea here is you know you you know user one's influenced by another user maybe they retweet their content maybe they say oh hey you know Derek gave a great presentation on elastic search and then we say oh there was some influence there and we and we take a look and we take a look at how important the people were right if you if you've uh influenced you know somebody really you know really important like if you influenced Martin Oderski to introduce a new feature into Scala that may be a little bit more influential than in influencing me to stop using camel So, you know, we need to try to take a look at here again. And if we take a look who's this big influencer here, unfortunately, it turns out it's Justin Bieber, which is sometimes the way of it with these, uh, online social networks, but it's all right. So, moving on. So, um, of course, when we wanted to build this search, we wanted to we didn't want to build everything from scratch
So, we actually built on top of this, um, engine called Elastic Search. And elastic search is actually building on a lot of technologies that have been built over the p over the past you know 10 years or so and and basically what it gives you is um it gives you a scalable search and it plugs into and it's based on lucine and what it gives you is sort of all the power of lucine in terms of you know having language analysis being able to do you know synonyms and other things and you can just build on top of that using the ent you know comp it's a query parser and yet you still get sharding for free so you don't have to be doing you know kind of manual sharding bills managing a solar cluster anything else and it it actually just um will you know autoconnect and autobalance shards autoreplicate and it makes makes your life a lot easier. And just you know for a little preview we just to give you a little bit of a peek into the kind of architecture. The whole thing that's really cool about it is this is a share nothing architecture and everything talks asynchronously. And so you know basically on each thing on um you basically have a cluster of servers and each cluster you know has underlying actually a just a normal lucine index and each cluster also has this sh a sort of shared metadata state. Every node actually knows about the state of every other node which means you can connect to any particular instance in this cluster and you get results across the whole cluster and it actually uses this um asynchronous framework from JBoss called Medi which basically allows it to query every shard in parallel because otherwise as your cluster goes up your performance is just going to keep on decreasing and this lets you keep on getting good performance across the cluster. But what we thought people would be more interested in here is how we use play to build the uh API on top of this architecture. um because um you know of course this is a Scala meetup not an elastic search meetup so we don't want to go into too much detail here and um play has actually given us really you know a lot of ability to build a scalable API very quickly and there's just kind of a few different uh different things we were able to do here um that'll that uh helped our uh helped our API development go pretty quickly um so one thing that you know we thought was really great is being able to use you know these DSLs um allows us to basically concentrate on our API and not concentrate on kind of um a lot of these details, let us just build these kind of high level constructs like if you you know if you um take a look here right for example instead of um having to go and you know man manage all your query parsing and manage and manage your results and turn them into JSON we and you know managing all these other aspects of the interface you know we're just like okay please JSONify and we do this and we have a a case object here that we can use to um represent our results and this just gets and this just gets translated into a result automatically and we again we don't have to think about JSON on when we're writing our search API, we just think about the search API
Um, and there's a few other things we um we kind of got um good uh good traction on here. So, um if we just move on, we um we also said um a couple of cool things. we had a couple of cool things here like um we we we um you know, one of the things that's really the bane of everyone's existence is writing configuration files just because you know, again, you have this repeated pattern. You have to do it again and again and again. And we um we actually used right there's this um uh a cake pattern in Scola. So you can um basically actually use code in the configuration of your server and then just having thing you know kind of automatically configured by code. So instead of having to like specify every server you can just say here's a list of servers read it out of the play config and then just apply a bunch of transforms to it. And so you can just again you can reduce the amount of code reduce the amount of configuration you're writing
And we also have a couple of constructs um we built on top of this that um that kind of um have helped us you know just enhance our service like um if we uh um we have this little construct we have built called cachify and this just lets us take any command and just automatically say you know hey we want to we want to memorize this result we don't want to execute it again we just want to store it in a cache and replay the results and again the nice thing is now we have this um this very compact trait that we can use and now we don't have to write it anymore we just specify that this is a cachable result and then we're, you know, we're good to go. Um, and we don't have to, you know, we can write the code once, we can make bug fixes once if there's something we want to change about the cache. If we wanted to go to distributed cache, it all lives in one place, which has been really a big win as we try to evolve this infrastructure. And the other thing, you know, that and um there's some other stuff that's uh kind of um still pending that we'd like to improve, which is um which is um sort of on um we actually have a federated search running here. So we actually um right now if you if you use um search oncloud.com it's actually giving you results for search results over users and search results over topics and sort of and you know the thing is right now we actually are getting this in serial so you have to wait for user results first and then you can and then you get your topic results and you know that's again it means that as you add more and more of these federated searches your performance just um keeps on decreasing and so and so we actually are you know in the process of going and moving this to basically um switch over to using actors which allows us to sort of um sort again easily multi-thread this so that we have one we can respond to user search in one thread um topic search in another thread and if we wanted to add other things like influencer search we could add an influencer surge in the third thread and again we wouldn't start seeing this linear performance hit um and again it's um it's been a it's been a really big win and so um really you know really actually that that's as simple as it's been again because we are we're leveraging kind of a lot of um powerful tools to um to get our search working and Um yeah, so just um to um wrap it up, we we can um actually take a little look into um some of the search results we got here. So um if you you know again, you know, because our whole goal here is allowing you to find influencers more quickly and just we have a couple of we have a little preview. We just Oh man, this is way too tiny. It's great because we have a search for play
Actually, could we just switch into a web browser here just so we can actually uh see the search running live? You can uh just go to cloud.com. Yeah, I'm used to having the laptop sitting over here so I can I can touch it. So I'm a little better I think we have to log in to search unfortunately. So, can you log in so we can get to the Cool. [clears throat] But yeah, so actually if you look over here and you can actually go I don't know if anyone has a search they want to they want to go try to search for influencers or topics or if not we can um we can always just search for um >> tacos. tacos. All right, let's search for tacos. And then you can go and you can pull pull this up and then immediately you can get both wow people who are relevant to tacos and you can also take a look at the topic
And again, you know, it just um it again it just it allows you to find find all this stuff a lot more quickly. And so it's um it's been a really big it's been a really big win. And again, we kind of have the ability to enhance this um more easily because of all the kind of advanced features that uh Scola gives us. So um >> yep Taco Bell I think is pretty reasonable influencer on tacos. I think it influences how a lot of people consume tacos in the modern world. So in any case um so we can just go back and now to the and we just had a few um resources on that you can um look for later on uh on basically everything we've done here. So we you know we actually as I said on Friday we posted the blog post on find your clout which just talks about same sort of things we talked about here about you know how to about how we did search at clout and how we've leveraged all these technologies to get it up and running fast and just some links for you know elastic search we didn't write that a we also didn't write and play when all these things that really you know helped us move move very fast and um just to give a shout out to the people who really did the biggest work on the search there's um uh we have uh Kenny who's in our analytics team who did a lot work on search as well as um Felipe who really was just very instrumental in getting everything together. So they're not um well Kenny Kenny is here but Felipe is not here to give a plug for himself
So I just want to give a plug for them and I all all I do is take credit for things so that's why I'm here. So um so yeah so um as I said hope everyone finds it exciting and again was good for us. >> People have some questions. >> No what questions? Questions? This is unquestioned. [laughter] >> Any questions? And while we're answering questions, maybe we can set up the next uh >> how are those people above Taco Bell more influential than Taco Bell and help the tacos? >> Well, for one thing, you may debate as to whether Taco Bell really makes tacos or just things called tacos, right? You know, and that's hard for us to distinguish, but you know, there are these two categories and they're just overlapping, you know? I mean you can imagine there may be other interpretations of tacos that we are also not using here but you know it's um >> it gets a little bit murky with English >> so simple words without any grammatical context >> are difficult to disambigulate >> yes so tacos happens to be a one word topic actually a lot of these other things are multi-word topics and the other thing is again you know you're not going to see that here but in fact we are using more than just the one word to match against tacos. So it's not that we're just saying tacos equals tacos or anything like that, you know. >> How do you get the list of topics? Is the elastic search pick them out for you or do you have something that >> So okay, so in there's an entire other back end here which I don't want to go into a lot of detail but we do other elastic search is just surfacing the topics. It's not doing the topic detection per se, right? um stuff that that um we are going to probably be interested in doing that it's not doing now is you know for example if you said you know tacos Melissa it would find you know it would find people named Melissa who were influential on topics that's not actually something we have now but you know but no we're not saying that you know yeah elastic search is just surfacing you know documents we put in trying to figure that sort of thing out may it may eventually figure out that hey when you search for tacos even though this is the most influential person on tacos this is the person you were always searching for and that may be stuff that's kind future work because search relevance is a that's a whole world of its own but um yeah >> cool any other questions >> so how do you get your topic list you get it just from people giving K plus >> plus well plusk actually is um plusk actually comes after the user has gotten the topic so no we actually get it both based on I'm kind of analyzing the uh influential content and we actually do have a feature now where users can suggest topics and yes that is driven by if you get enough plays plus K's we'll actually say that yes that topic is valid cool and and again you know there's if if you want to talk afterwards there's people who can tell you a lot more about the topic system than I can like I said all I can do is take credit for things so >> cool all right and let me um just hand this off and so we can hear about Spark which I here was the main headline event thanks >> [applause] and cheering] >> So please uh committer on on park and metas which is the uh cluster management system
So you see he's really driving this project and it's it's a great pleasure to come here. >> Yeah, thanks a lot for having me here. Um, so I'll talk a little bit about Spark and actually we have two folks that are using Spark here also. Uh, Dillip from Conviva and Tim Hunter from Berkeley. And so they'll talk a little bit about the applications they are doing um, using it so that you don't just see the toy examples that I have. Um, so yeah, so Spark is a is a parallel computing framework that we're developing at UC Berkeley. We're part of this lab called the AMP lab. That stands for algorithms, machines, and people, which are all things that we're trying to apply to the big data problem
So, we're trying to to sort of combine these things and let you deal with um big data. Um and in Spark, we kind of had two highle goals. So, the first was to give you a programming model for data analytics that's sort of as as nice to use or or nicer to use than map reduce um but is better for more types of applications. So in particular, we looked at two applications that are actually pretty important analytic apps when you get down to it, but that map just isn't very good for. Um, one of them is iterative algorithms, things like uh machine learning, optimization, graph processing, page rank, that kind of stuff. And the other thing is interactive data mining. So this is where you're a user sitting at a console and you want to interactively ask questions about the data. So we we allow you to do both of these much faster than you can with Hadoop
Um the second thing we wanted is to make this really easy to program and that's where Scala comes in. So we have an API in Scola that lets you write functional code that will automatically run in parallel on the cluster and you can use Spark interactively from the Scola interpreter. So let me talk first about why those two applications are a problem um for uh map agents. Um so when you look at map educe and actually other parallel programming frameworks as well that are out there today um a lot of them are based on this model of data flow