Devreal

BDS Alexy Intro

BDS Alexy Intro

Recording: BDS Alexy Intro

are you ready for the first-ever big data scholar conference alright so I'm electric rubber off when she says nitra and also the founder office of scala and scallop by the by serious this year we have the longest scala sequence in human history and we're in this final third part right which is completely new this year so first of all we have Wi-Fi and here's a quiz if the network name is scholar what's the password exactly so it's it's good it should be available in all areas let us know if it's not something wants to tell me ok it's good so that's the Wi-Fi all right I want to thank several people who made this conference reality right it started as just a kind of an idea a glimpse of collaboration the result engineer his name is Anatoly Fomenko his soft engineer cloud era and he's really enthusiastic about scala so he approached typesafe with an idea that closer would host Martin turski and type say for this to me to connect the sub Scala with closer and closer at that time could offer you know pretty small room you know 80 people like we regularly gather hundreds so one thing led to the other and here we are with a full today big data scala conference and cold air is a partner just ingested is the director of developer relations he is manning the booth out there say hi to him so basically we finalized and develop this idea with justin michael martin darsky we're early commuters to this idea and basically promised to deliver the keynotes and I think if you know if you think of two people who together for this phrase big data Scala in human form I think these guys are right we have some very interesting keynote speakers today and I also want to say thank you to close early nitra for the partnerships I want to thank nitra CTO teja bhai for community stewardship vision that basically followed from last year sponsorship of nitra of last year's gala by the bay is basically from adjoining nitra and neither food supports our community initiatives our meetups we expanded the scope and areas of the mid off so that was a really great year in natural just hosted scala by the bay which which was held here last week how many people here were scaled by the way how many people here are the first time at at our conference okay so we have a pretty pretty good about 5050 I would say so you're in for some serious fun so what what did this whole thing about right so there is the rise of big data scala happening in across the bay area and in the world I think apache spark is really what drives this and so there's this facetious acronym smack I wish you know for the lack of a better one we can use and this in one instance eec we have a spark masses acha Cassandra and Kafka pipeline we did a training on this yesterday so this is a fairly standard data pipeline built by companies like Pinterest we need web scale AP is and who basically have millions of users work in real time who's actually should be processed with some kind of predictive analytics some kind of recommendations and kind of there is another facetious in operation of this it's my card which is called available resilient and distributed smack which is an enterprise version right but we're not going to do this right now so I've seen recently what happens with spark you know there are there is a meet-up called by area spark me top which has started in 2012 as an offshoot of sf's column and what I've seen recently happening it basically child grew scholar by a factor of two so a sub Scala is the biggest meant up in the world it has about 2,700 members the spark met up at this point have about 5,000 members I mean there is shoter c/p ratio is much lower so I don't know how many of these guys are actually coming to the group but it's an indication that there is a tremendous amount of interest Ian spark and what often happens is that these guys who come to the spark met up they realize that spark is between scholars so a few days later they join the skull amitabh so there is this interesting sequence house park is driving scala and obviously there is a reason because park is written scala so because it's written scala and there are multiple other systems written scala we were at the point where we can basically pass around objects retaining their types which is really interesting which is really important right so the the way traditional big data is done a bunch of texas dumped in hdfs and then report back and you waste compute you waste basically your constraints if you stay within a type full pipeline you don't need to do this and I think this is one of the key ideas we want you guys to remember and take home from this conference the other is that data science on the JVM is a reality although the majority of data scientists are now working in languages like Python and are we believe that it is fully feasible today to do everything you need in data science while being in GBM and working in scala and we'll have some good examples of this throughout the conference so i have a question how much did is big data can anybody venture a guess or a definition 20 gigabytes and Johnson John anybody anybody else going once twice 22 go by 20 gigabytes on a hex abide but I mean like it is going by size like what is the a good way to define this so we had this discussion i put it actual to a class in berkeley at last fall and somebody said like all the data in the universe on the grants are really interesting all right let's say all the hit on my laptop all the data in the universe the number of grains of sand people in the industry sets somebody said big date is enough data to make money with it sounds like this right so and now the person said big data is the amount of data which you cannot afford right so so there was clearly second ohmic definition but you know what I'm driving at we need to think about it we can just take it at face value right big big data is certainly hype term and what's interesting about this group of people that we have certain values and you guys half of you who were just coming to this color conference I will kind of repeat what I said that skala by the bay but I think it's worth reiterating so there are certain shared values which we share in the in upholding the scholar community first of all functional programming with everything which comes around it and specific shapes of this user refresh on transparency then we really care about abstractions so folks are really working hard to to make things very concise very reusable and also not to repeat themselves so clarity of code is a value reuse and documentation is becoming an even higher priorities we've seen that the talks like cats that people are really striving to make the documentation compilable they they want the beginners to feel welcome and things like skull easy you know where to complex of cats and another library is trying to start from the very beginning with abundant newbie friendly documentation we should somatically compile so all the examples work right so i think this is a very good direction we're going in another thing this community does very naturally it basically formalizes various concepts and frameworks when we do streaming we don't just do it we think about the proper way to do it and compare various ways to do this and we have very good formulas we can express them as dsl's we have a very expressive type system so we can basically formalize various ways of doing it as code and then we can run it and it can compare it and we're doing the same with things like parallelism and all kinds of data processing so I think this is the value of each this committee brings to Big Data community and overall we can just formulate this as thoughtfulness so we know one just accept big data on his face value will ask what does it mean we did an example of this yesterday so yesterday was a fairly unprecedented training complete pipeline training how many people were at the training yesterday so a good amount right so we basically 300 people on the whole smack stack and opposite was the first time we've done it we've had a docker nitrous own boundaries kofsky who's the docker guru helped us really put together a wonderful set up and essentially most people were able to go through the whole thing and play with machine learning in a spark notebook basically bringing it all the way from ingestion through our Chi and took Afghan to spark and Cassandra and look at it with a notebook we have a great events coming next year following this year so scholar and beta the Scala coming back so we'll keep this sequence in April who held text by the bay which is the first NLP conference for startups bringing together advanced users of text and next year we're expanding it to four days one day will be text by the bay another artificial intelligence which basically everything which is not taxed which is vision and speech IOT smart home and things like that genomics we have a sizable group of folks here applying sparkin and scholar to genomic problems and democracy we have actually good group of startups such as Brigade and signal labs who are measuring voter sentiment cop posing various electoral issues to the voters knob viously this is an election year we have the chief data census of GOP presented strata in February and it apparently all the parties and all the candidates looking data mining as a tool in elections and so we're going to have a deal on that open source is going really strong one thing we don't see yet is the emergence of reference tax if you want to put together your own smack stack you just have to piece it together and generally you know it may take a year and the million-dollar investment for a company like Pinterest and maybe five engineers to put it together and the question is why why do we have to do this all over again it's generally ending up varies in a very similar fashion but what what we need to do we need to add a divorce component and we need to add the glue right basically we need to operationalize this Mac stack and make an open-source version which a company like nitro can take and use so we started this alliance of like-minded project we called framework foundation if you're interested in let's go to this this actual URL and we'll start with an alliance of like-minded projects so if you're interested you can see the list of projects you can see the folks in the steering committee and you can see some contact information another initiative is no ETA lat org in the basically the site for you two to submit your horror story about etl so you know if you if you have an interesting story how etl went wrong please go to knowledge of the torque and click Submit and right up open genomics is another volunteer group centered around bringing skills for scotland's park to help defeat cancer so we'll you'll hear more about this but we are very excited about this this this is a very good way for our community to give back to humankind at large right we have the skills required to process genomics data we have the skills required to match cancer mutations to novel treatments and we want to ask genomics scientists to teach us basically to write a big readme file so you can go to github you can go on it and kind of read the readme and work on the problems and will partner with Cagle to run some competitions around this and basically we will kind of try to bring folks in both fields together so you know if you have some spare cycles you can actually help solve this very important problem I hope favor to ask you will have a closing panel tomorrow with some great folks from the industry on the panel and we want some really good questions about everything relating to Big Data and with a scholar angle and kind of customer focus so if you have a question if you have an observation please go to this URL bit ly / ask there's big there's data as big data and ask a good question we really need the questions I want to thank again everybody who made this conference possible there are speakers obviously without speakers we cannot do this so there are some amazing talks had great feedback from scholar by the bay think of Sabrina talk next year trainers who did a great job yesterday all the sponsors volunteers there in orange shirts like this I thank you to them they're mining the all the breakfast lines and basically a registration everything else if something is something you need ask one of the volunteers Jeanie is there indefatigable organizer and Jason is kind of my partner in crime at the sub Scala and of course all that Indies and all the students we gave some tickets to students this year two women hook all through General Assembly we want everybody to who wants to come here to come here if somebody cannot afford it somebody's not the way please let them know and put put them in touch with us and we'll get them a ticket schedule is available online and schedule by the beta Toyota that is an iphone app it basically the website which tells you to save a linked as a little iphone app to your iPhone the Twitter account for this conference is big scholar and the hashtag is also big scholar with a pound sign please use this for all the tweets all right and basically that's all I have we're a little bit running behind there are no breaks between can also there are bounds of the schedule the keynotes will run as a consecutive sequence if you are leaving or enter and during the talks in this auditorium please do not use the middle door so use the side doors and also we should not bring any food or drink into this auditorium however is fine in the other suite 250 thank you very much and I hope you enjoy this conference you inspire and get inspired and have fun