SF Scala: Enhancing Spark's Power with ZIO, Qubism and NLP at Scale, Using Nix for Haskell
okay all right guys let's get this show started welcome everybody hopefully you know me I'm Alexia I'm the organizer of SF scholar here with salah who is the organizer and huge thanks to sell art for taking charge he organized this meetup with demon base who is our full of our long-term friends it's a no company with many connections in the community using the technologies we're talking about and it was supposed to be the physical meet up we said we you know month ago so this is our first virtual meet up we have very grateful for info demon base if they provided the zoom channel enterprise zoom they they feel that speakers and among the CTO will introduce the man base shortly at first I just want to kind of take care of a few things how we're gonna do it Salar we'll go over details basically there is a slack SF Scala slack channel at SM where I take that XYZ work space so the links should be on the mid top you should be able to request invitation by clicking there you'll then get an email you send off as a slack and you should see their subscribe channel if you don't see this of Scala channel you should command K go to it right so the way we're gonna run this if people have issues connecting to the slack they should basically ask questions use the chat of this zoom right there is a chart the chart should be used only for technical issues let's make sure that all the questions go into a service college channel on the slack feel free to drop questions for speakers during their talk speakers can be it's up to the speakers how if they want to pick questions during the talk right they might push it to the back at the end of the talk we'll do the regular voice Q&A and so so if you want some questions asked you know answer during the talk drop them in just a scholar on slack otherwise if you want to ask them during the Q&A wait until they're right there we'll see how this goes this is first of all this our first meetup like this so we're gonna just drive a bunch of kind of details how this is gonna run the main questions during the talk we want the speaker to be able to present right so we don't want distractions so the questions will be basically seen in in in the slag and the slag is persistent so so it keeps being there which is very useful there will be a second channel called jobs some people asked how do we hire how do we get hire there is a different Channel they're called jobs and I wanted to also add I think I've sent an email in February I am helping companies to hire people so I have worked kind of as a broker as a matchmaker so if you want to hire people let me know and especially if you want to get hired let me know I've been doing this behind the scenes if people tell me they're movable they want to move they want a job once I know you know I know a bunch of different companies I can help you to do that so so that there are basically little bit jobs channel feel free to advertise jobs for yourselves there let me know I have now more time to help personally so we have three talks today right and so this is a very packed plan so the next solar do you want to speak now or do should we tell a month is not on the sequence should we say Ryan on this so I'll quickly took what it's like to you then Ryan and then Amman and then sounds good yes I'm good that's right yeah great great thanks everyone for coming to Def scholar and we're actually doing a joint meetup with the Bay Area Haskell user group if any of you want to speak at a future event please let me know you don't have to live in San Francisco anymore apparently was within our minds today's Alexis said on the bottom of the slide you see the link to go and join our slack when he joined please join us a smaller channel for scholar and SF Haskell channel for Haskell and today's agenda with ragazza talks that was going to do we have three talks lined up for you there there on your screen so I'm not going to read it for you Google presented presenters as we go along as Alexa mentioned any job advertising or looking for job please use the slack jobs channel as well so as a scholar so that we can keep it relevant to the talks the conversations that goes on as a normal channel and before we introduce the city of demand base can I hand over to Ryan who's the organizer for the Bay Area Haskell user group yeah hey guys so there's a few organizers so it's not just me it's me Nick and also Stefan who organized the barrier house for meetup we're always looking in this interesting time we're looking for exactly what types of events to do we have a weather event planned for late April but if anyone wants to help out suggest be a speaker be a host or have any other type of event that they want to try to run hackathon or something like that please let me know you know it's great that you guys could be here thanks Ryan demand base we were supposed to have this meetup at their office but because of the current world situation and they're still hosting it for us it is their Zuma cancer I can't thank them enough for this and to just amount their CTO was gonna tell everyone a bit about the man base hey thanks a lot I'm on so yeah like like we said you know I was hoping to host you guys with beer and food but this is what we got so we're gonna maximize this opportunity I'm just gonna like page two man base for just a minute and then I'll introduce the speaker Leo so we are a ki data first company that's late stage startup about 350 people you know over 100 million in revenue we're basically we process the world's business data to help marketing and sales organizations at hundred plus of the fortune 500 with their marketing needs right and it's an AI application probably the earliest in our industry in the enterprise space what else can I say about it we process massive amounts of data right about five petabytes of data probably ten billion events a day have 50 plus machine learning models most of it in Scala Emma live some bigquery we have unbelievable you know data engineering data science team about a third of my organization are data scientists and engineers and unlike most companies we don't divide between them we we truly believe that the best organization is where an engineering product and data science work together and sometimes and many times we like to cross train people who are engineers but we we have you know like a bunch of data scientists with PhDs across train engineers into data science and conversely we train data scientists into engineering right because we believe to build an extra generation of applications we need people who are engineers and have data science skills so you know that's one of the ways we are very different than most organizations because being an AI data company you know the data engineers and data scientists is is really core of what we do so you know again you know Eric Leo Jerome the apart organization you know if anybody wants to chat more about the company please feel free to do that I'm doing doing this talk so I'm gonna hand this over to Leo who's our senior machine learning engineer giving the next talk thanks guys hello everyone or Leo starts can I just say that back that we zo anything that uses zero is very close to our heart ideas of Scala so it will very excited to hear on demand base are using 0ly over to you all right so let me try to share my screen does everybody slides yes all right awesome so today we're gonna talk about an open-source project I've been working on I was trying to use zio cuisines I mean we spark and it was not straightforward so I built like a boilerplate library to do that and I'm gonna talk about it today so my name is Leo bum kale I'm a senior design engineer at the main base I've been working there for about two years my handle and link team and github is double kill super8 is fine so today what are we gonna talk about so first I'm just gonna introduce calendar for programming very quickly just for the audience but spark futures CIO and then finally we're gonna dive into your easy path tiger alright so skyline for programming nor is it won't be like a five hours class it's very quick so just Scylla is a programming language based on the GBM most of you knows but about it here functional programming in mind inspired by a skull and just that's the basics of it you're gonna see more about it today um functional programming is all about transformation of types based on categories here in mathematics you don't want side-effects you don't want me to ability that will help you build better unit tests and better quality card which is easier to read and easier to maintain so most of the time what you hear is all those keywords I won't go into it but just because we're gonna need this function later so it's so you know what we're talking about so most of the time you use Matt which Turner type a to type B and you wrap that into a container and that allow you to do transformation without having internet variables and you can chain operation like this for instance and the compiler will help you so if the type don't match the compiler will scream and you won't be able to make mistakes and transforming you know the wrong type into the wrong other and that we're gonna leverage this with the IO and that will help us write better spot cut I know there's things that you're going to need to understand the rest of this slide is for conventions so basically is just syntactic sugar for Scala to be able to write to about poles and maps into a very easy to read like step by step operation okay you used sir an earlier in the case of Spock would be the data set and the type can is air and you will use map and the same construction we just talked about turn data set of a ticket I start be another aspect that's important to know is that the driver will wait until each we're in bag one so the each operation are similar easy is not lazy completely because some operation will materialize the operation you know so each of us build a job and driver wait for subscribe to be complete before sending it to the next executives what do you do is you will build ETL pipeline this infrastructure is important you understand because that's where you can find where the i/o can be useful so first you load the data from different sources you can load like from s3 and plus brass and be query if you don't do anything specific you will have to wait for the bigquery query to return or the Postgres query to return before being able to start the next one and then you do the transformation you reduce and then you wait you're gonna write your result somewhere in Postgres or s3 and if you write duplicate data if you write in bus location you're gonna have to wait pause the first right to be over before starting the second one you can see that here so read from database one and read from database do they don't interact with each other they could be done in parallel but by default or not so you will just read one and then wait for the next one to be over before you can merge all your types together and write out so about a year ago I went to use the spots to meet 2019 and I went to this amazing talk from Anna should I'm very sorry if I'd read Forbes name it was about how to paralyze a party spot in unnatural ways put in quotes so I took screenshot of a slide just to see more about it also the links here I'm gonna share the slide afterwards you're gonna be able to click on it and watch it for yourself but I just copied back screen shot from representation so you can see here that's Zeus reoperation are not related to each other you don't have to get back to them oh and Anna how it works for targets as well you see the target logo on the side each other's operation are completely dependent so you could if you don't do anything they would run one after the other but if you ripe them in future you will end up with something like that where they start and and and run in at the same time so last year I took one of our of the project that was I was working on and I wrapped everything in futures just to try and we dropped the processing time from 40 minute to 20 minutes is it but future trigger a lot of side effects for instance if you're doing a weight just to make sure that there is a time out you go now if the time I would run out the future will keep going and there you and you might end up keep writing to the database for instance which can be really disastrous so there is like and also implementing a retry for instance can be very tricky because the future as soon as you build it will execute itself you can't take the future objecting and do that we try that doesn't exist so the next step is to use z io z io you will be able to develop that so that the links you can so remember the maps earlier with yeah you can wrap everything into Maps into operation you don't have to differentiate between what you want synchronous and what you want a synchronous with the future you you end up with like a punch-out of the quad is not in the future and those are hot easing the features and so you need to like bridge the gaps with awaits everywhere and it's really dirty but we see IO everything is just chainable into like a full comprehension it makes it cut really easy to read and really really easy to maintain and see I always feel lazy it's nothing absolutely nothing will happen until you materialize the wall application at the edge of your program so that makes it much easier to debug and we will see in the next slide how's that yes it can be even more powerful we spot so zio just for basic knowledge plenty of talk out there it's basically Z start with three inputs three RJ next three parameters you will observe governments the first plug which are all the requirement to exist execute these tasks the second one will be the air and the last one will be is output if everything goes well so the futures are terrible but we CIO all of these problems are solved if you want to be to make your tasks as synchronous you just calls about phone if you want to cancel it you can just cancel it if you want retry you just folder to retry if you want to timeout you just called that timeout super easy if you want to wrap your you read from Postgres and you know that sometimes the database doesn't work which is not a big deal you can just do that we try and wait like a minute and we try you don't have to be like this crazy loops of operations that you would have to otherwise or if you want to have this touch make sure it won't last more than 20 minutes you can just pull the timeout and it will be cancelled it won't run under you without you knowing also because of the environment you can pass all the arguments that you're gonna need within this function through the environment you don't have to pass just bucks ation over and over through all your function so now before diving into the park at you I made a funny slide I'm gonna take opportunity to drink all right Zapata you now so thanks for my dad image animation nasty papa attempt I have to say it in French cause always you won't be able to understand in the recording that's a link to the github and you're gonna have to slide afterward you can just click on it so what is it I try to made like a simple library can just import and use spark and zio seamlessly within your card and you don't have to write any of the tedious boilerplate and only teach you spell up like you will have to write will be enforced to you by the compiler by the type built by Z papaya also you know that in your sparkle sometimes you have like imports parking places and you have to copy paste it everywhere to do all this transformation so I also try to hide some of Zeus there's a lot of work to do like this like I build this library along the US here and it's it's just working for races a lot more to do but it works it works and you get all the advantages of zio doing so retrials time as pollenization in the next slides i'm gonna show you how you use the libraries Moloch is almost like a tutorial and all the snippets of code are copied from Zeus to modules within the library and so you can read the code for yourself afterward or even in parallel if you want you to just go and and open the github repo you'll be able to see where is it cut fit and if you have even more question you can just send me a message on LinkedIn or you can send me two nichiren's into github repo so how do you use it you just do that kidding but almost the advantage of using CIO is that you can put all your application in at rate and because everything is in traits you can unit test your entire application you can just basically call the main into unit test and because all the side effects are inside the other environments you will just have to re-implement the cio environment for your test environment so for instance you can just make a mark database and you that all you to call your man you don't have to start a cluster or do anything or you don't need to do a sec connection to a local database you can just press play any friends so what is the application so I build the zip pocket your ad which is an extension of the zio app but there is a little bit more things you have to do because we need to build spot for you so first things to do is to implement the configuration so you take arguments from command line and that's how it can be build so you see how the type would force you to do the right side you can't miss that then you need to build a spark station and then you build the environment at the cio environment that will be used to to run your application and that content of you of your application that's where you put audio pad if we dive into the detail of the see Pocoyo AB the first argument is what come online is that's necessary because otherwise you won't be able to access old fields a bit the second one is an environment which extends the party's CIO environment and the last one is the output most of the time it will just be you know units because you don't reach on anything since it will be part of the main but you can override it if you for instance want to build in a library that would be called in another software now let's go one at a time so the environment will probably look something like that you can add as many services as you want for instance here we can see that all the CIO 1 so you don't have to do anything for that just cause the CIO ones since you come online arguments so it's being provide it's being provided for you by the O's amidst such a to implement it so everything is like assembled for you just do one step at a time into compiler guide with you sends a logger so you can do anything you want here and spot module and so same thing here is a spot service is building on the outside and if you follow the compiler it's easy so the configuration I made the decision to use calop because I find it really really also I mean easy to use you can submit an issue if you disagree and we can work together to use a mob strike one but I really like this one I've been using it for years and you see how you can just add as many arguments as you want and then on the on the companion object you just make the helper and to help you allow you to do this so within the CIO for comprehension you can access your parameters from everywhere in your software you don't have to pass you know a config object over and over in every single one of your method or that can increase its which is in every single see you your message becomes much much simpler and we're gonna see the same for holes or helper things so now spark within CIO so that's how you assemble the spark building so the spike builder allow you to override this but config based on your arguments there is already a message which have assumed a set already as the default implementation that's just an example if you were to override it I find it pretty useful for instance with some bigquery library sometimes they want you to override sequential and things like that and but you can just reach an Spock builders that's what the default library does and so you don't have you don't even have to implement it if you don't want and then that allow you to desert same as log Yuman you can just if you need spark you just go spark and it's there accessible for you all across your program without doing anything you don't have to pass the sparcstation everywhere in every single message over and over and we're gonna see that I build helper functions that sometimes you don't even need to do that at all because you're ready of this with its you already have a CIA operation which contains from the get-go so let's dive into all the help function so the first one is the zds it's like zio data bits are data set so you you you take the spark is part of the function so you don't have to do what we saw earlier you don't have to call the spot module and then do something with it it's already building and then you can do your import Sparky increase its blah blah so that's super easy another one for instance if you want to paralyze if you have like an array of case class and you want to turn it into data set you don't have to do all the boilerplate of like importing in 2d s and that polarization there's already a helper for that if you want to do it transformation in its transformation involve a zio region object so you see here it's written a task of the output case plus and so internally our I I builds runtime for CI hero and and turn and do the transformation in imports Falcon so you don't have anything to do and so you can read your code like very simply like it's everything is just zio Spock is becoming one thing instead of finding Z's to world if you want to broadcast its already built in as well you don't have to you know to do like spark and then get to spark context and then do the broadcast it it's already done and there is like so much more to write so if you want more things submit an issue or you can even like participate and like see me to PR and we'll add more things I was also thinking you know if you have like an application which is not done or that which already exists you want to convert it like one bits of time they all helper function to turn to convert from futures to zds so for instance if you have like a future which produce a data set you can just call to zio and if we make GDS output it will pass it will give the execution context from the zio environment so you don't have to do any of that either enjoy is like a little bit more helper function if you look at the a jay-z feature silent in the repo so I've been converting the same application as earlier but it grew a lot in the last year so as it used to take like 20 minutes it ended up like now taking 3 hours in the recent months because of adding more data and more features and everything so I try to wrap everything and it sounds like a full project that while using futures everywhere as you remember in the early slides and so we took a while but like from the three L's that you used to take it now it take like about 12 games it only tests are done and it's it's faster that's less error because it's easy to retry like if one of the read fail you don't have to kill the cluster and we started on tyrosine because now zio does a rich white automatically I mean not automatically but it's easy to use which why of zio when they are like minor failure you can differentiate easily between minor and major failure now you can also cut off on some operation if you have like several feature constriction and you cannot like okay if I have a lot of time I'm gonna do this more expensive things but if it takes too long I'm just gonna take this shortcut and you can do that with the time I'll Kim it's built with the i/o and also you have a better arrow log when something fell with luck sometime it can be very tricky to know where coming from but because everything is fibers now you get the exact stack trace of everything that was happening is happening in as field and what will even you even get what will happen if that we don't have here so it's really powerful to know exactly where the error is and the cut is easier to read and maintain because now everything is part of the i/o there is no more cards that is bang out of rather so everything can be chained together so tightly strong so you know exactly what if I don't see you out for each function so it's it's it's different world it's really really fun so you remember what we Adams a little example from the spot submit presentation we had like three consecutive tasks but that was you know like an example so let me show you what it looks like in prediction application there you go everything running parallel it's amazing you can unlike thousands and thousand you can max out every executors forever all the time they never have to wait they always lock something to do uh because sometimes you know if you do a lot of little things you might have so much don't even have in a row to partition across all the executors blame Zeus cases you can just run everything up all the tasks at the same time and you never ever executors doing nothing it's amazing so what next so the library is just like we sent it's not even in I didn't even push like person one so it's a milestone I I posted myself some issues if you want to help me I'd love to you I'd love to have like a Gatorade so also example Asher earlier will just be rolling for you right now I heart could use a spot version of the one we use a demand base but that would be awesome to to have like a plugin to build for all the also stock version there is also a didn't putting this aside but I'm still on zio LC 16 and I have been made too big move to Susie layers yet so if someone wants with help as well that would be also a member and I walk on it and I think that's it I don't know where we're in time but I'm ready for questions I have one sorry I didn't raise my hands I didn't find the feature so so since we were talking about Zeo I have to ask this question what was your design decision between CEO and and cats effect I've never used gas effect I've been using zero for a long time so I couldn't tell you what advantage versus inconvenience I just went to a lot of talks about the i/o and I really liked the easy of use of zeusie I've never used a spec so I can't tell you okay actually so in your first example you showed you went from 40 minutes to 20 minutes and actually all all what you've you've showed here I actually tried also to submit in parallel well I used cats effect but that's the same concept and I didn't actually gain anything so I'm wondering if don't you think you were I think you actually kind of half answer it at the end but I suspect that a new first example your first job was not utilizing the whole cluster and so you were able to actually submit twice and utilize like more like basically you were utilizing half of the cluster in each of the job and you've been able to run them in parallel right also in my use case they were I was reading from a lot of different sources like bigquery s3 phosphorus and so for each was really do you sends Aquarians and you do nothing while you wait for the result and so being able to send all the query at once that speed up a lot and then there were a lot of little tasks outside you know like tons of little tables that you rub from everywhere that you transform to just merge into this big thing so is that like a pure rose and so is entire cluster we're just waiting for two executor to finish tweeting Zeus little tables and so with that you can do all the little operation in follow and then use a big operation on your utilizing the fruit cluster okay see ya because I guess I was so the first that I was submitting well basically submitting them in parallel the first one was already taking the food cluster and then the rest was like just waiting so it was already parallel pipes park yeah but so you if you see that Lester is fully utilized but you I mean you can try to like divide the politician number by two and see everything magazine run bus at the same time and it might take so exact same time but sometime it's faster you never know and also if you just read from one source and have one transformation in one right yes that won't be very useful this is this transformation this will just be a lot of work not much game but if you have like this type of thing like I was showing into the first slides those guys you know where you have like several sources and several rights you can even do all the rights in parallel as well like if you want to write first question right where you can do both at the same time but if you if your job is very linear which like one source one output in one transformation that that won't help you cool thanks good you have a few questions from slack I can read them out for you one of the questions is Ludovic Claude has asked whether you know about a another sparks your project by a github user Univ Alan's I don't know I'm terrible with name but I I submitted this project when I first started working on it and read it and someone contacting me I did like your helper file which does like if you have a function and we talked about merging the product together but then we talked for a while and then the communication stopped and they never came back to me I don't know it was there where the same people okay great then Mark hamster is asking how do you handle Spock job and task cancellation relationship to CIO cancellation semantics I my belief is that if the task if the CIO task get cancelled I don't think this spark step will be if they all submitted to the executor those things they will just finish before you can do anything else but I haven't done heavy testing on that time I'm not really sure great and the last question in the slack channel is if you could post the URL for Z spark in the slack Channel yeah that's a good idea let me do it right now any other questions from anyone by voice or that's very good all right you showed an example where a lot of tasks were paralyzed and you said that they wouldn't be without CIO because so because the it would be not possible to split in the data but so how is it that CIO magically is able to split the data when spark cannot so this is a little strange to see that what exactly is the magic here the the driver will have these queue of job to do and it will just take one and submit it will exist there why do you think that could be done and get a result when you die or when he calls it for you will basically tell the driver just submit everything to the executor at once you don't care for if they're done or not and so hope you do that since executors will be able to finalize that that has nothing to do with each other like out of the cluster will do one part and those are how we do something else so executor is running CIO fibers that's right yeah the driver that is running five fibers the executors okay you can with with zmapp helper you can make zio tasks within the map of the data set so then juice fibers will run on the executors but z-pak io is focusing on the macro on the higher layer above spot matters distributions operations but right any more questions so I have one actually on the on your slide when you said you switch from future future to Z you say oh and you went from three hours to two hours yeah that sounds big how come I mean I know futures are slower than Zeo but you know you're paralyzing only the the submission of the work so I'm a little a little bit surprised on those numbers though it's I mean it must be specific to my to my very use case that we have a lot of like you see on this slide the little transformation like C to D and the A to B we have a lot of small one of zoos we're in we wheat from like I don't know dozens of table and data source and each of zoos sauce is neat little transformation and so by exams they ought to be done like bonus time and with and also they were like for instance if you do one specific query for one specific ID this idea might take forever and you might not want you might want to give up on them and so with zio you can do this time out but with futures you would have to wait until they're done and you don't have a choice so that's all those tiny little like addition you get with the iOS that probably ended up to this game so I see so it's not really the pure performance of desire as zo but more like features that consolation stuff an anti-communist that you don't do like very fine testing unlike each of the aspects of parent but overall you see a pretty good game even in like development a development time like it's much easier to add something and not make mistakes and because everything is just enforced for you by the compiler cool thanks and if you have any questions this is Erica I'm hosting this meeting if you unmute yourself you'll move to the top of the participant list because everybody else is muted so if you just unmute yourself you can kind of create a little queue and then I can call on you or or Meucci if you start making any noise of course but try doing that or use the slack channel for questions as well any more questions about notes great well if you think of any questions afterwards please feel free to continue conversations in the slack as that's our hallway track racing click on support io you can just go on the on the repo and you can also add me on LinkedIn if you need to help and yeah thank you everyone great thank you Aleksey are you available to introduce the next speaker absolutely so our next speaker is Jerome Banks Jerome is really as og as they come in the world of data right so we you were colleague says if some of you remember that the social the social startup who brought Justin Bieber to the pod top of the world in terms of cloud score and so you know it was actually an early kind of preparation and so Jerome was doing all the you know crazy things with Hadoop just you know if anybody remembers Ozzie that forever plays scars on our collective souls but we survived it so and Jerome started open source projects there I think Brickhouse was one and it's I did follow closely but I think cubism is an inheritor of that so really happy to have Jerome in the community and super awesome to have Jerome speak and being hosted by his company welcome to Rome take it to take it okay yeah thank you thank you Alexi yeah thank you so much okay so okay can you silly King this lights no slides yet we can see you ourselves see you so what you need to present here yes and now all right now what guess it when you can see it present okay can you see that yep looks good okay great great so so this is gonna be a little more a little more abstract a little more philosophical than than Leo's talk and talk about yeah like the old times what Lexi was saying and then sort of the new times and sort of how we approach problems you know just talk about sort of like what the problems are at demand base like a mom was saying you know the challenges we have here the special sort of you know approach that we have because we have to deal with data at scale we want to do you know advanced AI data science you know what the problems we run into what sort of the approach is and then this library we've developed called cubism as sort of like a solution at it and you know how philosophically it fits in and then and then pragmatically what what are the tools that we have in there and and and how we apply them and how we apply a to to the intent problem and here's sort of a marketing slide that I sold I stole from another slide deck this is sort of this is not a sales presentation you guys aren't gonna by demand based product but but basically you know the idea is like if you're a b2b company and you sell construction equipment you have a website and people from IBM come to your website and then people from IBM are interested in buying construction equipment and if we if you know by scouring the the you know by scouring the web and and crawling lots of web pages and getting lots of signals from you know various secret sources and you know various places we fear about that people at IBM have the intention of by construction equipment and if they have the intention of buying construction equipment then then your salespeople should should go to IBM and then find those people you know trying to sell them something so that's that's sort of what intent is that sort of what you know account-based Marting is like like um I was saying and and and you know we have we have we crawl tons of web pages and and we have various signals from various different sources and and and you know all this data and and the trick is how to figure out you know they want to buy construction equipment and yeah here's here's just an example there's a lot of keywords so a lot a lot of the sources is that you know if your your company construction equipment what are keywords associated with it you know we have we're sort of you know at some level keyword based right so so then you know a lot of the problems we run into is generating bags of words bags of keywords and this is you know a classical technique natural language processing it's been around for decades probably like since the 50s and sort of like a way approaching natural language like and and basically you know all it is is is is keeping you know a bag of words as a sparse vector so so imagine that you know each keyword it is a dimension in some very highly dimensional space and then maybe how many times add that keyword is mentioning is is like the magnitude of a vector and you know like sort of in skull you can sort of just imagine you know modeling you know thus far as keyword at least that you know enough check well as a map of strings doubles right so that's it's a map each string is your dimension each double is the magnitude of that dimension right so so so the trick is you know okay that's a dying a words but you know we want to generate lots and lots of these bag of keywords bag of words right so we want to calculate like all all the key words are the things IBM from IBM are saying you know all the key words of all the pages of you know construction weekly you know the website you know what that is you know and then then maybe you want to keep track of key words you know globally because maybe you want to normalize those scores you know and then maybe we want to like combine it with you know other sorts of dimensions like we wanted oh the people of France they say these words for construction equipment you know they're you know in French they these different construction equipment words and you know or maybe like these words are used in this industry or these words are using another ministry so there's there's there's lots of tricks we can play and lots of games you know we want to do one you know well but the thing is like we have to create lots and lots of keywords for its you know and and and that that's sort of the our modus operandi right and then now now you know sort of think back you know how do we approach this what what what is the actual problem we're trying to solve and and and I would say you know you know we're actually doing lots of aggregation and you know we're generating features for for machine learning models and aggregation is essentially you know extracting features front from events you know and then and then feature extraction is basically aggregation because feature extraction is reducing dimensionality at some set you know in rag Gration is dimension area dimensionality reduction right the world consists of events you know you have all these events the worlds and and and you you know you that's there to many events to to put into you know your your machine learning models so you need to reduce the dimension somehow and what is that really that's aggregation right and then you know aggregation you know really is just you know we want to you know you know when we're aggregating we really want to take features and you know we want to take the world events that generates something which is useful which you can use later which are features and then you know this is a little controversial you know but but you know so you know but people think like oh you know you're doing aggregation all you hydration people think oh you're creating a dashboard you know oh you're you know Ukraine tableau you know you're creating these fancy charts and graphs you know you're you're doing you're doing what Tricia has been said you know as analytics or business intelligence and and you know that's you know that's that's nice you know that's very pretty you know you spend a lot of time you know looking at these fancy graphs and charts and and and doing you know these things but you know that's nice but you know eventually you like you have a human who sits in front of you know some dashboard you know they look at tableau and they spent a lot of time there but they might not necessarily come up with sort of any sort of actionable you know actual actions to perform or you know they they they they might get an intuition but it doesn't really go across the organization so so what I'm saying is that you know instead of just being a dashboard you know you should see aggregation as the source of all the features that you need for your models right you know generate lots of features via aggregation and then from that that those are features you drive into to perform all the things you need to do with AI and you know machine learning so that then you know you take those features you can do things like you know use them for for your model right that's that's the model you develop from those features right you know you can do things like clustering you can find you know different these guys who talk about I be you know IBM and Apple there are computer companies they're all similar to each other because there's some sort of distance between them you know you get that from the features you know from the aggregate so you can do outliers and index say you know people who you know people by construction equipment are also going to buy concrete you know things things like that or you know maybe this this this guy is by much more concrete than the other guy you know you know if the guys in Germany are buying more concrete's and the other guys so so so all those sort of data science the things you can do is is buy models which you get from generate features and little controversial but but you know and now to go back sort of ancient history you know like like Alexi was saying you know so so in the beginning you know was brick house right so this is this is almost you know almost 10 years ago at this point about eight years ago you know like what Alexi was saying you know we worked on something called the cloud score and and and while developing the clouds score you know we develop breakouts and you know what what was the class score Klout score was how influenced you were and it was basically you know linear regression based on different features from you know social media you know like Justin be you know how many times people retweeted me how many how many times people retweeted you in the last seven days how many times people retweeting the last 30 days how many how many unique individuals retweeted you how many Facebook Likes you got in the last seven days you know how many Foursquare you know mentions did you get in the last 30 days and so you know that basically you know the feature engineering of that was basically producing these aggregates you know you're the aggregates of the social media events you know you have all the social media events the the tweet you know share is the tweak the tweets the the likes you know the the reshare it's you know all the different social media events were the events you know we aggregated them we had you know different yet unique we had counts you know some other aggregations we had a Gration so their time periods but essentially that I was the model right we generate aggregates we put it into you know we use weeka to generate this you know this linear regression model you know and so that you know if you had more retweets your class score was higher and some linear combination of that occurred to you know according to the model and and the way it was implemented was you know this ancient history I don't know if you guys remember this there's you know high tide is sort of sequel based you know sequel based development and what Brickhouse is is a set of high VDS and a high video apps you know user defined functions and user-defined aggregate functions that you know allowed us to generate those features in an efficient manner right so so it's it's it's you know not not only you know is it you know it's you know a single development will get you halfway there you know so it's good it's a DSL you know you can write things quickly but you know the Machine especially at the time Mitch the machine wasn't there in high so so we need to be prodded and and also you know the design patterns as well you know and you know that that in the in in in the sequel queries themself and this is sort of the other aspect of cubism is that you know it's a library uh you know so so you know data engineering you know is a thing you know there there's there's good Bay engineering there's bad data engineering but like you know you can do sequel you can do sequel based development and their design patterns of data engineering in the sequel right so things like salting a query or doing you know doing a join or doing doing a Maps I joined versus you know a self join right so so there's it's you know even though it's just you know looks like very simple sequel there data engineering design patterns and and the data engineering is in the shape of the sequel queries themselves right so so if you wrote like this Bayesian analysis this data engineering or so you have some design pattern or something like that in sequel you know how do you use that how do how do you capture the fact that you know here here's an element of data engineering and and that's something which is hard to do it's you know pure sequel based development live with hive and you know brick house and please please visit the site by the way and star the project it's still inactive it's still being actively used even though it's not really being maintained you know but but you know the answer that question is the next generation is a library developed developed very similar versions library at several different places but but you know here at demand base called cubism and we used in production on one project which is more more web analytics and we're applying it to to the intent project which is more keyword generation and you know natural language processing and and what what cubism is it's a Scala SPARC library and and and and the idea is that it has reusable transformation and and and and this is this goes back to you know first of all data engineering is a thing right you know what is that thing right so you have some logic in a sequel query which is data engineering because you're you you do the query in a smart way rather than a dumb way and how do you you know reuse that knowledge or how do you reuse that that's sort of a piece of work like you know if you're if you're writing software you you know you you have subroutines you have object libraries what is that at the data engineering level you know and and my philosophy is that it's you know that's what Scala and functional is all about is is that you know the data engineering you know you're coming up with transformations which is you crummy with functions that give it a data frame outputs another data right so all the good you know you guys are the Scala guys you understand that you know everything is functional [Music] so so so what what cubism provides is says the transformations you know for doing feature generation for doing you know various data data science type type transformations so that's what you know that's that's the purpose of the library and and it's focusing on swear these future generation aggregation and introduces you know several several different concepts and and it uses a concept which first sort of developed by my colleagues at quad cast and and and something this this is gonna get little little abstract a little little a little crazy here pretty soon so don't get you know you might get a headache soon and usually at this point you know I would look into this crowd and see the their eyes crossover but we have we have the notion of X units and Y paths and explain more detail later and and this is a way to represent you know multiple dimensional features in a simple in a simple way and and and I'll go to details later and then also what cubism provides on a more tactical level people are from you guys have probably heard you guys are the SF Scala guys you know about algebra and algebra is a famous library from the guys at Twitter it was used in a project called summing bird and it's sort of a model of you know it's a model category theory of abstract algebra and applied to Scala and they have have about reusable components and we'll go to that there but but but you know spoiler alert um what my cube isn't provide is that if you have an algebra aggregator we have a way to convert that to a spark user defined aggregate function and spark has you know if you ever try to use user defined a group function in spark it's it's it's not pretty it's not easy to use so so this provides a very seamless way to take your algebra Die graders and using spark easily and then and then on that note also cubism is a set of exotic I graders you know so so so you know we've already said you know aggregation is sort of the way we approach you know future generation we want to go from lots of lots of events - small small aggregates and then like we'd already said that you know aggregator soar the way to go because that's that's the category three way to do it rather than sort of sort of some weird you know arbitrary definition and by the by the guys at the spark spark level you know but then it's like you know how do I approach these other problems and make a greetings for them and you know things like you know finding the Arg max and doing the fish away you haven't do cardinality explanation vectors in time series so that's that's basically what cubism is it's a set of it's a library which is intended to be used for aggregation under several different use cases several different domains right and it uses X units and Y pads for handling arbitrary features and in arbitrary dimensions and then it provides you a way to create aggregators easily and then it has sort of off-the-shelf aggregates for doing these fancy jobs okay and now let's let's go back you know X units and white pads and a next unit is a string which can represent multiple slice and dice combination of segments of combination do you know of dimensions and that these these dimensions are called white hats and the idea is you know if you want to find all the people in Germany who go to your website all the people in Germany who go to a single page on your website all the people in Germany who are male who go to your website you know you have different dimensions you know geo you know gender you know the site you know you know maybe have like a browser you know you know you know different you know give an event it has different you know feature attributes and different combinations feature attributes and you know you might be you know so one person might be interested in one one of those dimensions you know you might want to interest it in several combinations in those those dimensions and the idea is that with actually in some white pads you're able generate aggregates you know for all possible combinations of your dimensions and the idea is that you know an event can be mapped on to multiple X units sososo say there's an event like you here here's the example like you know that somebody goes to the main D be calm and and they go to the page you know home.html and they have to not be you know service account one two three four and you know they're in San Francisco and you know that you know this person's in the ad tech industry so like that or the accounts associate with without tech right you know well this event corresponds to the segment you know oh they're visiting main D be calm you know it also corresponds to you know the the segment it counts you know I one two three four you know and then also it's you know the accounts you know ad tech you know then then geo happens to be us you know but huh so has to be geo the people from you know who went to be be calm in the US have to be these people from at Tech who went to be be calm so all the different combinations right so so that's that that's sort of the idea is that that any one event maps on to multiple exits right and then you know how do we process that you know so what ideally you know so so I want to come up with you know given sex you know I want to come up with aggregates for for each of those accidents I want to know how many people you know can't you know on my side you know were from the US how many people on my site we're from a tech how many for you know on my site were went to the home HTML page and how do we do that and just sort of an abstract MapReduce type way the technology's not important you know doesn't matter if it's park or hive or MapReduce or you know data flow or whatever whatever it is but in terms like the the general MapReduce concepts you know you would take an event row you would explode it to multiple X units right so given this one event you know you know it corresponds to all these different X units right and then like in a MapReduce job you know you know we so so whatever one row you know gets exposed to multiple rows so in the map phase you actually map on to multiple rows and you generate one row with the event plus an X unit and any and there multiple muscle words generate right so then then in MapReduce you have this phase in between which is sort of the shuffle and sort right so so so it goes from the map and then it sort of gets shuffled and then you group it by that X unit right so all the all the records which were saying you know geo you know country equals us all the rows which have the X unit you know attack you know that then they would get pushed together at some reducer right and then so on the reduce phase you know you just take all those you know you're grouping by the next unit so by that that segment that you want to aggregate and then you do your aggregation in that phase so so now you just count them up you're getting all the roads those event row so you do you your a great operation you're counting them up so now in that one reducer you can count all the people came from the US and the other reducer you can count all the people in attacking the other reducer you can count all the people you know in demand acecomm who you know if you know in the US so all those combinations of X units that you want you know interested in generating segments for and trying to keep track of this is a way to do it and and the important thing is like you're you're creating you're doing it in one pass of the data right you know so it's not like you're doing multiple passes data for each of these different dimensions you have all possible combinations and dimensions you just do it all in one pass there's there's a big explode but but you know this is big data you're able to handle that and there's other data engineering techniques you can do that to trick that you know but that's that's that that's basically what is so so the idea is that we want to generate X unit keyed aggregates from from event rows and this this I'm going to pause there take a drink and you guys just just sort of let that sink in for for a few seconds it's just sort of an ID in a way I'm going to look at sort of like the aggregation you've been doing and and this is a technique like I said we first first used it at you know quad caste and other companies generate jump shot their their pipeline is sort of based on this that's sort of e-commerce you know tag we did for our web analytics and it's good it's good when you have all these different dimensions which you want to add and get get the arbitrary you know combination with and it doesn't really matter on your use case or your domain this this is you know we're doing it for ad tech we're doing it for NLP you know but but you could say ply the same technique for you know ecommerce or you know other industries right so it's a generic data engineering technique applied applied in this manner so so let's move on yeah and then and then what is the good good thing about it it's just single single string that we can have have to for that segment we could add dimensions arbitrarily like if you didn't have Gia before suddenly we can add Gio without changing a lot of code and then what does cubism provide cubism is a toolkit for her being able to generate X units in a lot you know there's a way of you know on the line you know provides a way to define X units you know given your row or given business logic of pulling it out of your records I wanted this is how you got the geo feel this is how you got the industry field there's there's transformations for exploding and aggregating and then some other subtle things you might want to do you might want to salt your queries and then like things you want to you know you want to find the top top ten you know you know do a collect you know art max and then maybe you want to do outliner detection or you want to say you know index affinity people from you know people from from from this country are more likely to buy concrete from somebody for another country you know and then then also you do sort of like Association matrix clustering so that's that's sort of what cubism provides you Xing some iPads and then going forward to sort of the next step you know going back you know aggregators there's also a lot of tooling in there of being able to support these aggregation functions you know as an aggregator as sort of an algebra concept as mass directly more onto category theory you guys are the you know the SF Scala you know category Theory studs right I don't have to explain to you you know why it's a good thing you know the fact that you define you know your aggregators as monoids right is is you know then you can do the computation and different you know different parts of the cluster you can distribute it you know across different machines you can take aggregates from you know previous days and merge them today so you can take you know yesterday's work and reuse it to tomorrow right so that's that's that's the category theory you guys I'm not going to explain that to you guys you guys understand that better than then I can explain it but but what cubism does it allows you just you know does all the work of creating the the user-defined Aggron function from your egg reader and and yeah and then it's stuff um the code is ugly because it goes back a few versions a skull of a spark and it's before types really serializable so there's a lot of tooling in there to try support serializable type type information right that's something which only came in the later versions of Scala which we had it and we had to work with at least in my previous companies so that's that's sort of like the tooling that kiba's it provides and then let's go into sort of like what are the what are the actual off-the-shelf aggregators that would provide which which cuba's it has which is which is very useful and then and like what we're saying before you know we'd be produce parse vectors and I don't you guys ever saw I guess I can't look at the audience or get feedback there's a plane there's a film called airplane say there's a scene you know what's what's the vector Victor wait what's the what's the Clarence Clarence and I can't hear any laughter obviously even if you were laughing that thing but but but what key was it provides is an aggregator for for doing very efficient you know vector operations and being able to aggregate efficiently with with with spark right so you can say you know group by you know say so so group by so and so dot AG and then you would put in an aggregator as AUD AF and you can create these these keyword vectors right so so it has like an aggregator to give in given a keyword you know and then agreeing over some rain create a vector a keyword vector you know with a given magnitude right and it outputs it puts a string outputs a map string double right and there's also aggregators in there like give it a bunch of vectors you know merge them together like so so if you want to aggregate you know over the last 30 days or something like that you know group them together you take bunch of other vectors and enry merge them together and then then also to support this it has you know since you're sort of going vertically you know going horizontally you know as an essence structure that's how your vector rather multiple rows then you then you can take that that necessary mastering double then apply them as UDF because there's all these you know standard vector operations you want to do like scalar multiply normalize dot product you know cosines of all these things and then and then sort of the secret sauce also in there you know you know you can you can you know writing those things are very you know simple you know in terms of just naive implementations but the problem when you run into running you know really large-scale spark jobs is you know being able to serialize back and forth and and and that's where your your cost is really an efficient memory usage so there's a notation of something called a vector buffer which is a fairly efficient data structure for for emerging vectors and then for you know it's tours them as you know as a buffer as a byte buffer so you're able to realize that very quickly back and forth between your spark jobs so that that's that's also turn the secret sauce behind that and then we're not we're not gonna go through all the different you know sort of but you know tools that keep us calm because it's so feature feature-rich but another one that I want to mention is another thing which probably take a whole speech talking about them as well there's aggregators to provide an implementation of the can be sketched set and and can be sketches is similar to hyper log logs in the sense that it's a probabilistic data structure for estimating cardinality right so you know Justin Bieber you know how two million retweets but maybe he was you know over over 30 days but maybe the same 30 people doing the retweets you want to find the cardinality you know number unique people rather than the number accounts right so that's a hard thing to do you know in general alright that's that that you know to come up with an exact number basically you have to have to resort and and send all the other values to a single node and you know otherwise you might read you know double count the value right so so you guys you everybody use you know you know to date engineering you know you understand that funny makes a hard problem every time when you say select count distinct you know a puppy dies you're killing puppies you're destroying lots of resources so so KB sketch sets is a way to do this and that's what that's what humans and provides you can read elsewhere about there but but I like I I happen to like sketch that's better than Piper log logs because you know hyper log logs you know might take you know less space technically and they might have higher you know a higher precision then you know in this case at some level but but for low reach values you know sketch sketch dresses you collect the K smallest hash values so so so since you have you have keep an array of about like ten thousand and then once you go over ten thousand you drop some some values but but the good thing about K V sketches at their exact for small values and then over a certain amount it has a certain fixed block of memory and then you don't need any more memory to represent the setna matter how big it is and and another another thing is that you know you can create you know you use it to do Jaccard similarity you know how similar you know they are you know how some are the guys who you know who looked up concrete as the guys who looked up construction equipment you know what's the overlap you know in that set and then as before the other secret the secret sauce of cubism also is that it has a very infant in you know efficient version of the other sketch right so so you're able to merge it very quickly be able to see her eyes it very quickly right so so it's technically just an array of Long's but then how you actually deal and how how to make it perform and spark very efficiently that's that's where cubism provides you this this sort of secret implementation so so you know a bunch of other stuff in there I can't really go through all the tooling you know and then and you know tell you how great it is but just sort of just you know just say you know how we're using this right and as I said before like we have one pipeline which which is already a production which does sort of more generic web analytics and we're using cubism for that and we generate it's and and that allows us to process you know different dimensions in an efficient manner and then we're also now applying the same library to this more NLP focused pipeline of generating keywords and bag of keywords and and this is how we apply it you know towards the intent problem so so the idea is that we generate X units based on you know different art attributes of you know parse documents right so we we parse all these documents and these dockings have different attributes it's you know from from a certain website Wall Street Journal and you know maybe somebody from IBM read you know the Wall Street Journal and maybe the person was in Germany and maybe they maybe this document is in Spanish you know so we have all these different dimensions of different you know documents that we have and a people and and different things you know we have different dimensions so we generate xcms from that and then you know without going to too much detail about you know the secret sauce or anything like that you know what we do is you know we generate keyword vectors per X unit right so so given those we need to use the keyword aggregator and we output these vectors and then we want to create other vectors also you know to match that maybe want to create global vectors because we want to normalize these values maybe we want to take those vectors and merge some overtime Rangers and then we do all sorts of data science C type things with with I right maybe we want to you know calculate different scores how you know you know calculate scores by by combining vectors you know use the profile and the dot product with with the guy with this other vector of the domain so given the profile the dot product of the IBM site means that IBM has a high score right so that's you know that that's where the game you know that the Cubism provides tools for us to approach today science problem and since and since we're doing arrogation at scale since the spark you know is it can be very efficient and since we're you know since we've have an efficient aggregation then we can you know do things at a much larger scale that we could you know otherwise right if it was just honestly your laptop okay we're still early on time I guess and and you know just just I'm you know you guys are probably you know phased out I can't see anybody's you know eyeballs from from here since we're remote you guys are probably all fall asleep already but but this is more philosophical you know take away and and about the problem we have and and and and the thing I'm gonna say and this this you know all this stuff is very controversial you probably stick with it but but but you know beam engineering you know at the end you know is really all about aggregation you know data science really is all about aggregation and aggregation I would say and you know it seems like a blunt hammer but but in none of your problems your your solving is going to be tractable unless you can reduce a to an aggregation because you can't sort the universe right you know you can't like take the data from from across the other side of the galaxy and bring it here and then you can't you know you have to be able to read yesterday you can't read you all the time when you did yesterday today so everything has to be you know an aggregation and and then the approach we have is generate you know generate as many aggregates and feature as possible you know you know since its make it an efficient way to generate all the features at once and and and and and do your data science by looking all the features in one big big big swoop right so that's where the philosophical approaches well don't don't hunt and pick you know you know case by case don't don't create a graph and have some human try to figure about what the important data is just generate everything and run it through an algorithm and then cube is is the reusable library too super for that right you know so cubism is data engineering this you know the view number use and data engineering is functional that's why you know the future data engineering is skull and SPARC and functional languages and you know you know it's not gonna be sequel based development at some point you know everybody writing sequel queries that's gonna crash and burn without you know some real you know Scala spark development you know nobody's gonna be able to maintain that code so so so the real future being engineering is just coming up with reusable components and and and functions is that and then generating axioms and y paths and able to create features that's a scalable way to do that because otherwise you have to change schemas all the time and then implement you know these efficient Aughra Gators you know so reduce everything to an aggregation problem and then make that aggregation problem very very efficient okay so that's I'm not sure where we are in time but that's that's all I sort of had any questions Eric Eric what what is the process now so yeah Thank You Jerome if you have any questions either slack or unmute yourself now and I will see that you are queued up we I can't see the slack Channel presenting that's I just less sure how to we did have a question from a high schooler here from Jeff is the next unit basically a row filtering function it's sort of the opposite it's like all the possible rows all the possible values that a row could could match so so the next unit you know they're they're more more X units per row so so so the idea the are tonality is is is one row produces many X units in fact that sort of answers this question you know one one thing you know what one one possible question is like you know if you're you know if you're producing all these actually minutes and there many X units per row you know aren't you creating some some monstrous giant job with you know which is completely intractable and once you add a new new you're doing a combination and once you add a new dimension don't you have a you know commentary explosion and and and yes that that's true and then you know one one if this is what you're saying one one approach to solving that is cubism has a notion called filter rules so like maybe you don't want to you know you don't want you're not interested in the industry with the geo for some reason because you already have that information so you can decide what what X seems to generate and what extreme it's not to generate there was another small question on some slack from Martin basically about basically the keywords you show that you can think of it because it's a multi-dimensional vector space where the keyword itself is the is the dimension and then the W is the value or the magnitude he wanted to know why the map was between a string and a double and not let's say a big int or a long I took in a crack at answering him but why don't you why don't you yeah so the hashing yeah so basically you know an md5 hash of the string for the dimension and and we do that in places like under the hood you know so yes so you could be you know so that's that the hashing trick and the sorry the the longer the big int was instead of a double like for the magnitude why isn't that an integer why is it a double as we're counting keywords fundamentally well because well it has to be a scalar so so you right now that's how we implement is as a double just because you can represent an int as a double we can't represent a double as an int so it's the you know just just representationally so I sort of responded to him by saying that well if you do a tf-idf or other transformations on the keywords then they're no longer yes right yeah I mean yeah because yeah for the stuff we're really interested in you know we because we yeah we want to normalize the values right so so we really want to deal with doubles and and and yeah a lot of stuff yeah yeah I mean I mean and the other point is though that if you have a different data structure you want to make it an int you know maybe you would write an angry which does that right and then they're like for example like maybe you want to do you know word embeddings maybe it's like a map of you know the dimension is is you know it's the actual values is is a vector is actually array of values which is like some word embedding or something like that so you could write aggregators to do that as well right and you know we but we might we may or may not be doing stuff like that as well right without giving away any just to keep going on slack here Ralph Mahajan has a question can X units carry some measure of salience on their components in other words not every row negative data is important for a specific intent that's um yeah so so I think I think that's the point like with the filter rules if I understand this question like so that you don't necessarily generate all the X units except for the ones you're interested in hmm right so so so if that's if you don't think it's important yeah if you know it doesn't really make sense that you know I'm combining this you know one come nature I might not necessarily want to count all the people you know by industry or by geo or something like that for some reason then then we have we have ways to turn that on and off depending upon your use case right and then and then and then one one case you know like you might want to generate the global aggregate and we have a special axiom it called /g because maybe you want to normalize against your entire data set you want to know how how big a size you're a greedy you know you counted 300 of Germany but you're only counting 500 people so you want to normalize by the global accident anyone else so this is not you know unfortunately this is not really a open source right now maybe if there was you know people really interested me I could twist some arms and get Eric and people at demand base to open source this this technology and it you know and other people could use it if people are interested you know maybe that could happen so I have a question so this X unit does like how similar is it to a query planner it's it's not all all it is is a string so so so it doesn't tell you how to cook you so every the idea is that if you're creating you know this is another maybe philosophical point and I'll to mention so so you're creating ideas you're precomputing everything you're precomputing aggregates ahead of time rather than going back and doing queries so that's also sort of the philosophical things like so normally you'd have a database and even you set it up and then you you know put some human in front of it and they they type in you know queries and and go back and forth and and do sort of sort of interactive ad hoc query and that's that's how the analysts you know gets information and this is sort of the antithesis right so this this the idea is not instead of doing business intelligence where a human is going back and forth and you're doing query plan you just take you just pre calculate everything you're creating aggregate ahead of time creating all the aggregates that wants and then taking that and then dumping it into some data science algorithm you know don't don't involve a human don't involve the the intelligence or the intuition of some some you know business executive or stuff like that just push it into you know pre capturing everything so it's it's not a query plan it's not like you're gonna go back and then do it later and and then and they're not you know another use case which i think is actually you know once you create these X units and the aggregates normally what you would do is stand it up in like a key value store like you put in an NH base or BigTable and that's where you do like scans over time ranges and then you can have fan you know charts and graphs you know and that's that's where you have the necessary analytics poor and when you're looking at you know fancy grass and everything like that and then you could do the Bisbee I and you select date ranges and aggregate the time series so so we built systems like that at like like like tagged so so it's really designed for key value stores okay thank you and along that line Jerome there's another question on slack when you were started talking about cubism and feature engineering you made a comment about how this is a bit controversial and someone wants to know more about what you meant by that you don't agree you're just that sir you want to start a fight no no I mean it's I mean it's not necessarily controversial I don't think it's I don't I don't think you know anybody would really disagree but it's it's not the way a lot of you know organizations a lot of people sort of approach problems right that's all I'm saying right so so you know like a lot of a lot of people just sort of have you know they create they create analytics and then they place an analyst in front of a door the create systems to do ad hoc analysis you know rather than sort of seeing the problem as Dana sighs algorithms right it's probably not controversial to this audience but it might be controversial to to other other groups right and if you agree you know you know that that's great you know see if you know I'm not I'm not you know I'm not trying to you know you know start a fight or anything like that okay see it so I just want to make a quick comment before we turn it over to our last speaker if Jerome and Leo and I are from the man base as you can see from the slides if you want to learn more about demand base about engineering about machine learning about data science please reach out to me for them to continue this conversation on slack or on LinkedIn or wherever we are always you know looking for smart people and happy to continue this conversation if anyone wants to learn one another Eric can I give you the link to the slides and you'll post them somewhere or yeah why don't we discuss that with Salar afterwards we have the video recording and then links to the slides that we can definitely share afterwards yeah great yeah we'll take we'll discuss it after the Meetup thank you great well if there's no more questions for Jerome over to Ryan to introduce our last speaker hi guys so our last talk is not actually actually Haskell specific but related to Nick's so our speaker John will be talking about how to use Nick's in the next package manager in order to set up a variety of different environments for development including Haskell Scala and a few others so take it away John Thank You Jerome I'll need you to lift the screen sharing there you go all right so do does everyone here see my Emacs window yeah all right so my name I'm John Wiggly I'm a Suzy a stick about many different open source technologies among them especially Nick's and Haskell and just recently getting into rust and we use hat we use Nick's every day at Definity the company I work for but I've also been using it personally for about six years now thanks to Ryan trinkle one day I was at NYC hack staying at his apartment one night and he was showing me this laptop he had set up where he was running Nick's OS and he just started describing some of the basic features that Nick's had that I'm gonna cover and after I saw his little present I was converted right then and there on the spot and I switched my machine over and haven't looked back and I've really liked seeing how the community has developed over those six years it was a much different experience jumping in into NYX then as it is now there are way more people waking up to what NYX can do for them nowadays then then and the package system for NYX is huge back then I had to run homebrew on the side for the recipes that weren't yet indexed and I would use homebrew as a source to crib how things should be built but nowadays there's never been anything I wanted to build that was if it wasn't a NYX it usually wasn't anywhere else either because it was that new I do have some repositories that are adjacent to this talk if you want to see my NYX configuration in all of its detail there is a repo of mine on github called next - config the this file by the way is in one of these repos so if you want to see this data you can easily for this talk I made a repo on github called hello and it contains material specifically designed for you to be able to cargo cult it just copy it and start running and not have to know a lot about Nix just to put it to use for you and then of course you'll see me do a lot of things in Emacs relating to nicks and relating to Haskell that you will find in my dot - Emacs repository so why what is it what is nicks to get that out of the way so Nix is Nix is actually a few different things Nix is first and foremost the next language that I'll cover a little bit it is also a data store manager so if you have nicks installed there will be a directory on your machine slash mix and under that will be a very quickly growing pool of data and the next tool is used to manage the contents of that store and there are nice tie-ins with the language to make it easy to reference that store Nick's is also a build tool driver so although Nix is a packager it really doesn't know anything about how to build the packages that it's constructing so it defers to the build system of whatever language or whatever tool is being built so if I pull down fetch mail with Nix it's going to use automate to do the building it also knows how to talk to see make or to cargo or to cabal and as time goes by it's learning more and more and more of the build systems that are out there so Nix is not trying to solve that part of the problem it's more just giving a you a declarative way to specify what the build instruction should be at the top level so if you already have a project set up to use auto make or to you see make usually your Nix recipe is going to be incredibly small maybe 10 lines because Nix is just gonna look for certain files in the build tree that it recognizes as relating to build systems it knows about and if it finds auto make files it kind of knows what to do it knows whether it needs to generate the make file dot in or whether it just needs to run Auto Kampf using the existing make file dot in then it knows to call make then it knows to invoke it using make check and make install etc so those are all Drive build builders that are part of the next packages repository and then finally Nix is a very very large ecosystem of package definitions called Nix packages Nix packages right now has I think about 1.2 million lines of Nix code in it and covers tens of thousands of packages a lot of ecosystems or language environments that have their own Suites of libraries Nix finds ways to import those whole ecosystems into Nix usually by automated tools so it's not replicating all the package definition work that's been done by other maintainer x' so hack is Haskell packages for example these days is imported almost directly in from stack egde and it's definition of packages there's another one for Ruby for Perl for Python for nodejs all of these different systems also Emacs they get imported in so that 1.2 million lines doesn't even always include all of that information so it's a huge it's a huge tree of details about how to build things and where to go to get the sources for them so why Nix what was it that Ryan showed me that won me over so much if you haven't experienced Nix yet or read much about it just some of the things he showed me that I was kind of blown away by and had always wanted the first was atomic builds so one thing that I had noticed happened to me at least 7 or 8 in homebrew over the time that I had been using homebrew and homebrew was a really excellent system it was the best I had used up until then is that I would go to do a homegroup upgrade to upgrade all the packages in my system it would start chugging along and it was for example upgrade say Zeeland and it would succeed and it would install the new zealand and then it would go on to install something else and that would break and because it broke it never got to building everything so my max was pointing to the old Z Lib that had been there before but not to now to the new zealand because it didn't get rebuilt and then i would be happily coding because i don't exit Emacs usually during the day and the next day in the morning i would go to run Emacs and it just wouldn't run at all and i would find out oh it's referring to a shared lobby library that wasn't there so now I have to go rebuild I have to figure out what happened how did my system break and how do I get myself back into a working state quickly so I can get back to doing my work with NYX if I ask NYX to upgrade the system there are times depending on what they've changed in the NYX packages repository that it may update 8,000 packages you know sometimes they'll change something very fundamental like a version of GCC and that just wants to sweep through and down and rebuild rebuild the world well either all of those builds succeed or they don't if they don't nothing changes from my user environments point of view the intermediaries may have gotten installed and they're now in the datastore that NYX is managing but they're not part of my user environment furthermore NYX has this off this philosophy of making all references to products that are in the store absolute always so if I build a project and that project creates a binary that refers to shared libraries part of the fix up phase of the next build will be to rewrite all those shared library references into absolute paths so that even if I do say instead of an upgrading everything and relying on that to be economic what if I just upgrade zeal in my Emacs will continue to refer to and use the Z Lib that it was built with and as so there's kind of this idea that if you build something a mix and you on your machine and it works it will always work there are certain ways you can violate that premise if you garbage collect for example but if you don't garbage collect over the course of a week then everything you have that's that's functioning will stay functioning it won't be broken by updating the system NYX also supports atomic rollbacks so if I build everything and that part worked but I go to run and they've updated some feature and now all of a sudden my my Emacs doesn't work like today I got a notification from Forge that file there was no final notification facility available and I don't want to have to debug what that problem is now I can just say Nix rollback and it will take me to the exact state that my user environment was in the last time an update succeeded assuming of course I have not garbage collected away past environments if you but garbage collection also is always manual unlike with yet so if I don't actually run Nix garbage collect it will leave all of my past environments past environments alone there's also deduplication without having to do anything to enable it since all of the contents in the next store get assigned to hash when they get put there and then that hash is derived from all the sources and dependencies that were used to build that thing those dependencies if another package needs them it will recognize that they're already in the store and it will use them so if two different things use e lib they will both point at the same Z live there may be multiple versions of Z live in my store but there's usually only going to be one Z lib like whatever 103 whatever is the latest and anyone using that version will refer to that Zeeland this includes if I do local project bills so my use of mix is generally to keep my user environment relatively free of development tools I just have systems tools there editors such like but in my project directories each project has its own mix definition that pulls in the version of Haskell I want to use for that project the version of rust the version of its libraries all of those projects share the same binaries pool in the next store so if rust 131 pulls in grep and and rust one for one pulls in grep and they pulling in the same version of Gref they will use the same version of breath another feature of Nix is transparent caching so in some systems I've heard Gen 2 is this way although I've never used it myself you do a lot of compiling you're waiting for a lot of stuff to build and in Nix if you started out with Nix and you had no access to a NYX cache like the default one that NYX itself the next project itself give gives you building your system would involve building everything it knows how to build everything it knows how to run all the compilers and you can tell NYX by a flag don't use caches in which case it will do this and it will take forever and if you track the master of Nyx packages and you're tracking it faster than the cache can populate you'll find yourself in this scenario a bit taking hours and hours and hours to rebuild the system but once something has built because it is hash addressed if it exists in a cache that you have a pointer to you will download that material from the cache rather than building it yourself so the NYX project provides a cache my company has its own cache there are my my in my house I have one machine that acts as the cache to all my laptops and there are public services like cache shakes that give that make caches available to open-source projects for free and that's usually how I get my Travis build times down for example is I use Kasich's to cache everything for the holo repository otherwise it builds so much stuff that it exceeds the time limit and Travis always says the builds are failing then there's reproducibility and since Nix is all based on a functional language it has this idea that the same inputs should produce the same outputs so the same source at the same version with the same dependencies should produce the same binaries that means that if I have a project definition like I have in the hello repository and you out there in the internet decide to download you decide to get clone the hello repository and you have Nix installed you go in there and you type next build I know what you're using I know what your environment is I know the version of every single library that you're going to be pulling down so if you have problems I can help you because I can reproduce your problems here and that way we rely upon this at work so that all the developers that we have spread across four different time zones were never we're never asking each other well what version are you using what version of the compiler you have we always know what everyone is using and that's Nyx's reproducibility the other nice feature that kind of gives you for free is distributed builds I have three machines in my house they are all running Nix they all want to build but my laptop my tiny little 13-inch laptop gets pretty hot if it starts needing to compile a lot of things so by adding one line to my Nix Darwin configuration I can say that my big eye Mac Pro is a builder and what will happen then is when the little laptop wants to build something for Nix he'll use SSH to go talk to the Builder and have the Nix build happen over there and then he'll SCP down the results that should go into the next door without copying down the develop for the development only products that need to go into the next door so when I type Nix build Nix will download the compiler it'll download the tarball it'll download all the dependencies needed for building it but when you copy a derivation that you've built to some other machine you're only copying what that derivation needs to execute for to run so that's a very handy way for me to save on battery life and to keep my laptop cool is that when I go out around to cafes I'm using my machine at home as my transparent distributed builder I used to do this back in my C++ days using this CC and loved it then but this CC was you was something you had to integrate into your build system whereas Nix distributed builds will work for any project it is completely orthogonal to what it is you're building the bill just simply happens on another machine or even a fleet of different machines in your distributed builds list you can list a whole bunch of machines that your computer should rely upon before doing builds on itself sandboxing is another feature that works much better on Linux that it does on Darwin there are certain key system libraries on Darwin that you need we need to reference buy pads outside of the next store but unlit exceed Enix build you will automatically be in a sandbox that does not let the build access the Internet that does not let the the build access any files outside of a true to jail in which the build happens the build only has access to the dependencies that it needs to do building it cannot see or be polluted by anything else and this is important to ensure reproducibility we want to know that all the tests for that project your project will run and pass in a sandbox so sandbox is on by default for Linux you can turn it off likewise you can turn it on for Darwin and it'll work most of the time but sometimes when a key core dependency needs to be rebuilt you may find yourself building things that won't work with sandboxing on but will work with it off and lastly is environment I said that when I install something with nicks or I update my machine in its updating my environment that's because by default you get a user environment that all your end results are similar in tubes that if I type say fine it will get it out of the new fine utils in my NIC store because fine will be first on the path and the path is pointing to this user environment a user environment is really just a treat of symlinks into the next store there's nothing stopping me from having multiple user environments I could have a user environment named work another one you named gains another one named Scala and that would give me a global environment that I could just jump into if I wanted to do certain things I can also have project level environments I can even create Nick's defined environments that will produce as their result a shell script called load env - name that whenever I run it drops me into a shell with everything that I chose to be installed in that environment so Nix is very very flexible about scoped localized environments to whatever degree you want and these are these are exportable if I create an environment function for me to jump into like I have one for I have one for working on one of my projects with Python - and another for working on it with Python 3 because I don't want the project itself to define one or the other and I don't want to have to manually switch between them so I just run a shell script before I start doing mix in that project well I could send that definition to somebody else out on the web and say hey if you want to try this out in the Python 3 dev environment just run the shell script build the NIC's environments and then run the shell script and you'll be in the environment I'm expecting another neat nixb a service that I found recently is called Nick Suri which is just a public service that lets you run docker and be put into a machine a container running any any defined next thing from the next packages set so if you want to just really quickly try out Garrett as built by Nix for example you can find yourself in a docker container that way now in talking about Nix it's easy to go into the weeds so I usually identify three different categories of users of Nix one is just users and I imagine many of you are gonna fall into this category as most of the people at my company do you're you're just somebody who wants to use Nix to get stuff done you really don't care about Nix you don't want to read got Nick's files you don't want to learn anything about Nix beyond what it takes to get to where you actually do work so for a Nix user there are very few commands that you typically end up using the most common two are gonna be either Nix build which builds the whole project down to its final output binary or Nick shell that will put you into a shell where all of the dependencies needed to do building with whatever the build tool is will be in scope so if you have cloned my hello repository and you CD into the Haskell directory and you type Nix now she'll well be prepared to wait for quite a while unless you set up using my my Kasich's cash for this repository it'll probably take I would imagine about 40 minutes for it to build GHC all the libraries all the dependencies it's going to build it for both with profiling without it'll make the Google database but it will in the end get you into a show where you can type G HCI and you'll be in the version that I used for that hello project that's what I do most of is Nick shell then there's Nick env what you use for manipulating your user environment so if you wanted to install he mucks for your user account then you would type Nixon V - IP max - just install it for your user and then finally Nick search Nick search lets us do like say is there something a Knicks packages named fetch mail and then it will give me it'll tell me whether it found fetch mail in the Knicks packages database and if so a the blurb that was defined along with fetch mail in that database and then for power users if you want to get a little deeper into it you'll find that Nick shell sometimes has a bit of a pause or sometimes reaches out to the internet when you didn't want it to I'm a big fan of coding when I don't having access to Internet so at cafes or on airplanes when I'm traveling I use dirty envy and dirty Envy for me is the solution to almost all Nicks interaction problems I have dirty NB and some rapper scripts that I've included in the hello repository what it does is it basically runs a nick shell calls env inside the Nick shell to see what that environment looks like exits that Nick shell and stuffs that into the durian be definition so now when I CD into that directory I don't have to go into Nick shell dirty and B will just change my environment to be what the environment would have been if I had run Nick shell and since that's always being based on just the contents of a cache file there will never be any rebuilding the worst case scenario is either I am NOT up to date and my dirty env cache needs to be updated or I foolishly garbage-collected before going on a trip and my dirty NB cache relates to things that are no longer present I updated my script to create GC roots to prevent my dirty and B contents from being garbage collected but I haven't yet tested it in anger so I won't about say for it just yet then in conjunction with dirty env the real magic is using a project called Emacs dirty Andy and this lets you hook Emacs into dirty and V so that as I change buffers if I'm in a buffer that's in one project and I change to a file that's in another project dirty env will or Emacs through env will change my global Emacs environments environment variables to be as if I had started Emacs from that other project and I even have it tied into e shell here so I can say CD source hello if I type CD that's cool you'll see in my mode line down here it's telling me all the changes it made to the environment going into Haskell and now I can say GHC - - version and it shows me that I am look I'm at version 8a - well if I go out and I take GHC - - version I don't have GHC available anymore or I can go into rust still know GHC but now I can say rusty what's your version and so this works for both being an e shell and C being around you'll need to look at my dot Emacs file to see how I have that hooked up or it also works for just editing files so if I go into the Haskell name dot HS file I will again see a little notification down here at the bottom it's a little noisy but I like it because it it reminds me that something's happening and now when I'm in here I can do medibang because as I said it changed my global Emacs environment and I can say J HD version and now I'm again at 882 and if I am in this email if I'm in this Haskell buffer and I say something like misspell and identify right I get instant syntax checking on that using the appropriate version tools back - so that's that's env and Emacs 3nv for for power users if you want a little more automation there's a project out there called Laurie and Laurie is essentially doing the same thing as Derian be with a bit more automation so it will run as a background daemon process and it will keep your during NB cache up-to-date as files change that might affect it where as I kind of go the other route I want all cache updates to be absolutely manual so that I'm never pausing at an opportune moment like say giving a demo um the third cat was the category of users the second category of Nick's people is maintainer x' so these are people who are also mixed users but they need to know enough to edit the dot Nick's files in their projects because they're going to be setting up build structure for their users but they're not really necessarily Knicks heads and they don't want to go into Nick's packages they don't want to really get deep into the next language but you kind of need to know enough about next to be dangerous and this is mainly gonna be done based on cargo culty and copying recipes from the web but I will say one of Nick's downsides right now I would say is discoverability there are a lot of things that in Nix it's it's common to want to know but hard to know how to ask it and harder even still to coax Google into giving you the answer that you're looking for the best resource in this sense is to go to the hash nicks OS channel on freenode probably the best place to ask live questions about nicks feel free to also reach out to me if you want so Nick's doesn't make finding these things really all that great and the documentation could be a little better but again it's unproven all the time using Nick's a few years ago is it a lot like you didn't get in the beginning gets gotten a ton better at documentation and I think Nick's will catch up with it as well so if you're a maintainer of Nick's code as opposed to someone just using it there are a few more commands that you're going to use a lot of one really really handy one is Nick's edit so if I'm in a situation or a scenario where I know the attribute that I am able to build like that Nick's package is fetch mail that I saw be running Nick's search fetch mail it told me that the attribute was Nick's packages dot fetch mail if I say Nick's edit Nick's packages dot fetch mail using that same attribute I will be jumped into the file that defines that attribute and now I can see exactly how a fetch mail gets built which because it's an auto make project there's no instructions on here and how to build it because Nick's can infer how to build it all but this file defines is what is the version of betch met fetch mail where are we going to get those sources from what's the hash on those sources which will prevent injection attack possibilities what are the dependencies that Nick's needs to make sure are in the are available in order to build bechamel and it makes these dependencies available to the met fetch Mel build within the sandbox what configuration flags might be needed in addition to what Auto make is going to discover and then some metadata that defines this project within the next packages repository so that's Nick's edit very very handy tool there's also Nick's log which I can as well use here let me release this this will let me see what was the what was the build log for when fetch mail built the version that I am looking at in my current nick store and this is available until it gets garbage collected so even if I roll back to a previous user environment I can see the logs for the build that I'm using in that user environment Nick's ping store is if you are using a other store to cache binary products for example at work you can use Nick's ping store to see if you could have access to it there's a very handy way to debug connection problems Nick's locate is a something you can separately install it gives you a command called or sorry the package is called Nick's index it gives you a command called Nick's locate and with Nick's locate I can say you know what I don't know which package provides bin sqlite3 but can you tell me which ones would build that and then it goes through and it finds all of the output target paths for the different Nick's packages that are known and will tell me which ones which ones will create an output path in my user environment called bin sqlite3 takes a while for Nick's index to build that information but once you have it it's pretty invaluable Nick's info just gives you information about your current Nick's machine always great if you have users you're trying to service you need to know what version of Nick's they're running and what system they're on and finally Nick's mash completions or Nick to zsh completions all very very useful for remembering all the various options the third class of Nick's users who I'm not really going to address in this talk today are developers so these are people who are going to do things like pop into the raffle to tryout expression evaluations use Nick's email to see why something their writing is getting a recursion failure asking nicks you know there's a runtime dependency of fetch mail and open SSL I can say Nick's wide depends let's see Nick's packages that Fletch mail on Nick's packages open SSL and it'll tell me see here it's telling me it doesn't depend upon it I saw that working before I don't know which invocation oh it sorry it's it's it's it's not out got out because right and so it tells me the reason why why OpenSSL is being dependent upon is because the Dutch male binary is referencing it it's referencing it in the in the the shared object table mixed copy as I said is copying bill products from one store to another and the Knicks diff is a tool a tool I'm not going to go into the I'm running short on time Knicks language basics Knicks is a pure functional language by pure I mean outputs follow through inputs I don't mean that it's free of side effects because some things you do in Knicks may change the contents of the next data store however when you run the function you'll always get the same output from the same input even if the store is changed on the way to getting you that output and it's functional because functions are first class they can be passed around as values and if you write Knicks code you're gonna end up creating lots and lots of functions and lambda abstractions and you're gonna be passing around functions and calling functions at its core NYX is just an untyped lambda calculus for Haskell people familiar with that so you've got your typical constructors of variables lambda abstractions and function application then we add to that a bunch of basic data types these are the seven essential data types that are used in Nick's the only really special ones are paths and attribute sets and Nick's has a lot of special functionality both in the language and in the library for dealing with adds and attribute sets paths are things that live in the store and attribute sets are key Val heterogeneous key value mappings which have a bunch of sugar and a lot of special meaning in different contexts then Nick's adds on top of those basic foundations some syntactic sugar so when I define a function I just say Arg colon and then I say I'm returning a key value map here and I'm just saying key equals value but if the key and the value have the same name I could just use the keyword inherit I can have default arguments on my function arguments I can take and capture the whole argument set as a key value map itself I can say dot dot dot because I don't care about arguments you extra arguments you pass me and then inherit can take in the argument which is the key value map to pull the named item out that's just a very brief example but among those three things that covers actually a lot of the syntax of Nyx most of the richness of Nicks comes when you start nesting in function definitions and you take advantage of the fact that it's a lazily evaluated language that allows a lot of magical things to happen and then Nick's also I won't cover that has a ton of built-ins us I think it has about 120 to 140 built-in functions to do basic things and even some fancy things like fetching from git and then on top of that it has a standard library defined as part of the next packages definition which defines a ton of library function so with that let me go into the hello project so the hello project is something I wrote for this talk on github and what it does is it defines a very very quick getting started definition of projects for a bunch of languages so I have C++ Haskell Scala Ruby rust a lot of these were cribbed from existing packages I'm working on so if you see any strange looking artifacts in there that may be where it evolved from I'll go into the Haskell definition first by going into that directory you can see at the bottom dirty env put me into the right environment looking at the NYX file NYX files in your project you will typically have one NYX file named default Nicks that's what NYX build and Nick shell will both look for by default that file contains a function that takes several arguments all of which are defaulted because I don't want my users they have to explicitly specify them usually I like to define I like to give the user the ability to vary which version of the compiler they're using so that they can just test it out and try the build on an older version I also for especially for Haskell I offer flags for turning on benchmarking profiling of this of the project itself and also strict which means all warnings or errors then there's this bit which is I define a revision of the nicks packages repository itself and you'll see this pattern recur through all the default mixes of all these and a Shah for that specific package set and that's because I want to pin the project to a very specific tree of dependencies so that I'm not having to deal with cross dependency bugs on github then pkgs what it does is it fetch it uses a built-in to fetch the tarball from github check that that tar ball matches the SHA and then alters the nicks packages definition to allow unfree packages if they're necessary but not allow broken packages unless that's necessary I have the line here just so that I can easily toggle that and then I allow overlays which in this case I'm doing a very fancy overlay to make sure that everything builds a Google database and then the remainder of this is something you can just copy not all of it is explicable in the context of time we have remaining but basically Haskell packages that develop package is a very very fancy function provided by the next packages repository that is very smart about Haskell projects it knows how to read your COBOL file to figure out what your Haskell dependencies are and turn that into a list of Nix dependencies on those Haskell libraries it also knows how to take an H PAC package ml file and auto-generate it to a COBOL file to get its dependency information it allows you to selectively override the build tools or what tools you need in scope when running tests as opposed to anything else and a pass through here that is used for magical reasons I have actually forgotten by now but this little stent this little snippet right here gets copied from Haskell project - Haskell project across my machine I must have maybe 20 instances or variations of this little snippet right here and then I'll go show you the rest one the rest one's not hugely different but instead of calling dot Haskell packages develop package it calls rust platform dot build rust package which has a lot of its own magic for figuring out from the cargo lock file what all of your rust dependencies are and what version of rust it should be using and then I do some I do I do some various things in here and I also included a test make sure that the executable when it runs in ever one of these sub-directories does what I expected it to do so now if I go into this what I end up with is by having my whole environment I can do things like ask Google to give me the definition the documentation for put string and I know that this documentation will be for the version a put string that I'm actually calling here mmm I could talk more about Haskell mode but I'm sure that's a whole discussion in itself if I go over to rusts there env will switch me over to the rust project and now I could ask for definition on the on the UTC now function for example or on the UTC structure and if I go over here and I do something that doesn't make that makes rust unhappy I will get immediate syntax highlighting using the correct version of RLS so it's it's a pretty seamless experience once you've set it up and you've got the right environment set in your Emacs or in your durian be you could be just using Nik shell and running GHC ID or you could be just running gh CI and doing a colon load every time you change your sources I know several good developers that that's their their mode but the nice thing is I just CD and that's my interface to Nix 99% of the time I only have to dive into dot nix files generally when I update stuff and and the very last thing I'll show is Scala simply because you guys there's a lot of scholars here I have no idea how to build Scala so all I did I look up the hello world example for Scala this is a plain Jane mix derivation so it's neat there's no Scala packages not build scarlet if there is there could be I don't know this is just make derivation which says how do I build something standard well I pull in Scala I run Scala C and then I just run Java on the resulting class file with my class path set up to what it needs to be to run and I have a I have a check phase here that makes sure that the output is exactly HelloWorld as I expect to see it this org file that I've been reading from is at the root of this hello repository so anything I mentioned you can find in there and sorry to not leave much time for questions but if anyone has questions I will take them now I can relays people's questions one of the questions that people had was can Knicks work with stack or is it as a replacement for stack they do a lot of the same jobs they can be used compatibly Knicks is typically orthogonal to anything else so you can use homebrew while you use Knicks they won't interfere with each other Knicks only cares about the contents of / Knicks / store and manages those contents it doesn't touch anything else so if you're using stack that can do its own downloads pull its own stuff but then when it wants to say get Z lib or use some external system tool it could be getting it from Knicks I haven't tried that combination myself good question for you John you mentioned about how you become a maintainer on the Knicks repo how do you move from being a maintainer to actually having a permission to merge with the next group is there a procedure or you know I actually don't know when I started contributing to Knicks I was one of the I was one of the few people really caring about Darwin support so by like my 40th pull request somebody said hey why don't we just give this guy commit rights and save our save ourselves some time and that's when it happened there was no official nothing formal ok ok so I just got a lot of PRN yeah I think there's a Knicks organization now and I would contact them ok cool cool hey John there was another question does Nick support proxy servers in different Network environments like a home environment versus a work environment we have had pretty good success at work for example we use octa in duo to give people access to our servers and things have been working fairly well I mean what NYX wants is it wants to use HTTP it wants to use git it wants to use SSH as long as you can get that access to Nix Nix doesn't really care about much else in term of network connectivity then I guess this is a related question but Kim Nick support caches or repose in a clean environment they're disconnected from the Internet absolutely yes you can run your own you can run your own Nick's caching server at any time like you can take any one developer's machine and you could just pre populate your cache on a given day from what was on their machine and then from that day forward disable your access to the to the global NICs OS community provided cache and only point your machine at that internal cache you set up also when you update the cache Nix Nix copy that command I showed you is smart enough to only copy the closure of what you're uploading and only the part of the closure that it doesn't already have the parts of so if I build something that's a tree of a thousand build products and I put it in my cache yeah that copy may take a while I did a copy to my laptop the other day that copied 90 gigabytes then I do a rebuild I changed some dependency and 20 things in that fill tree change only those 20 go over so it's a incremental update of the cache each time you do a build so if you and it doesn't need any special build system either you could have Jenkins as its last build step as a post build hook just run Nick's copy as long as it has SSH access any more questions so all right I ask a question and the Eric kinda answered and you kind of answered you new with Scala how's it compared to like some tool like SBT or Gradle like what would be the advantage of using Nick's over over those tools depends on who you ask I've heard people the Mon the use of Nix at my company and want to use basil or use Gradle there is I think there's competing mindshare in that space I prefer Nix only because the breadth of exposure it has in the open-source community and the amount of popularity it's gaining and how how flexible it is for all the different uses I want to put it too but I could see those other teams satisfying corporate needs just as well so I'm not gonna say Nyx's the be-all end-all but it is a growing system with a large community of volunteers and it's a fun project to be a part of so that's that's a side advantage but technically there are there there are comparable app aspects to those projects okay there's another question about basically what this guy sis disk size would you recommend if you're gonna do Haskell development with Nix and that might also extend to kind of just any other kind of mix how the next door grows that is a great question if you are gonna keep to one version of GHC you could get away with a nix store size of about 100 gigabytes pretty easily and then garbage collect on the weekends I have a 2 terabyte drive I usually collect my NIC store when it gets to be about 3 to 400 gigabytes in size it takes months to get that big so we're not talking about using up terabytes of disk space in any short order of time I also have a ridiculous number of packages and different versions of different compilers installed as well as their entire dependency trees so it's almost unfair what I'm putting my disk through I think the typical mix user is going to have store contents on the range of 15 to 15 gigabytes most of the time and Nix wants an SSD let me tell you this there are so many inodes in that Nix store that the AI ops are what you want to optimize for Nix does not allow itself to be placed on an external device unfortunately the slash next directory has to be well-lit Linux users can do it you can do it with bind mounts but Darwin users can't we do it we do it a different way in Catalina with a shadowed volume and a PFS volume but typically mixes the Nick store is going to be located on whatever your boot SSD is I just wanted to comment on Knicks versus SBT and Riedl and my impression correct me but my impression is that Knicks is very holistic in its approach you have to go all in and all of your dependencies and all other packages have to be in mix and SPT and Gradle just build what you want you built one package at a time and it doesn't install it it doesn't check other things it doesn't make your development environment it just runs the build commands and and if it's no if some dependencies are not there too bad it will fail it will not fit everything and we compile your kernel before you do any any kind of build but so NYX is more for kind of an approach where a lot of dependencies are cross-platform difficult to maintain you have a dependency on GCC and you have a dependency on Lib C and you have a dependency on the rust Haskell and so on learn SPT and Gradle would wouldn't be able to handle that that's my impression so please let me know that's right um it depends on how you set it up because for example there's nothing stopping you from having a nix derivation defined that uses Gradle so you can draw a dividing line where you want to as to how much should be managed by NYX because maybe you want next to manager Colonel I mean that's what Nick so s is all about but maybe you don't want Nick's to manage the build of your project then don't use an X for your project just tell Nick stew use whatever you do want to use for your project one note of old Scala project that does use Nix for local development there's scholar native which is probably not a bad idea because it does record some extra OS specific dependencies so by there by having a mixed shell defined in the repo provides it all for you I thought I made sure that just is anyone from there as a scholar side wants to look at an example project using it and I think they're just using it as said as before like John said setting up the environment not necessarily as are made in replacement yeah if you build Java projects Nix we use maven it does not try to reproduce that yeah hey you know what's the Cross building story between Scala versions that's one thing that SBT does well as opposed to others I don't know it's not about Scala to know exactly what that question entails I would I would just say you need kind of a variable or or parameter in Nix that defines your style versions and then you just run your build function and derivation with different arguments in the you know dot map and we map over a sequence of scalar versions to some point but old dependencies you need to okay I guess okay I don't know if you saw it in my in the hello project I do not share I have a default Nix at the top level and what this does is it if you were to type Nix build at the talk of the hello project instead of in one of the language of directories it would build all of the language of directories in all of the versions that are currently available so it would build using Scala two ten eleven twelve and thirteen and it would make sure that the build works and passes tests for every one of those versions so so basically if you need cross version you just have an array of versions and you go over that array and run this whole package they're done yep so it's a very easy thing to do everything is parent tries you know everything is alarmed that so you call that lambda on the arguments you need and you know all the dependencies are guaranteed but one thing that I understand that Nix is different from pretty much every other tool is that it is very meticulous about dependencies and so you you can make it depend on everything up to the last config file if you want to and it will it will work so it's kind of you know when you do when you do a or Denari build tool like SBT in Gradle that is nowhere near that degree of reliability and reproducibility so a lot of things have to be installed pre-installed and nobody knows where they are this is not the case with with mix if you want it can manage your entire computer so it's a it's kind of another real really correct comparison you know what is mixed versus Gradle the Gradle is a glorified make file and SBT is a glorified make file and so you don't necessarily need nix if befall you want us to build some things in a fixed environment but it's a very good example in my view but you have Scala native that depends on so many other things other than Java other than spell and SP T wouldn't be able to handle that you know Scala native depends on LLVM it depends on maybe some other things so scale native really needs this kind of cross-platform compatibility and dependency handling which most stellar projects don't need so that's anyway that's what I'm that's what I think pretty cool thanks the answer I also use next to manage all of the Emacs packages that I have installed and I load into my Emacs right because I I never got into using any of the Emacs package managers that are out there so I used to just it was manual I would download them and and just reference them in a directory tree and one day I realized wait a minute I'm doing exactly what NYX is created to solve so I spent I spent a weekend and now I have it just brought so much sanity to my Emacs configuration so I have right now 422 Emacs packages that get compiled built installed and tests run and this is all being drawn from Melba so when I update at the end of the week I'm getting all the latest versions of all of these and seeing that they pass tests and it's an atomic upgrade and it's rollback it can be rolled back and that's just light-years beyond anything I ever had in the past for managing my Emacs dependency I have another one you know there is a brew for more like I mean as far as I understand and stuff installing for example some like epic UI application of stuff there is a brew cask install you can do the same right is what is the equivalent of for example instant IntelliJ for now when configuring telework job apps so are you talking Linux or Darwin Darwin so Darwin does not have the best native application support usually Java apps do install pretty well with Nix so I installed J disc report using Nix on on a Mac on Mac OS but if it's truly a native app meaning it needs to link in to the cocoa libraries that I have had really spotty success with that I think currently I have Wireshark is the only application I'm building with Nix usually for a native app I will just download the dmg and install it manually I would love not to have to do that but Nyx's darwin support is not quite as first-class as their support on linux fine it's evolving hey John there was a question about has Nix changed the way that you do testing as Nick's cheese the way I do testing not really because when I work on code I Nick shell and then and then from that point on the rest of the day I'm using either cargo tests or cobol test and and now I'm just doing what I would have done is a Haskell developer in my pre next days I really don't think about Nick's very much in my day-to-day work eight any more questions and I'll be on the slack channel for i'll stay in that slack for a couple days and i'll answer questions as they're asked there - awesome awesome thank you John really appreciate your insight the next as the next time I suffer it was very enjoyable bollocks since you talk so thank you thank you for the great talk great great great well thank you all for joining our first ever online meetup just so that we can have another one please contact myself Alexia Ryan if you have a subject you want to talk about our future online Meetup also feel free to continue using the slack if you want to use your slack which one is there is the slack we have to talk to non work colleagues feel free to do so it's not just limited to using it for this meetup we are hoping to have the video of tonight's available on a functional dot TV when that gets published we will announce it on the mean it relevant meetup ages I don't Twitter if you follow me Aleksey or Scala Alexia is there anything you want to add before we wrap up and also anything Ryan yeah I just wanted to reiterate that if anyone wants to give a talk on really any topic kind of even related to Haskell so as you saw John's talk was related to Nick's feel free to contact me or if you have any concepts as to other events that you'd like to see please let me know this and also once again a big thank you to demand base and everyone from there for hosting us with a zoom and all the help they did to make this event go well it's a fantastic job of being mute police amongst many others thank you enough yeah the man makes enough and their kids let us know how we can help them base if there's anything more you want to say my demand base please do so before we all log off yeah well first of all it's my pleasure and thanks to you guys for organizing a great Meetup and also if people know about other meetups you know I'm a fan of a lot of meetups that involve machine learning data science and out the reach out maybe we can host another one because you know we we pay for the enterprise version soon it went pretty well today and I'd love to keep doing this if I can so please feel free to reach out to me for other Minho's no thank you very good I would love to do that and consider we had a lot of content what was it what was impressive for first time running and online event wasn't at the peak of they you know we had about seven zero and these on Meetup and but by the time John finished it only went down to 55 which means I'd like that everyone enjoyed having their three talks and right right to the very end so that was that was that was that was breaks that cycle yeah so that's great populated I think we've got 55 people at slack and so you know I was worried a bit that people will not be able to login just like about the thing everybody here is really cause I think it testifies to the level of scholars husky people they can they can they can masters like for sure so everybody everybody was in so really really cool I think actually I wanted to ask you guys I feel free to you know raise your hand or is there anything the participant the audience the speaker's want to see improved can you give us any tips for the next time we do it anybody I think and feel free to comment on slack as well sorry god I think joining the slack was a little challenging for me just finding the URL but yeah research from Google made it okay this is in the description I couldn't find the link yeah maybe three yeah and let me just I could swear I put it there but I made a and I did send an email out to the mailing list with her on it as well it is on the very bottom of the description yeah but it works in the end okay so some eggs make slack URL more permanent what else was everybody able to ask the question should be should we allow voice question did the slack questions work okay did you guys feel like you were able to ask questions when you wanted to okay looked like everybody is in agreement you know in Russian whatever saying silence means agreement one thing to consider is that we didn't make use of the chat in zoom and so I mean is that something I think the problem with that is if you put links there or like tell everybody okay ask your questions in the chat like people who come in later they don't see the history of the chat right it's a little challenging and I think something like slack although I agree can be a little challenging to join a new workspace but you know something like that definitely works better than the than the chat here although maybe I should enable it for like just speaking to the host so if you have a specific problem about slang or something else I can address that but you know I really appreciated Eric I think it was a good call because you know I think it really wasn't clear how to use chat on Zoom because there was nothing so if people had to use slack so actually you know my initial idea was to kind of use it as a triage and tool but I think that was best I want to ask anybody did anybody try to use chat on Zoom to raise concerns or was it not necessary is there anybody here who kind of wanted to chat on Zoom itself maybe it's me again but I just can't find it like I have known I don't know maybe I'm getting old but I've been enabled when the meeting started but I went into zoom settings after it started and made it so that you could chat and I think because the meeting had already started it was too late it didn't like appear for people at that point yeah yeah I try to do the same I wanted to make sure that chad is managed properly I couldn't find it and when I found it final it's a chat dissent was just for things perfect so that's kind of which is great but yeah I we will make the joining the slack thing definitely more prominent and advantage slack has over that chat is that once this zoom is closed we won't lose the conversation people can continue talking for the whole next one week about the what was discussed here which is quite nice yeah one question I have I don't know as possible you know when you are in slack there is a notion of topic which is a message on top of the window so is there anything like ideally we would place and mess at another lay on zoom' go to slack for questions right like if you joined in the middle and didn't read anything like a lot of people do alright then basically an overlay would help to direct people to resources so I don't know if this is possible but I think you know that could solve all the problems if during like anybody else's presentation if the host can can put an overlay with basic URLs for this Meetup you know that would answer like a fake use for this Meetup right like if you came late and didn't read anything go to this FAQ if you experience any problems first and then you know raise your hand and do stop right like that'll solve a lot of problems I think I don't know if it's possible or not I'll look into that but also just one of us if it's our screen share it's just that like slack messes here join here then between speakers you can just go back to that right all right to make it more often did that that's actually a good point the slides I have done did have that on it so maybe between speakers we always go back to the our set of our holding slide right that would be a constant reminder right yeah I think that we yeah the other thing we could explore is that we could have all speakers use a template that we create for the Meetup and then that will always have it as the product footer right another thing we can do is you know if you guys listen to NPR you know if Terry Gross write fresh air so every time she goes to break in comes back said you know in case it's just join us we're talking to Barack Obama about his latest book right like basically we can keep keep pinging speakers to remind it do you really like but we'll just yeah we'll just have to ask folks to remind new listeners and viewers where to find resources and we can add in inserts twice at the beginning and end of talks right and like we can screen share the slides between the talks so we'll do that next basic basically we've come up with a new business idea new AI BOTS presume right exactly yeah interesting point too that means like you can app spreads so if someone post a question you can all swear wheezing the spread so the main channel stake claim that can be interesting to try to make people do that yeah yeah yeah making people do stuff as always an interesting undertaking yes cool all right so in case you guys think of something later right so please contact organizers on meetup we just click contact organizers it's creates a message which goes to all of us so if you can think of anything else we need to improve or if you want to give a talk right propose a topic now we're more flexible because we're basically can do this meetups as frequently as our families in current teen allows right especially the next week a school break so you know or two weeks or whatever so we can do more so please ping us with the topics then we'll we'll we'll do more of this we I think it turned out well so we can do between several of us we can do this actually more easy than easily than you yeah I next time I won't bring a glass of wine with you next time I'm going to bottle exactly you know my mind that's what's next yeah on that topic you had asked about hosting another meet up at the man base I was little hesitant because there is a cost associated with it but now that is yeah do you actually have some speakers in mind for something else I did definitely I would like us to have a meet-up every month so if you guys are happy to help us with their hosting this thing again with your zoom account I can work more more more actively on finding someone for April when we first spoke I did have a speaker for I did have a speaker for April but my the speaker I had pulled out because he was flying in from the east coast to here and because of the whole virus situation he canceled this flight but now that we can do it online I may be able to get a speaker even if they're not based locally to do a talk so I think if you guys are happy to help us again I'll definitely proactively we'll find a speaker for April and we can work on a date that suits you guys and works for the speaker yeah sure I mean happy - awesome awesome oh I'll jump on it straightaway tomorrow and as soon as I find a speaker I'll get some dates - I see what what we can agree on okay sounds good thank you so much Eric thank you so much you guys have been awesome thank you guys on this great note let's reconvene soon let us know if anything and thanks again everybody for the amazing first and lime it up Cheers stay safe thank you thanks everyone thank you all ladies and thank you everyone of demand-based really thank you thank you let's carry on chatting on slack leave goodbye