DBTB INT Alyssa Morrow r
Recording: DBTB INT Alyssa Morrow r
yeah I'm Melissa Mauro and I'm from university of california berkeley from the amp lab and i'm a graduate student researcher so my favorite thing about speaking here is a lot of the talks today um we're inherently technical but people really gave light into like the creativity in the thought process behind them there was a talk earlier at SQL and it was just a hilarious talk about the guy you know relating everything to transportation and very clever so it's really nice to meet people who do research and have talks that can really abstract the technical details to give really intuitive and inspiring our conversations so why is data so cool so that's extremely uh I guess general question I guess in particular I'm more interested in genomics so I I was a computer I used to do computer science and I thought it's a you know there's a lot of interesting problems but like how can we actually apply this technical knowledge to real problems and I think that's like really where data science comes into play so you have all this data sitting in a warehouse somewhere so the question is what what do you do with it how do you do it how do you deal with it and how you make it useful and so I think that problem is really exciting especially in the medical industry where that's an unsolved problem so there's a lot of really exciting new problems that we can try to explore in that area I think one of the main insights is I'm working on distributed visualization for genomics so I think the main insight is is that existing tools although people know and love them they're just not going to work for today's amount of data so if we have you know three petabytes of genomic data we can't really we can't really rely on current tools to support that amount of data especially for visualization so I understand that distributed visualization is a new concept and it might be a little scary because people who do use visualization tools Tech usually aren't the most technical people but I think what we're trying to what we're really trying to provide is like a core reliable set of API is that people can use even if they're not super technical sure so to become a data scientist I would definitely say take a computer science class you don't have to pay for it they're online they're free definitely keep up with current technologies there's a lot of really great you know computational infrastructure that's coming out every day so keeping up with that is extremely important one of the most important things i would say which people might disagree with is to not be a computer scientist you know if you come from the field of biology or chemistry or you know humanities you can get such a broader perspective on what data is and that can really be a key player and like how successful you can be as a data scientist so i definitely think that if you passionate about any data set on the technical skills can come through time but to have good understanding of the data you're dealing with is primary importance