Devreal

An Innovative Approach to Labeling Groun...

Event: Data by the Bay

data.bythebay.io: Jill Desmond, An Innovative Approach to Labeling Ground Truth in Speech

Recording: data.bythebay.io: Jill Desmond, An Innovative Approach to Labeling Ground Truth in Speech

thank you it's versity we're very with a difficult name to pronounce our company from so to give you a little bit of a background of what Bursa means working on and what our motivation is it's widely accepted now that the more parents talk to interact with their children the higher their children's IQ will be so we can see that sort of in this little image here that just represents that as a child ages into adulthood the brain's ability to change in response to experiences decreases and the effort to make such change increases so it's really vital that we that we make these changes early on there's an issue now pretty big in the media the the word gap so it's the idea that you can see here in this plot we have cumulative words and millions and the ages try out the child in months so from a higher interactive household we might expect this orange line here and then a lower interactive household at this blue line and so what this is illustrating is the difference between the amount of words that a child hears if they come from a household with more interaction versus if they come from a household with less interaction and we would see this 30 million word gap by the child by the time the child which is 48 months so as the child's younger what really matters is quantity of words but as the child ages quality really begins to matter the child's understanding more quantity we can just measure by counting words but how do we measure quality there's different ways of doing so one is vocabulary richness one is conversational turns so if I talk you talk I talk it's that back and forth another is the quantity of interrogatives because interrogatives questions they're more engaging than just talking to the child and also the emotional state of the speaker so it's more beneficial to you know to be happy to sort of be interactive versus don't do that stop come here so that can really affect the quality of interaction so the reason we're interested in machine learning is to be able to detect these metrics now to focus more on quality so to do this machine learning we really needed to start off by just labeling our audio so we needed ground truth so that we could train our machine learning algorithms so I don't know if anyone went to the talk actually earlier today on how to get good data but we've gone through a lot of these sort of steps that he was actually talking about starting with initial tagging methods we started out with one method of seeing if we could label our ground truth we took a pause we now we analyzed our consistency and we realized we needed to really step back and make sure we were getting the accuracy and the consistency in the labels that we needed so this step back we had to really figure out what we wanted from each label what we wanted an utterance to be even so that each person is defining a segment of speech the same way we also needed to really simplify how the person is doing the transcription so present it in a way that it's very black-and-white as to what we're asking of them so we developed a system we call Bertha tag for audio labeling and then we analyzed our new results to see if we were improving so initially we use software called clan which is used a lot in the speech processing world so in clan you have your audio signal shown on the bottom here and we're collecting audio from households so we have beta testers who volunteer to submit audio and these audio samples can be on the order of hours so you can imagine you need to scroll through hours of audio it's really not user friendly if you're gonna be sitting there labeling the audio so we have the transcribers to is highlight segments of speech and then they would manually enter the speaker what the speaker said and then whatever information we were interested in so they'd have to enter for interested in background noise or emotion or different things like this so it really was not optimized for speed accuracy or consistency so after doing a couple rounds of this that's when we stopped and we looked at how we were doing and we realized you really needed to make some changes so we for each piece of audio we had to transcribers label the audio because we wanted to make sure that we were getting this accuracy so on the Left we have two different plots since there's no truth aside from what our transcribers are writing we have one where we assumed that transcriber was one was the truth and then one where we assume that transcriber 2 was the truth and what we looked at here was when you're labeling speech how many Mis detections or false alarms were there so for instance on the left here the blue represents Mis detection so in that case label r1 would have said some segments of audio is speech label label or she would have said it's non speech and vice versa for the orange here so that's showing is where there misaligning and where there even saying that speech is occurring so even in addition to that there's sort of different ways you can have these kinds of errors one being that different transcriber is actually thought that segments links could be different so I'm up here I'm talking to you for about 20 minutes and someone might listen to that and say that is 20 minutes of speech and they can just write the whole thing if they want to on one extreme someone else could listen to it and say each thought is a piece of speech so I'm going to mark all of these thoughts as different segments so neither ones really wrong they're both marking speech they're both saying when I'm saying but the problem is we're not having that consistency so now if we give our algorithm these two pieces of information you know it's even though they're not wrong they're not if they're not conveying the information in the same way another source of error that we found was the way people figured out their end points so we could have people who are very generous with their end points and they are really capturing a lot of the nonspeech after the speech ends we could have people that are doing the opposite and cutting off the actual speech or what we hope for is this bottom plot we're really just getting the speech itself and this one it shows that I already went into a little bit where again we're not really accuracy is definitely one issue but consistency is sort of a bigger problem that we didn't even realize at the time where you could have 10 seconds of speech that one person says well there's one person talking so that is one segment or you could have another person say well I hear each of these different thoughts so each of these is a different segment and again nothing's really wrong but there are different ways of interpreting the data so we had it first take a big step back and just say one of the nutter ins what are we actually looking for here are we looking for 20 minutes of speech are we looking for each thought so we stopped and we will first had a look at our audio and our audio in itself is really difficult because we're in a household we're not doing read read audio where you know there's pretty good punctuation pretty much set thoughts there's chaos there's a dog barking there's the child there's interruptions and overlapping speakers so we really had to have a hard definition of what we're looking for so it started with defining and utterance and so we really wanted to make sure that people understood that wherever there is punctuation we want that to be a segment of speech but again we're not dealing with read speech there's not always a nice point of punctuation so wherever there's an you thought we also wanted that to me it marked so if someone gets interrupted it starts something new well there's no punctuation but that's not the same thought so we really went back to the drawing board and created a hard set of instructions for our transcribers another issue that we found was that I showed you guys clan before and again if we have two hours of audio it's a little bit difficult to go through using that clan software so we developed a GUI for tagging audio that we call a birth attack and the way this works is first we send the audio through a speech activity detector and so we try to segment out the audio as best we can into sections that acoustically resemble an utterance so we try to detect where their speech and then we see sort of these acoustic endpoints and present that to the tagger of sort of one other ends at a time or one acoustic what resembles an utterance the user then has to determine whether that is in fact one speaker one utterance and there's sort of a different box for each one of these points or whether it's one speaker with multiple utterances whether there's multiple speakers or whether there was TV and it's not actually a person talking and then because we're not a new word our first step is not manual there's also the chance that this could not even have speech in it at all so we do have that option too so you really wanted to just boil it down to something a lot more digestible by the transcriber from there then the first step is just saying how many others is is this even speech so the next step is the actual labeling of the audio so we first needed to know gender age group and then speaker ID on the transcription and then we care also if words are mispronounced because again it's not read speech there's a lot really going on whether they're singing and then we added drop-down menus for our language and for our noise and the reason we did this is because it enables us to be more consistent in what we call obeying a bump or these kinds of things we don't have different people writing a bunch of different words for these noises for TV electronics down we have a separate pop-up window so you really wanted it so that the transcriber is presented with only what they need to be presented with so if there's no TV we don't want them needs to deal with you know rummaging through the TV section so for TV electronic sound all we care is where it happens and if their speech in it so that later on we can try to figure out if we're detecting human speech or this TV speech and this again helps not only the accuracy but by not forcing them to transcribe the TV speech we're also trying to speed up the process again know another thing we really focus on was were the instructions because the the software itself was really important but how we presented what the user is supposed to do is really important so we we came up with a training program that each transcriber has to go through a series of videos and then also a reference sheet so they can look at what to do if they reach things that they're unfamiliar with so there's non speech sounds pauses fragments and intelligible an intelligible speech we have all of that so they can do it like a quick reference on that and that for we're hopefully again decreasing where our our discrepancies are coming between transcribers but the big question is did we actually do anything other that's of course when we care about so now we're at our second pause or we're now revisiting our data and seeing if we actually have seen improvements and we have so that's good what we're looking at here is the performances our initial methods are on the top and our optimized tagging methods are on the bottom and again similar to the plots I showed before since the truth is really dependent on transcribers right we have an instance where the first tagger is considered to be the crash tagger so their labels are true and then we have instances where the second tagger is the correct tagger and they are at the truth so what we want to see is really big yellow and green sections so the yellow is correct rejection so that's when both of the transcribers are saying this is not speech the green is our correct detection so that's or both of the transcribers are saying that a segment contains speech and then the bad segments are the the blue and the purple so in blue is in this detection so in this case on the Left we have transcriber one said that there was speech transcriber two said there was not speech so it's considered a miss and then again the opposite and purple we're now transcriber one said there was non speech transcriber two said there were speech so we have a false alarm so we do see now as we went from our initial methods to our to our versa tag and our more rigorous definition of utterances and instructions that we were able to decrease our miss detection errors from about seven and fourteen percent to about five percent for both and our false alarms who decreased some about twelve and six percent to about four and five percent so we're not saying that we're done we're still iterating through but this really is a big step in the direction that we want to go and for this whole thing really the way we've been integrating through the process is we've been working with um different students from local high schools universities and they've been actually doing the tagging for us and first of all we now have their data to analyze and they've been as project coming back to us with their recommendations with their concerns what's confusing and just that has sort of helped us get to a point of more clarity for the transcribers so basically our big goal again is to sum up today is to create a wearable that really encourages parents interact with their children so the big thing you know if a parents interacting more with their children we consider that a win but we also want to do really cool things with the audio along the way so our first release that's coming out now is just counting words more interaction more quantity our next release is really gonna focus on the data science and the quality and that's where we're now starting our data collection so that we will have the resources necessary to do the machine learning we found unsurprisingly that labeling ground truth for speech is not right wrong not easy whatsoever so we were inspired to develop versa tag and to develop a set of rigorous set of instructions and we really sort of learn the importance of just being very clear very black-and-white about what we want and so we're really excited that we've already seen these improvements we're still gonna continue this back and forth for a while until we really get to a point where we're getting that data right on but we're already very excited with the results we've been seeing so thank you everyone out enough we have time for questions or if anyone has any we have one minute was yes so the first thing we really noticed was that the utterance problem that we had people just labeling speech in every way possible basically so that's something that was our first issue that we really feel like we narrow down now we're entering sort of a new problem where the transcription itself is widely different you know because you hear different things some people care more about it and that's a little bit actually more difficult to combine because it's not just time-based you can't just either take the average or throw it out if it's too far apart you really have to figure out you know what these words are and how to combine those so that's sort of our newest headache but the next hurdle that were they were entering those two would be big ones then also noises so people would hear noise in the background and one's a bump one's a bang one you know and when we're gonna be parsing this later it's possible to do if we you know but it's a lot easier if we have a set of words that we can really look for so little things like that along the way we've tried to really Zone in on to get them as consistent as possible I'm sorry Oakland yet with software it's um it's so use a lot for sort of this kind of thing from visualizing audio there's a database the childís or Charlie's database and I know they use it quite quite regularly but yeah it's it's open it's free to use it's I don't really know who developed it to be honest but yeah no problem great that man that was like it's sort of a thing that we didn't really expect to be as difficult as it was we wanted to really focus instead of acoustically what's an utterance of what sort of to a human is an utterance Oh more thought based other ances so we have places you know you could do a huge pause and you could still be within an utterance so now we have framework for saying that there's a pause but saying that you're still in an utterance so we really wanted to focus on it being a thought so not necessary conversation with speech you're not really guaranteed to have like proper grammar and everything so really I mean ideally yet there's a period of there's punctuation then you're there but if you know mom starts talking and then baby's doing something and she stops and yells at the baby well that's a new thought because that's she's changed even though we haven't ended a sentence so we really wanted it to be what like based on the what is said and not necessarily how it was said I'll be wonderful I'd be very happy if the pain to do the manual tagging but it would be very very far down the road so far that I can't quite see it yet but that would be wonderful I mean the idea is once we start to get these probably one step in that direction would first be to use that to suggest to a person say this is what we think you know and then they can validate that and then and then further down the road once to get to a point we're doing really well then potentially but I think that as a good multiple step process to get there yeah yeah so right right now we're doing the word count for a version one for version two we want the wearable to be able to report back to the user sort of a report cards they can see not only the quantity but also the quality of interaction they're having correct yeah so this is the first step in getting there so it's it's that's always see as the future is being able to do that so it's what we wanted to be as as real life as possible so we have beta testers who volunteered and at this the first step was they were using their phone it was an app and we really encourage them to do we were reporting back to them on on what on how many words we said in the household but we also wanted them to really sort of set it and forget it so that natural life can happen and we're currently working on developing we have what we called the Starling which is the device clips to the child and I will report right now on the word count we're developing our research Starling so that we can gather audio now an even more realistic environment people our device doesn't record and when very like that's big to us we don't want it to record but our research device will so that we can collect the audio and it'll be real life as the child in the car as the child is you know have his sister soccer practice or or what have you yep okay thank you