Statistics even when you're missing data
Audience:
Tags: statistics
Analytics
Comments
I enjoyed the pace and narration of this video. You offer a simple example motivating the usefulness of that algorithm. Overall, the motivation was easy to follow, but there were numerous points where I questioned a few of the premises and got a bit distracted thinking about what statistical assumptions were being made, such as the model used, the sampling procedure, the ability to make confidence intervals, etc.
I think the mathematical goal of your video was only made clear by around 4:20 into the video, which might be slightly long in terms of clearly motivating the namesake of the video.
Your target audience consisted of undergraduate and graduate students. I personally feel that this video falls closer to the undergraduate level.
Your style is indeed unique, despite MLE and EM being pretty standard topics in statistics and data science. I enjoyed your slide-design and ability to emphasize/point to specific portions of text.
I think what I’ll remember best as my “aha” moment is the idea that EM is an “expansion pack” to MLE.
Nice explanation of the algorithm and illustration.
Very nice and useful! The animations helped keep everything grounded. I liked the way that you explained parametric distributions and likelihood. Some of it was a bit hard to follow, I assume because I don’t know that much statistics, but I feel like I learned a few stats ideas from watching. I appreciate that you point to some other references.
I express my vote as a an average of the scores I have given to the individual principles specified in the guidelines (I have considered for each of the principles a score from 1 to 9, as for the final vote).
-
motivation: 8.5/9, I think it the importance of the topic is made quite clear from the beginning to the end. A few more cases explaining other instances and possible meanings of “missing data” would have been helpful;
-
clarity: 7/9, I honestly feel that quite some time is devoted to the explanation / “recap” of the maximum likelihood method, and then we are a bit rushed into the explanation of the EM algorithm. In fact, the explanation of the EM method is relegated to the last 7 minutes of a 17 minutes video (less than half). There are pieces of the EM formula that remain not very clear to me after rewatching the video. Especially the conditional expectations of the missing data, i.e. the term at minute 13:00. It’s not very clear how we get at 13:55 to (and the other term with the exponential) for the expectation values;
-
novelty: 7/9 I watched the first result on my youtube search for the same topic, to have a term of comparison. What I appreciate in this other video is that the topic was built form the ground up. Meaning that if some time is taken to build up from an example (two coins in this other video) the statistical meaning of the EM method and formulation, then the algorithm becomes much more digestible. Here I have the feeling we went from up to the ground. That is, we didn’t build up from the example to the general principle. We rather went: example / case study; pause; general principle / formula; application to the case study. I think it’s a perfectly valid way of presenting a topic, and as a matter of fact the approach in many textbooks. But the scope of the SoME was to create “something that would be useful in a class-room”. If I were to imagine myself as a data science student, I wouldn’t have the feeling that I “grasped” the topic and that now it’s mine;
Memorability: 9/9, I think the example chosen (average permanence of a subscription) makes the video very relatable, engaging and easy to remember. The second experiment actually made me think about proton decay, where a similar issue is encountered (an event too long for anyone to wait to occur).
Note (average of principles’ marks) = (8.5 + 7 + 7 + 9) / 4 = 7.875
I enjoyed the video and appreciate the effort that went into its preparation!
La motivación del vídeo fue buena y estuvo justificada. El vídeo trataba un tema verdaderamente curioso y novedoso, que pudieras estimar una media de un conjunto de datos sin precisamente tener todos esos datos. La única pega que le pongo al vídeo es que era algo complicado de entender, parecía que pasaba de idea en idea demasiado deprisa, y algunos cálculos matemáticos parecían “salir de la nada”. Me gustó mucho que todo girara alrededor de un problema del mundo real.
Very clear and interesting. My advice would be to make it more memorable, for now it is not very much so.
This was great! I knew about the maximum likelihood estimation method, but not about the EM method.
I really liked that you didn’t go through every calculation step and instead focused more on the intuition behind the method.
This was a fantastic video explaining every step of the algorithm and how it relates to the example. However, it was not explained why there was a logarithm in the E-step, why the second experiment ended after six months, or why the second experiment did not keep track of when the customers left for those who left before the six month deadline. The first two seem arbitrary, and the latter seems like it would have given more data to work with. (Also, it’s hard to suspend my disbelief that a streaming company wouldn’t have a spreadsheet somewhere of every account’s join and leave date, which could easily figure out the average lifetime of everyone who has left. Plus, streaming service memberships get paid for once a month, so the CEO might not care so much about fractions of months.)
I had never heard of the EM algorithm, and after a quick skim of the Youtube results I can see why there is still room for good material on it despite its importance.
While the video is visually very appealing, I have some gripes with it. First of all, I perceived the narration as too fast. I’d simply have needed more time to think whereas here, I hardly had time to pause the video in between sections. There are some points where the structure makes sense, but isn’t followed through: For example, around 10:30 you say you’ll quickly give an intuition before moving into details. But what follows isn’t an intuition, it’s just a verbalisation of the formulae. Your final point of Q approaching the complete likelihood at the end is much more intuitive and would have, in my opinion, made more sense here.
The motivation at the start basically comes down to an appeal to popularity: It is important because many people use it. I didn’t find that very convincing. I believe something else such as illustrating the problem of missing data and asking how that could be fixed would have been more effective.
I think this video can help people who have heard of the EM algorithm and are looking for additional material on it. Without prior motivation, it might be difficult to pay enough attention to keep up with the pace.