Summer of Math Exposition

Presented by 3Blue1Brown 3blue1brown

The Math Behind AI Sanity

A motivation for the use of a certain type of loss function in machine learning and a theorem about an surprising result of using such a loss function. Some knowledge of how neural networks are trained might be helpful.


Analytics

4.71 Overall score*
92 Rank
10 Votes
5 Comments

Comments

4.8
I really like your approch, even if i wasn't really on the mood for learning about this specific topic, it was clear and easy understable, and i apprecied how to manage to show the probleme before giving a solution, it make it was more natural. Saddenly i wasn't able to watch all of your video (but i will later). You were talking not too fast or too slowly, which is pretty good. Now for the least favorable thing, i will said that for what i saw (60% of the video), there is a lack of focus on the most important points. What i mean is that there were a lot of information and an overall view or an enlightment of the most fondamental points was a bit missing, at least it is my feeling :). I wound said, maybe making more emphase on the core suject (also with the rythm and voice) could probably give more power and enlightment for some points. Also a summary at the end can be a great way to remember the core of your video. I still appreciate the last diagram that kind of do a little conclusion but i feel like you could give it a bit more time to fully sumerize the ideas of your video.
5.8
Fascinating, but a little hard to follow at times. When you introduced the equation for the proper sorting rule, it would've been helpful to explain what each of the symbols appearing in the formula means, and not just s. Still, nice work!
3.8
The topic is really interesting and definitely fits in the goals of the SoME. However I think this presentation is a bit lacking. Firstly, I think there should be a brief recall of the basics of what is the process of training a model, and the definition of loss function should be accompanied by a brief description of the role that it plays in the training. The definition of proper scoring rule is given too quickly, and the variables that appear in the equation aren't well explained. Moreover, it is not very clear that the equations shown are the definition of proper score. Also, the viewer is left wondering what exactly is a score as opposed to a penalty/loss, and how the two notions are connected. At 3:42 it is not clear that the statement is a theorem (only later does the author say that it is a theorem, and that he is going to present a proof of it): it may appear as a property that the viewer is expected to agree with. At 5:50 the formula that writes φ as a maximum is not clear, at least to me; I believe the author should spend more time justifying and explaining it. In the definition of Bregham Divergence, the comment about the gradient sounds out of place and unhelpful. But most importantly, no motivation is given for that definition or for why it indeed represents the extra loss. I believe the author should at least briefly motivate/explain why it is not just the difference φ(Y) - φ(X) but there needs to be a term involving derivatives. (It looks like the remainder of the first order Taylor series of φ(Y) centered at X; maybe there is an easily apparent reason why this is the case? If so, it could make sense to tell the viewer). Also, the result that d_φ(X, Y)>=0 with equality if and only if X=Y deserves a little more motivation. And the viewer is also left with the thought that Bregham Divergence may actually be a metric (a distance): I think the author should have talked about this (whether or not it actually is a distance). At 7:53: the notion of "coherent" output was never defined in a precise mathematical way (only the example made at the beginning of the presentation serves as an explanation to what coherent means, but here there is a need for a precise definition). And also the fact that the set of coherent outputs is convex seems to be central to the proof, but it is only mentioned en passant, as if obvious; this is also a consequence of the lack of definition of coherence, because in order to prove that the set is convex there first needs to be a definition of the set. Here there is also another major hole in the presentation: the narrator says (too quickly) that there is a unique coherent element that minimises the Bregham Divergence, since that's a convex function. However I believe that this result about convex functions holds true only under certain conditions on the domain, like the domain should be compact (and possibly also conditions on the function (like continuity or similar)). This also highlights the need to state that the set of possible outputs has a certain structure: we already have a structure of vector space (which is not mentioned, but in my opinion should be mentioned), but now we also need at least a topology. Here is where I think the Bregham Divergence may be a distance, thus defining a metric on the space of outputs. (If that is the case, then we may also need the metric space of possible outputs to be complete with respect to this metric). So a lot needs to be added to justify the existence of a minimiser. When taking the limit as t goes to zero, I cannot see how this gives the derivative. Perhaps we need to divide by t first, and then take the limit, but even then I can't figure out how the formula is deduced. I think the author should explain these steps more carefully and with more intermediate steps. And the next equation is even more obscure to me, I can't understand what the equal sign refers to from the line above: is it a continuation from the previous line that is supposed to be "0 = [...]"? (And in any case, why/how does the previous line reduce to this new line?) These steps lack explanation, and are only shown on screen without further comment. Further comments: When giving definitions or stating theorems, there should be a title denoting that the sentencs is a definition/theorem. Also, it should be clear where a proof begins and where it ends. When talking about convex functions, instead of "to start, let's take a quick detour into the land of convex functions" it would have been better, in my opinion, to phrase it as "let's recall some properties of convex functions".
1.5
Voice was hard to understand, explanation was confusing with lots of poorly defined concepts (I'm pretty familiar with machine learning and optimization and had trouble following)
4.7
Video okay, sound a little less good than average. You only show the equations on screen, but not concrete examples of what you're talking about so the whole video feels a bit too 'abstract', I feel. A few pictures of cat, dog, goat, table, neat 'a' and scribbled 'a' and the classifier's beliefs as to what they are might have helped.