Behavioural analysis
A meltdown is not one signal, it is several at once. The vision model rates each visible sign; a fixed weighting and threshold decide the state.
Score a clip?The vision model reads frames sampled across the whole clip and rates each distress indicator 0–1 from what it can see. It scores; the verdict is arithmetic.
Reads ~16 frames across the clip. Audio-only signs (like vocal distress) score low unless the face shows them.
Pick a clip and assess it.