Research

Multi-Moments in Time: Learning and Interpreting Models for Multi-Action Video Understanding

Authors

Mathew Monfort
Bowen Pan
Kandan Ramakrishnan
Alex Andonian
Barry A. McNamara
Alex Lascelles
Quanfu Fan
Dan Gutfreund
Rogerio Feris
Aude Oliva

Cite

Research

Multi-Moments in Time: Learning and Interpreting Models for Multi-Action Video Understanding

Computer Vision

Cite Paper Project Page

Authors

Mathew Monfort
Bowen Pan
Kandan Ramakrishnan
Alex Andonian
Barry A. McNamara
Alex Lascelles
Quanfu Fan
Dan Gutfreund
Rogerio Feris
Aude Oliva

Published on

11/09/2021

Categories

Computational neuroscience Computer Vision

Videos capture events that typically contain multiple sequential, and simultaneous, actions even in the span of only a few seconds. However, most large-scale datasets built to train models for action recognition in video only provide a single label per video. Consequently, models can be incorrectly penalized for classifying actions that exist in the videos but are not explicitly labeled and do not learn the full spectrum of information present in each video in training. Towards this goal, we present the Multi-Moments in Time dataset (M-MiT) which includes over two million action labels for over one million three second videos. This multi-label dataset introduces novel challenges on how to train and analyze models for multi-action detection. Here, we present baseline results for multi-action recognition using loss functions adapted for long tail multi-label learning, provide improved methods for visualizing and interpreting models trained for multi-label action detection and show the strength of transferring models trained on M-MiT to smaller datasets.

This work was presented in IEEE Transactions on Pattern Analysis and Machine Intelligence in 2021.

Please cite our work using the BibTeX below.

@ARTICLE{9609554,
  author={Monfort, Mathew and Pan, Bowen and Ramakrishnan, Kandan and Andonian, Alex and McNamara, Barry A. and Lascelles, Alex and Fan, Quanfu and Gutfreund, Dan and Feris, Rogerio Schmidt and Oliva, Aude},
  journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, 
  title={Multi-Moments in Time: Learning and Interpreting Models for Multi-Action Video Understanding}, 
  year={2022},
  volume={44},
  number={12},
  pages={9434-9445},
  doi={10.1109/TPAMI.2021.3126682}}