The Kinetics Human Action Video Dataset

We describe the DeepMind Kinetics human action video dataset. The dataset contains 400 human action classes, with at least 400 video clips for each action. Each clip lasts around 10s and is taken from a different YouTube video. The actions are human focussed and cover a broad range of classes including human-object interactions such as playing instruments, as well as human-human interactions such as shaking hands. We describe the statistics of the dataset, how it was collected, and give some baseline performance figures for neural network architectures trained and tested for human action classification on this dataset. We also carry out a preliminary analysis of whether imbalance in the dataset leads to bias in the classifiers.

UCF101: A Dataset of 101Human Actions Classes…UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild2D Human PoseEstimation: New…2D Human Pose Estimation: New Benchmark and State of the Art AnalysisActivityNet: Alarge-scale video…ActivityNet: A large-scale video benchmark for human activity understandingBatch Normalization:Accelerating Deep…Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate ShiftTensorFlow: Large-ScaleMachine Learning on…TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed SystemsSemantics derivedautomatically from…Semantics derived automatically from language corpora contain human-like biasesRecurrent BatchNormalizationRecurrent Batch NormalizationActional-StructuralGraph Convolutional…Actional-Structural Graph Convolutional Networks for Skeleton-Based Action RecognitionDense Dilated Networkfor Video Action…Dense Dilated Network for Video Action RecognitionBreakingWinner-Takes-All…Breaking Winner-Takes-All: Iterative-Winners-Out Networks for Weakly Supervised Temporal Action LocalizationListen to Look: ActionRecognition by…Listen to Look: Action Recognition by Previewing AudioSpatiotemporal Fusion in3D CNNs: A Probabilisti…Spatiotemporal Fusion in 3D CNNs: A Probabilistic ViewNo Frame Left Behind:Full Video Action…No Frame Left Behind: Full Video Action RecognitionSelf-SupervisedRepresentation Learning…Self-Supervised Representation Learning: Introduction, advances, and challengesSelf-Supervised Learningfor Videos: A SurveySelf-Supervised Learning for Videos: A SurveySelf-SupervisedSpatiotemporal…Self-Supervised Spatiotemporal Representation Learning by Exploiting Video ContinuityPIVOT: Prompting forVideo Continual LearningPIVOT: Prompting for Video Continual LearningVideoChat-Flash:Hierarchical Compressio…VideoChat-Flash: Hierarchical Compression for Long-Context Video ModelingVideoGPT+: IntegratingImage and Video Encoder…VideoGPT+: Integrating Image and Video Encoders for Enhanced Video UnderstandingThe Kinetics HumanAction Video DatasetThe Kinetics Human Action Video DatasetEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.