2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) · 2022 · 25 citations · 35 references
EngineeringMachine LearningAffective NeuroscienceMultimedia AnalysisMultimodal Temporal-aware FeaturesMultimodal Sentiment AnalysisArousal EstimationSocial SciencesTemporal Context InformationImage AnalysisData SciencePattern RecognitionAffective ComputingAffective Behavior AnalysisValence-arousal Estimation ChallengeBehavioral SciencesCognitive ScienceMultimodal Signal ProcessingVideo UnderstandingDeep LearningComputer VisionFacial Expression RecognitionEmotionEmotion Recognition
This paper presents our submission to the Valence-Arousal Estimation Challenge of the 3rd Affective Behavior Analysis in-the-wild (ABAW) competition. Based on multimodal feature representations that fuse the visual and aural information, we utilize two types of temporal encoder to capture the temporal context information in the video, including the transformer based encoder and LSTM based encoder. With the temporal context-aware representations, we employ fully-connected layers to predict the valence and arousal values of the video frames. In addition, smoothing processing is applied to refine the initial predictions, and a model ensemble strategy is used to combine multiple results from different model setups. Our system achieves the performance in Concordance Correlation Coefficients (ccc) of 0.606 for valence, 0.602 for arousal, and mean ccc of 0.601, which ranks the first place in the challenge.
35
DROPS (Schloss Dagstuhl – Leibniz Center for Informatics) · 2023 · 73.5K citations · Full text
Peter Salovey, John D. Mayer · Imagination Cognition and Personality · 1990 · 8.7K citations
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey et al. · 2015 · 5.7K citations
Audio Set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman et al. · 2017 · 2.9K citations
Music, Engineering, Machine Learning +22