ACM MM 2021: 'Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection'
-
Updated
Oct 23, 2023 - Python
ACM MM 2021: 'Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection'
Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection (ECCV 2022)
AnnoTheia is a data annotation toolkit that identifies when a person speaks in a scene and transcribes their speech, also offering flexibility to replace modules for different languages.
Accepted by TMM 2022
Open-source Opus Clip alternative. Turn long videos into vertical clips with AI-picked moments, virality scoring, animated captions, auto zoom and speaker-aware reframing. Free desktop app for macOS, Windows and Linux.
Active speaker detection: two-stream audio-visual CNN for AVA Active Speaker format data, with a live in-browser demo
Turns long podcasts and interviews into vertical short clips. WhisperX transcription, Gemini picks the moments, active speaker detection reframes to 9:16 with burned-in captions.
SyncNet based on Meta's Perception Encoder Audio-Visual (PE-AV)
Audio-visual synchronization and active-speaker detection in Python via cross-modal correlation
RNN-based Active Speaker Detection from facial landmarks (UniTalk-ASD + AVA + WASD)
Click a person in a video and listen to their voice.
视听多模态目标说话人增强:Light-ASD 身份先验、SRP-PHAT 与八通道 SBL-INCM-MVDR 波束形成
Proactive AI for smart glasses
Add a description, image, and links to the active-speaker-detection topic page so that developers can more easily learn about it.
To associate your repository with the active-speaker-detection topic, visit your repo's landing page and select "manage topics."