DataoceanAI/American_English_Male_Speech_Synthesis_Corpus_Gentle_and_Mature_Aged_30_40 Updated Jan 10, 2025 • 85 • 9
Native Action-Prior Learning from Videos for World Action Models Paper • 2610.03391 • Published 5 days ago • 84
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs Paper • 2609.26796 • Published 15 days ago • 37
StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training Paper • 2609.26774 • Published 15 days ago • 56
ehcalabres/wav2vec2-lg-xlsr-en-speech-emotion-recognition Audio Classification • 0.3B • Updated Oct 24, 2024 • 16.8k • 258
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue Paper • 2609.21465 • Published 19 days ago • 151
TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection Paper • 2609.18748 • Published 20 days ago • 10
Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control Paper • 2609.17909 • Published 22 days ago • 47
xmj2002/hubert-base-ch-speech-emotion-recognition Audio Classification • Updated May 16, 2023 • 516 • 57
r-f/wav2vec-english-speech-emotion-recognition Automatic Speech Recognition • Updated Jan 2, 2025 • 4.05k • 41
speechbrain/emotion-recognition-wav2vec2-IEMOCAP Audio Classification • Updated Jul 23, 2024 • 48.1k • 197
Building a Production Greek-English Speech Recognizer Paper • 2609.13498 • Published 26 days ago • 9
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement Paper • 2609.13406 • Published 26 days ago • 85