2026
T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
ICML 2026poster
Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped benchmarks that fail to capture cross-modal alignment, instruction…