← Search

Liting Wen

3 accepted papers

2026

MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX

AAAI 2026technical

We introduce MAVERIX (Multimodal Audio-Visual Evaluation and Recognition IndeX), a unified benchmark to probe video understanding in multimodal LLMs, encompassing video, audio, and text inputs with human performance baselines. Although recent advancements in audiovisual models have shown substantial

Cited by 0SourcePDFScholar
2026

OnlineHMR: Video-based Online World-Grounded Human Mesh Recovery

CVPR 2026

Human mesh recovery (HMR) models 3D human body from monocular videos, with recent works extending it to world-coordinate human trajectory and motion reconstruction. However, most existing methods remain offline, relying on future frames or global optimization, which limits their applicability in int

Cited by 0SourcecodeScholar
2025

FreeDance: Towards Harmonic Free-Number Group Dance Generation via a Unified Framework

ICCV 2025poster

Generating harmonic and diverse human motions from music signals, especially for a group of dancers, is a practical yet challenging task in virtual avatar creation. Existing methods merely model a fixed number of dancers, lacking the flexibility for arbitrary individuals. To fulfill this goal, we pr…