← Search

Shuo Ma

3 accepted papers

2026

Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models

AAAI 2026technical

Recent Large Audio-Language Models (LALMs) exhibit impressive capabilities in understanding audio content for conversational QA tasks. However, these models struggle to accurately understand timestamps for temporal localization (e.g., Temporal Audio Grounding) and are restricted to short audio perce

Cited by 0SourcePDFScholar
2025

SleepSMC: Ubiquitous Sleep Staging via Supervised Multimodal Coordination

ICLR 2025poster

Sleep staging is critical for assessing sleep quality and tracking health. Polysomnography (PSG) provides comprehensive multimodal sleep-related information, but its complexity and impracticality limit its practical use in daily and ubiquitous monitoring. Conversely, unimodal devices offer more conv…

Cited by 0SourcePDFScholar