← Search

Yuanchao Li

9 accepted papers

2026

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

ICML 2026poster

Recent advances in Omni-Multimodal Large Language Models (Omni-MLLMs) have enabled strong integration of vision, audio, and language. However, their audio-visual intelligence (AVI) remains insufficiently evaluated due to the lack of systematic and comprehensive benchmarks. We introduce AVI-Bench, a …

Cited by 0SourceScholar
2025

Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models

ICASSP 2025accepted

Utilizing Self-Supervised Learning (SSL) models for Speech Emotion Recognition (SER) has proven effective, yet limited research has explored cross-lingual scenarios. This study presents a comparative analysis between human performance and SSL models, beginning with a layer-wise analysis and an explo…

Cited by 0SourceScholar
2025

Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations

ICASSP 2025accepted

Emotion recognition from speech and music shares similarities due to their acoustic overlap, which has led to interest in transferring knowledge between these domains. However, the shared acoustic cues between speech and music, particularly those encoded by Self-Supervised Learning (SSL) models, rem…

Cited by 0SourceScholar
2025

Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction

ICASSP 2025accepted

Annotating and recognizing speech emotion using prompt engineering has recently emerged with the advancement of Large Language Models (LLMs), yet its efficacy and reliability remain questionable. In this paper, we conduct a systematic study on this topic, beginning with the proposal of novel prompts…

Cited by 0SourceScholar
2025

Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling

ICASSP 2025accepted

The lack of labeled data is a common challenge in speech classification tasks, particularly those requiring extensive subjective assessment, such as cognitive state classification. In this work, we propose a Semi-Supervised Learning (SSL) framework, introducing a novel multi-view pseudo-labeling met…

Cited by 0SourceScholar
2024

Can Textual Semantics Mitigate Sounding Object Segmentation Preference?

ECCV 2024poster

"The Audio-Visual Segmentation (AVS) task aims to segment sounding objects in the visual space using audio cues. However, in this work, it is recognized that previous AVS methods show a heavy reliance on detrimental segmentation preferences related to audible objects, rather than precise audio guida…

2023

Multimodal Dyadic Impression Recognition via Listener Adaptive Cross-Domain Fusion

ICASSP 2023accepted

As a sub-branch of affective computing, impression recognition, e.g., perception of speaker characteristics such as warmth or competence, is potentially a critical part of both human-human conversations and spoken dialogue systems. Most research has studied impressions only from the behaviors expres…

Cited by 0SourceScholar
2020

Cooperative Comfortable-Driving at Signalized Intersections for Connected and Automated Vehicles

RA-L 2020

This letter proposes a control framework for Connected and Automated Vehicles(CAVs) to approach the signalized intersections with good driving-comfortability. Both the velocity plan and longitudinal dynamics control are concerned in this study. Regarding the velocity plan problem, a two-layer framew

Cited by 36SourceScholar