← Search

Chunlei Meng

5 accepted papers

2026

CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning

CVPR 2026

Multimodal learning aims to capture both shared and private information from multiple modalities. However, existing methods that project all modalities into a single latent space for fusion often overlook the asynchronous, multi-level semantic structure of multimodal data. This oversight induces sem

Cited by 0SourceScholar
2026

CareBot-H: Enhancing Patient Transfer with Biomimetic Design and Trajectory Deformation Algorithm

ICRA 2026poster

This paper introduces the CareBot-H Robot, a humanoid nursing robot designed to perform patient transfer tasks in confined environments. The robot is equipped with biomimetic arms that replicate human arm size and function, and distributed tactile sensors that enhance operational safety during physi…

Cited by 0Scholar
2026

Group Cognition Learning: Making Everything Better Through Controlled Two-Stage Agents Collaboration

ICML 2026poster

Centralized multimodal learning commonly compresses language, acoustic, and visual signals into a single fused representation for prediction. While effective, this paradigm suffers from two limitations: modality dominance, where optimization gravitates towards the path of least resistance, ignoring …

Cited by 0SourceScholar
2026

Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis

ICASSP 2026oral

Multimodal Sentiment Analysis integrates Linguistic, Visual, and Acoustic. Mainstream approaches based on modality-invariant and modality-specific factorization or on complex fusion still rely on spatiotemporal mixed modeling. This ignores spatiotemporal heterogeneity, leading to spatiotemporal info…

Cited by 0SourcePDFScholar
2026

Tri-Subspaces Disentanglement for Multimodal Sentiment Analysis

CVPR 2026

Multimodal Sentiment Analysis (MSA) integrates language, visual, and acoustic modalities to infer human sentiment. Most existing methods either focus on globally shared representations or modality-specific features, while overlooking signals that are shared only by certain modality pairs. This limit

Cited by 0SourceScholar