← Search

Chenglizhao Chen

14 accepted papers

2026

AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios

AAAI 2026technical

Referring Multi-Object Tracking (RMOT) aims to achieve precise object detection and tracking through natural language instructions, representing a fundamental capability for intelligent robotic systems. However, current RMOT research remains mostly confined to ground-level scenarios, which constrain

Cited by 0SourcePDFScholar
2026

Controlling Decision Drift in Multimodal Sentiment Analysis with Missing Modalities

IJCAI 2026

Multimodal sentiment analysis relies on textual, acoustic, and visual signals, yet real-world data often suffer from modality missing and quality imbalance. Existing methods generate features for modality missing from available ones, but differences in expression mechanisms and sentiment dynamics ac

Cited by 0Scholar
2026

MoCoDiff: A Controllable Autoregressive Diffusion Model for Expressive Motion Generation

CVPR 2026

Diffusion-based motion generation has advanced rapidly, but current methods still struggle with long-horizon consistency, style control, and multi-condition guidance. A major reason is the fused-conditioning design, where semantic, stylistic, and temporal signals share a single pathway, causing inte

Cited by 0SourceScholar
2025

Fine-Grained Perception in Panoramic Scenes: A Novel Task, Dataset, and Method for Object Importance Ranking

AAAI 2025technical

Existing Salient Object Ranking (SOR) aims to infer ranking of salient objects based on their saliency degree. However, it tends to only focus on salient objects while neglecting non-salient ones. This coarse-grained ranking limits the performance of downstream tasks. For instance, in image retrieva…

2024

Arbitrary Motion Style Transfer with Multi-condition Motion Latent Diffusion Model

CVPR 2024poster

Computer animation's quest to bridge content and style has historically been a challenging venture with previous efforts often leaning toward one at the expense of the other. This paper tackles the inherent challenge of content-style duality ensuring a harmonious fusion where the core narrative of t…

2024

Cross-View Diversity Embedded Consensus Learning for Multi-View Clustering

IJCAI 2024poster

Multi-view clustering (MVC) has garnered significant attention in recent studies. In this paper, we propose a novel MVC method, named CCL-MVC. The novel method constructs a cross-order neighbor tensor of multi-view data to recover a low-rank essential tensor, preserves noise-free, comprehensive, and…

Cited by 0SourcePDFScholar
2024

Fine-Grained Bipartite Concept Factorization for Clustering

CVPR 2024poster

In this paper we propose a novel concept factorization method that seeks factor matrices using a cross-order positive semi-definite neighbor graph which provides comprehensive and complementary neighbor information of the data. The factor matrices are learned with bipartite graph partitioning which…

Cited by 2SourcePDFScholar
2024

HOIAnimator: Generating Text-prompt Human-object Animations using Novel Perceptive Diffusion Models

CVPR 2024poster

To date the quest to rapidly and effectively produce human-object interaction (HOI) animations directly from textual descriptions stands at the forefront of computer vision research. The underlying challenge demands both a discriminating interpretation of language and a comprehensive physics-centric…

Cited by 11SourcePDFScholar
2023

Lightweight Portrait Segmentation Via Edge-Optimized Attention

ICASSP 2023accepted

With the outbreak of COVID-19 around the world, the frequency of video conferencing at home is increasing. Therefore, a segmentation architecture that can quickly carry out close-range portrait segmentation has become a current need. However, the current portrait segmentation architectures cannot me…

Cited by 0SourceScholar
2023

Pixel Is All You Need: Adversarial Trajectory-Ensemble Active Learning for Salient Object Detection

AAAI 2023technical

Although weakly-supervised techniques can reduce the labeling effort, it is unclear whether a saliency model trained with weakly-supervised data (e.g., point annotation) can achieve the equivalent performance of its fully-supervised version. This paper attempts to answer this unexplored question by…

Cited by 10SourcePDFScholar
2021

From Semantic Categories to Fixations: A Novel Weakly-Supervised Visual-Auditory Saliency Detection Approach

CVPR 2021poster

Thanks to the rapid advances in the deep learning techniques and the wide availability of large-scale training sets, the performances of video saliency detection models have been improving steadily and significantly. However, the deep learning based visual-audio fixation prediction is still in its i…

Cited by 47PDFcodeScholar
2021

Mutual Graph Learning for Camouflaged Object Detection

CVPR 2021poster

Automatically detecting/segmenting object(s) that blend in with their surroundings is difficult for current models. A major challenge is that the intrinsic similarities between such foreground objects and background surroundings make the features extracted by deep model indistinguishable. To overcom…

Cited by 338PDFcodeScholar
2019

RES-PCA: A Scalable Approach to Recovering Low-Rank Matrices

CVPR 2019poster

Robust principal component analysis (RPCA) has drawn significant attentions due to its powerful capability in recovering low-rank matrices as well as successful appplications in various real world problems. The current state-of-the-art algorithms usually need to solve singular value decomposition of…

Cited by 30PDFScholar