← Search

Chenglong Li

31 accepted papers

2026

Apo2Mol: 3D Molecule Generation via Dynamic Pocket-Aware Diffusion Models

AAAI 2026technical

Deep generative models are rapidly advancing structure-based drug design, offering substantial promise for generating small molecule ligands that bind to specific protein targets. However, most current approaches assume a rigid protein binding pocket, neglecting the intrinsic flexibility of proteins

Cited by 3SourcePDFScholar
2026

Chain-of-Thought Guided Multi-Modal Object Re-Identification

CVPR 2026

With the rise of visual-language models, multi-modal ReID retrieves specific targets by integrating different spectra and textual descriptions. Existing methods merely adopt descriptive representation learning for image-text, ignoring the relationships among the intrinsic logical hierarchies of sema

Cited by 0SourcecodeScholar
2026

CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product

AAAI 2026technical

Human-defined creativity is highly abstract, posing a challenge for multimodal large language models (MLLMs) to comprehend and assess creativity that aligns with human judgments. The absence of an existing benchmark further exacerbates this dilemma. To this end, we propose CreBench, which consists o

Cited by 0SourcePDFScholar
2026

Cross-modal Fuzzy Alignment Network for Text-Aerial Person Retrieval and A Large-scale Benchmark

CVPR 2026

Text-aerial person retrieval aims to identify targets in UAV-captured images from eyewitness descriptions, supporting intelligent transportation and public security applications. Compared to ground-view text-image person retrieval, UAV-captured images often suffer from degraded visual information du

Cited by 0SourcecodeScholar
2026

Dual-Teacher Interactive Knowledge Distillation Network for Text-to-Visible & Infrared Person Retrieval

AAAI 2026technical

Text-to-visible & infrared person retrieval aims to retrieve the corresponding visible (RGB) and thermal infrared (TIR) images given the text descriptions. Existing methods perform semantic decoupling by aligning RGB and TIR features separately to different attributes, thereby facilitating the align

Cited by 0SourcePDFScholar
2026

GO-PRE:Goal-Oriented Next-Best-View Selection via Predictive Rendering Entropy for Active 3D Reconstruction

ICML 2026poster

Active 3D reconstruction relies on active view selection to maximize reconstruction fidelity under limited capture budgets. However, most existing methods rely on surrogate signals—such as parameter uncertainty or geometric heuristics—which are often misaligned with the ultimate goal: the fidelity o…

Cited by 0SourceScholar
2026

Progressive Multi-cue Alignment for Unaligned RGBT Tracking

CVPR 2026

Unaligned RGBT tracking aims to achieve robust target localization across spatially misaligned RGB and thermal infrared (TIR) videos, a crucial challenge for deploying RGBT tracking in real-world scenarios. Existing methods often calculate all cross-modal alignment parameters (i.e., spatial shift an

Cited by 0SourcecodeScholar
2026

ProxyTTT: Proxy-driven Test-Time Training for Multi-modal Re-identification

AAAI 2026technical

Multi-modal object re-identification (ReID) aims to retrieve specific targets by leveraging complementary cues from different sensing modalities. Despite recent progress, two key challenges remain: (1) the limited ability to jointly address both modality and viewpoint discrepancies, and (2) the diff

Cited by 0SourcePDFScholar
2026

RGB-Event based Pedestrian Attribute Recognition: A Benchmark Dataset and An Asymmetric RWKV Fusion Framework

CVPR 2026

Existing pedestrian attribute recognition methods are generally developed based on RGB frame cameras. However, these approaches are constrained by the limitations of RGB cameras, such as sensitivity to lighting conditions and motion blur, which hinder their performance. Furthermore, current attribut

Cited by 0SourcecodeScholar
2026

Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

CVPR 2026

Missing modalities in RGBT tracking often lead to incomplete and unstable multimodal feature representations that greatly degrade the performance. Existing methods typically attempt to recover missing modalities from available ones, but the quality of data generated in challenging scenarios might be

Cited by 0SourceScholar
2026

Unaligned UAV RGBT Tracking: A Largescale Benchmark and a Novel Approach

AAAI 2026technical

With the rapid development of the low-altitude economy, multimodal visual tracking in UAV scenarios has attracted extensive attention. UAVs are typically equipped with independent visible (RGB) and thermal infrared (TIR) sensors, resulting in an inherent spatial misalignment between the two modaliti

Cited by 0SourcePDFScholar
2025

Alignment-Free RGB-T Salient Object Detection: A Large-Scale Dataset and Progressive Correlation Network

AAAI 2025technical

Alignment-free RGB-Thermal (RGB-T) salient object detection (SOD) aims to achieve robust performance in complex scenes by directly leveraging the complementary information from unaligned visible-thermal image pairs, without requiring manual alignment. However, the labor-intensive process of collecti…

2025

Cross-modulated Attention Transformer for RGBT Tracking

AAAI 2025technical

Existing Transformer-based RGBT trackers achieve remarkable performance benefits by leveraging self-attention to extract uni-modal features and cross-attention to enhance multi-modal feature interaction and search-template correlation. Nevertheless, the independent search-template correlation calcul…

2025

DecoyDB: A Dataset for Graph Contrastive Learning in Protein-Ligand Binding Affinity Prediction

NeurIPS 2025poster

Predicting the binding affinity of protein-ligand complexes plays a vital role in drug discovery. Unfortunately, progress has been hindered by the lack of large-scale and high-quality binding affinity labels. The widely used PDBbind dataset has fewer than 20K labeled complexes. Self-supervised learn…

Cited by 0SourceScholar
2025

Pedestrian Attribute Recognition: A New Benchmark Dataset and a Large Language Model Augmented Framework

AAAI 2025technical

Pedestrian Attribute Recognition (PAR) is one of the indispensable tasks in human-centered research. However, existing datasets neglect different domains (e.g., environments, times, populations, and data sources), only conducting simple random splits, and the performance of these datasets has alread…

2025

RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba

AAAI 2025technical

Existing RGBT tracking methods often design various interaction models to perform cross-modal fusion of each layer, but can not execute the feature interactions among all layers, which plays a critical role in robust multimodal representation, due to large computational burden. To address this issue…

Cited by 2SourcePDFScholar
2025

Robo-GS: A Physics Consistent Spatial-Temporal Model for Robotic Arm with Hybrid Representation

ICRA 2025

The Real2Sim2Real (R2S2R) paradigm is critical for advancing robotic learning. Existing methods lack a comprehensive solution to accurately reconstruct real-world objects with both spatial representations and their associated physics attributes in the Real2Sim stage. We propose a Real2Sim pipeline t

Cited by 73SourceScholar
2025

Template-based Uncertainty Multimodal Fusion Network for RGBT Tracking

IJCAI 2025

RGBT tracking is to localize the predefined targets in video sequences by effectively leveraging the information from both visible light (RGB) and thermal infrared (TIR) modalities. However, the quality of different modalities changes dynamically in complex scenes, and effectively perceiving modal q

2025

UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-Identification

NeurIPS 2025poster

Multi-modal object Re-IDentification (ReID) has gained considerable attention with the goal of retrieving specific targets across cameras using heterogeneous visual data sources. At present, multi-modal object ReID faces two core challenges: (1) learning robust features under fine-grained local nois…

Cited by 0SourcecodeScholar
2024

DI-MVS: Learning Efficient Multi-View Stereo With Depth-Aware Iterations

ICASSP 2024accepted

Learning-based Multi-View Stereo (MVS) methods aim to reconstruct 3D scenes from a set of 2D calibrated images. However, existing learning-based MVS methods often overlook depth maps that include the geometric shapes of the scene when constructing the cost volume. This can result in suboptimal recon…

Cited by 0SourceScholar
2024

Parallel Augmentation and Dual Enhancement for Occluded Person Re-Identification

ICASSP 2024accepted

Occluded person re-identification (Re-ID), the task of searching for the same person’s images in occluded environments, has attracted lots of attention in the past decades. Recent approaches concentrate on improving performance on occluded data by data/feature augmentation or using extra models to p…

Cited by 0SourceScholar
2024

Structural Information Guided Multimodal Pre-training for Vehicle-Centric Perception

AAAI 2024technical

Understanding vehicles in images is important for various applications such as intelligent transportation and self-driving system. Existing vehicle-centric works typically pre-train models on large-scale classification datasets and then fine-tune them for specific downstream tasks. However, they neg…

2022

Attribute-Based Progressive Fusion Network for RGBT Tracking

AAAI 2022technical

RGBT tracking usually suffers from various challenge factors, such as fast motion, scale variation, illumination variation, thermal crossover and occlusion, to name a few. Existing works often study fusion models to solve all challenges simultaneously, and it requires fusion models complex enough an…

Cited by 171SourcePDFScholar
2022

Cross-Modal Object Tracking: Modality-Aware Representations and a Unified Benchmark

AAAI 2022technical

In many visual systems, visual tracking often bases on RGB image sequences, in which some targets are invalid in low-light conditions, and tracking performance is thus affected significantly. Introducing other modalities such as depth and infrared data is an effective way to handle imaging limitatio…

2022

Interact, Embed, and EnlargE: Boosting Modality-Specific Representations for Multi-Modal Person Re-identification

AAAI 2022technical

Multi-modal person Re-ID introduces more complementary information to assist the traditional Re-ID task. Existing multi-modal methods ignore the importance of modality-specific information in the feature fusion stage. To this end, we propose a novel method to boost modality-specific representations…

Cited by 51SourcePDFScholar
2021

Cross-Modal Collaborative Representation Learning and a Large-Scale RGBT Benchmark for Crowd Counting

CVPR 2021poster

Crowd counting is a fundamental yet challenging task, which desires rich information to generate pixel-wise crowd density maps. However, most previous methods only used the limited information of RGB images and cannot well discover potential pedestrians in unconstrained scenarios. In this work, we f…

Cited by 165PDFcodeScholar
2018

Cross-Modal Ranking with Soft Consistency and Noisy Labels for Robust RGB-T Tracking

ECCV 2018poster

Due to the complementary benefits of visible (RGB) and thermal infrared (T) data, RGB-T object tracking attracts more and more attention recently for boosting the performance under adverse illumination conditions. Existing RGB-T tracking methods usually localize a target object with a bounding box,…

Cited by 163SourcePDFScholar
2018

SINT++: Robust Visual Tracking via Adversarial Positive Instance Generation

CVPR 2018poster

Existing visual trackers are easily disturbed by occlusion,blurandlargedeformation. Inthechallengesofocclusion, motion blur and large object deformation, the performance of existing visual trackers may be limited due to the followingissues: i)Adoptingthedensesamplingstrategyto generate positive exam…

Cited by 152SourcePDFScholar
2015

SOLD: Sub-Optimal Low-rank Decomposition for Efficient Video Segmentation

CVPR 2015poster

This paper investigates how to perform robust and efficient unsupervised video segmentation while suppressing the effects of data noises and/or corruptions. We propose a general algorithm, called Sub-Optimal Low-rank Decomposition (SOLD), which pursues the low-rank representation for video segmentat…

Cited by 54SourcePDFScholar