← Search

Hong Chang

34 accepted papers

2026

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes

ICLR 2026poster

The aspiration for artificial general intelligence, fueled by the rapid progress of multimodal understanding, demands models to understand humans in diverse and complex scenarios, as humans manifests intelligence and embody the world. We propose HumanPCR, an evaluation suite for probing MLLMs’ capac…

Cited by 0SourceScholar
2026

LensWalk: Agentic Video Understanding by Planning How You See in Videos

CVPR 2026

The dense, temporal nature of video presents a profound challenge for automated analysis. Despite the use of powerful Vision-Language Models, prevailing methods for video understanding are limited by the inherent disconnect between reasoning and perception: they rely on static, pre-processed informa

Cited by 0SourceScholar
2026

MAST: Motif-Augmented Diffusion with Search Tree for Spectroscopic Molecular Structure Elucidation

ICML 2026poster

Elucidating molecular structures from spectra is a foundational problem in chemical and materials characterization, yet remains challenging due to spectral ambiguity and the vast molecular space. Although recent diffusion-based generators show strong promise for spectra-conditioned elucidation, exis…

Cited by 0SourceScholar
2026

MindFlow: Mind Supernet Powered Thinking Flows for Research Idea Innovation

ICML 2026poster

Research idea innovation is a fundamental engine of scientific progress, yet it remains difficult to generate and evaluate in a scalable and controllable way. This challenge lies in its inherently open-ended and multi-objective nature, where ideas should balance novelty, plausibility and feasibility…

Cited by 0SourceScholar
2026

Revisiting Multimodal Positional Encoding in Vision–Language Models

ICLR 2026poster

Multimodal position encoding is essential for vision-language models, yet there has been little systematic investigation into multimodal position encoding. We conduct a comprehensive analysis of multimodal Rotary Positional Embedding (RoPE) by examining its two core components: position design and f…

Cited by 0SourcecodeScholar
2025

G2PDiffusion: Cross-Species Genotype-to-Phenotype Prediction via Evolutionary Diffusion

ICCV 2025poster

Understanding how genes influence phenotype across species is a fundamental challenge in genetic engineering, which will facilitate advances in various fields such as crop breeding, conservation biology, and personalized medicine. However, current phenotype prediction models are limited to individua…

Cited by 0SourcePDFScholar
2025

HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding

ICCV 2025poster

We propose a new task to benchmark human-in-scene understanding for embodied agents: Human-In-Scene Question Answering (HIS-QA). Given a human motion within a 3D scene, HIS-QA requires the agent to comprehend human states and behaviors, reason about its surrounding environment, and answer human-rela…

2025

KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge

NeurIPS 2025poster

The molecular large language models have garnered widespread attention due to their promising potential on molecular applications. However, current molecular large language models face significant limitations in understanding molecules due to inadequate textual descriptions and suboptimal molecular…

Cited by 0SourceScholar
2025

MATS: An Audio Language Model under Text-only Supervision

ICML 2025poster

Large audio-language models (LALMs), built upon powerful Large Language Models (LLMs), have exhibited remarkable audio comprehension and reasoning capabilities. However, the training of LALMs demands a large corpus of audio-language pairs, which requires substantial costs in both data collection an…

2025

ProtInvTree: Deliberate Protein Inverse Folding with Reward-guided Tree Search

NeurIPS 2025spotlight

Designing protein sequences that fold into a target 3D structure—known as protein inverse folding—is a fundamental challenge in protein engineering. While recent deep learning methods have achieved impressive performance by recovering native sequences, they often overlook the one-to-many nature of t…

Cited by 0SourcecodeScholar
2025

Revisiting Logit Distributions for Reliable Out-of-Distribution Detection

NeurIPS 2025poster

Out-of-distribution (OOD) detection is critical for ensuring the reliability of deep learning models in open-world applications. While post-hoc methods are favored for their efficiency and ease of deployment, existing approaches often underexploit the rich information embedded in the model’s logits…

Cited by 0SourcecodeScholar
2025

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing

CVPR 2025highlight

Human pose plays a crucial role in the digital age. While recent works have achieved impressive progress in understanding and generating human poses, they often support only a single modality of control signals and operate in isolation, limiting their application in real-world scenarios. This paper…

2025

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP

NeurIPS 2025poster

Contrastive Language-Image Pre-training (CLIP) has become a foundation model and has been applied to various vision and multimodal tasks. However, recent works indicate that CLIP falls short in distinguishing detailed differences in images and shows suboptimal performance on dense-prediction and vis…

Cited by 0SourcecodeScholar
2024

M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation

NeurIPS 2024poster

This paper presents M$^3$GPT, an advanced $\textbf{M}$ultimodal, $\textbf{M}$ultitask framework for $\textbf{M}$otion comprehension and generation. M$^3$GPT operates on three fundamental principles. The first focuses on creating a unified representation space for various motion-relevant modalities…

2024

Scalable Modular Network: A Framework for Adaptive Learning via Agreement Routing

ICLR 2024poster

In this paper, we propose a novel modular network framework, called Scalable Modular Network (SMN), which enables adaptive learning capability and supports integration of new modules after pre-training for better adaptation. This adaptive capability comes from a novel design of router within SMN, na…

2024

UMFC: Unsupervised Multi-Domain Feature Calibration for Vision-Language Models

NeurIPS 2024poster

Pre-trained vision-language models (e.g., CLIP) have shown powerful zero-shot transfer capabilities. But they still struggle with domain shifts and typically require labeled data to adapt to downstream tasks, which could be costly. In this work, we aim to leverage unlabeled data that naturally span…

2023

Generalized Semi-Supervised Learning via Self-Supervised Feature Adaptation

NeurIPS 2023poster

Traditional semi-supervised learning (SSL) assumes that the feature distributions of labeled and unlabeled data are consistent which rarely holds in realistic scenarios. In this paper, we propose a novel SSL setting, where unlabeled samples are drawn from a mixed distribution that deviates from the…

Cited by 5SourcePDFScholar
2023

Understanding Few-Shot Learning: Measuring Task Relatedness and Adaptation Difficulty via Attributes

NeurIPS 2023poster

Few-shot learning (FSL) aims to learn novel tasks with very few labeled samples by leveraging experience from \emph{related} training tasks. In this paper, we try to understand FSL by exploring two key questions: (1) How to quantify the relationship between \emph{ training} and \emph{novel}…

2022

Clothes-Changing Person Re-Identification With RGB Modality Only

CVPR 2022poster

The key to address clothes-changing person re-identification (re-id) is to extract clothes-irrelevant features, e.g., face, hairstyle, body shape, and gait. Most current works mainly focus on modeling body shape from multi-modality information (e.g., silhouettes and sketches), but do not make full u…

Cited by 225PDFcodeScholar
2022

Joint Feature Learning and Relation Modeling for Tracking: A One-Stream Framework

ECCV 2022poster

"The current popular two-stream, two-stage tracking framework extracts the template and the search region features separately and then performs relation modeling, thus the extracted features lack the awareness of the target and have limited target-background discriminability. To tackle the above iss…

2022

Learning Continuous Graph Structure with Bilevel Programming for Graph Neural Networks

IJCAI 2022poster

Learning graph structure for graph neural networks (GNNs) is crucial to facilitate the GNN-based downstream learning tasks. It is challenging due to the non-differentiable discrete graph structure and lack of ground-truth. In this paper, we address these problems and propose a novel graph structure…

2022

Optimal Positive Generation via Latent Transformation for Contrastive Learning

NeurIPS 2022accept

Contrastive learning, which learns to contrast positive with negative pairs of samples, has been popular for self-supervised visual representation learning. Although great effort has been made to design proper positive pairs through data augmentation, few works attempt to generate optimal positives…

Cited by 10SourcePDFScholar
2022

Salient-to-Broad Transition for Video Person Re-Identification

CVPR 2022poster

Due to the limited utilization of temporal relations in video re-id, the frame-level attention regions of mainstream methods are partial and highly similar. To address this problem, we propose a Salient-to-Broad Module (SBM) to enlarge the attention regions gradually. Specifically, in SBM, while the…

Cited by 69PDFcodeScholar
2021

BiCnet-TKS: Learning Efficient Spatial-Temporal Representation for Video Person Re-Identification

CVPR 2021poster

In this paper, we present an efficient spatial-temporal representation for video person re-identification (reID). Firstly, we propose a Bilateral Complementary Network (BiCnet) for spatial complementarity modeling. Specifically, BiCnet contains two branches. Detail Branch processes frames at origina…

Cited by 126PDFcodeScholar
2020

Appearance-Preserving 3D Convolution for Video-based Person Re-identification

ECCV 2020poster

Due to the imperfect person detection results and posture changes, temporal appearance misalignment is unavoidable in video-based person re-identification (ReID). In this case, 3D convolution may destroy the appearance representation of person video clips, thus it is harmful to ReID. To address this…

2020

Dynamic R-CNN: Towards High Quality Object Detection via Dynamic Training

ECCV 2020poster

Although two-stage object detectors have continuously advanced the state-of-the-art performance in recent years, the training process itself is far from crystal. In this work, we first point out the inconsistency problem between the fixed network settings and the dynamic training procedure, which gr…

2020

TCTS: A Task-Consistent Two-Stage Framework for Person Search

CVPR 2020poster

The state of the art person search methods separate person search into detection and re-ID stages, but ignore the consistency between these two stages. The general person detector has no special attention on the query target; The re-ID model is trained on hand-drawn bounding boxes which are not avai…

Cited by 139PDFScholar
2020

Temporal Complementary Learning for Video Person Re-Identification

ECCV 2020poster

This paper proposes a Temporal Complementary Learning Network that extracts complementary features of consecutive video frames for video person re-identification. Firstly, we introduce a Temporal Saliency Erasing (TSE) module including a saliency erasing operation and a series of ordered learners. S…

2019

Cross Attention Network for Few-shot Classification

NeurIPS 2019poster

Few-shot classification aims to recognize unlabeled samples from unseen classes given only few labeled samples. The unseen classes and low-data problem make few-shot classification very challenging. Many existing approaches extracted features from labeled and unlabeled samples independently, as a re…

2019

Interaction-And-Aggregation Network for Person Re-Identification

CVPR 2019poster

Person re-identification (reID) benefits greatly from deep convolutional neural networks (CNNs) which learn robust feature embeddings. However, CNNs are inherently limited in modeling the large variations in person pose and scale due to their fixed geometric structures. In this paper, we propose a n…

Cited by 466PDFScholar
2019

Temporal Knowledge Propagation for Image-to-Video Person Re-Identification

ICCV 2019poster

In many scenarios of Person Re-identification (Re-ID), the gallery set consists of lots of surveillance videos and the query is just an image, thus Re-ID has to be conducted between image and videos. Compared with videos, still person images lack temporal information. Besides, the information asymme…

Cited by 85PDFcodeScholar
2019

VRSTC: Occlusion-Free Video Person Re-Identification

CVPR 2019poster

Video person re-identification (re-ID) plays an important role in surveillance video analysis. However, the performance of video re-ID degenerates severely under partial occlusion. In this paper, we propose a novel network, called Spatio-Temporal Completion network (STCnet), to explicitly handle par…

Cited by 271PDFScholar