← Search

Henghui Ding

75 accepted papers

2026

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

ICML 2026poster

Recent advances in Omni-Multimodal Large Language Models (Omni-MLLMs) have enabled strong integration of vision, audio, and language. However, their audio-visual intelligence (AVI) remains insufficiently evaluated due to the lack of systematic and comprehensive benchmarks. We introduce AVI-Bench, a …

Cited by 0SourceScholar
2026

AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios

AAAI 2026technical

Referring Multi-Object Tracking (RMOT) aims to achieve precise object detection and tracking through natural language instructions, representing a fundamental capability for intelligent robotic systems. However, current RMOT research remains mostly confined to ground-level scenarios, which constrain

Cited by 0SourcePDFScholar
2026

AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild

ICLR 2026poster

Vision-language navigation (VLN) requires intelligent agents to navigate environments by interpreting linguistic instructions alongside visual observations, serving as a cornerstone task in Embodied AI. Current VLN research for unmanned aerial vehicles (UAVs) relies on detailed, pre-specified instru…

Cited by 0SourceScholar
2026

EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect Erasing

CVPR 2026

Video object removal aims to eliminate dynamic target objects and their visual effects, such as deformation, shadows, and reflections, while restoring seamless backgrounds. Recent diffusion-based video inpainting and object removal methods can remove the objects but often struggle to erase these eff

Cited by 0SourcecodeScholar
2026

Fastcar: Cache Attentive Replay for Fast Auto-Regressive Video Generation on the Edge

ICLR 2026poster

Auto-regressive (AR) models, initially successful in language generation, have recently shown promise in visual generation tasks due to their superior sampling efficiency. Unlike image generation, video generation requires a substantially larger number of tokens to produce coherent temporal frames,…

Cited by 0SourcecodeScholar
2026

Free-Form Scene Editor: Enabling Multi-Round Object Manipulation Like in a 3D Engine

AAAI 2026technical

Recent advances in text-to-image (T2I) diffusion models have significantly improved semantic image editing, yet most methods fall short in performing 3D-aware object manipulation. In this work, we present FFSE, a 3D-aware autoregressive framework designed to enable intuitive, physically-consistent o

Cited by 0SourcePDFScholar
2026

GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering

CVPR 2026

Generating accurate glyphs for visual text rendering is essential yet challenging. Existing methods typically enhance text rendering by training on a large amount of high-quality scene text images, but the limited coverage of glyph variations and excessive stylization often compromise glyph accuracy

Cited by 0SourcecodeScholar
2026

PSDesigner: Automated Graphic Design with a Human-Like Creative Workflow

CVPR 2026

Graphic design is a creative and innovative process that plays a crucial role in applications such as e-commerce and advertising. However, developing an automated design system that can faithfully translate user intentions into editable design files remains an open challenge. Although recent studies

Cited by 0SourcecodeScholar
2026

Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation

ICML 2026poster

Unified multimodal models capable of both understanding and generation have achieved remarkable strides. However, despite their unified designs, existing evaluations typically assess understanding and generation capabilities in isolation, overlooking the synergy between comprehension and generation.…

Cited by 0SourceScholar
2025

CharaConsist: Fine-Grained Consistent Character Generation

ICCV 2025poster

In text-to-image generation, producing a series of consistent contents that preserve the same identity is highly valuable for real-world applications. Although a few works have explored training-free methods to enhance the consistency of generated subjects, we observe that they suffer from the follo…

2025

DViN: Dynamic Visual Routing Network for Weakly Supervised Referring Expression Comprehension

CVPR 2025poster

In this paper, we focus on weakly supervised referring expression comprehension (REC), and identify that the lack of fine-grained visual capability greatly limits the upper performance bound of existing methods. To address this issue, we propose a novel framework for weakly supervised REC, namely Dy…

2025

Exploiting Temporal State Space Sharing for Video Semantic Segmentation

CVPR 2025poster

Video semantic segmentation (VSS) plays a vital role in understanding the temporal evolution of scenes. Traditional methods often segment videos frame-by-frame or in a short temporal window, leading to limited temporal context, redundant computations, and heavy memory requirements. To this end, we i…

2025

Explore In-Context Segmentation via Latent Diffusion Models

AAAI 2025technical

In-context segmentation has drawn increasing attention with the advent of vision foundation models. Its goal is to segment objects using given reference images. Most existing approaches adopt metric learning or masked image modeling to build the correlation between visual prompts and input image que…

Cited by 10SourcePDFScholar
2025

Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation

ICCV 2025poster

Controlling the movements of dynamic objects and the camera within generated videos is a meaningful yet challenging task. Due to the lack of datasets with comprehensive 6D pose annotations, existing text-to-video methods can not simultaneously control the motions of both camera and objects in 3D-awa…

Cited by 0SourcePDFScholar
2025

Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension

AAAI 2025technical

In this work, we address the challenging task of Generalized Referring Expression Comprehension (GREC). Compared to the classic Referring Expression Comprehension (REC) that focuses on single-target expressions, GREC extends the scope to a more practical setting by further encompassing no-target and…

Cited by 1SourcePDFScholar
2025

Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation

ICCV 2025poster

Video instance segmentation (VIS) has gained significant attention for its capability in tracking and segmenting object instances across video frames. However, most of the existing VIS approaches unrealistically assume that the categories of object instances remain fixed over time. Moreover, they ex…

2025

LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers

AAAI 2025technical

Diffusion Transformers have emerged as the preeminent models for a wide array of generative tasks, demonstrating superior performance and efficacy across various applications. The promising results come at the cost of slow inference, as each denoising step requires running the whole transformer mode…

2025

QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the Edge

CVPR 2025poster

Monocular Depth Estimation (MDE) has emerged as a pivotal task in computer vision, supporting numerous real-world applications. However, deploying accurate depth estimation models on resource-limited edge devices, especially Application-Specific Integrated Circuits (ASICs), is challenging due to the…

2025

ReferSplat: Referring Segmentation in 3D Gaussian Splatting

ICML 2025oral

We introduce Referring 3D Gaussian Splatting Segmentation (R3DGS), a new task that aims to segment target objects in a 3D Gaussian scene based on natural language descriptions, which often contain spatial relationships or object attributes. This task requires the model to identify newly described o…

2025

SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models

NeurIPS 2025poster

Achieving fine-grained spatio-temporal understanding in videos remains a major challenge for current Video Large Multimodal Models (Video LMMs). Addressing this challenge requires mastering two core capabilities: video referring understanding, which captures the semantics of video regions, and video…

Cited by 0SourceScholar
2025

SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose Manipulation

NeurIPS 2025spotlight

Controllable image generation has attracted increasing attention in recent years, enabling users to manipulate visual content such as identity and style. However, achieving simultaneous control over the 9D poses (location, size, and orientation) of multiple objects remains an open challenge. Despite…

Cited by 0SourceScholar
2025

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation

ICCV 2025poster

Referring audio-visual segmentation (RAVS) has recently seen significant advancements, yet challenges remain in integrating multimodal information and deeply understanding and reasoning about audiovisual content. To extend the boundaries of RAVS and facilitate future research in this field, we propo…

Cited by 0SourcePDFScholar
2024

Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation

CVPR 2024poster

Referring video segmentation relies on natural language expressions to identify and segment objects often emphasizing motion clues. Previous works treat a sentence as a whole and directly perform identification at the video-level mixing up static image-level cues with temporal motion cues. However i…

2024

Duolando: Follower GPT with Off-Policy Reinforcement Learning for Dance Accompaniment

ICLR 2024poster

We introduce a novel task within the field of human motion generation, termed dance accompaniment, which necessitates the generation of responsive movements from a dance partner, the "follower", synchronized with the lead dancer’s movements and the underlying musical rhythm. Unlike existing solo or…

Cited by 19SourcePDFScholar
2024

How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?

NeurIPS 2024poster

Custom diffusion models (CDMs) have attracted widespread attention due to their astonishing generative ability for personalized concepts. However, most existing CDMs unreasonably assume that personalized concepts are fixed and cannot change over time. Moreover, they heavily suffer from catastrophic…

2024

Mitigating the Curse of Dimensionality for Certified Robustness via Dual Randomized Smoothing

ICLR 2024poster

Randomized Smoothing (RS) has been proven a promising method for endowing an arbitrary image classifier with certified robustness. However, the substantial uncertainty inherent in the high-dimensional isotropic Gaussian noise imposes the curse of dimensionality on RS. Specifically, the upper bound o…

2024

OMG-Seg: Is One Model Good Enough For All Segmentation?

CVPR 2024poster

In this work we address various segmentation tasks each traditionally tackled by distinct or partially unified models. We propose OMG-Seg One Model that is Good enough to efficiently and effectively handle all the segmentation tasks including image semantic instance and panoptic segmentation as well…

2024

PointCVaR: Risk-Optimized Outlier Removal for Robust 3D Point Cloud Classification

AAAI 2024technical

With the growth of 3D sensing technology, the deep learning system for 3D point clouds has become increasingly important, especially in applications such as autonomous vehicles where safety is a primary concern. However, there are growing concerns about the reliability of these systems when they enc…

2024

Referring Image Editing: Object-level Image Editing via Referring Expressions

CVPR 2024poster

Significant advancements have been made in image editing with the recent advance of the Diffusion model. However most of the current methods primarily focus on global or subject-level modifications and often face limitations when it comes to editing specific objects when there are other objects coex…

Cited by 14SourcePDFScholar
2024

Region-Native Visual Tokenization

ECCV 2024poster

"We explore an innovative region-based visual token representation and present the REgion-native AutoencoDER (Reader). In contrast to the majority of previous methods, which represent each image as a grid-shaped tokens map, Reader perceives each image into sequential region-based tokens, with each t…

2024

SemFlow: Binding Semantic Segmentation and Image Synthesis via Rectified Flow

NeurIPS 2024poster

Semantic segmentation and semantic image synthesis are two representative tasks in visual perception and generation. While existing methods consider them as two distinct tasks, we propose a unified framework (SemFlow) and model them as a pair of reverse problems. Specifically, motivated by rectified…

2024

Transferable Adversarial Attacks on SAM and Its Downstream Models

NeurIPS 2024poster

The utilization of large foundational models has a dilemma: while fine-tuning downstream tasks from them holds promise for making use of the well-generalized knowledge in practical applications, their open accessibility also poses threats of adverse usage. This paper, for the first time, explores th…

2023

Betrayed by Captions: Joint Caption Grounding and Generation for Open Vocabulary Instance Segmentation

ICCV 2023poster

In this work, we focus on open vocabulary instance segmentation to expand a segmentation model to classify and segment instance-level novel categories. Previous approaches have relied on massive caption datasets and complex pipelines to establish one-to-one mappings between image regions and words i…

Cited by 36PDFcodeScholar
2023

Decoupling with Entropy-based Equalization for Semi-Supervised Semantic Segmentation

IJCAI 2023poster

Semi-supervised semantic segmentation methods are the main solution to alleviate the problem of high annotation consumption in semantic segmentation. However, the class imbalance problem makes the model favor the head classes with sufficient training samples, resulting in poor performance of the tai…

Cited by 3SourcePDFScholar
2023

Deep Geometrized Cartoon Line Inbetweening

ICCV 2023poster

We aim to address a significant but understudied problem in the anime industry, namely the inbetweening of cartoon line drawings. Inbetweening involves generating intermediate frames between two black-and-white line drawings and is a time-consuming and expensive process that can benefit from automat…

Cited by 17PDFcodeScholar
2023

Federated Incremental Semantic Segmentation

CVPR 2023poster

Federated learning-based semantic segmentation (FSS) has drawn widespread attention via decentralized training on local clients. However, most FSS models assume categories are fxed in advance, thus heavily undergoing forgetting on old categories in practical applications where local clients receive…

2023

Global Knowledge Calibration for Fast Open-Vocabulary Segmentation

ICCV 2023poster

Recent advancements in pre-trained vision-language models, such as CLIP, have enabled the segmentation of arbitrary concepts solely from textual inputs, a process commonly referred to as open-vocabulary semantic segmentation (OVS). However, existing OVS techniques confront a fundamental challenge: t…

Cited by 55PDFScholar
2023

MOSE: A New Dataset for Video Object Segmentation in Complex Scenes

ICCV 2023poster

Video object segmentation (VOS) aims at segmenting a particular object throughout the entire video clip sequence. The state-of-the-art VOS methods have achieved excellent performance (e.g., 90+% J&F) on existing datasets. However, since the target objects in these existing datasets are usually relat…

Cited by 148PDFcodeScholar
2023

Mask-Free Video Instance Segmentation

CVPR 2023poster

The recent advancement in Video Instance Segmentation (VIS) has largely been driven by the use of deeper and increasingly data-hungry transformer-based models. However, video masks are tedious and expensive to annotate, limiting the scale and diversity of existing VIS datasets. In this work, we aim…

2023

MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions

ICCV 2023poster

This paper strives for motion expressions guided video segmentation, which focuses on segmenting objects in video content based on a sentence describing the motion of the objects. Existing referring video object datasets typically focus on salient objects and use language expressions that contain ex…

Cited by 110PDFcodeScholar
2023

OVTrack: Open-Vocabulary Multiple Object Tracking

CVPR 2023poster

The ability to recognize, localize and track dynamic objects in a scene is fundamental to many real-world applications, such as self-driving and robotic systems. Yet, traditional multiple object tracking (MOT) benchmarks rely only on a few object categories that hardly represent the multitude of pos…

Cited by 66SourcePDFScholar
2023

Primitive Generation and Semantic-Related Alignment for Universal Zero-Shot Segmentation

CVPR 2023poster

We study universal zero-shot segmentation in this work to achieve panoptic, instance, and semantic segmentation for novel categories without any training samples. Such zero-shot segmentation ability relies on inter-class relationships in semantic space to transfer the visual knowledge learned from s…

2023

SegRefiner: Towards Model-Agnostic Segmentation Refinement with Discrete Diffusion Process

NeurIPS 2023poster

In this paper, we explore a principal way to enhance the quality of object masks produced by different segmentation models. We propose a model-agnostic solution called SegRefiner, which offers a novel perspective on this problem by interpreting segmentation refinement as a data generation process. A…

2023

Semantic-Promoted Debiasing and Background Disambiguation for Zero-Shot Instance Segmentation

CVPR 2023poster

Zero-shot instance segmentation aims to detect and precisely segment objects of unseen categories without any training samples. Since the model is trained on seen categories, there is a strong bias that the model tends to classify all the objects into seen categories. Besides, there is a natural con…

2022

Coarse-To-Fine Feature Mining for Video Semantic Segmentation

CVPR 2022poster

The contextual information plays a core role in semantic segmentation. As for video semantic segmentation, the contexts include static contexts and motional contexts, corresponding to static content and moving content in a video clip, respectively. The static contexts are well exploited in image sem…

Cited by 75PDFcodeScholar
2022

Degradation-Aware Unfolding Half-Shuffle Transformer for Spectral Compressive Imaging

NeurIPS 2022accept

In coded aperture snapshot spectral compressive imaging (CASSI) systems, hyperspectral image (HSI) reconstruction methods are employed to recover the spatial-spectral signal from a compressed measurement. Among these algorithms, deep unfolding methods demonstrate promising performance but suffer fro…

2022

Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning

NeurIPS 2022accept

Vision transformers have recently achieved competitive results across various vision tasks but still suffer from heavy computation costs when processing a large number of tokens. Many advanced approaches have been developed to reduce the total number of tokens in the large-scale vision transformers,…

Cited by 32SourcePDFScholar
2022

Flow-Guided Sparse Transformer for Video Deblurring

ICML 2022spotlight

Exploiting similar and sharper scene patches in spatio-temporal neighborhoods is critical for video deblurring. However, CNN-based methods show limitations in capturing long-range dependencies and modeling non-local self-similarity. In this paper, we propose a novel framework, Flow-Guided Sparse Tra…

2022

Learning Transferable Human-Object Interaction Detector With Natural Language Supervision

CVPR 2022poster

It is difficult to construct a data collection including all possible combinations of human actions and interacting objects due to the combinatorial nature of human-object interactions (HOI). In this work, we aim to develop a transferable HOI detector for unseen interactions. Existing HOI detectors…

Cited by 66PDFcodeScholar
2022

Primitive3D: 3D Object Dataset Synthesis From Randomly Assembled Primitives

CVPR 2022poster

Numerous advancements of deep learning can be attributed to access to large-scale and well-annotated datasets. However, such a dataset is prohibitively expensive in 3D computer vision due to the substantial collection cost. To alleviate this issue, we propose a cost-effective method for automaticall…

Cited by 5PDFScholar
2022

Video Mask Transfiner for High-Quality Video Instance Segmentation

ECCV 2022poster

"While Video Instance Segmentation (VIS) has seen rapid progress, current approaches struggle to predict high-quality masks with accurate boundary details. Moreover, the predicted segmentations often fluctuate over time, suggesting that temporal consistency cues are neglected or not fully utilized.…

Cited by 38SourcePDFScholar
2021

A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto-Encoder

ICCV 2021poster

We present a unified and flexible framework to address the generalized problem of 3D motion synthesis that covers the tasks of motion prediction, completion, interpolation, and spatial-temporal recovery. Since these tasks have different input constraints and various fidelity and diversity requiremen…

Cited by 80PDFScholar
2021

Adaptive Data Augmentation on Temporal Graphs

NeurIPS 2021poster

Temporal Graph Networks (TGNs) are powerful on modeling temporal graph data based on their increased complexity. Higher complexity carries with it a higher risk of overfitting, which makes TGNs capture random noise instead of essential semantic information. To address this issue, our idea is to tran…

Cited by 66SourcePDFScholar
2021

Directed Graph Contrastive Learning

NeurIPS 2021poster

Graph Contrastive Learning (GCL) has emerged to learn generalizable representations from contrastive views. However, it is still in its infancy with two concerns: 1) changing the graph structure through data augmentation to generate contrastive views may mislead the message passing scheme, as such g…

2021

Discovering Human Interactions With Large-Vocabulary Objects via Query and Multi-Scale Detection

ICCV 2021poster

In this work, we study the problem of human-object interaction (HOI) detection with large vocabulary object categories. Previous HOI studies are mainly conducted in the regime of limit object categories (e.g., 80 categories). Their solutions may face new difficulties in both object detection and int…

Cited by 32PDFScholar
2021

Else-Net: Elastic Semantic Network for Continual Action Recognition From Skeleton Data

ICCV 2021poster

We address continual action recognition from skeleton sequence, which aims to learn a recognition model over time from a continuous stream of skeleton data. This task is very important in changing environment. Due to catastrophic forgetting problems of deep neural networks and large discrepancies be…

Cited by 56PDFScholar
2021

Interaction via Bi-Directional Graph of Semantic Region Affinity for Scene Parsing

ICCV 2021poster

In this work, we devote to address the challenging problem of scene parsing. Previous methods, though capture context to exploit global clues, handle scene parsing as a pixel-independent task. However, it is well known that pixels in an image are highly correlated with each other, especially those f…

Cited by 19PDFScholar
2021

MINE: Towards Continuous Depth MPI With NeRF for Novel View Synthesis

ICCV 2021poster

In this paper, we propose MINE to perform novel view synthesis and depth estimation via dense 3D reconstruction from a single image. Our approach is a continuous depth generalization of the Multiplane Images (MPI) by introducing the NEural radiance fields (NeRF). Given a single image as input, MINE…

Cited by 171PDFcodeScholar
2021

Meta Navigator: Search for a Good Adaptation Policy for Few-Shot Learning

ICCV 2021poster

Few-shot learning aims to adapt knowledge learned from previous tasks to novel tasks with only a limited amount of labeled data. Research literature on few-shot learning exhibits great diversity, while different algorithms often excel at different few-shot learning scenarios. It is therefore tricky…

Cited by 60PDFScholar
2021

Vision-Language Transformer and Query Generation for Referring Segmentation

ICCV 2021poster

In this work, we address the challenging task of referring segmentation. The query expression in referring segmentation typically indicates the target object by describing its relationship with others. Therefore, to find the target one among all instances in the image, the model must have a holistic…

Cited by 297PDFcodeScholar
2020

PhraseClick: Toward Achieving Flexible Interactive Segmentation by Phrase and Click

ECCV 2020poster

Existing interactive object segmentation methods mainly take spatial interactions such as bounding boxes or clicks as input. However, these interactions do not contain information about explicit attributes of the target-of-interest and thus cannot quickly specify what the selected object exactly is,…

Cited by 64SourcePDFScholar
2019

Boundary-Aware Feature Propagation for Scene Segmentation

ICCV 2019poster

In this work, we address the challenging issue of scene segmentation. To increase the feature similarity of the same object while keeping the feature discrimination of different objects, we explore to propagate information throughout the image under the control of objects' boundaries. To this end, w…

Cited by 286PDFcodeScholar
2019

Semantic Correlation Promoted Shape-Variant Context for Segmentation

CVPR 2019oral

Context is essential for semantic segmentation. Due to the diverse shapes of objects and their complex layout in various scene images, the spatial scales and shapes of contexts for different objects have very large variation. It is thus ineffective or inefficient to aggregate various context informa…

Cited by 215PDFcodeScholar
2018

Context Contrasted Feature and Gated Multi-Scale Aggregation for Scene Segmentation

CVPR 2018poster

Scene segmentation is a challenging task as it need label every pixel in the image. It is crucial to exploit discriminative context and aggregate multi-scale features to achieve better segmentation. In this paper, we first propose a novel context contrasted local feature that not only leverages the…

Cited by 450SourcePDFScholar