← Search

Zhibo Chen

66 accepted papers

2026

4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models

CVPR 2026

World Generation Models are emerging as a cornerstone of next-generation multimodal intelligence systems. Unlike traditional 2D visual generation, World Models aim to construct realistic, dynamic, and physically consistent 3D/4D worlds from images, videos, or text. These models not only need to prod

Cited by 0SourcecodeScholar
2026

AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation

ICML 2026poster

Vision-Language Navigation (VLN) requires agents to follow natural language instructions by grounding them in sequential visual observations over long horizons. Explicit reasoning could enhance temporal consistency and perception–action alignment, but reasoning at fixed steps often leads to suboptim…

Cited by 0SourceScholar
2026

ColorFLUX: A Structure-Color Decoupling Framework for Old Photo Colorization

CVPR 2026

Old photos preserve invaluable historical memories, making their restoration and colorization highly desirable. While existing restoration models can address some degradation issues like denoising and scratch removal, they often struggle with accurate colorization.This limitation arises from the uni

Cited by 0SourcecodeScholar
2026

Embodied Navigation Foundation Model

ICLR 2026poster

Navigation is a fundamental capability in embodied AI, representing the intelligence required to perceive and interact within physical environments. To achieve such intelligence, recent advanced works leverage Vision-Language Models (VLMs), which demonstrate strong generalizability and possess a wel…

Cited by 0SourcecodeScholar
2026

IQA-Spider: Unifying Reasoning, Grounding, and Referring for Multi-Granularity Image Quality Assessment

ICML 2026poster

We present IQA-Spider, the first image quality assessment (IQA) framework that unifies reasoning, grounding, and referring within a LMM-based system for multi-granularity quality understanding. Existing LMM-based IQA methods typically support only partial perception dimensions, \egno, quality descri…

Cited by 0SourceScholar
2026

Sonic4D: Spatial Audio Generation for Immersive 4D Scene Exploration

AAAI 2026technical

Recent advancements in 4D generation have demonstrated its remarkable capability in synthesizing photorealistic renderings of dynamic 3D scenes. However, despite achieving impressive visual performance, almost all existing methods overlook the generation of spatial audio aligned with the correspondi

Cited by 0SourcePDFScholar
2026

Test-Time Preference Optimization for Image Restoration

AAAI 2026technical

Image restoration (IR) models are typically trained to recover high-quality images using L1 or LPIPS loss. To handle diverse unknown degradations, zero-shot IR methods have also been introduced. However, existing pre-trained and zero-shot IR approaches often fail to align with human preferences, res

Cited by 0SourcePDFScholar
2026

TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking

ICRA 2026poster

Embodied Visual Tracking (EVT) is a fundamental ability that underpins practical applications, such as companion robots, guidance robots and service assistants, where continuously following moving targets is essential. Recent advances have enabled language-guided tracking in complex and unstructured…

2026

Training-Free Sparse Attention for Fast Video Generation via Offline Layer-Wise Sparsity Profiling and Online Bidirectional Co-Clustering

ICML 2026poster

Diffusion Transformers (DiTs) achieve strong video generation quality but suffer from high inference cost due to dense 3D attention, leading to the development of sparse attention technologies to improve efficiency. However, existing training-free sparse attention methods in video generation still f…

Cited by 0SourceScholar
2026

UrbanVLA: A Vision-Language-Action Model for Urban Micromobility

ICRA 2026poster

Urban micromobility applications, such as delivery robots, demand reliable navigation across large-scale urban environments while following long-horizon route instructions. This task is particularly challenging due to the dynamic and unstructured nature of real-world city areas, yet most existing na…

2026

VisualScore: Learning Holistic Visual Quality Scores via Multi-Task Reasoning

ICML 2026poster

Image quality assessment (IQA) is inherently multi-mage quality assessment (IQA) is inherently multi-dimensional, yet existing reward models are typically limited to a single task and become unstable when extended to multi-task settings. In particular, heterogeneous reward scales and variances acros…

Cited by 0SourceScholar
2025

Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative Prior

NeurIPS 2025poster

Image compression methods are usually optimized isolatedly for human perception or machine analysis tasks. We reveal fundamental commonalities between these objectives: preserving accurate semantic information is paramount, as it directly dictates the integrity of critical information for intelligen…

Cited by 0SourceScholar
2025

Na Vid-4D: Unleashing Spatial Intelligence in Egocentric RGB-D Videos for Vision-and-Language Navigation

ICRA 2025

Understanding and reasoning about the 4D space-time is crucial for Vision-and-Language Navigation (VLN). However, previous works lack in-depth exploration in this aspect, resulting in bottlenecked spatial perception and action precision of VLN agents. In this work, we introduce NaVid-4D, a Vision La

Cited by 4SourceScholar
2025

QMamba: On First Exploration of Vision Mamba for Image Quality Assessment

ICML 2025poster

In this work, we take the first exploration of the recently popular foundation model, *i.e.,* State Space Model/Mamba, in image quality assessment (IQA), aiming at observing and excavating the perception potential in vision Mamba. A series of works on Mamba has shown its significant potential in va…

2025

RAGD: Regional-Aware Diffusion Model for Text-to-Image Generation

ICCV 2025poster

Regional prompting, or compositional generation, which enables fine-grained spatial control, has gained increasing attention for its practicality in real-world applications. However, previous methods either introduce additional trainable modules, thus only applicable to specific models, or manipulat…

2025

RaSS: Improving Denoising Diffusion Samplers with Reinforced Active Sampling Scheduler

CVPR 2025poster

Recent years have witnessed the great success of denoising diffusion samplers in improving the generative capability and sampling efficiency given a pre-trained diffusion model. However, most sampling schedulers in diffusion models lack the sampling dynamics and planning capability for future genera…

Cited by 0SourcePDFScholar
2025

StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition

ICCV 2025poster

With the rise of real-world human-AI interaction applications, such as AI assistants, the need for Streaming Video Dialogue is critical. To address this need, we introduce StreamMind, a video LLM framework that achieves ultra-FPS streaming video processing (100 fps on a single A100) and enables proa…

Cited by 0SourcePDFScholar
2025

TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation

AAAI 2025technical

Text-driven Image to Video Generation (TI2V) aims to generate controllable video given the first frame and corresponding textual description. The primary challenges of this task lie in two parts: (i) how to identify the target objects and ensure the consistency between the movement trajectory and th…

Cited by 1SourcePDFScholar
2024

CCEdit: Creative and Controllable Video Editing via Diffusion Models

CVPR 2024poster

In this paper we present CCEdit a versatile generative video editing framework based on diffusion models. Our approach employs a novel trident network structure that separates structure and appearance control ensuring precise and creative editing capabilities. Utilizing the foundational ControlNet a…

Cited by 43SourcePDFScholar
2024

Compositional 3D-aware Video Generation with LLM Director

NeurIPS 2024poster

Significant progress has been made in text-to-video generation through the use of powerful generative models and large-scale internet data. However, substantial challenges remain in precisely controlling individual elements within the generated video, such as the movement and appearance of specific…

Cited by 10SourcePDFScholar
2024

End-to-End Rate-Distortion Optimized 3D Gaussian Representation

ECCV 2024poster

"3D Gaussian Splatting (3DGS) has become an emerging technique with remarkable potential in 3D representation and image rendering. However, the substantial storage overhead of 3DGS significantly impedes its practical applications. In this work, we formulate the compact 3D Gaussian learning as an end…

2024

Is Vanilla MLP in Neural Radiance Field Enough for Few-shot View Synthesis?

CVPR 2024poster

Neural Radiance Field (NeRF) has achieved superior performance for novel view synthesis by modeling the scene with a Multi-Layer Perception (MLP) and a volume rendering procedure however when fewer known views are given (i.e. few-shot view synthesis) the model is prone to overfit the given views. To…

Cited by 15SourcePDFScholar
2024

KVQ: Kwai Video Quality Assessment for Short-form Videos

CVPR 2024poster

Short-form UGC video platforms like Kwai and TikTok have been an emerging and irreplaceable mainstream media form thriving on user-friendly engagement and kaleidoscope creation etc. However the advancing content generation modes e.g. special effects and sophisticated processing workflows e.g. de-art…

2024

Learn from Failure: Fine-tuning LLMs with Trial-and-Error Data for Intuitionistic Propositional Logic Proving

ACL 2024long

Recent advances in Automated Theorem Proving have shown the effectiveness of leveraging a (large) language model that generates tactics (i.e. proof steps) to search through proof states. The current model, while trained solely on successful proof paths, faces a discrepancy at the inference stage, as…

2024

Rate-Distortion-Cognition Controllable Versatile Neural Image Compression

ECCV 2024poster

"Recently, the field of Image Coding for Machines (ICM) has garnered heightened interest and significant advances thanks to the rapid progress of learning-based techniques for image compression and analysis. Previous studies often require training separate codecs to support various bitrate levels, m…

Cited by 4SourcePDFScholar
2024

SeD: Semantic-Aware Discriminator for Image Super-Resolution

CVPR 2024poster

Generative Adversarial Networks (GANs) have been widely used to recover vivid textures in image super-resolution (SR) tasks. In particular one discriminator is utilized to enable the SR network to learn the distribution of real-world high-quality images in an adversarial training manner. However the…

2024

UCIP: A Universal Framework for Compressed Image Super-Resolution using Dynamic Prompt

ECCV 2024poster

"Compressed Image Super-resolution (CSR) aims to simultaneously super-resolve the compressed images and tackle the challenging hybrid distortions caused by compression. However, existing works on CSR usually focus on single compression codec, , JPEG, ignoring the diverse traditional or learning-base…

2023

Compression with Bayesian Implicit Neural Representations

NeurIPS 2023spotlight

Many common types of data can be represented as functions that map coordinates to signal values, such as pixel locations to RGB values in the case of an image. Based on this view, data can be compressed by overfitting a compact neural network to its functional representation and then encoding the ne…

2023

Deep Frequency Filtering for Domain Generalization

CVPR 2023poster

Improving the generalization ability of Deep Neural Networks (DNNs) is critical for their practical uses, which has been a longstanding challenge. Some theoretical studies have uncovered that DNNs have preferences for some frequency components in the learning process and indicated that this may affe…

Cited by 63SourcePDFScholar
2023

GraphAdapter: Tuning Vision-Language Models With Dual Knowledge Graph

NeurIPS 2023poster

Adapter-style efficient transfer learning (ETL) has shown excellent performance in the tuning of vision-language models (VLMs) under the low-data regime, where only a few additional parameters are introduced to excavate the task-specific knowledge based on the general and powerful representation of…

2023

Learning Distortion Invariant Representation for Image Restoration From a Causality Perspective

CVPR 2023poster

In recent years, we have witnessed the great advancement of Deep neural networks (DNNs) in image restoration. However, a critical limitation is that they cannot generalize well to real-world degradations with different degrees or types. In this paper, we are the first to propose a novel training str…

2023

SelectAugment: Hierarchical Deterministic Sample Selection for Data Augmentation

AAAI 2023technical

Data augmentation (DA) has been extensively studied to facilitate model optimization in many tasks. Prior DA works focus on designing augmentation operations themselves, while leaving selecting suitable samples for augmentation out of consideration. This might incur visual ambiguities and further in…

Cited by 7SourcePDFScholar
2023

Semantically Structured Image Compression via Irregular Group-Based Decoupling

ICCV 2023poster

Image compression techniques typically focus on compressing rectangular images for human consumption, however, resulting in transmitting redundant content for downstream applications. To overcome this limitation, some previous works propose to semantically structure the bitstream, which can meet spe…

Cited by 13PDFScholar
2023

Versatile Neural Processes for Learning Implicit Neural Representations

ICLR 2023poster

Representing a signal as a continuous function parameterized by neural network (a.k.a. Implicit Neural Representations, INRs) has attracted increasing attention in recent years. Neural Processes (NPs), which model the distributions over functions conditioned on partial observations (context set), pr…

2022

Cloth-Changing Person Re-Identification From a Single Image With Gait Prediction and Regularization

CVPR 2022poster

Cloth-Changing person re-identification (CC-ReID) aims at matching the same person across different locations over a long-duration, e.g., over days, and therefore inevitably has cases of changing clothing. In this paper, we focus on handling well the CC-ReID problem under a more challenging setting,…

Cited by 179PDFcodeScholar
2022

Image Coding for Machines with Omnipotent Feature Learning

ECCV 2022poster

"Image Coding for Machines (ICM) aims to compress images for AI tasks analysis rather than meeting human perception. Learning a kind of feature that is both general (for AI tasks) and compact (for compression) is pivotal for its success. In this paper, we attempt to develop an ICM framework by learn…

2022

Mask-based Latent Reconstruction for Reinforcement Learning

NeurIPS 2022accept

For deep reinforcement learning (RL) from pixels, learning effective state representations is crucial for achieving high performance. However, in practice, limited experience and high-dimensional inputs prevent effective representation learning. To address this, motivated by the success of mask-base…

2021

Dense Interaction Learning for Video-Based Person Re-Identification

ICCV 2021poster

Video-based person re-identification (re-ID) aims at matching the same person across video clips. Efficiently exploiting multi-scale fine-grained features while building the structural interaction among them is pivotal for its success. In this paper, we propose a hybrid framework, Dense Interaction…

Cited by 67PDFcodeScholar
2021

Learning Omni-Frequency Region-adaptive Representations for Real Image Super-Resolution

AAAI 2021technical

Traditional single image super-resolution (SISR) methods that focus on solving single and uniform degradation (i.e., bicubic down-sampling), typically suffer from poor performance when applied into real-world low-resolution (LR) images due to the complicated realistic degradations. The key to solvin…

Cited by 44SourcePDFScholar
2021

MetaAlign: Coordinating Domain Alignment and Classification for Unsupervised Domain Adaptation

CVPR 2021poster

For unsupervised domain adaptation (UDA), to alleviate the effect of domain shift, many approaches align the source and target domains in the feature space by adversarial learning or by explicitly aligning their statistics. However, the optimization objective of such domain alignment is generally no…

Cited by 137PDFScholar
2021

Partial Person Re-Identification With Part-Part Correspondence Learning

CVPR 2021poster

Driven by the success of deep learning, the last decade has seen rapid advances in person re-identification (re-ID). Nonetheless, most of approaches assume that the input is given with the fulfillment of expectations, while imperfect input remains rarely explored to date, which is a non-trivial prob…

Cited by 52PDFScholar
2021

PlayVirtual: Augmenting Cycle-Consistent Virtual Trajectories for Reinforcement Learning

NeurIPS 2021poster

Learning good feature representations is important for deep reinforcement learning (RL). However, with limited experience, RL often suffers from data inefficiency for training. For un-experienced or less-experienced trajectories (i.e., state-action sequences), the lack of data limits the use of them…

2021

Re-Energizing Domain Discriminator With Sample Relabeling for Adversarial Domain Adaptation

ICCV 2021poster

Many unsupervised domain adaptation (UDA) methods exploit domain adversarial training to align the features to reduce domain gap, where a feature extractor is trained to fool a domain discriminator in order to have aligned feature distributions. The discrimination capability of the domain classifier…

Cited by 18PDFScholar
2021

Soft then Hard: Rethinking the Quantization in Neural Image Compression

ICML 2021spotlight

Quantization is one of the core components in lossy image compression. For neural image compression, end-to-end optimization requires differentiable approximations of quantization, which can generally be grouped into three categories: additive uniform noise, straight-through estimator and soft-to-ha…

Cited by 92SourcePDFScholar
2021

ToAlign: Task-Oriented Alignment for Unsupervised Domain Adaptation

NeurIPS 2021poster

Unsupervised domain adaptive classifcation intends to improve the classifcation performance on unlabeled target domain. To alleviate the adverse effect of domain shift, many approaches align the source and target domains in the feature space. However, a feature is usually taken as a whole for alignm…

2021

Uncertainty-Aware Few-Shot Image Classification

IJCAI 2021poster

Few-shot image classification learns to recognize new categories from limited labelled data. Metric learning based approaches have been widely investigated, where a query sample is classified by finding the nearest prototype from the support set based on their feature similarities. A neural network…

Cited by 30SourcePDFScholar
2020

Global Distance-distributions Separation for Unsupervised Person Re-identification

ECCV 2020poster

Supervised person re-identification (ReID) often has poor scalability and usability in real-world deployments due to domain gaps and the lack of annotations for the target domain data. Unsupervised person ReID through domain adaptation is attractive yet challenging. Existing unsupervised ReID approa…

Cited by 90SourcePDFScholar
2020

LIRA: Lifelong Image Restoration from Unknown Blended Distortions

ECCV 2020poster

Most existing image restoration networks are designed in a disposable way and catastrophically forget previously learned distortions when trained on a new distortion removal task. To alleviate this problem, we raise the novel lifelong image restoration problem for blended distortions. We first desig…

Cited by 20SourcePDFScholar
2020

Learning Disentangled Feature Representation for Hybrid-distorted Image Restoration

ECCV 2020poster

Hybrid-distorted image restoration (HD-IR) is dedicated to restore real distorted image that is degraded by multiple distortions. Existing HD-IR approaches usually ignore the inherent interference among hybrid distortions which compromises the restoration performance. To decompose such interference,…

Cited by 55SourcePDFScholar
2020

Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-Based Person Re-Identification

CVPR 2020poster

Video-based person re-identification (reID) aims at matching the same person across video clips. It is a challenging task due to the existence of redundancy among frames, newly revealed appearance, occlusion, and motion blurs. In this paper, we propose an attentive feature aggregation module, namely…

Cited by 138PDFScholar
2020

Relation-Aware Global Attention for Person Re-Identification

CVPR 2020poster

For person re-identification (re-id), attention mechanisms have become attractive as they aim at strengthening discriminative features and suppressing irrelevant ones, which matches well the key of re-id, i.e., discriminative feature learning. Previous approaches typically learn attention using loca…

Cited by 713PDFcodeScholar
2020

Style Normalization and Restitution for Generalizable Person Re-Identification

CVPR 2020poster

Existing fully-supervised person re-identification (ReID) methods usually suffer from poor generalization capability caused by domain gaps. The key to solving this problem lies in filtering out identity-irrelevant interference and learning domain-invariant person representations. In this paper, we a…

Cited by 448PDFcodeScholar
2020

TuiGAN: Learning Versatile Image-to-Image Translation with Two Unpaired Images

ECCV 2020poster

An unsupervised image-to-image translation (UI2I) task deals with learning a mapping between two domains without paired images. While existing UI2I methods usually require numerous unpaired images from different domains for training, there are many scenarios where training data is quite limited. In…

2019

Towards Accurate One-Stage Object Detection With AP-Loss

CVPR 2019poster

One-stage object detectors are trained by optimizing classification-loss and localization-loss simultaneously, with the former suffering much from extreme foreground-background class imbalance issue due to the large number of anchors. This paper alleviates this issue by proposing a novel framework t…

Cited by 173PDFcodeScholar
2018

Layer-Wise Coordination between Encoder and Decoder for Neural Machine Translation

NeurIPS 2018poster

Neural Machine Translation (NMT) has achieved remarkable progress with the quick evolvement of model structures. In this paper, we propose the concept of layer-wise coordination for NMT, which explicitly coordinates the learning of hidden representations of the encoder and decoder together layer by…