← Search

Liang Zheng

70 accepted papers

2026

Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls

ICML 2026poster

The ability to use tools is fundamental for large language model (LLM) agents. Given a task, existing systems use LLMs to plan and generate tool calls, which are executed by real-world tools to complete the task. However, tool calls are prone to errors because they are derived merely from LLM intrin…

Cited by 0SourceScholar
2026

LeapAlign: Post-training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories

CVPR 2026

This paper focuses on the alignment of flow-matching models with human preference. A promising way is fine-tuning by directly backpropagating reward signals through the differentiable generation process of flow matching. However, backpropagating through long trajectories results in prohibitive memor

Cited by 0SourcecodeScholar
2026

Mosaic: Unlocking Over 30$\times$ Context Length for Diffusion LLMs Inference via Global Memory Planning and Dynamic Peak Taming

ICML 2026poster

Diffusion-based large language models (dLLMs) have emerged as a promising alternative to autoregressive models, leveraging simultaneous denoising to enable global planning and iterative refinement. These properties make dLLMs particularly attractive for long-context generation. However, deploying dL…

Cited by 0SourceScholar
2026

What matters for Representation Alignment: Global Information or Spatial Structure?

ICLR 2026poster

Representation alignment helps generation by distilling representations from a pretrained vision encoder to intermediate diffusion features. We investigate a fundamental question - `what aspect of the target representation matters for generation, its global information (measured by Imagenet1K accura…

Cited by 0SourcecodeScholar
2025

Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization

CVPR 2025poster

Generating visually appealing images is fundamental to modern text-to-image generation models. A potential solution to better aesthetics is direct preference optimization (DPO), which has been applied to diffusion models to improve general image quality including prompt alignment and aesthetics. Pop…

2025

Can We Predict Performance of Large Models across Vision-Language Tasks?

ICML 2025poster

Evaluating large vision-language models (LVLMs) is very expensive, due to high computational cost and the wide variety of tasks. The good news is that if we already have some observed performance scores, we may be able to infer unknown ones. In this study, we propose a new framework for predicting u…

2025

DDPA-3DVG: Vision-Language Dual-Decoupling and Progressive Alignment for 3D Visual Grounding

IJCAI 2025

3D visual grounding aims to localize target objects in point clouds based on free-form natural language, which often describes both target and reference objects. Effective alignment between visual and text features is crucial for this task. However, existing two-stage methods that rely solely on obj

2025

Effective Training Data Synthesis for Improving MLLM Chart Understanding

ICCV 2025poster

Being able to effectively read scientific plots, or chart understanding, is a central part toward building effective agents for science. However, existing multimodal large language models (MLLMs), especially open-source ones, are still falling behind with a typical success rate of 30%-50% on challen…

2025

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation

NeurIPS 2025spotlight

This paper presents JavisGPT, the first unified multimodal large language model (MLLM) for Joint Audio-Video (JAV) comprehension and generation. JavisGPT adopts a concise encoder–LLM–decoder architecture, featuring a SyncFusion module for spatio-temporal audio- video fusion and synchrony-aware learn…

Cited by 0SourceScholar
2025

REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion Transformers

ICCV 2025poster

In this paper we tackle a fundamental question: "Can we train latent diffusion models together with the variational auto-encoder (VAE) tokenizer in an end-to-end manner?" Traditional deep-learning wisdom dictates that end-to-end training is often preferable when possible. However, for latent diffusi…

2025

Think on Your Feet: Seamless Transition Between Human-Like Locomotion in Response to Changing Commands

ICRA 2025

While it is relatively easier to train humanoid robots to mimic specific locomotion skills, it is more challenging to learn from various motions and adhere to continuously changing commands. These robots must accurately track motion instructions, seamlessly transition between a variety of movements,

Cited by 2SourceScholar
2025

Vec2Face: Scaling Face Dataset Generation with Loosely Constrained Vectors

ICLR 2025poster

This paper studies how to synthesize face images of non-existent persons, to create a dataset that allows effective training of face recognition (FR) models. Besides generating realistic face images, two other important goals are: 1) the ability to generate a large number of distinct identities (int…

Cited by 4SourcePDFScholar
2024

Adapting Humanoid Locomotion over Challenging Terrain via Two-Phase Training

CoRL 2024poster

Humanoid robots are a key focus in robotics, with their capacity to navigate tough terrains being essential for many uses. While strides have been made, creating adaptable locomotion for complex environments is still tough. Recent progress in learning-based systems offers hope for robust legged loco…

Cited by 4SourceScholar
2024

Alice Benchmarks: Connecting Real World Re-Identification with the Synthetic

ICLR 2024poster

For object re-identification (re-ID), learning from synthetic data has become a promising strategy to cheaply acquire large-scale annotated datasets and effective models, with few privacy concerns. Many interesting research problems arise from this strategy, e.g., how to reduce the domain gap betwee…

Cited by 0SourcePDFScholar
2024

Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments

ICLR 2024spotlight

Bounding boxes uniquely characterize object detection, where a good detector gives accurate bounding boxes of categories of interest. However, in the real-world where test ground truths are not provided, it is non-trivial to find out whether bounding boxes are accurate, thus preventing us from asses…

2024

CIFAR-10-Warehouse: Broad and More Realistic Testbeds in Model Generalization Analysis

ICLR 2024poster

Analyzing model performance in various unseen environments is a critical research problem in the machine learning community. To study this problem, it is important to construct a testbed with out-of-distribution test sets that have broad coverage of environmental discrepancies. However, existing tes…

Cited by 7SourcePDFScholar
2024

SmartMask: Context Aware High-Fidelity Mask Generation for Fine-grained Object Insertion and Layout Control

CVPR 2024poster

The field of generative image inpainting and object insertion has made significant progress with the recent advent of latent diffusion models. Utilizing a precise object mask can greatly enhance these applications. However due to the challenges users encounter in creating high-fidelity masks there i…

Cited by 9SourcePDFScholar
2024

The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision-Language Models?

ECCV 2024poster

"Large vision-language models (LVLMs), designed to interpret and respond to human instructions, occasionally generate hallucinated or harmful content due to inappropriate instructions. This study uses linear probing to shed light on the hidden knowledge at the output layers of LVLMs. We demonstrate…

2024

Towards Optimal Feature-Shaping Methods for Out-of-Distribution Detection

ICLR 2024poster

Feature shaping refers to a family of methods that exhibit state-of-the-art performance for out-of-distribution (OOD) detection. These approaches manipulate the feature representation, typically from the penultimate layer of a pre-trained deep learning model, so as to better differentiate between in…

2023

A Bag-of-Prototypes Representation for Dataset-Level Applications

CVPR 2023poster

This work investigates dataset vectorization for two dataset-level tasks: assessing training set suitability and test set difficulty. The former measures how suitable a training set is for a target domain, while the latter studies how challenging a test set is for a learned model. Central of the two…

Cited by 12SourcePDFScholar
2023

Adaptive Calibrator Ensemble: Navigating Test Set Difficulty in Out-of-Distribution Scenarios

ICCV 2023poster

Model calibration usually requires optimizing some parameters (e.g., temperature) w.r.t an objective function like negative log-likelihood. This work uncovers a significant aspect often overlooked that the objective function is influenced by calibration set difficulty: the ratio of misclassified to…

Cited by 8PDFcodeScholar
2023

CircNet: Meshing 3D Point Clouds with Circumcenter Detection

ICLR 2023poster

Reconstructing 3D point clouds into triangle meshes is a key problem in computational geometry and surface reconstruction. Point cloud triangulation solves this problem by providing edge information to the input points. Since no vertex interpolation is involved, it is beneficial to preserve sharp de…

2023

Confidence and Dispersity Speak: Characterizing Prediction Matrix for Unsupervised Accuracy Estimation

ICML 2023poster

This work aims to assess how well a model performs under distribution shifts without using labels. While recent methods study prediction confidence, this work reports prediction dispersity is another informative cue. Confidence reflects whether the individual prediction is certain; dispersity indica…

Cited by 17SourcePDFScholar
2023

Divide, Evaluate, and Refine: Evaluating and Improving Text-to-Image Alignment with Iterative VQA Feedback

NeurIPS 2023poster

The field of text-conditioned image generation has made unparalleled progress with the recent advent of latent diffusion models. While revolutionary, as the complexity of given text input increases, the current state of art diffusion models may still fail in generating images that accurately convey…

Cited by 23SourcePDFScholar
2023

High-Fidelity Guided Image Synthesis With Latent Diffusion Models

CVPR 2023poster

Controllable image synthesis with user scribbles has gained huge public interest with the recent advent of text-conditioned latent diffusion models. The user scribbles control the color composition while the text prompt provides control over the overall image semantics. However, we find that prior w…

2023

How Far Pre-trained Models Are from Neural Collapse on the Target Dataset Informs their Transferability

ICCV 2023poster

This paper focuses on model transferability estimation, i.e., assessing the performance of pre-trained models on a downstream task without performing fine-tuning. Motivated by the neural collapse (NC) that reveals the feature geometry at the terminal stage of training, our method considers the model…

Cited by 23PDFScholar
2023

Privacy Assessment on Reconstructed Images: Are Existing Evaluation Metrics Faithful to Human Perception?

NeurIPS 2023spotlight

Hand-crafted image quality metrics, such as PSNR and SSIM, are commonly used to evaluate model privacy risk under reconstruction attacks. Under these metrics, reconstructed images that are determined to resemble the original one generally indicate more privacy leakage. Images determined as overall d…

Cited by 8SourcePDFScholar
2022

How to Synthesize a Large-Scale and Trainable Micro-Expression Dataset?

ECCV 2022poster

"This paper does not contain technical novelty but introduces our key discoveries in a data generation protocol, a database and insights. We aim to address the lack of large-scale datasets in micro-expression (MiE) recognition due to the prohibitive cost of data collection, which renders large-scale…

2022

Intelli-Paint: Towards Developing More Human-Intelligible Painting Agents

ECCV 2022poster

"Stroke based rendering methods have recently become a popular solution for the generation of stylized paintings. However, the current research in this direction is focused mainly on the improvement of final canvas quality, and thus often fails to consider the intelligibility of the generated painti…

Cited by 18SourcePDFScholar
2022

On the Strong Correlation Between Model Invariance and Generalization

NeurIPS 2022accept

Generalization and invariance are two essential properties of machine learning models. Generalization captures a model's ability to classify unseen data while invariance measures consistency of model predictions on transformations of the data. Existing research suggests a positive relationship: a m…

Cited by 19SourcePDFScholar
2022

Paint2Pix: Interactive Painting Based Progressive Image Synthesis and Editing

ECCV 2022poster

"Controllable image synthesis with user scribbles is a topic of keen interest in the computer vision community. In this paper, for the first time we study the problem of photorealistic image synthesis from incomplete and primitive human paintings. In particular, we propose a novel approach paint2pix…

2021

Combining Semantic Guidance and Deep Reinforcement Learning for Generating Human Level Paintings

CVPR 2021poster

Generation of stroke-based non-photorealistic imagery, is an important problem in the computer vision community. As an endeavor in this direction, substantial recent research efforts have been focused on teaching machines "how to paint", in a manner similar to a human painter. However, the applicabi…

Cited by 32PDFcodeScholar
2021

Positive Sample Propagation Along the Audio-Visual Event Line

CVPR 2021poster

Visual and audio signals often coexist in natural environments, forming audio-visual events (AVEs). Given a video, we aim to localize video segments containing an AVE and identify its category. In order to learn discriminative features for a classifier, it is pivotal to identify the helpful (or posi…

Cited by 127PDFcodeScholar
2021

What Does Rotation Prediction Tell Us about Classifier Accuracy under Varying Testing Environments?

ICML 2021spotlight

Understanding classifier decision under novel environments is central to the community, and a common practice is evaluating it on labeled test sets. However, in real-world testing, image annotations are difficult and expensive to obtain, especially when the test environment is changing. A natural qu…

Cited by 84SourcePDFScholar
2020

Circle Loss: A Unified Perspective of Pair Similarity Optimization

CVPR 2020oral

This paper provides a pair similarity optimization viewpoint on deep feature learning, aiming to maximize the within-class similarity s_p and minimize the between-class similarity s_n. We find a majority of loss functions, including the triplet loss and the softmax cross-entropy loss, embed s_n and…

Cited by 1174PDFScholar
2020

CycAs: Self-supervised Cycle Association for Learning Re-identifiable Descriptions

ECCV 2020poster

This paper proposes a self-supervised learning method for the person re-identification (re-ID) problem, where existing unsupervised methods usually rely on pseudo labels, such as those from video tracklets or clustering. A potential drawback of using pseudo labels is that errors may accumulate and i…

Cited by 116SourcePDFScholar
2020

From Depth What Can You See? Depth Completion via Auxiliary Image Reconstruction

CVPR 2020poster

Depth completion recovers dense depth from sparse measurements, e.g., LiDAR. Existing depth-only methods use sparse depth as the only input. However, these methods may fail to recover semantics consistent boundaries, or small/thin objects due to 1) the sparse nature of depth points and 2) the lack o…

Cited by 100PDFScholar
2020

Learning Object Relation Graph and Tentative Policy for Visual Navigation

ECCV 2020poster

Target-driven visual navigation aims at navigating an agent towards a given target based on the observation of the agent. In this task, it is critical to learn informative visual representation and robust navigation policy. Aiming to improve these two components, this paper proposes three complement…

2020

Simulating Content Consistent Vehicle Datasets with Attribute Descent

ECCV 2020poster

This paper uses a graphic engine to simulate a large amount of training data with free annotations. Between synthetic and real data, there is a two-level domain gap, i.e., content level and appearance level. While the latter has been widely studied, we focus on reducing the content gap in attributes…

2019

Invariance Matters: Exemplar Memory for Domain Adaptive Person Re-Identification

CVPR 2019poster

This paper considers the domain adaptive person re-identification (re-ID) problem: learning a re-ID model from a labeled source domain and an unlabeled target domain. Conventional methods are mainly to reduce feature distribution gap between the source and target domains. However, these studies larg…

Cited by 780PDFcodeScholar
2019

Joint Discriminative and Generative Learning for Person Re-Identification

CVPR 2019oral

Person re-identification (re-id) remains challenging due to significant intra-class variations across different cameras. Recently, there has been a growing interest in using generative models to augment training data and enhance the invariance to input changes. The generative pipelines in existing m…

Cited by 1005PDFScholar
2019

Taking a Closer Look at Domain Shift: Category-Level Adversaries for Semantics Consistent Domain Adaptation

CVPR 2019oral

We consider the problem of unsupervised domain adaptation in semantic segmentation. The key in this campaign consists in reducing the domain shift, i.e., enforcing the data distributions of the two domains to be similar. A popular strategy is to align the marginal distribution in the feature space t…

Cited by 933PDFcodeScholar
2018

Beyond Part Models: Person Retrieval with Refined Part Pooling (and A Strong Convolutional Baseline)

ECCV 2018poster

Employing part-level features offers fine-grained information for pedestrian image description. A prerequisite of part discovery is that each part should be well located. Instead of using external resources like pose estimator, we consider content consistency within each part for precise part locati…

2018

Camera Style Adaptation for Person Re-Identification

CVPR 2018poster

Being a cross-camera retrieval task, person re-identification suffers from image style variations caused by different cameras. The art implicitly addresses this problem by learning a camera-invariant descriptor subspace. In this paper, we explicitly consider this challenge by introducing camera styl…

2018

Deep Adversarial Attention Alignment for Unsupervised Domain Adaptation: the Benefit of Target Expectation Maximization

ECCV 2018poster

In this paper, we make two contributions to unsupervised domain adaptation (UDA) using the convolutional neural network (CNN). First, our approach transfers knowledge in all the convolutional layers through attention alignment. Most previous methods align high-level representations, e.g., activation…

Cited by 162SourcePDFScholar
2018

Generalizing A Person Retrieval Model Hetero- and Homogeneously

ECCV 2018poster

Person re-identification (re-ID) poses unique challenges for unsupervised domain adaptation (UDA) in that classes in the source and target sets (domains) are entirely different and that image variations are largely caused by cameras. Given a labeled source training set and an unlabeled target traini…

2018

Image-Image Domain Adaptation With Preserved Self-Similarity and Domain-Dissimilarity for Person Re-Identification

CVPR 2018poster

Person re-identification (re-ID) models trained on one domain often fail to generalize well to another. In our attempt, we present a ``learning via translation'' framework. In the baseline, we translate the labeled images from source to target domain in an unsupervised manner. We then train re-ID mo…

Cited by 1224SourcePDFScholar
2018

Macro-Micro Adversarial Network for Human Parsing

ECCV 2018poster

In human parsing, the pixel-wise classification loss has drawbacks in its low-level local inconsistency and high-level semantic inconsistency. The introduction of the adversarial network tackles the two problems using a single discriminator. However, the two types of parsing inconsistency are genera…

2017

Dynamic Label Graph Matching for Unsupervised Video Re-Identification

ICCV 2017poster

Label estimation is an important component in an unsupervised person re-identification (re-ID) system. This paper focuses on cross-camera label estimation, which can be subsequently used in feature learning to learn robust re-ID models. Specifically, we propose to construct a graph for samples in ea…

Cited by 224PDFScholar
2017

Person Re-Identification in the Wild

CVPR 2017spotlight

This paper presents a novel large-scale dataset and comprehensive baselines for end-to-end pedestrian detection and person recognition in raw video frames. Our baselines address three issues: the performance of various combinations of detectors and recognizers, mechanisms for pedestrian detection to…

Cited by 1011PDFcodeScholar
2017

Unlabeled Samples Generated by GAN Improve the Person Re-Identification Baseline in Vitro

ICCV 2017spotlight

The main contribution of this paper is a simple semi-supervised pipeline that only uses the original training set without collecting extra data. It is challenging in 1) how to obtain more training data only from the training set and 2) how to use the newly generated data. In this work, the generativ…

Cited by 2056PDFcodeScholar
2015

Query-Adaptive Late Fusion for Image Search and Person Re-Identification

CVPR 2015poster

Feature fusion has been proven effective [31, 32] in image search. Typically, it is assumed that the to-be-fused heterogeneous features work well by themselves for the query. However, in a more realistic situation, one does not know in advance whether a feature is effective or not for a given query.…

Cited by 383SourcePDFScholar