← Search

Feng Zhu

60 accepted papers

2026

GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models

ICLR 2026poster

Recently spatial-temporal intelligence of Visual-Language Models (VLMs) has attracted much attention due to its importance for Autonomous Driving, Embodied AI and General Artificial Intelligence. Existing spatial-temporal benchmarks mainly focus on egocentric perspective reasoning with images/video…

Cited by 0SourcecodeScholar
2026

Gram2Token: Enabling Run-time GPU-Native Grammar-Constrained Decoding for LLMs

ICML 2026poster

Grammar-constrained decoding is essential for enabling large language models (LLMs) to efficiently generate structured outputs in applications, such as JSON objects for parameter passing. Existing approaches typically execute grammar constraint masking on the CPU, while LLM inference is performed on…

Cited by 0SourceScholar
2026

Hybrid-Adaptive Thread Tuning to Mitigate Simulation Execution Bottlenecks in High-Performance Reinforcement Learning Inference

IJCAI 2026

In simulation-in-the-loop decision-making systems, reinforcement learning (RL) inference is often constrained by simulator-side execution overhead, where workloads are highly dynamic and sensitive to runtime thread configurations. Existing multithreaded strategies struggle to match thread resources

Cited by 0Scholar
2026

PO-GVINS: A Tightly Coupled GNSS-Visual-Inertial Navigation Framework Using Pose-Only Representation

ICRA 2026poster

Accurate and reliable positioning is essential for perception, decision-making, and other high-level applications in autonomous driving, autonomous aerial vehicles, and intelligent robotics. Due to the inherent limitations of standalone sensors, integrating heterogeneous sensors with complementary c…

Cited by 0SourceScholar
2026

UNITE-NBV: Uncertainty-Driven and Information-Enhanced Gain Estimation for Next Best View

RA-L 2026

Next Best View (NBV) algorithms are a critical area of research in 3D reconstruction. They aim to efficiently reconstruct 3D scenes by maximizing information gain from the next optimal viewpoint. However, current NBV methods often neglect the importance of high-quality candidate view sampling, leadi

Cited by 0SourceScholar
2025

Adaptive Variance Inflation in Thompson Sampling: Efficiency, Safety, Robustness, and Beyond

NeurIPS 2025poster

Thompson Sampling (TS) has emerged as a powerful algorithm for sequential decision-making, with strong empirical success and theoretical guarantees. However, it has been shown that its behavior under stringent safety and robustness criteria --- such as safety of cumulative regret distribution and ro…

Cited by 0SourceScholar
2025

ChemActor: Enhancing Automated Extraction of Chemical Synthesis Actions with LLM-Generated Data

ACL 2025long

With the increasing interest in robotic synthesis in the context of organic chemistry, the automated extraction of chemical procedures from literature is critical. However, this task remains challenging due to the inherent ambiguity of chemical language and the high cost of human annotation required…

2025

Diffusion Transformer meets Multi-level Wavelet Spectrum for Single Image Super-Resolution

ICCV 2025poster

Discrete Wavelet Transform (DWT) has been widely explored to enhance the performance of image super-resolution (SR). Despite some DWT-based methods improving SR by capturing fine-grained frequency signals, most existing approaches neglect the interrelations among multi-scale frequency sub-bands, res…

Cited by 0SourcePDFScholar
2025

Enhance Pose Accuracy of GNSS/INS Integration by Fusing LiDAR Structure Features Based on Continuous-Time State Representation

RA-L 2025

Accurate and reliable reference poses are a crucial foundation for numerous scientific and engineering applications. This study introduces a two-stage reference pose generation method to produce a more reliable trajectory for large-scale outdoor environments. The first stage is a tightly coupled GNS

Cited by 1SourceScholar
2025

Graph Pooling via Dropping Task-Irrelevant Nodes

ICASSP 2025accepted

Graph neural networks (GNNs) face scalability challenges. While recent approaches have adopted pooling strategies inspired by convolutional neural networks (CNNs) to reduce graph size and improve efficiency, these methods often focus on local information and are optimized for single graph-level task…

Cited by 0SourceScholar
2025

Re-Aligning Language to Visual Objects with an Agentic Workflow

ICLR 2025poster

Language-based object detection (LOD) aims to align visual objects with language expressions. A large amount of paired data is utilized to improve LOD model generalizations. During the training process, recent studies leverage vision-language models (VLMs) to automatically generate human-like expres…

Cited by 0SourcePDFScholar
2025

RoboHanger: Learning Generalizable Robotic Hanger Insertion for Diverse Garments

RA-L 2025

For the task of hanging clothes, learning how to insert a hanger into a garment is a crucial step, but has rarely been explored in robotics. In this work, we address the problem of inserting a hanger into various unseen garments that are initially laid flat on a table. This task is challenging due t

Cited by 5SourceScholar
2024

Dynamic Service Fee Pricing under Strategic Behavior: Actions as Instruments and Phase Transition

NeurIPS 2024poster

We study a dynamic pricing problem for third-party platform service fees under strategic, far-sighted customers. In each time period, the platform sets a service fee based on historical data, observes the resulting transaction quantities, and collects revenue. The platform also monitors equilibrium…

Cited by 0SourcePDFScholar
2024

Instruct-ReID: A Multi-purpose Person Re-identification Task with Instructions

CVPR 2024poster

Human intelligence can retrieve any person according to both visual and language descriptions. However the current computer vision community studies specific person re-identification (ReID) tasks in different scenarios separately which limits the applications in the real world. This paper strives to…

2024

InstructDET: Diversifying Referring Object Detection with Generalized Instructions

ICLR 2024poster

We propose InstructDET, a data-centric method for referring object detection (ROD) that localizes target objects based on user instructions. While deriving from referring expressions (REC), the instructions we leverage are greatly diversified to encompass common user intentions related to object det…

2024

LinK3D: Linear Keypoints Representation for 3D LiDAR Point Cloud

RA-L 2024

Feature extraction and matching are the basic parts of many robotic vision tasks, such as 2D or 3D object detection, recognition, and registration. As is known, 2D feature extraction and matching have already achieved great success. Unfortunately, in the field of 3D, the current methods may fail to

Cited by 54SourcecodeScholar
2024

PaReNeRF: Toward Fast Large-scale Dynamic NeRF with Patch-based Reference

CVPR 2024poster

With photo-realistic image generation Neural Radiance Field (NeRF) is widely used for large-scale dynamic scene reconstruction as autonomous driving simulator. However large-scale scene reconstruction still suffers from extremely long training time and rendering time. Low-resolution (LR) rendering c…

Cited by 1SourcePDFScholar
2024

ScissorBot: Learning Generalizable Scissor Skill for Paper Cutting via Simulation, Imitation, and Sim2Real

CoRL 2024poster

This paper tackles the challenging robotic task of generalizable paper cutting using scissors. In this task, scissors attached to a robot arm are driven to accurately cut curves drawn on the paper, which is hung with the top edge fixed. Due to the frequent paper-scissor contact and consequent frac…

Cited by 5SourceScholar
2023

BoW3D: Bag of Words for Real-Time Loop Closing in 3D LiDAR SLAM

RA-L 2023

Loop closing is a fundamental part of simultaneous localization and mapping (SLAM) for autonomous mobile systems. In the field of visual SLAM, bag of words (BoW) has achieved great success in loop closure. The BoW features for loop searching can also be used in the subsequent 6-DoF loop correction.

Cited by 99SourcecodeScholar
2023

CORA: Adapting CLIP for Open-Vocabulary Detection With Region Prompting and Anchor Pre-Matching

CVPR 2023poster

Open-vocabulary detection (OVD) is an object detection task aiming at detecting objects from novel categories beyond the base categories on which the detector is trained. Recent OVD methods rely on large-scale visual-language pre-trained models, such as CLIP, for recognizing novel objects. We identi…

2023

CORE: Co-planarity Regularized Monocular Geometry Estimation with Weak Supervision

ICCV 2023poster

The ill-posed nature of monocular 3D geometry (depth map and surface normals) estimation makes it rely mostly on data-driven approaches such as Deep Neural Networks (DNN). However, data acquisition of surface normals, especially the reliable normals, is acknowledged difficult. Commonly, reconstructi…

Cited by 0PDFScholar
2023

Contextual Image Masking Modeling via Synergized Contrasting without View Augmentation for Faster and Better Visual Pretraining

ICLR 2023poster

We propose a new contextual masking image modeling (MIM) approach called contrasting-aided contextual MIM (ccMIM), under the MIM paradigm for visual pretraining. Specifically, we adopt importance sampling to select the masked patches with richer semantic information for reconstruction, instead of ra…

Cited by 22SourcePDFScholar
2023

Cycle-consistent Masked AutoEncoder for Unsupervised Domain Generalization

ICLR 2023poster

Self-supervised learning methods undergo undesirable performance drops when there exists a significant domain gap between training and testing scenarios. Therefore, unsupervised domain generalization (UDG) is proposed to tackle the problem, which requires the model to be trained on several different…

Cited by 7SourcePDFScholar
2023

Described Object Detection: Liberating Object Detection with Flexible Expressions

NeurIPS 2023poster

Detecting objects based on language information is a popular task that includes Open-Vocabulary object Detection (OVD) and Referring Expression Comprehension (REC). In this paper, we advance them to a more practical setting called *Described Object Detection* (DOD) by expanding category names to fle…

2023

Human Preference Score: Better Aligning Text-to-Image Models with Human Preference

ICCV 2023poster

Recent years have witnessed a rapid growth of deep generative models, with text-to-image models gaining significant attention from the public. However, existing models often generate images that do not align well with human preferences, such as awkward combinations of limbs and facial expressions. T…

Cited by 122PDFcodeScholar
2023

HumanBench: Towards General Human-Centric Perception With Projector Assisted Pretraining

CVPR 2023poster

Human-centric perceptions include a variety of vision tasks, which have widespread industrial applications, including surveillance, autonomous driving, and the metaverse. It is desirable to have a general pretrain model for versatile human-centric downstream tasks. This paper forges ahead along this…

2023

Locating Noise is Halfway Denoising for Semi-Supervised Segmentation

ICCV 2023poster

We investigate semi-supervised semantic segmentation with self-training, where a teacher model generates pseudo masks to exploit the benefits of a large amount of unlabeled images. We notice that the noisy label from the generated pseudo masks is the major obstacle to achieving good performance. Pre…

Cited by 13PDFScholar
2023

Patch-Level Contrasting without Patch Correspondence for Accurate and Dense Contrastive Representation Learning

ICLR 2023poster

We propose ADCLR: \underline{A}ccurate and \underline{D}ense \underline{C}ontrastive \underline{R}epresentation \underline{L}earning, a novel self-supervised learning framework for learning accurate and dense vision representation. To extract spatial-sensitive information, ADCLR introduces query pat…

Cited by 19SourcePDFScholar
2023

Stochastic Multi-armed Bandits: Optimal Trade-off among Optimality, Consistency, and Tail Risk

NeurIPS 2023spotlight

We consider the stochastic multi-armed bandit problem and fully characterize the interplays among three desired properties for policy design: worst-case optimality, instance-dependent consistency, and light-tailed risk. We show how the order of expected regret exactly affects the decaying rate of th…

Cited by 3SourcePDFScholar
2023

Trust Your Partner's Friends: Hierarchical Cross-Modal Contrastive Pre-Training for Video-Text Retrieval

ICASSP 2023accepted

Video-text retrieval has greatly benefited from the massive web video in recent years, while the performance is still limited to the weak supervision from the uncurated data. In this work, we propose to leverage the well-represented information of each original modality and exploit complementary inf…

Cited by 0SourceScholar
2023

UniHCP: A Unified Model for Human-Centric Perceptions

CVPR 2023poster

Human-centric perceptions (e.g., pose estimation, human parsing, pedestrian detection, person re-identification, etc.) play a key role in industrial applications of visual models. While specific human-centric tasks have their own relevant semantic aspect to focus on, they also share the same underly…

2022

A Simple and Optimal Policy Design for Online Learning with Safety against Heavy-tailed Risk

NeurIPS 2022accept

We consider the classical multi-armed bandit problem and design simple-to-implement new policies that simultaneously enjoy two properties: worst-case optimality for the expected regret, and safety against heavy-tailed risk for the regret distribution. Recently, Fan and Glynn (2021) showed that infor…

Cited by 3SourcePDFScholar
2022

Align Representations With Base: A New Approach to Self-Supervised Learning

CVPR 2022poster

Existing symmetric contrastive learning methods suffer from collapses (complete and dimensional) or quadratic complexity of objectives. Departure from these methods which maximize mutual information of two generated views, along either instance or feature dimension, the proposed paradigm introduces…

Cited by 30PDFScholar
2022

Counterfactual Intervention Feature Transfer for Visible-Infrared Person Re-identification

ECCV 2022poster

"Graph-based models have achieved great success in person re-identification tasks recently, which compute the graph topology structure (affinities) among different people first and then pass the information across them to achieve stronger features. But we find existing graph-based methods in the vis…

Cited by 51SourcePDFScholar
2022

Debiased Causal Tree: Heterogeneous Treatment Effects Estimation with Unmeasured Confounding

NeurIPS 2022accept

Unmeasured confounding poses a significant threat to the validity of causal inference. Despite that various ad hoc methods are developed to remove confounding effects, they are subject to certain fairly strong assumptions. In this work, we consider the estimation of conditional causal effects in the…

Cited by 13SourcePDFScholar
2022

Domain Invariant Masked Autoencoders for Self-Supervised Learning from Multi-Domains

ECCV 2022poster

"Generalizing learned representations across significantly different visual domains is a fundamental yet crucial ability of the human visual system. While recent self-supervised learning methods have achieved good performances with evaluation set on the same domain as the training set, they will hav…

Cited by 19SourcePDFScholar
2022

Feature Erasing and Diffusion Network for Occluded Person Re-Identification

CVPR 2022poster

Occluded person re-identification (ReID) aims at matching occluded person images to holistic ones across different camera views. Target Pedestrians (TP) are often disturbed by Non-Pedestrian Occlusions (NPO) and Non-Target Pedestrians (NTP). Previous methods mainly focus on increasing the model's ro…

Cited by 179PDFcodeScholar
2022

Instance As Identity: A Generic Online Paradigm for Video Instance Segmentation

ECCV 2022poster

"Modeling temporal information for both detection and tracking in a unified framework has been proved a promising solution to video instance segmentation (VIS). However, how to effectively incorporate the temporal information into an online model remains an open problem. In this work, we propose a n…

2022

Learning Memory-Augmented Unidirectional Metrics for Cross-Modality Person Re-Identification

CVPR 2022poster

This paper tackles the cross-modality person re-identification (re-ID) problem by suppressing the modality discrepancy. In cross-modality re-ID, the query and gallery images are in different modalities. Given a training identity, the popular deep classification baseline shares the same proxy (i.e.,…

Cited by 178PDFScholar
2022

Relative Contrastive Loss for Unsupervised Representation Learning

ECCV 2022poster

"Defining positive and negative samples is critical for learning visual variations of the semantic classes in an unsupervised manner. Previous methods either construct positive sample pairs as different data augmentations on the same image (i.e., single-instance-positive) or estimate a class prototy…

Cited by 3SourcePDFScholar
2022

Revisiting the Transferability of Supervised Pretraining: An MLP Perspective

CVPR 2022poster

The pretrain-finetune paradigm is a classical pipeline in visual learning. Recent progress on unsupervised pretraining methods shows superior transfer performance to their supervised counterparts. This paper revisits this phenomenon and sheds new light on understanding the transferability gap betwee…

Cited by 72PDFScholar
2022

Unifying Visual Contrastive Learning for Object Recognition from a Graph Perspective

ECCV 2022poster

"Recent contrastive based unsupervised object recognition methods leverage a Siamese architecture, which has two branches composed of a backbone, a projector layer, and an optional predictor layer in each branch. To learn the parameters of the backbone, existing methods have a similar projector laye…

Cited by 8SourcePDFScholar
2022

Unsupervised Object Detection Pretraining with Joint Object Priors Generation and Detector Learning

NeurIPS 2022accept

Unsupervised pretraining methods for object detection aim to learn object discrimination and localization ability from large amounts of images. Typically, recent works design pretext tasks that supervise the detector to predict the defined object priors. They normally leverage heuristic methods to p…

Cited by 5SourcePDFScholar
2022

Zero-CL: Instance and Feature decorrelation for negative-free symmetric contrastive learning

ICLR 2022poster

For self-supervised contrastive learning, models can easily collapse and generate trivial constant solutions. The issue has been mitigated by recent improvement on objective design, which however often requires square complexity either for the size of instances ($\mathcal{O}(N^{2})$) or feature dime…

Cited by 48SourcePDFScholar
2021

Cross-Domain Recommendation: Challenges, Progress, and Prospects

IJCAI 2021poster

To address the long-standing data sparsity problem in recommender systems (RSs), cross-domain recommendation (CDR) has been proposed to leverage the relatively richer information from a richer domain to improve the recommendation performance in a sparser domain. Although CDR has been extensively stu…

2021

Deep inference of latent dynamics with spatio-temporal super-resolution using selective backpropagation through time

NeurIPS 2021poster

Modern neural interfaces allow access to the activity of up to a million neurons within brain circuits. However, bandwidth limits often create a trade-off between greater spatial sampling (more channels or pixels) and the temporal frequency of sampling. Here we demonstrate that it is possible to obt…

2021

MixMix: All You Need for Data-Free Compression Are Feature and Data Mixing

ICCV 2021poster

User data confidentiality protection is becoming a rising challenge in the present deep learning research. Without access to data, conventional data-driven model compression faces a higher risk of performance degradation. Recently, some works propose to generate images from a specific pretrained mod…

Cited by 40PDFScholar
2021

Progressive Correspondence Pruning by Consensus Learning

ICCV 2021poster

Correspondence pruning aims to correctly remove false matches (outliers) from an initial set of putative correspondences. The selection is challenging since putative matches are typically extremely unbalanced, largely dominated by outliers, and the random distribution of such outliers further compli…

Cited by 91PDFScholar
2021

Temporal ROI Align for Video Object Recognition

AAAI 2021technical

Video object detection is challenging in the presence of appearance deterioration in certain video frames. Therefore, it is a natural choice to aggregate temporal information from other frames of the same video into the current frame. However, ROI Align, as one of the most core procedures of video d…

2020

A Graphical and Attentional Framework for Dual-Target Cross-Domain Recommendation

IJCAI 2020poster

The conventional single-target Cross-Domain Recommendation (CDR) only improves the recommendation accuracy on a target domain with the help of a source domain (with relatively richer information). In contrast, the novel dual-target CDR has been proposed to improve the recommendation accuracies on bo…

Cited by 0SourcePDFScholar
2020

Self-paced Contrastive Learning with Hybrid Memory for Domain Adaptive Object Re-ID

NeurIPS 2020poster

Domain adaptive object re-ID aims to transfer the learned knowledge from the labeled source domain to the unlabeled target domain to tackle the open-class re-identification problems. Although state-of-the-art pseudo-label-based methods have achieved great success, they did not make full use of all v…

2020

Self-supervising Fine-grained Region Similarities for Large-scale Image Localization

ECCV 2020poster

The task of large-scale retrieval-based image localization is to estimate the geographical location of a query image by recognizing its nearest reference images from a city-scale dataset. However, the general public benchmarks only provide noisy GPS labels associated with the training images, which…

2020

Towards Unified INT8 Training for Convolutional Neural Network

CVPR 2020poster

Recently low-bit (e.g., 8-bit) network quantization has been extensively studied to accelerate the inference. Besides inference, low-bit training with quantized gradients can further bring more considerable acceleration, since the backward process is often computation-intensive. Unfortunately, the i…

Cited by 218PDFScholar
2018

Attention-Aware Compositional Network for Person Re-Identification

CVPR 2018poster

Person re-identification (ReID) is to identify pedestrians observed from different camera views based on visual appearance. It is a challenging task due to large pose variations, complex background clutters and severe occlusions. Recently, human pose estimation by predicting joint locations was larg…

Cited by 565SourcePDFScholar
2017

Learning Spatial Regularization With Image-Level Supervisions for Multi-Label Image Classification

CVPR 2017poster

Multi-label image classification is a fundamental but challenging task in computer vision. Great progress has been achieved by exploiting semantic relations between labels in recent years. However, conventional approaches are unable to model the underlying spatial relations between labels in multi-l…

Cited by 476PDFcodeScholar