← Search

Zhe Zhang

47 accepted papers

2026

Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction

ICLR 2026poster

The evaluation of code-generating Large Language Models (LLMs) is fundamentally constrained by two intertwined challenges: a reliance on static, easily contaminated problem sources and the use of superficial, low-rigor testing. This paper introduces a new benchmark construction philosophy, Dual Scal…

Cited by 0SourcecodeScholar
2026

Cooperative Informed Tree (CoIT*): Cooperative Bi-Directional Multi-Resolution Motion Planning with Adaptive Edge Screening

ICRA 2026poster

In informed search-based path planning, heuristic functions that incorporate problem knowledge are essential for guiding the search and improving efficiency. The accuracy and computational cost of these heuristics are therefore critical to performance. However, accuracy and computational efficiency …

Cited by 0Scholar
2026

DCFold: Efficient Protein Structure Generation with Single Forward Pass

ICLR 2026oral

AlphaFold3 introduces a diffusion-based architecture that elevates protein structure prediction to all-atom resolution with improved accuracy. This state-of-the-art performance has established AlphaFold3 as a foundation model for diverse generation and design tasks. However, its iterative design sub…

Cited by 0SourceScholar
2026

Decoupling Variance and Scale-Invariant Updates in Adaptive Gradient Descent for Unified Vector and Matrix Optimization

ICML 2026poster

Adaptive methods like Adam have become the *de facto* standard for large-scale vector and Euclidean optimization due to their coordinate-wise adaptation with a second-order nature. More recently, matrix-based spectral optimizers like Muon (Jordan et al., 2024b) show the power of treating weight matr…

Cited by 0SourceScholar
2026

FlexProtein: Joint Sequence and Structure Pretraining for Protein Modeling

ICLR 2026poster

Protein foundation models have advanced rapidly, with most approaches falling into two dominant paradigms. Sequence-only language models (e.g., ESM-2) capture sequence semantics at scale but lack structural grounding. MSA-based predictors (e.g., AlphaFold 2/3) achieve accurate folding by exploiting…

Cited by 0SourceScholar
2026

PGS: Effective LLM Code Refinement via Property-Oriented and Structurally Minimal Feedback

ICML 2026poster

LLMs excel at code generation, yet ensuring the functional correctness of their outputs remains a persistent challenge. While recent studies have applied Test-Driven Development (TDD) to refine code, these methods are often undermined by poor feedback quality, stemming from the scarcity of high-qual…

Cited by 0SourceScholar
2026

Syllogism-Inspired TableQA: Evidentialization Makes Decomposition Reasoning and Answer Verification More Reliable

AAAI 2026technical

Existing large language model (LLM)-based table question answering (TableQA) methods primarily involve decomposition reasoning and answer verification processes. However, decomposing questions solely at the semantic level, without considering the factual evidence in tables, fails to significantly re

Cited by 0SourcePDFScholar
2025

CWNet: Causal Wavelet Network for Low-Light Image Enhancement

ICCV 2025poster

Traditional Low-Light Image Enhancement (LLIE) methods primarily focus on uniform brightness adjustment, often neglecting instance-level semantic information and the inherent characteristics of different features. To address these limitations, we propose CWNet (Causal Wavelet Network), a novel archi…

2025

Co-Progression Knowledge Distillation with Knowledge Prototype for Industrial Anomaly Detection

AAAI 2025technical

Unsupervised anomaly detection has emerged as a powerful technique for identifying abnormal patterns in images without relying on pre-labeled defective samples. Many unsupervised methods use pre-trained feature extractors from large datasets, with knowledge distillation between teacher and student m…

Cited by 0SourcePDFScholar
2025

CostFilter-AD: Enhancing Anomaly Detection through Matching Cost Filtering

ICML 2025poster

Unsupervised anomaly detection (UAD) seeks to localize the anomaly mask of an input image with respect to normal samples. Either by reconstructing normal counterparts (reconstruction-based) or by learning an image feature embedding space (embedding-based), existing approaches fundamentally rely on i…

2025

MAPS: Advancing Multi-Modal Reasoning in Expert-Level Physical Science

ICLR 2025poster

Pre-trained on extensive text and image corpora, current Multi-Modal Large Language Models (MLLM) have shown strong capabilities in general visual reasoning tasks. However, their performance is still lacking in physical domains that require understanding diagrams with complex physical structures an…

Cited by 0SourcePDFScholar
2025

PGDGS: Improving Few-shot 3D Gaussian Splatting with Progressive Gaussian Densification

ICASSP 2025accepted

Synthesizing novel views from sparse input images is a significant and challenging problem in neural rendering. As an innovative 3D representation, 3D Gaussian Splatting (3DGS) has demonstrated exceptional performance and real-time rendering capabilities. However, rendering novel views from few-shot…

Cited by 0SourceScholar
2025

Piloting Structure-Based Drug Design via Modality-Specific Optimal Schedule

ICML 2025poster

Structure-Based Drug Design (SBDD) is crucial for identifying bioactive molecules. Recent deep generative models are faced with challenges in geometric structure modeling. A major bottleneck lies in the twisted probability path of multi-modalities—continuous 3D positions and discrete 2D topologies—w…

2025

Rationalized All-Atom Protein Design with Unified Multi-Modal Bayesian Flow

NeurIPS 2025poster

Designing functional proteins is a critical yet challenging problem due to the intricate interplay between backbone structures, sequences, and side-chains. Current approaches often decompose protein design into separate tasks, which can lead to accumulated errors, while recent efforts increasingly f…

Cited by 0SourceScholar
2025

ShortListing Model: A Streamlined Simplex Diffusion for Discrete Variable Generation

NeurIPS 2025poster

Generative modeling of discrete variables is challenging yet crucial for applications in natural language processing and biological sequence design. We introduce the Shortlisting Model (SLM), a novel simplex-based diffusion model inspired by progressive candidate pruning. SLM operates on simplex cen…

Cited by 0SourcecodeScholar
2025

Steering Protein Family Design through Profile Bayesian Flow

ICLR 2025oral

Protein family design emerges as a promising alternative by combining the advantages of de novo protein design and mutation-based directed evolution.In this paper, we propose ProfileBFN, the Profile Bayesian Flow Networks, for specifically generative modeling of protein families. ProfileBFN extends…

Cited by 0SourcePDFScholar
2025

UICOMPASS: UI Map Guided Mobile Task Automation via Adaptive Action Generation

EMNLP 2025

Mobile task automation is an emerging technology that leverages AI to automatically execute routine tasks by users’ commands on mobile devices like Android, thus enhancing efficiency and productivity. While large language models (LLMs) excel at general mobile tasks through training on massive datase

2024

An Implicit Trust Region Approach to Behavior Regularized Offline Reinforcement Learning

AAAI 2024technical

We revisit behavior regularization, a popular approach to mitigate the extrapolation error in offline reinforcement learning (RL), showing that current behavior regularization may suffer from unstable learning and hinder policy improvement. Motivated by this, a novel reward shaping-based behavior re…

Cited by 6SourcePDFScholar
2024

FDC-NeRF: Learning Pose-Free Neural Radiance Fields with Flow-Depth Consistency

ICASSP 2024accepted

Learning neural radiance fields (NeRF) without camera poses has been widely studied. However, recent methods lack explicit and effective supervision for pose estimation, resulting in ambiguous optimization of camera pose and NeRF geometry during joint training, particularly in scenarios involving la…

Cited by 0SourceScholar
2024

First-Order Methods for Linearly Constrained Bilevel Optimization

NeurIPS 2024poster

Algorithms for bilevel optimization often encounter Hessian computations, which are prohibitive in high dimensions. While recent works offer first-order methods for unconstrained bilevel problems, the constrained setting remains relatively underexplored. We present first-order linearly constrained…

Cited by 18SourcePDFScholar
2024

Learned Lossless Image Compression based on Bit Plane Slicing

CVPR 2024poster

Autoregressive Initial Bits (ArIB) a framework that combines subimage autoregression and latent variable models has shown its advantages in lossless image compression. However in current methods the image splitting makes the information of latent variables being uniformly distributed in each subimag…

2024

LiDAR-Camera Extrinsic Calibration with Hierachical and Iterative Feature Matching

ICRA 2024poster

In autonomous driving, the LiDAR-Camera system plays a crucial role in a vehicle’s perception of 3D environments. To effectively fuse information from both camera and LiDAR, extrinsic calibration is indispensable. Recently, some researchers have proposed deep learning-based methods that utilize conv…

Cited by 1SourceScholar
2024

Multi-Robot Path Planning With Boolean Specification Tasks Under Motion Uncertainties

IROS 2024poster

This paper studies the path planning problem of multi-robot systems under motion uncertainties with high-level tasks that are expressed as Boolean specifications. The specification imposes logical constraints on robot trajectories and final states. First, a global Markov decision process model of th…

Cited by 1SourceScholar
2023

CL-MVSNet: Unsupervised Multi-View Stereo with Dual-Level Contrastive Learning

ICCV 2023poster

Unsupervised Multi-View Stereo (MVS) methods have achieved promising progress recently. However, previous methods primarily depend on the photometric consistency assumption, which may suffer from two limitations: indistinguishable regions and view-dependent effects, e.g., low-textured areas and refl…

Cited by 16PDFcodeScholar
2023

D2Former: Jointly Learning Hierarchical Detectors and Contextual Descriptors via Agent-Based Transformers

CVPR 2023highlight

Establishing pixel-level matches between image pairs is vital for a variety of computer vision applications. However, achieving robust image matching remains challenging because CNN extracted descriptors usually lack discriminative ability in texture-less regions and keypoint detectors are only good…

Cited by 10SourcePDFScholar
2023

Event-Centric Query Expansion in Web Search

ACL 2023industry

In search engines, query expansion (QE) is a crucial technique to improve search experience. Previous studies often rely on long-term search log mining, which leads to slow updates and is sub-optimal for time-sensitive news searches. In this work, we present Event-Centric Query Expansion (EQE), the…

Cited by 2SourcePDFScholar
2023

GeoMVSNet: Learning Multi-View Stereo With Geometry Perception

CVPR 2023poster

Recent cascade Multi-View Stereo (MVS) methods can efficiently estimate high-resolution depth maps through narrowing hypothesis ranges. However, previous methods ignored the vital geometric information embedded in coarse stages, leading to vulnerable cost matching and sub-optimal reconstruction resu…

2023

SE-ORNet: Self-Ensembling Orientation-Aware Network for Unsupervised Point Cloud Shape Correspondence

CVPR 2023poster

Unsupervised point cloud shape correspondence aims to obtain dense point-to-point correspondences between point clouds without manually annotated pairs. However, humans and some animals have bilateral symmetry and various orientations, which leads to severe mispredictions of symmetrical parts. Besid…

2023

Towards Accurate and Real-Time End-of-Speech Estimation

ICASSP 2023accepted

We introduce a variant of the endpoint (EP) detection problem in automatic speech recognition (ASR), which we call the end-of-speech (EOS) estimation. Given an utterance, EOS estimation aims to identify the timestamp when the utterance waveform has fully decayed and is then used to measure the EP la…

Cited by 0SourceScholar
2022

Motion-Modulated Temporal Fragment Alignment Network for Few-Shot Action Recognition

CVPR 2022poster

While the majority of FSL models focus on image classification, the extension to action recognition is rather challenging due to the additional temporal dimension in videos. To address this issue, we propose an end-to-end Motion-modulated Temporal Fragment Alignment Network (MTFAN) by jointly explor…

Cited by 79PDFScholar
2021

Finite Sample Analysis of Average-Reward TD Learning and $Q$-Learning

NeurIPS 2021poster

The focus of this paper is on sample complexity guarantees of average-reward reinforcement learning algorithms, which are known to be more challenging to study than their discounted-reward counterparts. To the best of our knowledge, we provide the first known finite sample guarantees using both cons…

Cited by 36SourcePDFScholar
2020

Fusing Wearable IMUs With Multi-View Images for Human Pose Estimation: A Geometric Approach

CVPR 2020poster

We propose to estimate 3D human pose from multi-view images and a few IMUs attached at person's limbs. It operates by firstly detecting 2D poses from the two signals, and then lifting them to the 3D space. We present a geometric approach to reinforce the visual features of each pair of joints based…

Cited by 80PDFcodeScholar
2018

PIRVS: An Advanced Visual-Inertial SLAM System with Flexible Sensor Fusion and Hardware Co-Design

ICRA 2018poster

In this paper, we present the PerceptIn Robotics Vision System (PIRVS), a visual-inertial computing hardware with embedded simultaneous localization and mapping (SLAM) algorithm. The PIRVS hardware is equipped with a multi-core processor, a global-shutter stereo camera, and an IMU with precise hardw…

Cited by 54SourceScholar
2018

Trifo-VIO: Robust and Efficient Stereo Visual Inertial Odometry Using Points and Lines

IROS 2018poster

In this paper, we present the Trifo Visual Inertial Odometry (Trifo-VIO), a tightly-coupled filtering-based stereo VIO system using both points and lines. Line features help improve system robustness in challenging scenarios when point features cannot be reliably detected or tracked, e.g. low-textur…

Cited by 70SourceScholar
2018

π-SoC: Heterogeneous SoC Architecture for Visual Inertial SLAM Applications

IROS 2018poster

In recent years, we have observed a clear trend in the rapid rise of autonomous vehicles and robotics. One of the core technologies enabling these applications, Simultaneous Localization And Mapping (SLAM), imposes two main challenges: first, these workloads are computationally intensive and they of…

Cited by 25SourceScholar
2017

Low-complexity optimization for two-dimensional direction-of-arrival estimation via decoupled atomic norm minimization

ICASSP 2017accepted

This paper presents an efficient optimization technique for super-resolution two-dimensional (2D) direction of arrival (DOA) estimation by introducing a new formulation of atomic norm minimization (ANM). ANM allows gridless angle estimation for correlated sources even when the number of snapshots is…

Cited by 0SourceScholar