← Search

Jiahui Zhang

43 accepted papers

2026

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation

ICML 2026poster

Video generation models offer a promising imagination mechanism for robot manipulation by predicting long-horizon future observations, but effectively exploiting these imagined futures for action execution remains challenging. Existing approaches either condition policies on predicted frames or dire…

Cited by 0SourceScholar
2026

I2Mole: Interaction-aware Invariant Molecular Learning For Generalizable Property Prediction

ICLR 2026poster

Molecular interactions are a common phenomenon in physical chemistry field, which could produce unexpected biochemical properties harmful to humans, such as drug-drug interactions. Machine learning has the potential to deliver rapid and accurate predictions. However, the complexity of molecular stru…

Cited by 0SourceScholar
2026

MILD: Tractable Terrain Modeling for Learning Improved Bipedal Locomotion on Deformable Surfaces

RA-L 2026

Enabling robots to walk on yielding terrain is vital for applications ranging from disaster response to planetary exploration. While bipedal robots hold immense potential, their locomotion on deformable surfaces remains limited as current simulators fail to capture the spatiotemporal heterogeneity o

Cited by 1SourceScholar
2026

MILD: Tractable Terrain Modeling for Learning Improved Bipedal Locomotion on Deformable Surfaces

ICRA 2026poster

Enabling robots to walk on yielding terrain is vital for applications ranging from disaster response to planetary exploration. While bipedal robots hold immense potential, their locomotion on deformable surfaces remains limited as current simulators fail to capture the spatiotemporal heterogeneity o…

Cited by 0SourceScholar
2026

MemoryART: Enhancing LLMs via Multi-Memory Models with Adaptive Resonance Theory for Healthcare Agents

AAAI 2026technical

Though promising in healthcare consultation applications, large language models (LLMs) face critical limitations in retaining and utilizing long-term memory across multi-turn interactions. In particular, existing memory enhancing paradigms are constrained by limited context windows and embedding-bas

Cited by 0SourcePDFScholar
2026

NeuroMamba: A Universal Spatiotemporal Module for Robust Perception in Degraded Sensory Streams

ICML 2026poster

In open-world intelligent systems, processing continuous sensory streams disrupted by heterogeneous degradation sources presents a fundamental challenge: reconciling the inherent tension between observational completeness and reconstruction fidelity. Methods that prioritize completeness by bridging …

Cited by 0SourceScholar
2026

OmniNet: Omnidirectional Jumping Neural Network with Height-Awareness for Quadrupedal Robots

ICRA 2026poster

In the robotics community, it has been a longstanding challenge for quadrupeds to achieve highly explosive movements similar to their biological counterparts. In this work, we introduce a novel training framework that achieves height-aware and omnidirectional jumping for quadrupedal robots. To facil…

Cited by 0SourceScholar
2026

PepBenchmark: A Standardized Benchmark for Peptide Machine Learning

ICLR 2026poster

Peptide therapeutics are widely regarded as the “third generation” of drugs, yet progress in peptide Machine Learning (ML) are hindered by the absence of standardized benchmarks. Here we present \textbf{PepBenchmark}, which standardizes datasets, preprocessing, and evaluation protocols for peptide d…

Cited by 0SourcecodeScholar
2026

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons

RSS 2026poster

General-purpose robot reward models are typically trained to predict absolute task progress from expert demonstrations, providing only local, frame-level supervision. While effective for expert demonstrations, this paradigm scales poorly to large scale real-world robotics datasets where failed and s…

Cited by 0SourceScholar
2026

UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding

ICLR 2026poster

Despite the impressive progress on understanding and generating images shown by the recent unified architectures, the integration of 3D tasks remains challenging and largely unexplored. In this paper, we introduce UniUGG, the first unified understanding and generation framework for 3D modalities. Ou…

Cited by 0SourcecodeScholar
2025

4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration

NeurIPS 2025poster

Leveraging diverse robotic data for pretraining remains a critical challenge. Existing methods typically model the dataset’s action distribution using simple observations as inputs. However, these inputs are often incomplete, resulting in a dispersed conditional action distribution—an issue we refer…

Cited by 0SourceScholar
2025

Augmenting Sequential Recommendation with Balanced Relevance and Diversity

AAAI 2025technical

By generating new yet effective data, data augmentation has become a promising method to mitigate the data sparsity problem in sequential recommendation. Existing works focus on augmenting the original data but rarely explore the issue of imbalanced relevance and diversity for augmented data, leadin…

2025

DO-CoLM: Dynamic 3D Conformation Relationships Capture with Self-Adaptive Ordering Molecular Relational Modeling in Language Models

IJCAI 2025

Molecular Relational Learning (MRL) aims to understand interactions between molecular pairs, playing a critical role in advancing biochemical research. Recently, Large Language Models (LLMs), with their extensive knowledge bases and advanced reasoning capabilities, have emerged as powerful tools for

Cited by 0SourcePDFScholar
2025

Dynamic and Chemical Constraints to Enhance the Molecular Masked Graph Autoencoders

NeurIPS 2025poster

Masked Graph Autoencoders (MGAEs) have gained significant attention recently. Their proxy tasks typically involve random corruption of input graphs followed by reconstruction. However, in the molecular domain, two main issues arise: the predetermined mask ratio and reconstruction objectives can lead…

Cited by 0SourcecodeScholar
2025

Enhancing the Maximum Effective Window for Long-Term Time Series Forecasting

NeurIPS 2025poster

Long-term time series forecasting (LTSF) aims to predict future trends based on historical data. While longer lookback windows theoretically offer more comprehensive insights, Transformer-based models often struggle with them. On one hand, longer windows introduce more noise and redundancy, hinderin…

Cited by 0SourcecodeScholar
2025

FR-Net: Learning Robust Quadrupedal Fall Recovery on Challenging Terrains through Mass-Contact Prediction

RA-L 2025

Fall recovery for legged robots remains challenging, particularly on complex terrains where traditional controllers fail due to incomplete terrain perception and uncertain interactions. We present FR-Net, a learning-based framework that enables quadrupedal robots to recover from arbitrary fall poses

Cited by 1SourceScholar
2025

From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D

NeurIPS 2025poster

Recent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unlike previous approaches that incorporate 3D representations into models to improve spatial understanding, we aim to unlo…

Cited by 0SourceScholar
2025

OmniNet: Omnidirectional Jumping Neural Network With Height-Awareness for Quadrupedal Robots

RA-L 2025

In the robotics community, it has been a longstanding challenge for quadrupeds to achieve highly explosive movements similar to their biological counterparts. In this work, we introduce a novel training framework that achieves height-aware and omnidirectional jumping for quadrupedal robots. To facil

Cited by 3SourceScholar
2025

PCR-GS: COLMAP-Free 3D Gaussian Splatting via Pose Co-Regularizations

ICCV 2025poster

COLMAP-free 3D Gaussian Splatting (3D-GS) has recently attracted increasing attention due to its remarkable performance in reconstructing high-quality 3D scenes from unposed images or videos. However, it often struggles to handle scenes with complex camera trajectories as featured by drastic rotatio…

Cited by 0SourcePDFScholar
2025

ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations

CoRL 2025oral

We introduce ReWiND, a framework for learning robot manipulation tasks solely from language instructions without per-task demonstrations. Standard reinforcement learning (RL) and imitation learning methods require expert supervision through human-designed reward functions or demonstrations for every…

Cited by 0SourceScholar
2025

Versatile Transition Generation with Image-to-Video Diffusion

ICCV 2025poster

Leveraging text, images, structure maps, or motion trajectories as conditional guidance, diffusion models have achieved great success in automated and high-quality video generation. However, generating smooth and rational transition videos given the first and last video frames as well as descriptive…

Cited by 0SourcePDFScholar
2024

An Ultralight Air-Ground Vehicle Capable of Sustained Amphibious Maneuverability and Bio-Inspired Modality Transition

RA-L 2024

This letter presents a 5.2g ultra-lightweight air-ground vehicle capable of passive stable flying and terrestrial cruising. Such a design proposes a passive stability layout, including a single-axis rotor, film dampers, and a stabilizer bar. The advantageous synergy between the rotor and the passive

Cited by 1SourceScholar
2024

FreGS: 3D Gaussian Splatting with Progressive Frequency Regularization

CVPR 2024poster

3D Gaussian splatting has achieved very impressive performance in real-time novel view synthesis. However it often suffers from over-reconstruction during Gaussian densification where high-variance image regions are covered by a few large Gaussians only leading to blur and artifacts in the rendered…

Cited by 55SourcePDFScholar
2024

SPRINT: Scalable Policy Pre-Training via Language Instruction Relabeling

ICRA 2024poster

Pre-training robots with a rich set of skills can substantially accelerate the learning of downstream tasks. Prior works have defined pre-training tasks via natural language instructions, but doing so requires tedious human annotation of hundreds of thousands of instructions. Thus, we propose SPRINT…

Cited by 19SourceScholar
2023

Bootstrap Your Own Skills: Learning to Solve New Tasks with Large Language Model Guidance

CoRL 2023oral

We propose BOSS, an approach that automatically learns to solve new long-horizon, complex, and meaningful tasks by growing a learned skill library with minimal supervision. Prior work in reinforcement learning require expert supervision, in the form of demonstrations or rich reward functions, to lea…

Cited by 80SourceScholar
2023

Pose-Free Neural Radiance Fields via Implicit Pose Regularization

ICCV 2023poster

Pose-free neural radiance fields (NeRF) aim to train NeRF with unposed multi-view images and it has achieved very impressive success in recent years. Most existing works share the pipeline of training a coarse pose estimator with rendered images at first, followed by a joint optimization of estimate…

Cited by 11PDFScholar
2023

Regularized Vector Quantization for Tokenized Image Synthesis

CVPR 2023poster

Quantizing images into discrete representations has been a fundamental problem in unified generative modeling. Predominant approaches learn the discrete representation either in a deterministic manner by selecting the best-matching token or in a stochastic manner by sampling from a predicted distrib…

2023

StyleRF: Zero-Shot 3D Style Transfer of Neural Radiance Fields

CVPR 2023poster

3D style transfer aims to render stylized novel views of a 3D scene with multi-view consistency. However, most existing work suffers from a three-way dilemma over accurate geometry reconstruction, high-quality stylization, and being generalizable to arbitrary new styles. We propose StyleRF (Style Ra…

2023

WaveNeRF: Wavelet-based Generalizable Neural Radiance Fields

ICCV 2023poster

Neural Radiance Field (NeRF) has shown impressive performance in novel view synthesis via implicit scene representation. However, it usually suffers from poor scalability as requiring densely sampled images for each new scene. Several studies have attempted to mitigate this problem by integrating Mu…

Cited by 16PDFScholar
2023

Weakly Supervised 3D Open-vocabulary Segmentation

NeurIPS 2023poster

Open-vocabulary segmentation of 3D scenes is a fundamental function of human perception and thus a crucial objective in computer vision research. However, this task is heavily impeded by the lack of large-scale and diverse 3D open-vocabulary segmentation datasets for training robust and generalizabl…

2022

Auto-Regressive Image Synthesis with Integrated Quantization

ECCV 2022poster

"Deep generative models have achieved conspicuous progress in realistic image synthesis with multifarious conditional inputs, while generating diverse yet high-fidelity images remains a grand challenge in conditional image generation. This paper presents a versatile framework for conditional image g…

2022

Bi-Level Feature Alignment for Versatile Image Translation and Manipulation

ECCV 2022poster

"Generative adversarial networks (GANs) have achieved great success in image translation and manipulation. However, high-fidelity image generation with faithful style control remains a grand challenge in computer vision. This paper presents a versatile image translation and manipulation framework th…

Cited by 52SourcePDFScholar
2022

Marginal Contrastive Correspondence for Guided Image Generation

CVPR 2022oral

Exemplar-based image translation establishes dense correspondences between a conditional input and an exemplar (from two different domains) for leveraging detailed exemplar styles to achieve realistic image translation. Existing work builds the cross-domain correspondences implicitly by minimizing f…

Cited by 78PDFScholar
2022

Modulated Contrast for Versatile Image Synthesis

CVPR 2022poster

Perceiving the similarity between images has been a long-standing and fundamental problem underlying various visual generation tasks. Predominant approaches measure the inter-image distance by computing pointwise absolute deviations, which tends to estimate the median of instance distributions and l…

Cited by 216PDFcodeScholar
2022

Parametric Path Optimization for Wheeled Robots Navigation

ICRA 2022poster

Collision risk and smoothness are the most important factors in global path planning. Currently, planning methods that reduce global path collision risk and improve its smoothness through numerical optimization have achieved good results. However, these methods cannot always optimize the path. The r…

Cited by 3SourceScholar
2021

Learning To Match Features With Seeded Graph Matching Network

ICCV 2021poster

Matching local features across images is a fundamental problem in computer vision. Targeting towards high accuracy and efficiency, we propose Seeded Graph Matching Network, a graph neural network with sparse structure to reduce redundant connectivity and learn compact representation. The network con…

Cited by 142PDFcodeScholar
2020

ASLFeat: Learning Local Features of Accurate Shape and Localization

CVPR 2020poster

This work focuses on mitigating two limitations in the joint learning of local feature detectors and descriptors. First, the ability to estimate the local shape (scale, orientation, etc.) of feature points is often neglected during dense feature extraction, while the shape-awareness is crucial to ac…

Cited by 379PDFcodeScholar
2020

KFNet: Learning Temporal Camera Relocalization Using Kalman Filtering

CVPR 2020oral

Temporal camera relocalization estimates the pose with respect to each video frame in sequence, as opposed to one-shot relocalization which focuses on a still image. Even though the time dependency has been taken into account, current temporal relocalization methods still generally underperform the…

Cited by 99PDFcodeScholar
2019

ContextDesc: Local Descriptor Augmentation With Cross-Modality Context

CVPR 2019oral

Most existing studies on learning local features focus on the patch-based descriptions of individual keypoints, whereas neglecting the spatial relations established from their keypoint locations. In this paper, we go beyond the local detail representation by introducing context awareness to augment…

Cited by 315PDFcodeScholar
2019

Learning Two-View Correspondences and Geometry Using Order-Aware Network

ICCV 2019poster

Establishing correspondences between two images requires both local and global spatial context. Given putative correspondences of feature points in two views, in this paper, we propose Order-Aware Network, which infers the probabilities of correspondences being inliers and regresses the relative pos…

Cited by 468PDFcodeScholar
2018

Efficient Semantic Scene Completion Network with Spatial Group Convolution

ECCV 2018poster

We introduce Spatial Group Convolution (SGC) for accelerating the computation of 3D dense prediction tasks. SGC is orthogonal to group convolution, which works on spatial dimensions rather than feature channel dimension. It divides input voxels into different groups, then conducts 3D sparse convolut…