← Search

Bin Zhao

64 accepted papers

2026

Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation

CVPR 2026

Reinforcement learning (RL), earlier proven to be effective in large language and multi-modal models, has been successfully extended to enhance 2D image generation recently. However, applying RL to 3D generation remains largely unexplored due to the higher spatial complexity of 3D objects, which req

Cited by 0SourcecodeScholar
2026

AutoMat: Physics-Guided Agentic Reasoning for Solving Ill-Posed Inverse Microscopy Problems

ICML 2026poster

Reconstructing atomistic crystal structures from a single noisy STEM projection is an ill-posed inverse problem: multiple lattices can explain similar contrast, and purely feed-forward models cannot verify physical validity. We present **AutoMat**, a failure-aware agentic *controller* that performs …

Cited by 0SourceScholar
2026

CLUHCS:Dual-View Contrastive Learning Enabled Unsupervised Heterogeneous Community Search with Meta-Path Behavior Modeling

AAAI 2026technical

Existing community search methods heavily rely on labeled data or predefined structures, thus fail to capture obscure and dynamic community boundaries in open-world heterogeneous networks, leading to poor adaptability. They also ignore modeling behavioral patterns, resulting in poor search performan

Cited by 0SourcePDFScholar
2026

Can VLMs Diagnose and Recover from VLA Manipulation Faults?

ICML 2026poster

Existing VLA models frequently fail in robotic manipulation tasks, with poorly structured fault types that often require expert diagnosis.While VLMs offer strong explanatory capabilities, their effectiveness in assisting VLAs is limited by their unclear role in diagnostics and inadequate collaborati…

Cited by 0SourceScholar
2026

Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion Policy

ICRA 2026poster

Diffusion-based policies have achieved remarkable results in robotic manipulation but often struggle to adapt rapidly in dynamic scenarios, leading to delayed responses or task failures. We present DCDP, a Dynamic Closed-Loop Diffusion Policy framework that integrates chunk-based action generation w…

2026

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

CVPR 2026

Vision-and-Language Navigation (VLN) requires the agent to navigate based on natural instructions. This task is challenging due to partial observability, which makes it difficult to align perception with language. Recent methods mitigate this by imagining future scenes, yet they rely on vision-based

Cited by 0SourcecodeScholar
2026

Exploring the Potential of Encoder-free Architectures in 3D LMMs

ICLR 2026poster

Encoder-free architectures have been preliminarily explored in the 2D Large Multimodal Models (LMMs), yet it remains an open question whether they can be effectively applied to 3D understanding scenarios. In this paper, we present the first comprehensive investigation into the potential of encoder-f…

Cited by 0SourcecodeScholar
2026

FM-Steer: Enhance Generalist Policies with Value-Guided Cascaded Denoising

CVPR 2026

Humans naturally allocate more time before acting when handling complex tasks in the physical world. This paradigm has recently led to remarkable advances in boosting Large Language Models (LLMs) on complex tasks in digital domains. However, the potential of test-time computing remains largely unexp

Cited by 0SourcecodeScholar
2026

FreeGaussian: Annotation-free Control of Articulated Objects via 3D Gaussian Splats with Flow Derivatives

AAAI 2026technical

Reconstructing controllable Gaussian splats for articulated objects from monocular video is especially challenging due to its inherently insufficient constraints. Existing methods address this by relying on dense masks and manually defined control signals, limiting their real-world applications. In

Cited by 0SourcePDFScholar
2026

MLM: Learning Multi-Task Loco-Manipulation Whole-Body Control for Quadruped Robot With Arm

RA-L 2026

Whole-body loco-manipulation for quadruped robots with arms remains a challenging problem, particularly in achieving multi-task control. To address this, we propose MLM, a reinforcement learning framework driven by both real-world and simulation data. It enables a six-DoF robotic arm–equipped quadru

Cited by 4SourceScholar
2026

MindSight: A Bio-Inspired Neural Architecture for Visual Restoration via Cortical Electrical Stimulation

AAAI 2026technical

Visual impairment is a common condition worldwide, and cortical electrical stimulation is one of the approaches to aid in visual restoration. However, existing methods suffer from limited precision, flexibility, and generalization in generating the desired visual perception. In this paper, we propos

Cited by 0SourcePDFScholar
2026

OpenFly: A COMPREHENSIVE PLATFORM FOR AERIAL VISION-LANGUAGE NAVIGATION

ICLR 2026poster

Aerial Vision-Language Navigation (VLN) seeks to guide UAVs by leveraging language instructions and visual cues, establishing a new paradigm for human-UAV interaction. However, the collection of VLN data demands extensive human effort to construct trajectories and corresponding instructions, hinderi…

Cited by 0SourcecodeScholar
2026

Trajectory Conditioned Cross-Embodiment Skill Transfer

ICRA 2026poster

Learning manipulation skills from human demonstration videos presents a promising yet challenging problem, primarily due to the significant embodiment gap between human body and robot manipulators. Existing methods rely on paired datasets or hand-crafted rewards, which limit scalability and generali…

2025

APA-BI: Adaptive Partition Aggregation and Bidirectional Integration for UAV-View Geo-Localization

ICRA 2025

The task of UAV-view geo-localization is to match a query image with database images to estimate the current geographic location of the query image. This is particularly useful in environments where GPS is not available or when the device fails. Although deep learning methods make sufficient progres

Cited by 1SourceScholar
2025

AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations

ICCV 2025poster

Visual grounding (VG) aims to localize target objects in an image based on natural language descriptions. In this paper, we propose AerialVG, a new task focusing on visual grounding from aerial views. Compared to traditional VG, AerialVG poses new challenges, e.g., appearance-based grounding is insu…

2025

AlignBot: Aligning VLM-Powered Customized Task Planning with User Reminders Through Fine-Tuning for Household Robots

ICRA 2025

This paper presents AlignBot, a novel framework designed to optimize VLM-powered customized task planning for household robots by effectively aligning with user reminders. In domestic settings, aligning task planning with user reminders poses significant challenges due to the limited quantity, diver

Cited by 9SourceScholar
2025

COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language Models

ICRA 2025

Leveraging the powerful reasoning capabilities of large language models (LLMs), recent LLM-based robot task planning methods yield promising results. However, they mainly focus on single or multiple homogeneous robots on simple tasks. Practically, complex long-horizon tasks always require collaborat

Cited by 41SourcecodeScholar
2025

Cocube: a Tabletop Modular Multi-Robot Platform for Education and Research

ICRA 2025

This paper presents CoCube a tabletop modular robotics platform designed for robotics education and multirobot algorithm research. CoCube is characterized by its low cost low floors high ceilings and wide walls offering flexibility and broad applicability across various use cases. The platform compr

Cited by 0SourceScholar
2025

Efficient Diffusion as Low Light Enhancer

CVPR 2025poster

The computational burden of the iterative sampling process remains a major challenge in diffusion-based Low-Light Image Enhancement (LLIE). Current acceleration methods, whether training-based or training-free, often lead to significant performance degradation, highlighting the trade-off between per…

Cited by 0SourcePDFScholar
2025

FastUMI: A Scalable and Hardware-Independent Universal Manipulation Interface with Dataset

CoRL 2025poster

Real-world manipulation datasets for robotic arms remain scarce due to the high costs, rigid hardware dependencies, and complex setup procedures associated with existing data collection methods. We introduce, a redesigned Universal Manipulation Interface (UMI) that addresses these challenges, enabli…

Cited by 0SourceScholar
2025

Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding

AAAI 2025technical

3D Object Affordance Grounding aims to predict the functional regions on a 3D object and has laid the foundation for a wide range of applications in robotics. Recent advances tackle this problem via learning a mapping between 3D regions and a single human-object interaction image. However, the geome…

2025

MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation

ICCV 2025poster

In mobile manipulation, navigation and manipulation are often treated as separate problems, resulting in a significant gap between merely approaching an object and engaging with it effectively. Many navigation approaches primarily define success by proximity to the target, often overlooking the nece…

Cited by 0SourcePDFScholar
2025

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Models

RSS 2025poster

In this paper, we claim that spatial understanding is the keypoint in robot manipulation, and propose SpatialVLA to explore effective spatial representations for the robot foundation model. Specifically, we propose Ego3D Position Encoding to inject 3D information into VLA’s input observations, and i…

Cited by 18PDFScholar
2025

Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation

CVPR 2025poster

Learning a generalist robot that can effectively leverage prior knowledge for continuous skill acquisition remains significantly challenging. Despite the success of experience replay and parameter-efficient methods in maintaining knowledge across skills, naively applying these methods causes a failu…

Cited by 0SourcePDFScholar
2025

VSS-SLAM: Voxelized Surfel Splatting for Geometally Accurate SLAM

ICRA 2025

[1] Visual Simultaneous Localization and Mapping (SLAM) helps robots estimate their poses and perceive the environment in unknown settings. Recent work has demonstrated that implicit neural radiance fields and 3D Gaussian Splatting (3DGS) offer higher fidelity scene representation than traditional m

Cited by 1SourceScholar
2024

Calibration-Free Vision-Assisted Container Loading of RTG Cranes

IROS 2024poster

Vision-assisted container loading of Rubber Tyred Gantry (RTG) cranes are facing two primary challenges. Firstly, the uncertainty inherent in Covolutional Neural Network (CNN) based detection hinders its direct application in the safety-critical operation of such heavy-duty machinery. Secondly, sens…

Cited by 0SourceScholar
2024

Color Event Enhanced Single-Exposure HDR Imaging

AAAI 2024technical

Single-exposure high dynamic range (HDR) imaging aims to reconstruct the wide-range intensities of a scene by using its single low dynamic range (LDR) image, thus providing significant efficiency. Existing methods pay high attention to restoring the luminance by inversing the tone-mapping process, w…

2024

Depth Helps: Improving Pre-trained RGB-based Policy with Depth Information Injection

IROS 2024poster

3D perception ability is crucial for generalizable robotic manipulation. While recent foundation models have made significant strides in perception and decision-making with RGB-based input, their lack of 3D perception limits their effectiveness in fine-grained robotic manipulation tasks. To address…

Cited by 2SourcecodeScholar
2024

GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting

CVPR 2024highlight

In this paper we introduce GS-SLAM that first utilizes 3D Gaussian representation in the Simultaneous Localization and Mapping (SLAM) system. It facilitates a better balance between efficiency and accuracy. Compared to recent SLAM methods employing neural implicit representations our method utilizes…

2024

HPL-ESS: Hybrid Pseudo-Labeling for Unsupervised Event-based Semantic Segmentation

CVPR 2024poster

Event-based semantic segmentation has gained popularity due to its capability to deal with scenarios under high-speed motion and extreme lighting conditions which cannot be addressed by conventional RGB cameras. Since it is hard to annotate event data previous approaches rely on event-to-image recon…

Cited by 5SourcePDFScholar
2024

HSS-SLAM: Human-in-the-Loop Semantic SLAM Represented by Superquadrics

IROS 2024poster

The advancement of object detection algorithms has catalyzed the development of object-level semantic SLAM. However, due to missed and false detections, object-level semantic SLAM fails to represent the objects within the scene adequately. Therefore, this paper proposes a novel object-level semantic…

Cited by 0SourceScholar
2024

KOI: Accelerating Online Imitation Learning via Hybrid Key-state Guidance

CoRL 2024poster

Online Imitation Learning methods struggle with the gap between extensive online exploration space and limited expert trajectories, which hinder efficient exploration due to inaccurate task-aware reward estimation. Inspired by the findings from cognitive neuroscience that task decomposition coul…

Cited by 0SourceScholar
2024

Kinematic-aware Prompting for Generalizable Articulated Object Manipulation with LLMs

ICRA 2024poster

Generalizable articulated object manipulation is essential for home-assistant robots. Recent efforts focus on imitation learning from demonstrations or reinforcement learning in simulation, however, due to the prohibitive costs of real-world data collection and precise object simulation, it still re…

Cited by 23SourcecodeScholar
2024

L-VIWO: Visual-Inertial-Wheel Odometry based on Lane Lines

ICRA 2024poster

To achieve precise localization for autonomous vehicles and mitigate the problem of accumulated drift error in odometry, this paper proposes L-VIWO, a Visual-Inertial-Wheel Odometry based on lane lines. This method effectively utilizes the lateral constraints provided by lane lines to eliminate and…

Cited by 2SourceScholar
2024

Learning Manipulation by Predicting Interaction

RSS 2024poster

Representation learning approaches for robotic manipulation have boomed in recent years. Due to the scarcity of in-domain robot data, prevailing methodologies tend to leverage large-scale human video datasets to extract generalizable features for visuomotor policy learning. Despite the progress achi…

2024

Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training

NeurIPS 2024poster

Learning a generalist embodied agent capable of completing multiple tasks poses challenges, primarily stemming from the scarcity of action-labeled robotic datasets. In contrast, a vast amount of human videos exist, capturing intricate tasks and interactions with the physical world. Promising prospec…

2024

LiveScene: Language Embedding Interactive Radiance Fields for Physical Scene Control and Rendering

NeurIPS 2024poster

This paper scales object-level reconstruction to complex scenes, advancing interactive scene reconstruction. We introduce two datasets, OmniSim and InterReal, featuring 28 scenes with multiple interactive objects. To tackle the challenge of inaccurate interactive motion recovery in complex scenes, w…

2024

Point-PEFT: Parameter-Efficient Fine-Tuning for 3D Pre-trained Models

AAAI 2024technical

The popularity of pre-trained large models has revolutionized downstream tasks across diverse fields, such as language, vision, and multi-modality. To minimize the adaption cost for downstream tasks, many Parameter-Efficient Fine-Tuning (PEFT) techniques are proposed for language and 2D image pre-tr…

2024

Robust Quadrupedal Locomotion via Risk-Averse Policy Learning

ICRA 2024poster

The robustness of legged locomotion is crucial for quadrupedal robots in challenging terrains. Recently, Reinforcement Learning (RL) has shown promising results in legged locomotion and various methods try to integrate privileged distillation, scene modeling, and external sensors to improve the gene…

Cited by 13SourceScholar
2024

SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation

ICML 2024poster

Acquiring a multi-task imitation policy in 3D manipulation poses challenges in terms of scene understanding and action prediction. Current methods employ both 3D representation and multi-view 2D representation to predict the poses of the robot’s end-effector. However, they still require a considerab…

Cited by 11SourcePDFScholar
2024

X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-Modal Knowledge Transfer

AAAI 2024technical

The field of 4D point cloud understanding is rapidly developing with the goal of analyzing dynamic 3D point cloud sequences. However, it remains a challenging task due to the sparsity and lack of texture in point clouds. Moreover, the irregularity of point cloud poses a difficulty in aligning tempo…

2023

Affordance-Driven Next-Best-View Planning for Robotic Grasping

CoRL 2023poster

Grasping occluded objects in cluttered environments is an essential component in complex robotic manipulation tasks. In this paper, we introduce an AffordanCE-driven Next-Best-View planning policy (ACE-NBV) that tries to find a feasible grasp for target object via continuously observing scenes from…

Cited by 14SourceScholar
2023

Behavior Contrastive Learning for Unsupervised Skill Discovery

ICML 2023poster

In reinforcement learning, unsupervised skill discovery aims to learn diverse skills without extrinsic rewards. Previous methods discover skills by maximizing the mutual information (MI) between states and skills. However, such an MI objective tends to learn simple and static skills and may hinder e…

2023

Brain Network Features Differentiate Intentions from Different Emotional Expressions of the Same Text

ICASSP 2023accepted

Intent differentiation in speech communication relies not only on linguistic information but also on paralinguistic information. The same textual content, when pronounced with different prosodies and emotions, may express totally different intentions. The true intentions in this condition can be eas…

Cited by 0SourceScholar
2023

Cross-Domain Policy Adaptation via Value-Guided Data Filtering

NeurIPS 2023poster

Generalizing policies across different domains with dynamics mismatch poses a significant challenge in reinforcement learning. For example, a robot learns the policy in a simulator, but when it is deployed in the real world, the dynamics of the environment may be different. Given the source and targ…

Cited by 19SourcePDFScholar
2023

Diffusion Model is an Effective Planner and Data Synthesizer for Multi-Task Reinforcement Learning

NeurIPS 2023poster

Diffusion models have demonstrated highly-expressive generative capabilities in vision and NLP. Recent studies in reinforcement learning (RL) have shown that diffusion models are also powerful in modeling complex policies or trajectories in offline datasets. However, these works have been limited to…

2023

Fully Self-Supervised Depth Estimation From Defocus Clue

CVPR 2023poster

Depth-from-defocus (DFD), modeling the relationship between depth and defocus pattern in images, has demonstrated promising performance in depth estimation. Recently, several self-supervised works try to overcome the difficulties in acquiring accurate depth ground-truth. However, they depend on the…

2023

Not All Features Matter: Enhancing Few-shot CLIP with Adaptive Prior Refinement

ICCV 2023poster

The popularity of Contrastive Language-Image Pre-training (CLIP) has propelled its application to diverse downstream vision tasks. To improve its capacity on downstream tasks, few-shot learning has become a widely-adopted technique. However, existing methods either exhibit limited performance or suf…

Cited by 89PDFcodeScholar
2023

One-Shot High-Fidelity Talking-Head Synthesis With Deformable Neural Radiance Field

CVPR 2023poster

Talking head generation aims to generate faces that maintain the identity information of the source image and imitate the motion of the driving image. Most pioneering methods rely primarily on 2D representations and thus will inevitably suffer from face distortion when large head rotations are encou…

Cited by 54SourcePDFScholar
2023

Propagate and Calibrate: Real-Time Passive Non-Line-of-Sight Tracking

CVPR 2023poster

Non-line-of-sight (NLOS) tracking has drawn increasing attention in recent years, due to its ability to detect object motion out of sight. Most previous works on NLOS tracking rely on active illumination, e.g., laser, and suffer from high cost and elaborate experimental conditions. Besides, these te…

2023

Towards Nonlinear-Motion-Aware and Occlusion-Robust Rolling Shutter Correction

ICCV 2023poster

This paper addresses the problem of rolling shutter correction in complex nonlinear and dynamic scenes with extreme occlusion. Existing methods suffer from two main drawbacks. Firstly, they face challenges in estimating the accurate correction field due to the uniform velocity assumption, leading t…

Cited by 8PDFcodeScholar
2023

ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding

ICCV 2023poster

Understanding 3D scenes from multi-view inputs has been proven to alleviate the view discrepancy issue in 3D visual grounding. However, existing methods normally neglect the view cues embedded in the text modality and fail to weigh the relative importance of different views. In this paper, we propos…

Cited by 64PDFScholar
2022

Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-training

NeurIPS 2022accept

Masked Autoencoders (MAE) have shown great potentials in self-supervised pre-training for language and 2D image transformers. However, it still remains an open question on how to exploit masked autoencoding for learning 3D representations of irregular point clouds. In this paper, we propose Point-M2…

2022

RCLane: Relay Chain Prediction for Lane Detection

ECCV 2022poster

"Lane detection is an important component of many real-world autonomous systems. Despite a wide variety of lane detection approaches have been proposed, reporting steady benchmark improvements over time, lane detection remains a largely unsolved problem. This is because most of the existing lane det…

Cited by 32SourcePDFScholar
2021

Generating Masks From Boxes by Mining Spatio-Temporal Consistencies in Videos

ICCV 2021poster

Segmenting objects in videos is a fundamental computer vision task. The current deep learning based paradigm offers a powerful, but data-hungry solution. However, current datasets are limited by the cost and human effort of annotating object masks in videos. This effectively limits the performance a…

Cited by 23PDFcodeScholar
2021

PSRR-MaxpoolNMS: Pyramid Shifted MaxpoolNMS With Relationship Recovery

CVPR 2021poster

Non-maximum Suppression (NMS) is an essential post-processing step in modern convolutional neural networks for object detection. Unlike convolutions which are inherently parallel, the de-facto standard for NMS, namely GreedyNMS, cannot be easily parallelized and thus could be the performance bottlen…

Cited by 12PDFcodeScholar
2020

A Continuum Manipulator with Closed-form Inverse Kinematics and Independently Tunable Stiffness

ICRA 2020poster

Continuum manipulators can accomplish various tasks in confined spaces, benefiting from their compliant structures and improved dexterity. Confined and unstructured spaces may require both enhanced stiffness of a continuum manipulator for precision and payload, as well as compliance for safe interac…

Cited by 11SourceScholar
2019

MaxpoolNMS: Getting Rid of NMS Bottlenecks in Two-Stage Object Detectors

CVPR 2019poster

Modern convolutional object detectors have improved the detection accuracy significantly, which in turn inspired the development of dedicated hardware accelerators to achieve real-time performance by exploiting inherent parallelism in the algorithm. Non-maximum suppression (NMS) is an indispensable…

Cited by 39PDFScholar
2018

Continuum Manipulator with Redundant Backbones and Constrained Bending Curvature for Continuously Variable Stiffness

IROS 2018poster

Snake-like manipulators can navigate and perform manipulation in confined spaces. Their recent implementations in surgical robots attracted a lot of attentions. These slender manipulators usually possess either a hyper-redundant articulated vertebrate structure or a continuum one. Primary design con…

Cited by 22SourceScholar