← Search

Junhao Cai

17 accepted papers

2026

Cross-Modal Dynamic Hypergraph Computation via Functional-Structural Brain Network for Brain Disorder Diagnosis

IJCAI 2026

Cross-modal brain networks characterize the complex connections between different brain regions from both functional and structural perspectives, which is of significant importance for brain network analysis and the diagnosis of brain diseases. However, existing methods have failed to fully exploit

Cited by 0Scholar
2026

HCF: Hierarchical Cascade Framework for Distributed Multi-Stage Image Compression

AAAI 2026technical

Distributed multi-stage image compression—where visual content traverses multiple processing nodes under varying quality requirements—poses challenges. Progressive methods enable bitstream truncation but underutilize available compute resources; successive compression repeats costly pixel-domain ope

Cited by 0SourcePDFScholar
2026

InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy

CVPR 2026

Recent work explores how real and synthetic data contribute to VLA model generalization. While the \pi-series model has shown the strong effectiveness of large-scale real-robot pre-training, synthetic data has not previously demonstrated comparable capability at scale.This paper provides the first e

Cited by 0SourceScholar
2024

GIC: Gaussian-Informed Continuum for Physical Property Identification and Simulation

NeurIPS 2024oral

This paper studies the problem of estimating physical properties (system identification) through visual observations. To facilitate geometry-aware guidance in physical property estimation, we introduce a novel hybrid framework that leverages 3D Gaussian representation to not only capture explicit sh…

2024

IPoD: Implicit Field Learning with Point Diffusion for Generalizable 3D Object Reconstruction from Single RGB-D Images

CVPR 2024highlight

Generalizable 3D object reconstruction from single-view RGB-D images remains a challenging task particularly with real-world data. Current state-of-the-art methods develop Transformer-based implicit field learning necessitating an intensive learning paradigm that requires dense query-supervision uni…

2024

Open-Vocabulary Category-Level Object Pose and Size Estimation

RA-L 2024

This letter studies a new open-set problem, the open-vocabulary category-level object pose and size estimation. Given human text descriptions of arbitrary novel object categories, the robot agent seeks to predict the position, orientation, and size of the target object in the observed scene image. T

Cited by 11SourceScholar
2023

ERRA: An Embodied Representation and Reasoning Architecture for Long-Horizon Language-Conditioned Manipulation Tasks

RA-L 2023

This letter introduces ERRA, an embodied learning architecture that enables robots to jointly obtain three fundamental capabilities (reasoning, planning, and interaction) for solving long-horizon language-conditioned manipulation tasks. ERRA is based on tightly-coupled probabilistic inferences at tw

Cited by 17SourceScholar
2023

Flipbot: Learning Continuous Paper Flipping via Coarse-to-Fine Exteroceptive-Proprioceptive Exploration

ICRA 2023poster

This paper tackles the task of singulating and grasping paper-like deformable objects. We refer to such tasks as paper-flipping. In contrast to manipulating deformable objects that lack compression strength (such as shirts and ropes), minor variations in the physical properties of the paper-like def…

Cited by 4SourcecodeScholar
2023

Learn to Grasp Via Intention Discovery and Its Application to Challenging Clutter

RA-L 2023

Humans excel in grasping objects through diverse and robust policies, many of which are so probabilistically rare that exploration-based learning methods hardly observe and learn. Inspired by the human learning process, we propose a method to extract and exploit latent intents from demonstrations, a

Cited by 1SourceScholar
2022

Open-World Semantic Segmentation for LIDAR Point Clouds

ECCV 2022poster

"Classical LIDAR semantic segmentation is not robust for real-world applications, e.g., autonomous driving, since it is closed-set and static. The closed-set network is only able to output labels of trained classes, even for objects never seen before, while a static network cannot update its knowled…

2022

Real-Time Collision-Free Grasp Pose Detection With Geometry-Aware Refinement Using High-Resolution Volume

RA-L 2022

In this letter, we proposea novel vision-based grasp system for closed-loop 6-degrees of freedom grasping of unknown objects in cluttered environments. The key factor in our system is that we make the most of a geometry-aware scene representation based on a truncated signed distance function (TSDF)

Cited by 22SourceScholar
2022

Uncertainty-based Exploring Strategy in Densely Cluttered Scenes for Vacuum Cup Grasping

ICRA 2022poster

Grasping a wide range of novel objects in densely cluttered scenes is difficult due to irregular shapes of objects and the uncertainty in sensing. In this paper, a novel vacuum cup grasping method, based on uncertainty modeling of perception data and grasp geometric heuristics, is proposed to grasp…

Cited by 8SourceScholar
2022

Volumetric-based Contact Point Detection for 7-DoF Grasping

CoRL 2022poster

In this paper, we propose a novel grasp pipeline based on contact point detection on the truncated signed distance function (TSDF) volume to achieve closed-loop 7-degree-of-freedom (7-DoF) grasping on cluttered environments. The key aspects of our method are that 1) the proposed pipeline exploits th…

Cited by 11SourcecodeScholar
2019

MetaGrasp: Data Efficient Grasping by Affordance Interpreter Network

ICRA 2019poster

Data-driven approach for grasping shows significant advance recently. But these approaches usually require much training data. To increase the efficiency of grasping data collection, this paper presents a novel grasp training system including the whole pipeline from data collection to model inferenc…

Cited by 57SourceScholar
2018

Fusing Object Context to Detect Functional Area for Cognitive Robots

ICRA 2018poster

A cognitive robot usually needs to perform multiple tasks in practice and needs to locate the desired area for each task. Since deep learning has achieved substantial progress in image recognition, to solve this area detection problem, it is straightforward to label a functional area (affordance) im…

Cited by 0SourceScholar