← Search

Jianhua Sun

16 accepted papers

2026

Empowering Precise Embodied Agents with Executable Analytic Concepts as Semantic-Physical Blueprints

IJCAI 2026

A core challenge for embodied agents is the ``semantic-to-physical gap"—the difficulty of mapping symbolic reasoning to precise execution. While Vision-Language Models (VLMs) enhance agent task planning, they often fail in problem classes requiring accurate alignment between functional geometry and

Cited by 0Scholar
2026

Learning Realistic Depth via Physics-Grounded Noise Disentanglement with Semantic-Geometric Collaboration

ICML 2026poster

Real-world physical sensing exhibits complex, heterogeneous noise patterns that deviate significantly from idealized simulation, posing a fundamental bottleneck for sim-to-real transfer. Existing sensor modelings typically treat depth noise as a monolithic black-box process, overlooking the distinct…

Cited by 0SourceScholar
2026

Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts

CVPR 2026

We humans rely on a wide range of commonsense knowledge to interact with an extensive number and categories of objects in the physical world. Likewise, such commonsense knowledge is also crucial for robots to successfully develop generalized object manipulation skills. While recent advancements in M

Cited by 0SourceScholar
2025

Arti-PG: A Toolbox for Procedurally Synthesizing Large-Scale and Diverse Articulated Objects with Rich Annotations

ICCV 2025poster

The acquisition of substantial volumes of 3D articulated object data is expensive and time-consuming, and consequently the scarcity of 3D articulated object data becomes an obstacle for deep learning methods to achieve remarkable performance in various articulated object understanding tasks. Meanwhi…

2025

Discovering Conceptual Knowledge with Analytic Ontology Templates for Articulated Objects

AAAI 2025technical

Human cognition can leverage fundamental conceptual knowledge, like geometry and kinematic ones, to appropriately perceive, comprehend and interact with novel objects. Motivated by this finding, we aim to endow machine intelligence with an analogous capability through performing at the conceptual le…

2025

Interactive Adjustment for Human Trajectory Prediction with Individual Feedback

ICLR 2025poster

Human trajectory prediction is fundamental for autonomous driving and service robot. The research community has studied various important aspects of this task and made remarkable progress recently. However, there is an essential perspective which is not well exploited in previous research all along,…

Cited by 0SourcePDFScholar
2024

ConceptFactory: Facilitate 3D Object Knowledge Annotation with Object Conceptualization

NeurIPS 2024poster

We present ConceptFactory, a novel scope to facilitate more efficient annotation of 3D object knowledge by recognizing 3D objects through generalized concepts (i.e. object conceptualization), aiming at promoting machine intelligence to learn comprehensive object knowledge from both vision and roboti…

2023

Stimulus Verification Is a Universal and Effective Sampler in Multi-Modal Human Trajectory Prediction

CVPR 2023poster

To comprehensively cover the uncertainty of the future, the common practice of multi-modal human trajectory prediction is to first generate a set/distribution of candidate future trajectories and then sample required numbers of trajectories from them as final predictions. Even though a large number…

Cited by 17SourcePDFScholar
2023

Symbol-LLM: Leverage Language Models for Symbolic System in Visual Human Activity Reasoning

NeurIPS 2023poster

Human reasoning can be understood as a cooperation between the intuitive, associative "System-1'' and the deliberative, logical "System-2''. For existing System-1-like methods in visual activity understanding, it is crucial to integrate System-2 processing to improve explainability, generalization,…

Cited by 15SourcePDFScholar
2022

Correlation Field for Boosting 3D Object Detection in Structured Scenes

AAAI 2022technical

Data augmentation is an efficient way to elevate 3D object detection performance. In this paper, we propose a simple but effective online crop-and-paste data augmentation pipeline for structured 3D point cloud scenes, named CorrelaBoost. Observing that 3D objects should have reasonable relative posi…

Cited by 10SourcePDFScholar
2022

Human Trajectory Prediction With Momentary Observation

CVPR 2022poster

Human trajectory prediction task aims to analyze human future movements given their past status, which is a crucial step for many autonomous systems such as self-driving cars and social robots. In real-world scenarios, it is unlikely to obtain sufficiently long observations at all times for predicti…

Cited by 39PDFScholar
2022

Unified and Fast Human Trajectory Prediction Via Conditionally Parameterized Normalizing Flow

RA-L 2022

Human trajectory prediction is crucial for service robots, autonomous driving and advanced driver assistant systems. Current top-performing methods mainly rely on intractable generative models to learn a distribution of future trajectories, and sample multiple plausible ones as prediction results. I

Cited by 14SourceScholar
2021

Three Steps to Multimodal Trajectory Prediction: Modality Clustering, Classification and Synthesis

ICCV 2021poster

Multimodal prediction results are essential for trajectory prediction task as there is no single correct answer for the future. Previous frameworks can be divided into three categories: regression, generation and classification frameworks. However, these frameworks have weaknesses in different aspec…

Cited by 89PDFScholar
2019

InstaBoost: Boosting Instance Segmentation via Probability Map Guided Copy-Pasting

ICCV 2019poster

Instance segmentation requires a large number of training samples to achieve satisfactory performance and benefits from proper data augmentation. To enlarge the training set and increase the diversity, previous methods have investigated using data annotation from other domain (e.g. bbox, point) in a…

Cited by 253PDFcodeScholar