← Search

Boyang Sun

22 accepted papers

2026

FUN REC * Reconstructing Functional 3D Scenes from Egocentric Interaction Videos

CVPR 2026

We present FunREC, a method for reconstructing functional 3D digital twins of indoor scenes directly from egocentric RGB-D interaction videos. Unlike existing methods on articulated reconstruction, which rely on controlled setups, multi-state captures, or CAD priors, FunREC operates directly on in-t

Cited by 0SourcecodeScholar
2026

Independence Test for Linear Non-Gaussian Data and Applications in Causal Discovery

ICLR 2026poster

Independence testing involves determining whether two variables are independent based on observed samples, which is a fundamental problem in statistics and machine learning. Existing testing methods, such as HSIC, can theoretically detect broad forms of dependence, but may sacrifice statistical powe…

Cited by 0SourceScholar
2026

Loop Closure From Two Views: Revisiting PGO for Scalable Trajectory Estimation Through Monocular Priors

RA-L 2026

(Visual) Simultaneous Localization and Mapping (SLAM) remains a fundamental challenge in enabling autonomous systems to navigate and understand large-scale environments. Traditional SLAM approaches struggle to balance efficiency and accuracy, particularly in large-scale settings where extensive comp

Cited by 2SourceScholar
2026

OpenFrontier: General Navigation with Visual-Language Grounded Frontiers

RSS 2026poster

Open-world navigation requires robots to make decisions in complex everyday environments while adapting to flexible task requirements. Conventional navigation approaches often rely on dense 3D reconstruction and hand-crafted goal metrics, which limits their generalization across tasks and environmen…

Cited by 0SourceScholar
2026

REACT3D: Recovering Articulations for Interactive Physical 3D Scenes

RA-L 2026

Interactive 3D scenes are increasingly vital for embodied intelligence, yet existing datasets remain limited due to the labor-intensive process of annotating part segmentation, kinematic types, and motion trajectories. We present REACT3D, a scalable zero-shot framework that converts static 3D scenes

Cited by 2SourceScholar
2026

Selection, Reflection and Self-Refinement: Revisit Reasoning Tasks via a Causal Lens

ICLR 2026poster

Due to their inherent complexity, reasoning tasks have long been regarded as rigorous benchmarks for assessing the capabilities of machine learning models, especially large language models (LLMs). Although humans can solve these tasks with ease, existing models, even after extensive pre-training and…

Cited by 0SourcecodeScholar
2026

Sight Over Site: Perception-Aware Reinforcement Learning for Efficient Robotic Inspection

ICRA 2026poster

Autonomous inspection is a central problem in robotics, with applications ranging from industrial monitoring to search-and-rescue. Traditionally, inspection has often been reduced to navigation tasks, where the objective is to reach a predefined location while avoiding obstacles. However, this formu…

2025

A Conditional Independence Test in the Presence of Discretization

ICLR 2025poster

Testing conditional independence (CI) has many important applications, such as Bayesian network learning and causal discovery. Although several approaches have been developed for learning CI structures for observed variables, those existing methods generally fail to work when the variables of intere…

2025

A Sample Efficient Conditional Independence Test in the Presence of Discretization

ICML 2025poster

Conditional independence (CI) test is a fundamental concept in statistics. In many real-world scenarios, some variables may be difficult to measure accurately, often leading to data being represented as discretized values. Applying CI tests directly to discretized data, however, can lead to incorrec…

2025

ActLoc: Learning to Localize on the Move via Active Viewpoint Selection

CoRL 2025poster

Reliable localization is critical for robot navigation, yet many existing systems assume that all viewpoints along a trajectory are equally informative. In practice, localization becomes unreliable when the robot observes unmapped, ambiguous, or uninformative regions. To address this, we present Act…

Cited by 0SourceScholar
2025

FrontierNet: Learning Visual Cues to Explore

RA-L 2025

Exploration of unknown environments is crucial for autonomous robots; it allows them to actively reason and decide on what new data to acquire for different tasks, such as mapping, object discovery, and environmental assessment. Existing solutions, such as frontier-based exploration approaches, rely

Cited by 12SourcecodeScholar
2025

Gene Regulatory Network Inference in the Presence of Selection Bias and Latent Confounders

NeurIPS 2025poster

Gene regulatory network inference (GRNI) aims to discover how genes causally regulate each other from gene expression data. It is well-known that statistical dependencies in observed data do not necessarily imply causation, as spurious dependencies may arise from *latent confounders*, such as non-co…

Cited by 0SourceScholar
2025

Permutation-based Rank Test in the Presence of Discretization and Application in Causal Discovery with Mixed Data

ICML 2025poster

Recent advances have shown that statistical tests for the rank of cross-covariance matrices play an important role in causal discovery. These rank tests include partial correlation tests as special cases and provide further graphical information about latent variables. Existing rank tests typically…

2025

SuperDec: 3D Scene Decomposition with Superquadrics Primitives

ICCV 2025poster

We present SuperDec, an approach for compact 3D scene representations based on geometric primitives, namely superquadrics.While most recent works leverage geometric primitives to obtain photorealistic 3D scene representations, we propose to leverage them to obtain a compact yet expressive representa…

Cited by 0SourcePDFScholar
2025

VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

CVPR 2025poster

Future robots are envisioned as versatile systems capable of performing a variety of household tasks. The big question remains, how can we bridge the embodiment gap while minimizing physical robot learning, which fundamentally does not scale well. We argue that learning from in-the-wild human videos…

Cited by 0SourcePDFScholar
2024

Active Visual Localization for Multi-Agent Collaboration: A Data-Driven Approach

ICRA 2024poster

Rather than having each newly deployed robot create its own map of its surroundings, the growing availability of SLAM-enabled devices provides the option of simply localizing in a map of another robot or device. In cases such as multi-robot or human-robot collaboration, localizing all agents in the…

Cited by 6SourceScholar
2024

Identifying Latent State-Transition Processes for Individualized Reinforcement Learning

NeurIPS 2024poster

The application of reinforcement learning (RL) involving interactions with individuals has grown significantly in recent years. These interactions, influenced by factors such as personal preferences and physiological differences, causally influence state transitions, ranging from health conditions i…

Cited by 3SourcePDFScholar
2024

Learning Where to Look: Self-supervised Viewpoint Selection for Active Localization using Geometrical Information

ECCV 2024poster

"Accurate localization in diverse environments is a fundamental challenge in computer vision and robotics. The task involves determining a sensor’s precise position and orientation, typically a camera, within a given space. Traditional localization methods often rely on passive sensing, which may st…

2024

NeRF On-the-go: Exploiting Uncertainty for Distractor-free NeRFs in the Wild

CVPR 2024poster

Neural Radiance Fields (NeRFs) have shown remarkable success in synthesizing photorealistic views from multi-view images of static scenes but face challenges in dynamic real-world environments with distractors like moving objects shadows and lighting changes. Existing methods manage controlled envir…

2023

Subspace Identification for Multi-Source Domain Adaptation

NeurIPS 2023spotlight

Multi-source domain adaptation (MSDA) methods aim to transfer knowledge from multiple labeled source domains to an unlabeled target domain. Although current methods achieve target joint distribution identifiability by enforcing minimal changes across domains, they often necessitate stringent conditi…

2022

See Yourself in Others: Attending Multiple Tasks for Own Failure Detection

ICRA 2022poster

Autonomous robots deal with unexpected scenarios in real environments. Given input images, various visual perception tasks can be performed, e.g., semantic segmentation, depth estimation and normal estimation. These different tasks provide rich information for the whole robotic perception system. Al…

Cited by 12SourcecodeScholar