← Search

Na Zhao

58 accepted papers

2026

Anatomical Domain Shifts: Test-time Heterogeneous Adaptation for 3D Human Pose Prediction

CVPR 2026

The research frontier in human pose prediction (HPP) is advancing toward continual test-time adaptation (TTA), where models must self-adapt to dynamic test distributions. To date, the homeostatic continual TTA remains the sole viable solution, which isolates the model parameters and update domain-se

Cited by 0SourceScholar
2026

Artemis: Structured Visual Reasoning for Perception Policy Learning

ICML 2026poster

Recent reinforcement-learning frameworks for visual perception policy usually incorporate intermediate reasoning chains expressed in natural language. Empirical observations indicate that such purely linguistic intermediate reasoning often reduces performance on perception tasks. We argue that the c…

Cited by 0SourceScholar
2026

CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection

CVPR 2026

Multi-modal fusion has emerged as a promising paradigm for accurate 3D object detection. However, performance degrades substantially when deployed in target domains different from training. In this work, focusing on dual-branch proposal-level detectors, we identify two factors that limit robust cros

Cited by 0SourcecodeScholar
2026

DuoCast: Duo-Probabilistic Diffusion for Precipitation Nowcasting

AAAI 2026technical

Accurate short-term precipitation forecasting is critical for weather-sensitive decision-making in agriculture, transportation, and disaster response. Existing deep learning approaches often struggle to balance global structural consistency with local detail preservation, especially under complex me

Cited by 0SourcePDFScholar
2026

Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception

CVPR 2026

Training Large Multimodality Models (LMMs) relies on descriptive image caption that connects image and language. Existing methods for generating such captions often rely on distilling the captions from pretrained LMMs, constructing them from publicly available internet images, or even generating the

Cited by 0SourcecodeScholar
2026

Few-Shot Incremental 3D Object Detection in Dynamic Indoor Environments

CVPR 2026

Incremental 3D object perception is a critical step toward embodied intelligence in dynamic indoor environments. However, existing incremental 3D detection methods rely on extensive annotations of novel classes for satisfactory performance. To address this limitation, we propose FI3Det, a Few-shot I

Cited by 0SourcecodeScholar
2026

Graph Smoothing for Enhanced Local Geometry Learning in Point Cloud Analysis

AAAI 2026technical

Graph-based methods have proven to be effective in capturing relationships among points for 3D point cloud analysis. However, these methods often suffer from suboptimal graph structures, particularly due to sparse connections at boundary points and noisy connections in junction areas. To address the

Cited by 0SourcePDFScholar
2026

Multi-label learning with contrastive cluster self-supervision for 3D hierarchical semantic segmentation

ICML 2026poster

3D hierarchical semantic segmentation (3DHS) is crucial for embodied intelligence that demands the coarse-to-fine grained and multi-hierarchy understanding of 3D scenes. 3DHS tasks can be addressed by multi-label learning, but facing two issues: I) learning multiple labels for each point with a shar…

Cited by 0SourceScholar
2026

PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving

CVPR 2026

This paper presents the first study on Unsupervised Domain Adaptation (UDA) for multimodal 3D panoptic segmentation (mm-3DPS), aiming to improve generalization under domain shifts commonly encountered in real-world autonomous driving. A straightforward solution is to employ a pseudo-labeling strateg

Cited by 0SourceScholar
2026

Path Tracking Control for a Transformable Wheel-Legged Robot Using Model Predictive Control

ICRA 2026poster

Transformable wheel-legged robots can adjust their configuration according to terrain conditions, enabling effective operation in harsh environments. While existing controllers based on preset commands have successfully demonstrated the feasibility of reconfigurable mechanisms, they still struggle t…

Cited by 0Scholar
2026

Time-Optimal Trajectory Planning and Model Predictive Control of Morphing Quadrotors

ICRA 2026poster

Morphing quadrotors offer enhanced maneuverability and adaptability in confined spaces, while their structural variations pose challenges to trajectory planning and control. This paper presents a time-optimal trajectory planning and model predictive control framework for the morphing quadrotor. The …

Cited by 0Scholar
2026

TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models

ICML 2026poster

Large vision-language models (LVLMs) excel at vision-language tasks but remain vulnerable to backdoor attacks. Most existing backdoor attacks on LVLMs force the model to generate predefined target patterns. However, these fixed-pattern attacks are easy to detect, as the model tends to memorize frequ…

Cited by 0SourceScholar
2026

ViLoMem: Agentic Learner with Grow-and-Refine Multimodal Semantic Memory

CVPR 2026

MLLMs exhibit strong reasoning on isolated queries, yet they operate de novo--solving each problem independently and often repeating the same mistakes. Existing memory-augmented agents mainly store past trajectories for reuse. However, trajectory-based memory suffers from brevity bias, gradually los

Cited by 0SourcecodeScholar
2025

AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language Models

NeurIPS 2025poster

Effective human-agent collaboration in physical environments requires understanding not only what to act upon, but also where the actionable elements are and how to interact with them. Existing approaches often operate at the object level or disjointedly handle fine-grained affordance reasoning, lac…

Cited by 0SourceScholar
2025

AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based Referring

AAAI 2025technical

3D visual grounding (3DVG), which aims to correlate a natural language description with the target object within a 3D scene, is a significant yet challenging task. Despite recent advancements in this domain, existing approaches commonly encounter a shortage: a limited amount and diversity of text-3D…

Cited by 1SourcePDFScholar
2025

Collaborative Tree Search for Enhancing Embodied Multi-Agent Collaboration

CVPR 2025poster

Embodied agents based on large language models (LLMs) face significant challenges in collaborative tasks, requiring effective communication and reasonable division of labor to ensure efficient and correct task completion. Previous approaches with simple communication patterns carry erroneous or inco…

Cited by 0SourcePDFScholar
2025

Data-Driven MPC for Attitude Control of Autonomous Underwater Robot

IROS 2025

High maneuverability is essential to the autonomous operation of underwater robots. To achieve real-time maneuvering motion, the control strategy must take into account nonlinear hydrodynamic effects, which are extremely difficult to accurately capture during motion and therefore a balance must be s

Cited by 1SourceScholar
2025

Efficient Cross-Boundary Grasping in Stacked Clutter with Single-Visual Mapping Multi-Step

ICRA 2025

In logistics applications, the vision-based technology for grasping target objects in the air is relatively mature. However, when operating across the air and water such as grasping marine products from the water, the visual information collected by the camera will be disturbed by ripples and bubble

Cited by 0SourceScholar
2025

GaussianBlock: Building Part-Aware Compositional and Editable 3D Scene by Primitives and Gaussians

ICLR 2025poster

Recently, with the development of Neural Radiance Fields and Gaussian Splatting, 3D reconstruction techniques have achieved remarkably high fidelity. However, the latent representations learnt by these methods are highly entangled and lack interpretability. In this paper, we propose a novel part-awa…

Cited by 0SourcePDFScholar
2025

Geometric Alignment and Prior Modulation for View-Guided Point Cloud Completion on Unseen Categories

ICCV 2025poster

View-Guided Point Cloud Completion (VG-PCC) aims to reconstruct complete point clouds from partial inputs by referencing single-view images. While existing VG-PCC models perform well on in-class predictions, they exhibit significant performance drops when generalizing to unseen categories. We identi…

Cited by 0SourcePDFScholar
2025

Graph Embedded Contrastive Learning for Multi-View Clustering

IJCAI 2025

Recently, numerous multi-view clustering (MVC) and multi-view graph clustering (MVGC) methods have been proposed. Despite significant progress, they still face two issues: I) MVC and MVGC are often developed independently for multi-view and multi-graph data. They have redundancy but lack a unified m

2025

How Do Images Align and Complement LiDAR? Towards a Harmonized Multi-modal 3D Panoptic Segmentation

ICML 2025poster

LiDAR-based 3D panoptic segmentation often struggles with the inherent sparsity of data from LiDAR sensors, which makes it challenging to accurately recognize distant or small objects. Recently, a few studies have sought to overcome this challenge by integrating LiDAR inputs with camera images, leve…

2025

MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion Paradigm

ICCV 2025poster

Human motion generation and editing are key components of computer vision. However, current approaches in this field tend to offer isolated solutions tailored to specific tasks, which can be inefficient and impractical for real-world applications. While some efforts have aimed to unify motion-relate…

2025

Provably Secure Image Robust Steganography via Cross-modal Error Correction

AAAI 2025technical

The rapid development of image generation models has facilitated the widespread dissemination of generated images on social networks, creating favorable conditions for provably secure image steganography. However, existing methods face issues such as low quality of generated images and lack of sema…

Cited by 0SourcePDFScholar
2025

Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated Perturbation

ICCV 2025poster

Recently, multi-view learning (MVL) has garnered significant attention due to its ability to fuse discriminative information from multiple views. However, real-world multi-view datasets are often heterogeneous and imperfect, which usually causes MVL methods designed for specific combinations of view…

2025

STEAD: Robust Provably Secure Linguistic Steganography with Diffusion Language Model

NeurIPS 2025poster

Recent provably secure linguistic steganography (PSLS) methods rely on mainstream autoregressive language models (ARMs) to address historically challenging tasks, that is, to disguise covert communication as ``innocuous'' natural language communication. However, due to the characteristic of sequen…

Cited by 0SourceScholar
2025

Uncertainty Meets Diversity: A Comprehensive Active Learning Framework for Indoor 3D Object Detection

CVPR 2025poster

Active learning has emerged as a promising approach to reduce the substantial annotation burden in 3D object detection tasks, spurring several initiatives in outdoor environments. However, its application in indoor environments remains unexplored. Compared to outdoor 3D datasets, indoor datasets fac…

2024

A Multi-modal Hybrid Robot with Enhanced Traversal Performance*

ICRA 2024poster

Current multi-modal hybrid robots with flight and wheeled modes have fallen into the dilemma that they can only avoid obstacles by re-taking off when encountering obstacles due to the poor performance of wheeled obstacle-crossing. To tackle this problem, this paper presents a novel multi-modal hybri…

Cited by 0SourceScholar
2024

Attitude Control for Morphing Quadrotor through Model Predictive Control with Constraints*

ICRA 2024poster

Morphing quadrotors that can be potentially applied to confined spaces such as warehouses, tanks, and pipelines have flourished in recent years. Most work has focused on the mechanical feasibility of the morphing systems and high-level flight controller design, with limited discussions on low-level…

Cited by 2SourceScholar
2024

Dual-Perspective Knowledge Enrichment for Semi-supervised 3D Object Detection

AAAI 2024technical

Semi-supervised 3D object detection is a promising yet under-explored direction to reduce data annotation costs, especially for cluttered indoor scenes. A few prior works, such as SESS and 3DIoUMatch, attempt to solve this task by utilizing a teacher model to generate pseudo-labels for unlabeled sam…

2024

LASO: Language-guided Affordance Segmentation on 3D Object

CVPR 2024poster

Segmenting affordance in 3D data is key for bridging perception and action in robots. Existing efforts mostly focus on the visual side and overlook the affordance knowledge from a semantic aspect. This oversight not only limits their generalization to unseen objects but more importantly hinders thei…

2024

Model Predictive Control for an Autonomous Underwater Robot with Fully Vectored Propulsion

ICRA 2024poster

Due to the low motion efficiency and maneuver-ability of underwater robots with six degrees of freedom, it is challenging for them to respond quickly to the attitude requirements during underwater autonomous manipulation. This paper presents a novel autonomous underwater robot with fully vectored pr…

Cited by 3SourceScholar
2024

View-Consistent 3D Editing with Gaussian Splatting

ECCV 2024poster

"The advent of 3D Gaussian Splatting (3DGS) has revolutionized 3D editing, offering efficient, high-fidelity rendering and enabling precise local manipulations. Currently, diffusion-based 2D editing models are harnessed to modify multi-view rendered images, which then guide the editing of 3DGS model…

Cited by 25SourcePDFScholar
2023

Refining 6-DoF Grasps with Context-Specific Classifiers

IROS 2023poster

In this work, we present GraspFlow, a refinement approach for generating context-specific grasps. We formulate the problem of grasp synthesis as a sampling problem: we seek to sample from a context-conditioned probability distribution of successful grasps. However, this target distribution is unknow…

Cited by 2SourcecodeScholar
2022

Patch Steganalysis: A Sampling Based Defense Against Adversarial Steganography

ICASSP 2022accepted

In recent years, the classification accuracy of CNN (convolutional neural network) steganalyzers has rapidly improved. However, as general CNN classifiers will misclassify adversarial samples, CNN steganalyzers can hardly detect adversarial steganography, which combines adversarial samples and stega…

Cited by 0SourceScholar
2022

Rethinking IoU-Based Optimization for Single-Stage 3D Object Detection

ECCV 2022poster

"Since Intersection-over-Union (IoU) based optimization maintains the consistency of the final IoU prediction metric and losses, it has been widely used in both regression and classification branches of single-stage 2D object detectors. Recently, several 3D object detection methods adopt IoU-based o…

2022

Style-Hallucinated Dual Consistency Learning for Domain Generalized Semantic Segmentation

ECCV 2022poster

"In this paper, we study the task of synthetic-to-real domain generalized semantic segmentation, which aims to learn a model that is robust to unseen real-world scenes using only synthetic data. The large domain shift between synthetic and real-world data, including the limited source environmental…

2022

Teaching with Soft Label Smoothing for Mitigating Noisy Labels in Facial Expressions

ECCV 2022poster

"Recent studies have highlighted the problem of noisy labels in large scale in-the-wild facial expressions datasets due to the uncertainties caused by ambiguous facial expressions, low-quality facial images, and the subjectiveness of annotators. To solve the problem of noisy labels, we propose Soft…

2021

Comparative Validation Study on Bioinspired Morphology-Adaptation Flight Performance of a Morphing Quad-Rotor

RA-L 2021

Inspired by the advantages from natural birds' shape-morphing adapted flight performance, this work investigates how the bioinspired adaptative morph induced inertia variation affects the flight performance of a morphing quad-rotor. Extensive numerical and experimental evaluations on a custom-built

Cited by 20SourceScholar
2021

Morphologically Adapatative Quad-Rotor Towards Acquiring High-Performance Flight: A Comparative Study and Validation

ICRA 2021poster

This paper presents our comparative study on how the flight performances of an in-flight morphing quad-rotor are affected by the morph induced inertia variation. A custom-built in-flight morphing quad-rotor was employed in numerical and experimental tests for the study and analysis. In these tests,…

Cited by 0SourceScholar
2020

Distributed Consensus Control of Multiple UAVs in a Constrained Environment

ICRA 2020poster

In this paper, we investigate the consensus problem of multiple unmanned aerial vehicles (UAVs) in the presence of environmental constraints under a general communication topology containing a directed spanning tree. First, based on a position transformation function, we propose a novel dynamic refe…

Cited by 15SourceScholar
2020

Faster Healthcare Time Series Classification for Boosting Mortality Early Warning System

IROS 2020poster

Electronic Health Record (EHR) and healthcare claim data provide rich clinical information for time series analysis. In this work, we provide a different angle of solving healthcare multivariate time series classification by turning it into a computer vision problem. We propose a Convolutional Featu…

Cited by 4SourceScholar
2019

An Approximation-Free Simple Control Scheme for Uncertain Quadrotor Systems: Theory and Validations

IROS 2019poster

In this paper, a simple tracking control scheme is proposed for quadrotor systems with uncertain dynamics. It precludes the necessity for prohibitive analytic computation of the derivatives of the desired (virtual) attitude that is typically employed in controlling quadrotor systems. Moreover, this…

Cited by 3SourceScholar
2018

Inchworm Locomotion Mechanism Inspired Self-Deformable Capsule-Like Robot: Design, Modeling, and Experimental Validation

ICRA 2018poster

Inspired by the inchworm locomotion mechanism, this paper presents our recently developed self-deformable capsule-like robot. The robot has the actuated deformation capability that relies on a novel rigid elements-based morphing structure (REMS) and its soft actuation mechanisms. When the robot defo…

Cited by 16SourceScholar
2018

The Deformable Quad-Rotor Enabled and Wasp-Pedal-Carrying Inspired Aerial Gripper

IROS 2018poster

The paper presents the development of a novel deformable quad-rotor enabled aerial gripper. The mechanism of our deformable quad-rotor is based on simultaneous expansion or contraction of the quad-rotor body, which is generated by controlling a rigid elements based morphing structure (REMS). Such de…

Cited by 25SourceScholar
2017

Design, modeling and experimental validation of a scissor mechanisms enabled compliant modular earthworm-like robot

IROS 2017poster

Inspired by natural earthworm locomotion behavior and segmental muscle motion mechanism, this paper presents our recently developed compliant modular earthwormlike robot with the novel segmental muscle-mimetic design unit that is capable of efficiently mimicking earthworms' segmental muscle contract…

Cited by 26SourceScholar
2017

The deformable quad-rotor: Design, kinematics and dynamics characterization, and flight performance validation

IROS 2017poster

To improve the obstacle surmounting performance of the quad-rotor vehicle, this paper focuses on designing, kinematically and dynamically characterizing a novel deformable quad-rotor that is based on the scissor-like foldable structures. The foldable structure allows that the volume of the quad-roto…

Cited by 72SourceScholar