← Search

Zhi Han

16 accepted papers

2026

All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker Adaptation

ICLR 2026poster

Deploying vision-and-language navigation (VLN) agents requires adaptation across diverse scenes and environments, but fine-tuning on a specific scenario often causes catastrophic forgetting in others, which severely limits flexible long-term deployment. We formalize this challenge as the all-day mul…

Cited by 0SourceScholar
2026

B*: Efficient and Optimal Base Placement for Fixed-Base Manipulators

ICRA 2026poster

B* is a novel optimization framework that addresses a critical challenge in fixed-base manipulator robotics: optimal base placement. Current methods rely on pre-computed kinematics databases generated through sampling to search for solutions. However, they face an inherent trade-off between solution…

2026

Bring Your Dreams to Life: Continual Text-to-Video Customization

AAAI 2026technical

Customized text-to-video generation (CTVG) has recently witnessed great progress in generating tailored videos from user-specific text. However, most CTVG methods assume that personalized concepts remain static and do not expand incrementally over time. Additionally, they struggle with forgetting an

Cited by 0SourcePDFScholar
2026

Lifelong Language-Conditioned Robotic Manipulation Learning

AAAI 2026technical

Traditional language-conditioned manipulation agent adaptation to new manipulation skills leads to catastrophic forgetting of old skills, limiting dynamic scene practical deployment. In this paper, we propose SkillsCrafter, a novel robotic manipulation framework designed to continually learn multipl

Cited by 0SourcePDFScholar
2026

SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical Planning

AAAI 2026technical

Sequential-Horizon Vision-and-Language Navigation (SH-VLN) presents a challenging scenario where agents should sequentially execute multi-task trajectory navigation guided by complex, long-horizon natural language instructions. Current vision-and-language navigation models exhibit significant perfor

Cited by 0SourcePDFScholar
2026

The Power of Small Initialization in Noisy Low-Tubal-Rank Tensor Recovery

ICLR 2026poster

We study the problem of recovering a low-tubal-rank tensor $\mathcal{X}\_\star\in \mathbb{R}^{n \times n \times k}$ from noisy linear measurements under the t-product framework. A widely adopted strategy involves factorizing the optimization variable as $\mathcal{U} * \mathcal{U}^\top$, where $\math…

Cited by 0SourceScholar
2026

Unleashing the Potential of Large Language Models for Text-to-Image Generation Through Autoregressive Representation Alignment

AAAI 2026technical

We present Autoregressive Representation Alignment (ARRA), a new training framework that unlocks global-coherent text-to-image generation in autoregressive LLMs without architectural modifications. Different from prior works that require complex architectural redesigns, ARRA aligns LLM

Cited by 0SourcePDFScholar
2026

Vi-TacMan: Articulated Object Manipulation Via Vision and Touch

ICRA 2026poster

Autonomous manipulation of articulated objects represents a basic skill for robots deployed in human environments. Current vision-based methods can infer object hidden kinematics, but their estimates are sometimes imprecise in driving reliable actions, especially on previously unseen objects. Tactil…

2025

GeoRecon: Geometric Coherence for Online 3D Scene Reconstruction From Monocular Video

RA-L 2025

Online 3D scene reconstruction from monocular video aims to incrementally recover 3D mesh from monocular RGB videos. It enables robots to accomplish tasks involving interactions with the environment. Due to the high memory consumption of 3D data, almost all existing methods adopt the coarse-to-fine

Cited by 3SourceScholar
2024

Discovering Syntactic Interaction Clues for Human-Object Interaction Detection

CVPR 2024poster

Recently Vision-Language Model (VLM) has greatly advanced the Human-Object Interaction (HOI) detection. The existing VLM-based HOI detectors typically adopt a hand-crafted template (e.g. a photo of a person [action] a/an [object]) to acquire text knowledge through the VLM text encoder. However such…

Cited by 5SourcePDFScholar
2024

Exploiting Multi-Modal Synergies for Enhancing 3D Multi-Object Tracking

RA-L 2024

3D Multi-Object Tracking (MOT) aims to establish and maintain consistent object trajectories in continuously dynamic environments. At present, the tracking-by-detection has emerged as a dominant paradigm for 3D MOT, due to its simplicity and efficiency. However, this paradigm depends heavily on the

Cited by 5SourceScholar
2017

Tensor RPCA by Bayesian CP Factorization With Complex Noise

ICCV 2017poster

The RPCA model has achieved good performances in various applications. However, two defects limit its effectiveness. Firstly, it is designed for dealing with data in matrix form, which fails to exploit the structure information of higher order tensor data in some pratical situations. Secondly, it ad…

Cited by 23PDFScholar
2017

Video Desnowing and Deraining Based on Matrix Decomposition

CVPR 2017poster

The existing snow/rain removal methods often fail for heavy snow/rain and dynamic scene. One reason for the failure is due to the assumption that all the snowflakes/rain streaks are sparse in snow/rain scenes. The other is that the existing methods often can not differentiate moving objects and snow…

Cited by 199PDFScholar