← Search

Jugang Fan

4 accepted papers

2026

NaVLA$^2$: A Vision-Language-Audio-Action Model for Multimodal Instruction Navigation

AAAI 2026technical

Embodied navigation is a fundamental capability for intelligent agents, yet remains challenging in partially observable environments where navigation instructions can be difficult to interpret. However, existing tasks only provide unimodal instructions, which are ambiguous in complex multimodal envi

Cited by 0SourcePDFScholar
2025

Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image Synthesis

ICCV 2025poster

Visual autoregressive modeling, based on the next-scale prediction paradigm, exhibits notable advantages in image quality and model scalability over traditional autoregressive and diffusion models. It generates images by progressively refining resolution across multiple stages. However, the computat…

2024

UBSoft: A Simulation Platform for Robotic Skill Learning in Unbounded Soft Environments

CoRL 2024poster

It is desired to equip robots with the capability of interacting with various soft materials as they are ubiquitous in the real world. While physics simulations are one of the predominant methods for data collection and robot training, simulating soft materials presents considerable challenges. Spec…

Cited by 1SourcecodeScholar
2023

FGPrompt: Fine-grained Goal Prompting for Image-goal Navigation

NeurIPS 2023poster

Learning to navigate to an image-specified goal is an important but challenging task for autonomous systems like household robots. The agent is required to well understand and reason the location of the navigation goal from a picture shot in the goal position. Existing methods try to solve this prob…

Cited by 13SourcePDFScholar