← Search

Fengda Zhu

12 accepted papers

2025

Goku: Flow Based Video Generative Foundation Models

CVPR 2025highlight

This paper introduces Goku, a state-of-the-art family of joint image-and-video generation models leveraging rectified flow Transformers to achieve industry-leading performance. We detail the foundational elements enabling high-quality visual generation, including the data curation pipeline, model ar…

Cited by 15SourcePDFScholar
2025

InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation

NeurIPS 2025oral

We introduce InfinityStar, a unified spacetime autoregressive framework for high-resolution image and dynamic video synthesis. Building on the recent success of autoregressive modeling in both vision and language, our purely discrete approach jointly captures spatial and temporal dependencies within…

Cited by 0SourceScholar
2023

Vision Language Navigation with Knowledge-driven Environmental Dreamer

IJCAI 2023poster

Vision-language navigation (VLN) requires an agent to perceive visual observation in a house scene and navigate step-by-step following natural language instruction. Due to the high cost of data annotation and data collection, current VLN datasets provide limited instruction-trajectory data samples.…

Cited by 2SourcePDFScholar
2022

Contrastive Instruction-Trajectory Learning for Vision-Language Navigation

AAAI 2022technical

The vision-language navigation (VLN) task requires an agent to reach a target with the guidance of natural language instruction. Previous works learn to navigate step-by-step following an instruction. However, these works may fail to discriminate the similarities and discrepancies across instruction…

2022

Visual-Language Navigation Pretraining via Prompt-based Environmental Self-exploration

ACL 2022long

Vision-language navigation (VLN) is a challenging task due to its large searching space in the environment. To address this problem, previous works have proposed some methods of fine-tuning a large model that pretrained on large-scale datasets. However, the conventional fine-tuning methods require e…

2021

SOON: Scenario Oriented Object Navigation With Graph-Based Exploration

CVPR 2021poster

The ability to navigate like a human towards a language-guided target from anywhere in a 3D embodied environment is one of the 'holy grail' goals of intelligent robots. Most visual navigation benchmarks, however, focus on navigating toward a target from a fixed starting point, guided by an elaborate…

Cited by 131PDFcodeScholar
2021

Self-Motivated Communication Agent for Real-World Vision-Dialog Navigation

ICCV 2021poster

Vision-Dialog Navigation (VDN) requires an agent to ask questions and navigate following the human responses to find target objects. Conventional approaches are only allowed to ask questions at predefined locations, which are built upon expensive dialogue annotations, and inconvenience the real-word…

Cited by 35PDFScholar
2021

UPDeT: Universal Multi-agent RL via Policy Decoupling with Transformers

ICLR 2021spotlight

Recent advances in multi-agent reinforcement learning have been largely limited in training one model from scratch for every new task. The limitation is due to the restricted model architecture related to fixed input and output dimensions. This hinders the experience accumulation and transfer of the…

Cited by 0SourcePDFScholar
2021

Vision-Language Navigation With Random Environmental Mixup

ICCV 2021poster

Vision-language Navigation (VLN) task requires an agent to perceive both the visual scene and natural language and navigate step-by-step. Large data bias makes the VLN task challenging, which is caused by the disparity ratio between small data scale and large navigation space. Previous works have pr…

Cited by 99PDFcodeScholar
2020

Vision-Dialog Navigation by Exploring Cross-Modal Memory

CVPR 2020poster

Vision-dialog navigation posed as a new holy-grail task in vision-language disciplinary targets at learning an agent endowed with the capability of constant conversation for help with natural language and navigating according to human responses. Besides the common challenges faced in visual language…

Cited by 55PDFcodeScholar