← Search

Zhaohuan Zhan

4 accepted papers

2026

HouseTune: Two-Stage Floorplan Generation with LLM Assistance

AAAI 2026technical

This paper proposes a two-stage text-to-floorplan generation framework that combines the reasoning capability of Large Language Models (LLMs) with the generative power of diffusion models. In the first stage, we leverage a Chain-of-Thought (CoT) prompting strategy to guide an LLM in generating an in

Cited by 0SourcePDFScholar
2026

UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model

AAAI 2026technical

Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions—remains highly challenging. Recent research on enhancing language-guided navigation reasoning using pre-trained large language models (LLMs) has show

Cited by 0SourcePDFScholar
2025

Personalized Subgraph Federated Learning with Differentiable Auxiliary Projections

NeurIPS 2025poster

Federated Learning (FL) on graph-structured data typically faces non-IID challenges, particularly in scenarios where each client holds a distinct subgraph sampled from a global graph. In this paper, we introduce **Fed**erated learning with **Aux**iliary projections (FedAux), a personalized subgraph…

Cited by 0SourcecodeScholar
2020

Vision-Dialog Navigation by Exploring Cross-Modal Memory

CVPR 2020poster

Vision-dialog navigation posed as a new holy-grail task in vision-language disciplinary targets at learning an agent endowed with the capability of constant conversation for help with natural language and navigating according to human responses. Besides the common challenges faced in visual language…

Cited by 55PDFcodeScholar