← Search

Junshen K. Chen

2 accepted papers

2025

Storybooth: Training-Free Multi-Subject Consistency for Improved Visual Storytelling

ICLR 2025poster

Consistent text-to-image generation depicting the *same* subjects across different images has gained significant recent attention due to its widespread applications in the fields of visual-storytelling and multiple-shot video generation. While remarkable, existing methods often require costly finet…

Cited by 0SourcePDFScholar
2021

Topological Planning With Transformers for Vision-and-Language Navigation

CVPR 2021poster

Conventional approaches to vision-and-language navigation (VLN) are trained end-to-end but struggle to perform well in freely traversable environments. Inspired by the robotics community, we propose a modular approach to VLN using topological maps. Given a natural language instruction and topologica…

Cited by 129PDFScholar