← Search

Xiao Chu

6 accepted papers

2026

WorldGen: From Text to Traversable and Interactive 3D Worlds

CVPR 2026

We introduce WorldGen, a method for generating large, fully formed, navigable 3D worlds from a single text prompt. Existing approaches to 3D scene generation often trade off scene diversity, completeness, and correctness in different ways. We push this envelope by producing large scenes explicitly d

Cited by 0SourceScholar
2018

Visual Question Generation as Dual Task of Visual Question Answering

CVPR 2018poster

Visual question answering (VQA) and visual question generation (VQG) are two trending topics in the computer vision, but they are usually explored separately despite their intrinsic complementary relationship. In this paper, we propose an end-to-end unified model, the Invertible Question Answering N…

Cited by 198SourcePDFScholar
2017

Multi-Context Attention for Human Pose Estimation

CVPR 2017poster

In this paper, we propose to incorporate convolutional neural networks with a multi-context attention mechanism into an end-to-end framework for human pose estimation. We adopt stacked hourglass networks to generate attention maps from features at multiple resolutions with various semantics. The Con…

Cited by 909PDFScholar
2016

CRF-CNN: Modeling Structured Information in Human Pose Estimation

NeurIPS 2016poster

Deep convolutional neural networks (CNN) have achieved great success. On the other hand, modeling structural information has been proved critical in many vision problems. It is of great interest to integrate them effectively. In a classical neural network, there is no message passing between neurons…

Cited by 97SourcePDFScholar