← Search

Tao Tu

7 accepted papers

2026

Pantheon360: Taming Digital Twin Generation via 3D-Aware 360deg Video Diffusion

CVPR 2026

Generating complete digital twins from videos requires precise camera control, global scene coverage, and strict spatial-temporal consistency--constraints that remain challenging for perspective video generators due to their limited field of view (FoV). Their narrow FoV forces long or multi-view tra

Cited by 0SourceScholar
2025

OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations

ICCV 2025poster

Open-vocabulary (OV) 3D object detection is an emerging field, yet its exploration through image-based methods remains limited compared to 3D point cloud-based methods. We introduce OpenM3D, a novel open-vocabulary multi-view indoor 3D object detector trained without human annotations. In particular…

Cited by 0SourcePDFScholar
2023

Disentangle then Parse: Night-time Semantic Segmentation with Illumination Disentanglement

ICCV 2023poster

Most prior semantic segmentation methods have been developed for day-time scenes, while typically underperforming in night-time scenes due to insufficient and complicated lighting conditions. In this work, we tackle this challenge by proposing a novel night-time semantic segmentation paradigm, i.e.,…

Cited by 11PDFcodeScholar
2023

ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object Detection

ICCV 2023poster

We propose ImGeoNet, a multi-view image-based 3D object detection framework that models a 3D space by an image-induced geometry-aware voxel representation. Unlike previous methods which aggregate 2D features into 3D voxels without considering geometry, ImGeoNet learns to induce geometry from multi-…

Cited by 11PDFcodeScholar
2021

Learning Better Visual Dialog Agents With Pretrained Visual-Linguistic Representation

CVPR 2021poster

GuessWhat?! is a visual dialog guessing game which incorporates a Questioner agent that generates a sequence of questions, while an Oracle agent answers the respective questions about a target object in an image. Based on this dialog history between the Questioner and the Oracle, a Guesser agent mak…

Cited by 25PDFcodeScholar
2020

Towards Unsupervised Speech Recognition and Synthesis with Quantized Speech Representation Learning

ICASSP 2020accepted

In this paper we propose a Sequential Representation Quantization AutoEncoder (SeqRQ-AE) to learn from primarily unpaired audio data and produce sequences of representations very close to phoneme sequences of speech utterances. This is achieved by proper temporal segmentation to make the representat…

Cited by 0SourceScholar
2019

A state-space model for inferring effective connectivity of latent neural dynamics from simultaneous EEG/fMRI

NeurIPS 2019poster

Inferring effective connectivity between spatially segregated brain regions is important for understanding human brain dynamics in health and disease. Non-invasive neuroimaging modalities, such as electroencephalography (EEG) and functional magnetic resonance imaging (fMRI), are often used to make m…