← Search

Sinan Tan

7 accepted papers

2025

A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation

ICLR 2025poster

This work tackles the information loss bottleneck of vector-quantization (VQ) autoregressive image generation by introducing a novel model architecture called the 2-Dimensional Autoregression (DnD) Transformer. The DnD-Transformer predicts more codes for an image by introducing a new direction, **mo…

2023

Embodied Referring Expression for Manipulation Question Answering in Interactive Environment

ICRA 2023poster

Embodied agents are expected to perform more complicated tasks in an interactive environment, with the progress of Embodied AI in recent years. Existing embodied tasks including Embodied Referring Expression (ERE) and other QA-form tasks mainly focuses on interaction in term of linguistic instructio…

Cited by 7SourceScholar
2023

Mixed Neural Voxels for Fast Multi-view Video Synthesis

ICCV 2023oral

Synthesizing high-fidelity videos from real-world multiview input is challenging due to the complexities of real-world environments and high-dynamic movements. Previous works based on neural radiance fields have demonstrated high-quality reconstructions of dynamic scenes. However, training such mode…

Cited by 72PDFcodeScholar
2022

Depth-Aware Vision-and-Language Navigation using Scene Query Attention Network

ICRA 2022poster

Vision-and-language navigation (VLN) has been an important task in the field of Robotics and Computer Vision. However, most existing vision-and-language navigation models only use features extracted from RGB observation as input, while robots can utilize depth sensors in the real world. Existing res…

Cited by 4SourceScholar
2022

Embodied Multi-Agent Task Planning from Ambiguous Instruction

RSS 2022poster

In human-robots collaboration scenarios, a human would give robots an instruction that is intuitive for the human himself to accomplish. However, the instruction given to robots is likely ambiguous for them to understand as some information is implicit in the instruction. Therefore, it is necessary…

Cited by 26SourcePDFScholar
2020

Multi-Agent Embodied Question Answering in Interactive Environments

ECCV 2020poster

We investigate a new AI task --- Multi-Agent Interactive Question Answering --- where several agents explore the scene jointly in interactive environments to answer a question. To cooperate efficiently and answer accurately, agents must be well-organized to have balanced work division and share know…

Cited by 38SourcePDFScholar