← Search

Tao Xu

15 accepted papers

2026

Refine3D: Scene-Adaptive Reference Point Refinement for Sparse 3D Object Detection

AAAI 2026technical

Sparse query-based detectors have emerged as the dominant paradigm in camera-only 3D object detection, owing to their exceptional performance and computational efficiency. A central component of these approaches is the use of reference points, which serve as learnable spatial anchors to guide queri

Cited by 0SourcePDFScholar
2025

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

ICML 2025poster

Visual tokenization via auto-encoding empowers state-of-the-art image and video generative models by compressing pixels into a latent space. However, questions remain about how auto-encoder design impacts reconstruction and downstream generative performance. This work explores scaling in auto-encode…

Cited by 6SourcePDFScholar
2025

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity

CVPR 2025poster

Text-to-video generation enhances content creation but is highly computationally intensive: The computational cost of Diffusion Transformers (DiTs) scales quadratically in the number of pixels. This makes minute-length video generation extremely expensive, limiting most existing models to generatin…

2024

Batch-ICL: Effective, Efficient, and Order-Agnostic In-Context Learning

ACL 2024findings

In this paper, by treating in-context learning (ICL) as a meta-optimization process, we explain why LLMs are sensitive to the order of ICL examples. This understanding leads us to the development of Batch-ICL, an effective, efficient, and order-agnostic inference algorithm for ICL. Differing from th…

2024

Saliency Prediction of Sports Videos: A Large-Scale Database and a Self-Adaptive Approach

ICASSP 2024accepted

Predicting video saliency is crucial for improving sports video processing efficiency, thereby providing an enriched viewing experience for a wide-ranging audience. However, there is a long-term absence of well-established eye-tracking database and learning-based approach, particularly tailored for…

Cited by 0SourceScholar
2023

Robust Speech Recognition via Large-Scale Weak Supervision

ICML 2023poster

We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet. When scaled to 680,000 hours of multilingual and multitask supervision, the resulting models generalize well to standard benchmarks and are often competitive with p…

2021

Autonomous object harvesting using synchronized optoelectronic microrobots

IROS 2021poster

Optoelectronic tweezer-driven microrobots (OETdMs) are a versatile micromanipulation technology based on the application of light induced dielectrophoresis to move small dielectric structures (microrobots) across a photoconductive substrate. The microrobots in turn can be used to exert forces on sec…

Cited by 7SourceScholar
2021

Principal component analysis in the stochastic differential privacy model

UAI 2021poster

In this paper, we study the differentially private Principal Component Analysis (PCA) problem in stochastic optimization settings. We first propose a new stochastic gradient perturbation PCA mechanism (DP-SPCA) for the calculation of the right singular subspace to achieve $(\epsilon,\delta)$-differe…

Cited by 6SourcePDFScholar
2020

FBNetV2: Differentiable Neural Architecture Search for Spatial and Channel Dimensions

CVPR 2020poster

Differentiable Neural Architecture Search (DNAS) has demonstrated great success in designing state-of-the-art, efficient neural networks. However, DARTS-based DNAS's search space is small when compared to other search methods', since all candidate network layers must be explicitly instantiated in me…

Cited by 383PDFcodeScholar
2018

AttnGAN: Fine-Grained Text to Image Generation With Attentional Generative Adversarial Networks

CVPR 2018poster

In this paper, we propose an Attentional Generative Adversarial Network (AttnGAN) that allows attention-driven, multi-stage refinement for fine-grained text-to-image generation. With a novel attentional generative network, the AttnGAN can synthesize fine-grained details at different sub-regions of…

2018

On the Discrimination-Generalization Tradeoff in GANs

ICLR 2018poster

Generative adversarial training can be generally understood as minimizing certain moment matching loss defined by a set of discriminator functions, typically neural networks. The discriminator set should be large enough to be able to uniquely identify the true distribution (discriminative), and als…

Cited by 0SourcePDFScholar
2017

StackGAN: Text to Photo-Realistic Image Synthesis With Stacked Generative Adversarial Networks

ICCV 2017oral

Synthesizing high-quality images from text descriptions is a challenging problem in computer vision and has many practical applications. Samples generated by existing text-to-image approaches can roughly reflect the meaning of the given descriptions, but they fail to contain necessary details and vi…

Cited by 2957PDFcodeScholar
2016

RSS-based sensor localization in underwater acoustic sensor networks

ICASSP 2016accepted

Since the global positioning system (GPS) is not applicable underwater, source localization using wireless sensor networks (WSNs) is gaining popularity in oceanographic applications. Unlike terrestrial WSNs (TWSNs) which uses electromagnetic signaling, underwater WSNs (UWSNs) require underwater acou…

Cited by 0SourceScholar
2016

SPDA-CNN: Unifying Semantic Part Detection and Abstraction for Fine-Grained Recognition

CVPR 2016poster

Most convolutional neural networks (CNNs) lack midlevel layers that model semantic parts of objects. This limits CNN-based methods from reaching their full potential in detecting and utilizing small semantic parts in recognition. Introducing such mid-level layers can facilitate the extraction of par…

Cited by 382PDFScholar