← Search

Tao Wu

40 accepted papers

2026

EES: A Data-Driven End-To-End Escorting System Via Spatiotemporal Feature Fusion

ICRA 2026poster

This letter presents a technique that allows unmanned vehicles to escort a human to their destinations. Current human-centered following methods depend solely on human movement, which presents significant limitations. The complexity of human movement during tactical maneuvers can lead to erratic veh…

Cited by 0SourceScholar
2026

GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation

CVPR 2026

Recent years have seen impressive advances in text-to-image generation, with image generative or unified models, generating high-quality images from text. Yet these models still struggle with fine-grained color control, often failing to accurately match colors specified in text prompts. While existi

Cited by 0SourcecodeScholar
2026

GeodesicNVS: Probability Density Geodesic Flow Matching for Novel View Synthesis

CVPR 2026

Recent advances in generative modeling have substantially enhanced novel view synthesis, yet maintaining consistency across viewpoints remains challenging. Diffusion-based models rely on stochastic noise-to-data transitions, which obscure deterministic structures and yield inconsistent view predicti

Cited by 0SourceScholar
2026

MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment

CVPR 2026

Multi-subject image generation aims to synthesize user-provided subjects in a single image while preserving subject fidelity, ensuring prompt consistency, and aligning with human aesthetic preferences. Existing In-Context-Learning based methods are limited by their highly coupled training paradigm.

Cited by 0SourceScholar
2026

SinGeo: Unlock Single Model's Potential for Robust Cross-View Geo-Localization

CVPR 2026

Robust cross-view geo-localization (CVGL) remains challenging despite the surge in recent progress. Existing methods still rely on field-of-view (FoV)-specific training paradigms, where models are optimized under a fixed FoV but collapse when tested on unseen FoVs and unknown orientations. This limi

Cited by 0SourcecodeScholar
2026

Spatially Generalizable Mobile Manipulation via Adaptive Experience Selection and Dynamic Imagination

IJCAI 2026

Mobile Manipulation (MM) involves long-horizon decision-making over multi-stage compositions of heterogeneous skills, such as navigation and picking up objects. Despite recent progress, existing MM methods still face two key limitations: (i) low sample efficiency, due to ineffective use of redundant

Cited by 0Scholar
2026

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

CVPR 2026

Enhancing temporal understanding of MLLMs is essential for long-form video analysis, supporting tasks such as temporal localization and time-sensitive question answering. While reinforcement learning (RL) has been explored for temporal reasoning, existing approaches are often limited to specific tas

Cited by 0SourceScholar
2026

When Robots Should Say ''I Don't Know'': Benchmarking Abstention in Embodied Question Answering

CVPR 2026

Embodied Question Answering (EQA) requires an agent to interpret language, perceive its environment, and navigate within 3D scenes to produce responses. Existing EQA benchmarks assume that every question must be answered, but embodied agents should know when they do not have sufficient information t

Cited by 0SourceScholar
2025

CoEvo: Coevolution of LLM and Retrieval Model for Domain-Specific Information Retrieval

EMNLP 2025

Information retrieval in specialized domains (e.g., legal and medical) faces challenges in aligning user queries, often expressed in colloquial language, with highly structured, terminology-rich documents. This discrepancy creates a distribution gap in the text representation. Recent methods aim to

2025

Cognitive-Level Adaptive Generation via Capability-Aware Retrieval and Style Adaptation

EMNLP 2025

Large Language Models (LLMs) have demonstrated strong performance in open-ended generation tasks. However, they often struggle to adapt content to users with differing cognitive capacities, leading to a phenomenon we term cognitive misalignment. This issue arises in two forms: knowledge-level misali

2025

CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities

AAAI 2025technical

Customized video generation aims to generate high-quality videos guided by text prompts and subject's reference images. However, since it is only trained on static images, the fine-tuning process of subject learning disrupts abilities of video diffusion models (VDMs) to combine concepts and generate…

2025

Do Current Video LLMs Have Strong OCR Abilities? A Preliminary Study

COLING 2025main

With the rise of multi-modal large language models, accurately extracting and understanding textual information from video content—referred to as video-based optical character recognition (Video OCR)—has become a crucial capability. This paper introduces a novel benchmark designed to evaluate the vi…

2025

Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents

ACL 2025long

Large language models (LLMs) are revolutionizing education, with LLM-based agents playing a key role in simulating student behavior. A major challenge in student simulation is modeling the diverse learning patterns of students at various cognitive levels. However, current LLMs, typically trained as…

Cited by 0SourcePDFScholar
2025

LithoSim: A Large, Holistic Lithography Simulation Benchmark for AI-Driven Semiconductor Manufacturing

NeurIPS 2025poster

Lithography orchestrates a symphony of light, mask and photochemicals to transfer the integrated circuit patterns onto the wafer. Lithography simulation serves as the critical nexus between circuit design and manufacturing, where its speed and accuracy fundamentally govern the optimization quality o…

Cited by 0SourcecodeScholar
2025

Online Video Understanding: OVBench and VideoChat-Online

CVPR 2025poster

Multimodal Large Language Models (MLLMs) have significantly progressed in offline video understanding. However, applying these models to real-world scenarios, such as autonomous driving and human-computer interaction, presents unique challenges due to the need for real-time processing of continuous…

Cited by 0SourcePDFScholar
2025

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control

ICCV 2025poster

Recent advancements in camera-trajectory-guided image-to-video generation offer higher precision and better support for complex camera control compared to text-based approaches. However, they also introduce significant usability challenges, as users often struggle to provide precise camera parameter…

Cited by 0SourcePDFScholar
2025

Simpler Is Better: Revisiting Doppler Velocity for Enhanced Moving Object Tracking with FMCW LiDAR

IROS 2025

Real-time and accurate perception of dynamic objects is crucial for autonomous driving. To better capture the motion information of objects, some methods now employ 4D Doppler point clouds collected by frequency-modulated continuous-wave (FMCW) LiDAR to enhance the detection and tracking of moving o

Cited by 1SourcecodeScholar
2025

TransiT: Transient Transformer for Non-line-of-sight Videography

ICCV 2025poster

High quality and high speed videography using Non-Line-of-Sight (NLOS) imaging benefit autonomous navigation, collision prevention, and post-disaster search and rescue tasks. Current solutions have to balance between the frame rate and image quality. High frame rates, for example, can be achieved by…

Cited by 0SourcePDFScholar
2025

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay

ICCV 2025poster

Despite the remarkable performance of multimodal large language models (MLLMs) across diverse tasks, the substantial training and inference costs impede their advancement. In this paper, we propose p-MoD, an efficient MLLM architecture that significantly reduces training and inference costs while ma…

2024

LRS: Enhancing Adversarial Transferability through Lipschitz Regularized Surrogate

AAAI 2024technical

The transferability of adversarial examples is of central importance to transfer-based black-box adversarial attacks. Previous works for generating transferable adversarial examples focus on attacking given pretrained surrogate models while the connections between surrogate models and adversarial tr…

2024

ReinforceNS: Reinforcement Learning-based Multi-start Neighborhood Search for Solving the Traveling Thief Problem

IJCAI 2024poster

The Traveling Thief Problem (TTP) is a challenging combinatorial optimization problem with broad practical applications. TTP combines two NP-hard problems: the Traveling Salesman Problem (TSP) and Knapsack Problem (KP). While a number of machine learning and deep learning based algorithms have been…

Cited by 0SourcePDFScholar
2024

SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion Model

AAAI 2024technical

Controllable spherical panoramic image generation holds substantial applicative potential across a variety of domains. However, it remains a challenging task due to the inherent spherical distortion and geometry characteristics, resulting in low-quality content generation. In this paper, we introduc…

Cited by 8SourcePDFScholar
2024

SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos

CVPR 2024poster

Video-based visual relation detection tasks such as video scene graph generation play important roles in fine-grained video understanding. However current video visual relation detection datasets have two main limitations that hinder the progress of research in this area. First they do not explore c…

2023

SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation

IJCAI 2023poster

As an important and challenging problem in computer vision, PAnoramic Semantic Segmentation (PASS) gives complete scene perception based on an ultra-wide angle of view. Usually, prevalent PASS methods with 2D panoramic image input focus on solving image distortions but lack consideration of the 3D p…

2023

SpatialFormer: Semantic and Target Aware Attentions for Few-Shot Learning

AAAI 2023technical

Recent Few-Shot Learning (FSL) methods put emphasis on generating a discriminative embedding features to precisely measure the similarity between support and query sets. Current CNN-based cross-attention approaches generate discriminative representations via enhancing the mutually semantic similar r…

2022

FormLM: Recommending Creation Ideas for Online Forms by Modelling Semantic and Structural Information

EMNLP 2022main

Online forms are widely used to collect data from human and have a multi-billion market. Many software products provide online services for creating semi-structured forms where questions and descriptions are organized by predefined structures. However, the design and creation process of forms is sti…

Cited by 1SourcePDFScholar
2022

Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding

AAAI 2022technical

Temporal grounding aims to localize a video moment which is semantically aligned with a given natural language query. Existing methods typically apply a detection or regression pipeline on the fused representation with the research focus on designing complicated prediction heads or fusion strategies…

2021

Fine-Grained Shape-Appearance Mutual Learning for Cloth-Changing Person Re-Identification

CVPR 2021poster

Recently, person re-identification (Re-ID) has achieved great progress. However, current methods largely depend on color appearance, which is not reliable when a person changes the clothes. Cloth-changing Re-ID is challenging since pedestrian images with clothes change exhibit large intra-class vari…

Cited by 206PDFScholar
2021

Multiple Contextual Cues Integrated Trajectory Prediction for Autonomous Driving

RA-L 2021

Trajectory prediction is an essential and challenging task for autonomous driving and mobile robots. The main difficulty is to model actor-actor interaction and actor-scene interaction. In addition, the different motion characteristics of each actor also increase the challenge of prediction. Most ex

Cited by 11SourceScholar
2020

Optimization of Graph Total Variation via Active-Set-based Combinatorial Reconditioning

AISTATS 2020poster

Structured convex optimization on weighted graphs finds numerous applications in machine learning and computer vision. In this work, we propose a novel adaptive preconditioning strategy for proximal algorithms on this problem class. Our preconditioner is driven by a sharp analysis of the local linea…

Cited by 4SourcePDFScholar
2019

Optimization of Inf-Convolution Regularized Nonconvex Composite Problems

AISTATS 2019poster

In this work, we consider nonconvex composite problems that involve inf-convolution with a Legendre function, which gives rise to an anisotropic generalization of the proximal mapping and Moreau-envelope. In a convex setting such problems can be solved via alternating minimization of a splitting for…

Cited by 7SourcePDFScholar
2019

Variational Uncalibrated Photometric Stereo Under General Lighting

ICCV 2019poster

Photometric stereo (PS) techniques nowadays remain constrained to an ideal laboratory setup where modeling and calibration of lighting is amenable. To eliminate such restrictions, we propose an efficient principled variational approach to uncalibrated PS under general illumination. To this end, the…

Cited by 44PDFcodeScholar
2018

A Nonconvex Proximal Splitting Algorithm under Moreau-Yosida Regularization

AISTATS 2018poster

We tackle highly nonconvex, nonsmooth composite optimization problems whose objectives comprise a Moreau-Yosida regularized term. Classical nonconvex proximal splitting algorithms, such as nonconvex ADMM, suffer from lack of convergence for such a problem class. To overcome this difficulty, in this…

Cited by 0SourcePDFScholar
2018

Combinatorial Preconditioners for Proximal Algorithms on Graphs

AISTATS 2018poster

We present a novel preconditioning technique for proximal optimization methods that relies on graph algorithms to construct effective preconditioners. Such combinatorial preconditioners arise from partitioning the graph into forests. We prove that certain decompositions lead to a theoretically optim…

Cited by 0SourcePDFScholar
2017

A Non-Convex Variational Approach to Photometric Stereo Under Inaccurate Lighting

CVPR 2017poster

This paper tackles the photometric stereo problem in the presence of inaccurate lighting, obtained either by calibration or by an uncalibrated photometric stereo method. Based on a precise modeling of noise and outliers, a robust variational approach is introduced. It explicitly accounts for self-sh…

Cited by 73PDFScholar