← Search

Xinxiao Wu

20 accepted papers

2026

AmbiRefer3D: 3D Visual Grounding with Referential Ambiguity

ICML 2026poster

Traditional 3D visual grounding typically assumes that natural language expressions unambiguously refer to target objects in a 3D scene. However, in practical applications, human instructions are often ambiguous or insufficient, which may lead existing models to associate the query with multiple pos…

Cited by 0SourceScholar
2026

Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Edit via In-Context Learning

ICML 2026poster

Scene text editing aims to modify text in a target region of an image while preserving its background style and texture. Existing methods rely solely on image background information while neglecting the visual details of target regions, which discards stylistic features in the original text and esse…

Cited by 0SourceScholar
2026

TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents

AAAI 2026technical

Building Graphical User Interface (GUI) agents is a promising research direction, which simulates human interaction with computers or mobile phones to perform diverse GUI tasks. However, a major challenge in developing generalized GUI agents is the lack of sufficient trajectory data across various o

Cited by 0SourcePDFScholar
2026

VUDG: A Dataset for Video Understanding Domain Generalization

ICLR 2026poster

Video understanding has made remarkable progress in recent years, largely driven by advances in deep models and the availability of large-scale annotated datasets. However, the robustness of these models to domain shifts encountered in real-world video applications remains a critical yet underexplor…

Cited by 0SourceScholar
2026

What to Trust? A Trust-aware Knowledge-guided Method for Zero-shot Object State Understanding in Videos

AAAI 2026technical

Object state understanding aims at recognizing the co-occurrence and transitions of multiple object states in videos. While learning from videos handles seen object states well, it struggles with novel ones. We address this task in a zero-shot setting by extracting state-specific knowledge from pre-

Cited by 0SourcePDFScholar
2025

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection

IJCAI 2025

Open-vocabulary video visual relationship detection aims to detect objects and their relationships in videos without being restricted by predefined object or relationship categories. Existing methods leverage the rich semantic knowledge of pre-trained vision-language models such as CLIP to identify

2025

Storyboard-guided Alignment for Fine-grained Video Action Recognition

NeurIPS 2025poster

Fine-grained video action recognition can be formulated as a video–text matching problem. Previous approaches primarily rely on global video semantics to consolidate video embeddings, often leading to misaligned video–text pairs due to inaccurate atomic-level action understanding. This inaccuracy ar…

Cited by 0SourceScholar
2025

Video Summarization Using Denoising Diffusion Probabilistic Model

AAAI 2025technical

Video summarization aims to eliminate visual redundancy while retaining key parts of video to construct concise and comprehensive synopses. Most existing methods use discriminative models to predict the importance scores of video frames. However, these methods are susceptible to annotation inconsist…

Cited by 0SourcePDFScholar
2024

Event-based Few-shot Fine-grained Human Action Recognition

IROS 2024poster

Few-shot fine-grained human (FGH) action recognition is crucial in the context of human-robot interaction within open-set real-world environments. Existing works mainly focus on features extracted from RGB frames. However, their performances are drastically impacted in challenging scenarios, such as…

Cited by 1SourceScholar
2024

Multi-Modal Prompting for Open-Vocabulary Video Visual Relationship Detection

AAAI 2024technical

Open-vocabulary video visual relationship detection aims to extend video visual relationship detection beyond annotated categories by detecting unseen relationships between objects in videos. Recent progresses in open-vocabulary perception, primarily driven by large-scale image-text pre-trained mod…

2024

Relational Distant Supervision for Image Captioning without Image-Text Pairs

AAAI 2024technical

Unsupervised image captioning aims to generate descriptions of images without relying on any image-sentence pairs for training. Most existing works use detected visual objects or concepts as bridge to connect images and texts. Considering that the relationship between objects carries more informatio…

Cited by 5SourcePDFScholar
2023

Teaching What You Should Teach: A Data-Based Distillation Method

IJCAI 2023poster

In real teaching scenarios, an excellent teacher always teaches what he (or she) is good at but the student is not. This gives the student the best assistance in making up for his (or her) weaknesses and becoming a good one overall. Enlightened by this, we introduce the "Teaching what you Should Tea…

Cited by 4SourcePDFScholar
2022

Adaptive Image-to-Video Scene Graph Generation via Knowledge Reasoning and Adversarial Learning

AAAI 2022technical

Scene graph in a video conveys a wealth of information about objects and their relationships in the scene, thus benefiting many downstream tasks such as video captioning and visual question answering. Existing methods of scene graph generation require large-scale training videos annotated with objec…

Cited by 1SourcePDFScholar
2022

Entity-aware and Motion-aware Transformers for Language-driven Action Localization

IJCAI 2022poster

Language-driven action localization in videos is a challenging task that involves not only visual-linguistic matching but also action boundary prediction. Recent progress has been achieved through aligning language queries to video segments, but estimating precise boundaries is still under-explored.…

2021

Spatial-temporal Causal Inference for Partial Image-to-video Adaptation

AAAI 2021technical

Image-to-video adaptation leverages off-the-shelf learned models in labeled images to help classification in unlabeled videos, thus alleviating the high computation overhead of training a video classifier from scratch. This task is very challenging since there exist two types of domain shifts betwee…

2019

Joint Syntax Representation Learning and Visual Cue Translation for Video Captioning

ICCV 2019poster

Video captioning is a challenging task that involves not only visual perception but also syntax representation learning. Recent progress in video captioning has been achieved through visual perception, but syntax representation learning is still under-explored. We propose a novel video captioning ap…

Cited by 112PDFScholar