← Search

Wenhao XU

13 accepted papers

2026

DialogueVPR: Towards Conversational Visual Place Recognition

CVPR 2026

Inspired by how humans communicate spatial information, language-guided geo-localization has gained significant traction for its intuitive and practical value. Despite this progress, most methods still rely on a static, one-shot retrieval paradigm, which fails to handle the ambiguity and incompleten

Cited by 0SourcecodeScholar
2026

Judging by the Rules: Compliance-Aligned Framework for Modern Slavery Statement Monitoring

AAAI 2026technical

Modern slavery affects millions of people worldwide, and regulatory frameworks such as Modern Slavery Acts now require companies to publish detailed disclosures. However, these statements are often vague and inconsistent, making manual review time-consuming and difficult to scale. While NLP offers a

Cited by 0SourcePDFScholar
2026

Modeling Cross-vision Synergy for Unified Large Vision Model

CVPR 2026

Recent advances in large vision models (LVMs) have shifted from modality-specific designs toward unified architectures that jointly process images, videos, and 3D data. However, existing unified LVMs primarily pursue functional integration, while overlooking the deeper goal of cross-vision synergy:

Cited by 0SourceScholar
2026

ReFAct: Empowering Multimodal Web Agents with Visual and Context Focusing

CVPR 2026

Multimodal Web Search Agents demonstrate a practically valuable capability by fusing information from diverse modalities (e.g., text and vision), retrieved iteratively from the internet, to address complex user queries. However, the visual modality is prone to information overload, and the noise con

Cited by 0SourceScholar
2026

SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition

ICLR 2026poster

Visual Place Recognition (VPR) requires robust retrieval of geotagged images despite large appearance, viewpoint, and environmental variation. Prior methods focus on descriptor fine-tuning or fixed sampling strategies yet neglect the dynamic interplay between spatial context and visual similarity d…

Cited by 0SourcecodeScholar
2025

Event-boosted Deformable 3D Gaussians for Dynamic Scene Reconstruction

ICCV 2025poster

Deformable 3D Gaussian Splatting (3D-GS) is limited by missing intermediate motion information due to the low temporal resolution of RGB cameras. To address this, we introduce the first approach combining event cameras, which capture high-temporal-resolution, continuous motion data, with deformable…

2025

Let's Chorus: Partner-aware Hybrid Song-Driven 3D Head Animation

CVPR 2025poster

Singing is a vital form of human emotional expression and social interaction, distinguished from speech by its richer emotional nuances and freer expressive style. Thus, investigating 3D facial animation driven by singing holds significant research value. Our work focuses on 3D singing facial animat…

Cited by 0SourcePDFScholar
2025

STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation

ACL 2025finding

Recent breakthroughs in singing voice synthesis (SVS) have heightened the demand for high-quality annotated datasets, yet manual annotation remains prohibitively labor-intensive and resource-intensive. Existing automatic singing annotation (ASA) methods, however, primarily tackle isolated aspects of…

2025

Towards Ship License Plate Recognition in the Wild: A Large Benchmark and Strong Baseline

AAAI 2025technical

The paper targets the challenging task of Ship License Plate (SLP) recognition. Existing methods for SLP recognition are hampered by the scarcity of large and publicly available datasets, leading to evaluations on small and non-representative datasets. To alleviate it, we have built a large dataset,…

2024

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks

NeurIPS 2024spotlight

The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages and singers, absence of multi-technique information and real…

2024

Spectral Prompt Tuning: Unveiling Unseen Classes for Zero-Shot Semantic Segmentation

AAAI 2024technical

Recently, CLIP has found practical utility in the domain of pixel-level zero-shot segmentation tasks. The present landscape features two-stage methodologies beset by issues such as intricate pipelines and elevated computational costs. While current one-stage approaches alleviate these concerns and…

2023

Regret Bounds for Markov Decision Processes with Recursive Optimized Certainty Equivalents

ICML 2023poster

The optimized certainty equivalent (OCE) is a family of risk measures that cover important examples such as entropic risk, conditional value-at-risk and mean-variance models. In this paper, we propose a new episodic risk-sensitive reinforcement learning formulation based on tabular Markov decision p…

Cited by 13SourcePDFScholar