← Search

Xu Tang

26 accepted papers

2026

ID-Splat: Propagating Object Identities for Segmenting 3D Aerial-view Scenes

AAAI 2026technical

High-resolution Earth Observation technologies present unprecedented opportunities for geospatial analysis, yet traditional 2D aerial-view semantic segmentation remains limited by its inability to model spatial relationships and handle object occlusions. While 3D Aerial-view Segmentation (3DAS) has

Cited by 0SourcePDFScholar
2026

IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning

ICLR 2026poster

Large Vision-Language Models (LVLMs) achieve impressive performance across multiple tasks. A significant challenge, however, is their prohibitive inference cost when processing high-resolution visual inputs. While visual token pruning has emerged as a promising solution, existing methods that primar…

Cited by 0SourceScholar
2026

PROMO: Promptable Outfitting for Efficient High-Fidelity Virtual Try-On

CVPR 2026

Virtual Try-on (VTON) has become a core capability for online retail, where realistic try-on results provide reliable fit guidance, reduce returns, and benefit both consumers and merchants. Diffusion-based VTON methods achieve photorealistic synthesis, yet often rely on intricate architectures such

Cited by 0SourceScholar
2026

ReMatch: Boosting Representation through Matching for Multimodal Retrieval

CVPR 2026

We present ReMatch, a framework that leverages the generative strength of MLLMs for multimodal retrieval. Previous approaches treated an MLLM as a simple encoder, ignoring its generative nature, and under-utilising its compositional reasoning and world knowledge. We train the embedding MLLM end-to-e

Cited by 0SourcecodeScholar
2026

SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMs

ICML 2026poster

Multimodal Large Language Models (MLLMs) possess intrinsic reasoning and world-knowledge capabilities, yet adapting them for dense retrieval remains challenging. Existing approaches typically rely on invasive parameter updates, such as full fine-tuning and LoRA, which risk disrupting the pre-trained…

Cited by 0SourceScholar
2026

Vision-Language Model Guided Source-Free Domain Adaptation via Optimal Transport

CVPR 2026

Unsupervised domain adaptation transfers knowledge from a labeled source domain to an unlabeled target domain. When source data cannot be accessed, source-free domain adaptation (SFDA) becomes a practical alternative. However, existing SFDA methods mainly rely on pseudo-label based self-training, wh

Cited by 0SourcecodeScholar
2025

Beyond Guilt: Legal Judgment Prediction with Trichotomous Reasoning

EMNLP 2025

In legal practice, judges apply the trichotomous dogmatics of criminal law, sequentially assessingthe elements of the offense, unlawfulness, and culpability to determine whether an individual’s conduct constitutes a crime. Although current legal large language models (LLMs) show promising accuracy i

2025

CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection

NeurIPS 2025poster

With the exponential growth of data, traditional object detection methods are increasingly struggling to handle vast vocabulary object detection tasks effectively. We analyze two key limitations of classification-based detectors: positive gradient dilution, where rare positive categories receive ins…

Cited by 0SourcecodeScholar
2025

Category-Specific Selective Feature Enhancement for Long-Tailed Multi-Label Image Classification

ICCV 2025poster

Since real-world multi-label data often exhibit significant label imbalance, long-tailed multi-label image classification has emerged as a prominent research area in computer vision. Traditionally, it is considered that deep neural networks' classifiers are vulnerable to long-tailed distributions, w…

Cited by 0SourcePDFScholar
2025

DynamicFace: High-Quality and Consistent Face Swapping for Image and Video using Composable 3D Facial Priors

ICCV 2025poster

Face swapping transfers the identity of a source face to a target face while retaining the attributes like expression, pose, hair, and background of the target face. Advanced face swapping methods have achieved attractive results. However, these methods often inadvertently transfer identity informat…

2025

InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention

NeurIPS 2025poster

Diffusion models have demonstrated remarkable capabilities in generating high-quality images. Recent advancements in Layout-to-Image (L2I) generation have leveraged positional conditions and textual descriptions to facilitate precise and controllable image synthesis. Despite overall progress, curren…

Cited by 0SourcecodeScholar
2025

Legal Mathematical Reasoning with LLMs: Procedural Alignment through Two-Stage Reinforcement Learning

EMNLP 2025

Legal mathematical reasoning is essential for applying large language models (LLMs) in high-stakes legal contexts, where outputs must be both mathematically accurate and procedurally compliant. However, existing legal LLMs lack structured numerical reasoning, and open-domain models, though capable o

2025

RegionMatch: Pixel-Region Collaboration for Semi-Supervised Semantic Segmentation in Remote Sensing Images

IJCAI 2025

Semi-supervised semantic segmentation (S4) has shown significant promise in reducing the burden of labor-intensive data annotation. However, existing methods mainly rely on pixel-level information, neglecting the strong region consistency inherent in remote sensing images (RSIs), which limits their

Cited by 0SourcePDFScholar
2024

Controllable Mind Visual Diffusion Model

AAAI 2024technical

Brain signal visualization has emerged as an active research area, serving as a critical interface between the human visual system and computer vision models. Diffusion-based methods have recently shown promise in analyzing functional magnetic resonance imaging (fMRI) data, including the reconstruct…

2024

Harmonizing Stochasticity and Determinism: Scene-responsive Diverse Human Motion Prediction

NeurIPS 2024poster

Diverse human motion prediction (HMP) is a fundamental application in computer vision that has recently attracted considerable interest. Prior methods primarily focus on the stochastic nature of human motion, while neglecting the specific impact of external environment, leading to the pronounced art…

Cited by 4SourcePDFScholar
2024

Multimodal Sense-Informed Forecasting of 3D Human Motions

CVPR 2024poster

Predicting future human pose is a fundamental application for machine intelligence which drives robots to plan their behavior and paths ahead of time to seamlessly accomplish human-robot collaboration in real-world 3D scenarios. Despite encouraging results existing approaches rarely consider the eff…

Cited by 6SourcePDFScholar
2024

SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven Generation

CVPR 2024poster

Recent advancements in subject-driven image generation have led to zero-shot generation yet precise selection and focus on crucial subject representations remain challenging. Addressing this we introduce the SSR-Encoder a novel architecture designed for selectively capturing any subject from single…

2024

ZONE: Zero-Shot Instruction-Guided Local Editing

CVPR 2024poster

Recent advances in vision-language models like Stable Diffusion have shown remarkable power in creative image synthesis and editing.However most existing text-to-image editing methods encounter two obstacles: First the text prompt needs to be carefully crafted to achieve good results which is not in…

2023

OvarNet: Towards Open-Vocabulary Object Attribute Recognition

CVPR 2023poster

In this paper, we consider the problem of simultaneously detecting objects and inferring their visual attributes in an image, even for those with no manual annotations provided at the training stage, resembling an open-vocabulary scenario. To achieve this goal, we make the following contributions: (…

2023

Towards Open-Vocabulary Video Instance Segmentation

ICCV 2023oral

Video Instance Segmentation (VIS) aims at segmenting and categorizing objects in videos from a closed set of training categories, lacking the generalization ability to handle novel categories in real-world videos. To address this limitation, we make the following three contributions. First, we intro…

Cited by 36PDFcodeScholar
2022

Absolute Wrong Makes Better: Boosting Weakly Supervised Object Detection via Negative Deterministic Information

IJCAI 2022poster

Weakly supervised object detection (WSOD) is a challenging task, in which image-level labels (e.g., categories of the instances in the whole image) are used to train an object detector. Many existing methods follow the standard multiple instance learning (MIL) paradigm and have achieved promising pe…

Cited by 16SourcePDFScholar
2022

SVIP: Sequence VerIfication for Procedures in Videos

CVPR 2022poster

In this paper, we propose a novel sequence verification task that aims to distinguish positive video pairs performing the same action sequence from negative ones with step-level transformations but still conducting the same task. Such a challenging task resides in an open-set setting without prior a…

Cited by 25PDFcodeScholar
2020

HAMBox: Delving Into Mining High-Quality Anchors on Face Detection

CVPR 2020poster

Current face detectors utilize anchors to frame a multi-task learning problem which combines classification and bounding box regression. Effective anchor design and anchor matching strategy enable face detectors to localize faces under large pose and scale variations. However, we observe that, more…

Cited by 50PDFScholar
2018

Face Aging With Identity-Preserved Conditional Generative Adversarial Networks

CVPR 2018poster

Face aging is of great importance for cross-age recognition and entertainment related applications. However, the lack of labeled faces of the same person across a long age range makes it challenging. Because of different aging speed of different persons, our face aging approach aims at synthesizing…

Cited by 287SourcePDFScholar
2018

PyramidBox: A Context-assisted Single Shot Face Detector

ECCV 2018poster

Face detection has been well studied for many years and one of remaining challenges is to detect small, blurred and partially occluded faces in uncontrolled environment. This paper proposes a novel context-assisted single shot face detector, named emph{PyramidBox} to handle the hard face detection p…