← Search

Yong Xie

7 accepted papers

2026

PC-CrossDiff: Point-Cluster Dual-Level Cross-Modal Differential Attention for Unified 3D Referring and Segmentation

AAAI 2026technical

3D Visual Grounding (3DVG) aims to localize the referent of natural language referring expressions through two core tasks: Referring Expression Comprehension (3DREC) and Segmentation (3DRES). While existing methods achieve high accuracy in simple, single-object scenes, they suffer from severe perfor

Cited by 0SourcePDFScholar
2026

RECOM: REALISTIC CO-SPEECH MOTION GENERATION WITH RECURRENT EMBEDDED TRANSFORMER

ICASSP 2026poster

We present ReCoM, an efficient framework for generating high-fidelity and generalizable human body motions synchronized with speech. The core innovation lies in the Recurrent Embedded Transformer (RET), which integrates Dynamic Embedding Regularization (DER) into a Vision Transformer (ViT) core arch…

Cited by 0SourcePDFScholar
2026

Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective

AAAI 2026technical

Open-vocabulary semantic segmentation (OVSS) employs pixel-level vision-language alignment to associate category-related prompts with corresponding pixels. A key challenge is enhancing the multimodal dense prediction capability, specifically this pixel-level multimodal alignment. Although existing m

Cited by 0SourcePDFScholar
2025

Dual-Population Watermark Vaccine: Efficient and Imperceptible Adversarial Attack for Watermarked Image Protection

ICASSP 2025accepted

The current watermark-removal neural networks (WRNNs) can effectively remove the watermarks from watermarked images without damaging their host images, which poses a significant threat to image copyright protection. As one of the most effective technologies of preventing watermarks from being remove…

Cited by 0SourceScholar
2025

Towards Million-Scale Adversarial Robustness Evaluation With Stronger Individual Attacks

CVPR 2025poster

As deep learning models are increasingly deployed in safety-critical applications, evaluating their vulnerabilities to adversarial perturbations is essential for ensuring their reliability and trustworthiness. Over the past decade, a large number of white-box adversarial robustness methods (i.e., at…

Cited by 1SourcePDFScholar
2024

Efficient Continual Pre-training for Building Domain Specific Large Language Models

ACL 2024findings

Large language models (LLMs) have demonstrated remarkable open-domain capabilities. LLMs tailored for a domain are typically trained entirely on domain corpus to excel at handling domain-specific tasks. In this work, we explore an alternative strategy of continual pre-training as a means to develop…

2022

A Word is Worth A Thousand Dollars: Adversarial Attack on Tweets Fools Stock Prediction

NAACL 2022long

More and more investors and machine learning models rely on social media (e.g., Twitter and Reddit) to gather information and predict movements stock prices. Although text-based models are known to be vulnerable to adversarial attacks, whether stock prediction models have similar vulnerability given…