← Search

Yong Wu

12 accepted papers

2026

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention

CVPR 2026

Vision-Language-Action (VLA) models have shown remarkable progress in embodied tasks recently, but most methods process visual observations independently at each timestep. This history-agnostic design treats robot manipulation as a Markov Decision Process, even though real-world robotic control is i

Cited by 0SourceScholar
2026

FGD-Align: Pluralistic Alignment for Large Language Models via Fuzzy Group Decision-Making

AAAI 2026technical

Ensuring alignment with human values is essential for modern large language models (LLMs), especially amid growing concerns around AI safety and social impact. Yet achieving such alignment remains challenging due to the limited, noisy, and often conflicting nature of human feedback from diverse anno

Cited by 0SourcePDFScholar
2026

GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMs

ICLR 2026poster

Geometric spatial reasoning forms the foundation of many applications in artificial intelligence, yet the ability of large language models (LLMs) to operate over geometric spatial information expressed in procedural code remains underexplored. In this paper, we address this gap by formalizing the \t…

Cited by 0SourcecodeScholar
2026

Selective Coupling of Decoupled Informative Regions: Masked Attention Alignment for Data-Free Quantization of Vision Transformers

ICML 2026poster

Data-Free Quantization (DFQ) addresses data security concerns by synthesizing fake samples, without accessing real data. It has garnered increasing attention in the context of Vision Transformers (ViTs), owing to the superiority of the self-attention mechanism compared to classical convolutional ope…

Cited by 0SourceScholar
2025

Bivariate Causal Discovery with Proxy Variables: Integral Solving and Beyond

ICML 2025poster

Bivariate causal discovery is challenging when unmeasured confounders exist. To adjust for the bias, previous methods employed the proxy variable (*i.e.*, negative control outcome (NCO)) to test the treatment-outcome relationship through integral equations -- and assumed that violation of this equat…

Cited by 0SourcePDFScholar
2025

OOD-Barrier: Build a Middle-Barrier for Open-Set Single-Image Test Time Adaptation via Vision Language Models

NeurIPS 2025poster

In real-world environments, a well-designed model must be capable of handling dynamically evolving distributions, where both in-distribution (ID) and out-of-distribution (OOD) samples appear unpredictably and individually, making real-time adaptation particularly challenging. While open-set test-tim…

Cited by 0SourceScholar
2025

Relation-aware Semantic Alignment Network for Text-to-Image Person Retrieval

ICASSP 2025accepted

Text-to-Image Person Retrieval (TIPR) aims to utilize natural language descriptions as queries to retrieve pedestrian images. However, existing methods only concentrated on aligning individual text-image pairs and ignored the specific self-representations within both visible images and textual descr…

Cited by 0SourceScholar
2024

AV4GAInsp: An Efficient Dual-Camera System for Identifying Defective Kernels of Cereal Grains

RA-L 2024

Grain Appearance Inspection (GAI) is a pre-requisite for grain quality determination, providing guidance for grain processing, storage, and trade. GAI is routinely performed by trained inspectors who are required to visually inspect cereal grains for each individual kernel. Since grain kernels (e.g.

Cited by 10SourceScholar
2024

Day-Night Cross-domain Vehicle Re-identification

CVPR 2024poster

Previous advances in vehicle re-identification (ReID) are mostly reported under favorable lighting conditions while cross-day-and-night performance is neglected which greatly hinders the development of related traffic intelligence applications. This work instead develops a novel Day-Night Dual-domai…

2024

Doubly Robust Proximal Causal Learning for Continuous Treatments

ICLR 2024poster

Proximal causal learning is a powerful framework for identifying the causal effect under the existence of unmeasured confounders. Within this framework, the doubly robust (DR) estimator was derived and has shown its effectiveness in estimation, especially when the model assumption is violated. Howev…

2024

ELF-UA: Efficient Label-Free User Adaptation in Gaze Estimation

IJCAI 2024poster

We consider the problem of user-adaptive 3D gaze estimation. The performance of person-independent gaze estimation is limited due to interpersonal anatomical differences. Our goal is to provide a personalized gaze estimation model specifically adapted to a target user. Previous work on user-adaptive…

Cited by 1SourcePDFScholar
2024

MAP: MAsk-Pruning for Source-Free Model Intellectual Property Protection

CVPR 2024poster

Deep learning has achieved remarkable progress in various applications heightening the importance of safeguarding the intellectual property (IP) of well-trained models. It entails not only authorizing usage but also ensuring the deployment of models in authorized data domains i.e. making models excl…