← Search

Yuxiao Wang

9 accepted papers

2026

QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection

AAAI 2026technical

Human-Object Interaction (HOI) detection aims to localize human-object pairs and recognize their interactions in images. Although DETR-based methods have recently emerged as the mainstream framework for HOI detection, they still suffer from a key limitation: Randomly initialized queries lack explici

Cited by 0SourcePDFScholar
2026

What-Meets-Where: Unified Learning of Action and Contact Localization in Images

AAAI 2026technical

People control their bodies to establish contact with the environment. To comprehensively understand actions across diverse visual contexts, it is essential to simultaneously consider what action is occurring and where it is happening. Current methodologies, however, often inadequately capture this

Cited by 0SourcePDFScholar
2025

LLM Agents Can Be Choice-Supportive Biased Evaluators: An Empirical Study

AAAI 2025technical

With Large Language Model (LLM) agents taking on more evaluation responsibilities in decision-making, it is essential to recognize their possible biases to guarantee fair and trustworthy AI-supported decisions. This study is the first to thoroughly examine the choice-supportive bias in LLM agents, a…

Cited by 0SourcePDFScholar
2025

Precision-Enhanced Human-Object Contact Detection via Depth-Aware Perspective Interaction and Object Texture Restoration

AAAI 2025technical

Human-object contact (HOT) is designed to accurately identify the areas where humans and objects come into contact. Current methods frequently fail to account for scenarios where objects are frequently blocking the view, resulting in inaccurate identification of contact areas. To tackle this problem…

Cited by 2SourcePDFScholar
2025

Prompt Guidance and Human Proximal Perception for HOT Prediction with Regional Joint Loss

ICCV 2025poster

The task of Human-Object conTact (HOT) detection involves identifying the specific areas of the human body that are touching objects. Nevertheless, current models are restricted to just one type of image, often leading to too much segmentation in areas with little interaction, and struggling to main…

2025

Pruning for Sparse Diffusion Models Based on Gradient Flow

ICASSP 2025accepted

Diffusion Models (DMs) have impressive capabilities among generation models, but are limited to slower inference speeds and higher computational costs. Previous works utilize one-shot structure pruning to derive lightweight DMs from pre-trained ones, but this approach often leads to a significant dr…

Cited by 0SourceScholar
2023

Improving Zero-Shot Generalization and Robustness of Multi-Modal Models

CVPR 2023poster

Multi-modal image-text models such as CLIP and LiT have demonstrated impressive performance on image classification benchmarks and their zero-shot generalization ability is particularly exciting. While the top-5 zero-shot accuracies of these models are very high, the top-1 accuracies are much lower…

2023

Learning To Generate Image Embeddings With User-Level Differential Privacy

CVPR 2023poster

Small on-device models have been successfully trained with user-level differential privacy (DP) for next word prediction and image classification tasks in the past. However, existing methods can fail when directly applied to learn embedding models using supervised training data with a large class sp…

2021

Learning View-Disentangled Human Pose Representation by Contrastive Cross-View Mutual Information Maximization

CVPR 2021poster

We introduce a novel representation learning method to disentangle pose-dependent as well as view-dependent factors from 2D human poses. The method trains a network using cross-view mutual information maximization (CV-MIM) which maximizes mutual information of the same pose performed from different…

Cited by 41PDFcodeScholar