← Search

Xiaoyuan Yu

4 accepted papers

2026

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention

CVPR 2026

Vision-Language-Action (VLA) models have shown remarkable progress in embodied tasks recently, but most methods process visual observations independently at each timestep. This history-agnostic design treats robot manipulation as a Markov Decision Process, even though real-world robotic control is i

Cited by 0SourceScholar
2023

Multimodal High-order Relation Transformer for Scene Boundary Detection

ICCV 2023poster

Scene boundary detection breaks down long videos into meaningful story-telling units and plays a crucial role in high-level video understanding. Despite significant advancements in this area, this task remains a challenging problem as it requires a comprehensive understanding of multimodal cues and…

Cited by 5PDFScholar
2022

TA2N: Two-Stage Action Alignment Network for Few-Shot Action Recognition

AAAI 2022technical

Few-shot action recognition aims to recognize novel action classes (query) using just a few samples (support). The majority of current approaches follow the metric learning paradigm, which learns to compare the similarity between videos. Recently, it has been observed that directly measuring this si…

2021

Uncertainty Guided Collaborative Training for Weakly Supervised Temporal Action Detection

CVPR 2021poster

Weakly supervised temporal action detection aims to localize temporal boundaries of actions and identify their categories simultaneously with only video-level category labels during training. Among existing methods, attention-based methods have achieved superior performance by separating action and…

Cited by 105PDFScholar