← Search

Xiaoshan Yang

13 accepted papers

2025

Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation

NeurIPS 2025poster

In recent years, Multimodal Large Language Models (MLLMs) have been extensively utilized for multimodal reasoning tasks, including Graphical User Interface (GUI) automation. Unlike general offline multimodal tasks, GUI automation is executed in online interactive environments, necessitating step-by-…

Cited by 0SourcecodeScholar
2025

Pilot: Building the Federated Multimodal Instruction Tuning Framework

AAAI 2025technical

In this paper, we explore a novel federated multimodal instruction tuning task(FedMIT), which is significant for collaboratively fine-tuning MLLMs on different types of multimodal instruction data on distributed devices. To solve the new task, we propose a federated multimodal instruction tuning fra…

Cited by 1SourcePDFScholar
2025

Pseudo Informative Episode Construction for Few-Shot Class-Incremental Learning

AAAI 2025technical

Few-Shot Class-Incremental Learning (FSCIL) studies how to empower the machine learning system to learn novel classes with only a few annotated examples continually. To tackle the FSCIL task, recent state-of-the-art methods propose to employ the meta-learning mechanism, which constructs the pseudo i…

Cited by 0SourcePDFScholar
2024

Libra: Building Decoupled Vision System on Large Language Models

ICML 2024poster

In this work, we introduce **Libra**, a prototype model with a decoupled vision system on a large language model (LLM). The decoupled vision system decouples inner-modal modeling and cross-modal interaction, yielding unique visual information modeling and effective cross-modal comprehension. Libra i…

2024

Modality-Collaborative Test-Time Adaptation for Action Recognition

CVPR 2024poster

Video-based Unsupervised Domain Adaptation (VUDA) method improves the generalization of the video model enabling it to be applied to action recognition tasks in different environments. However these methods require continuous access to source data during the adaptation process which are impractical…

Cited by 6SourcePDFScholar
2024

OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling

NeurIPS 2024poster

Constrained by the separate encoding of vision and language, existing grounding and referring segmentation works heavily rely on bulky Transformer-based fusion en-/decoders and a variety of early-stage interaction technologies. Simultaneously, the current mask visual language modeling (MVLM) fails t…

2023

Active Exploration of Multimodal Complementarity for Few-Shot Action Recognition

CVPR 2023poster

Recently, few-shot action recognition receives increasing attention and achieves remarkable progress. However, previous methods mainly rely on limited unimodal data (e.g., RGB frames) while the multimodal information remains relatively underexplored. In this paper, we propose a novel Active Multimod…

Cited by 42SourcePDFScholar
2023

Multi-modal Queried Object Detection in the Wild

NeurIPS 2023poster

We introduce MQ-Det, an efficient architecture and pre-training strategy design to utilize both textual description with open-set generalization and visual exemplars with rich description granularity as category queries, namely, Multi-modal Queried object Detection, for real-world detection with bot…

2022

Cross-Modal Federated Human Activity Recognition via Modality-Agnostic and Modality-Specific Representation Learning

AAAI 2022technical

In this paper, we propose a new task of cross-modal federated human activity recognition (CMF-HAR), which is conducive to promote the large-scale use of the HAR model on more local devices. To address the new task, we propose a feature-disentangled activity recognition network (FDARN), which has fiv…

Cited by 30SourcePDFScholar
2022

Shifting More Attention to Visual Backbone: Query-Modulated Refinement Networks for End-to-End Visual Grounding

CVPR 2022poster

Visual grounding focuses on establishing fine-grained alignment between vision and natural language, which has essential applications in multimodal reasoning systems. Existing methods use pre-trained query-agnostic visual backbones to extract visual feature maps independently without considering the…

Cited by 88PDFcodeScholar
2021

ECKPN: Explicit Class Knowledge Propagation Network for Transductive Few-Shot Learning

CVPR 2021poster

Recently, the transductive graph-based methods have achieved great success in the few-shot classification task. However, most existing methods ignore exploring the class-level knowledge that can be easily learned by humans from just a handful of samples. In this paper, we propose an Explicit Class K…

Cited by 77PDFScholar