← Search

Shuyong Gao

10 accepted papers

2026

Commonality in Few: Few-Shot Multimodal Anomaly Detection via Hypergraph-Enhanced Memory

AAAI 2026technical

Few-shot multimodal industrial anomaly detection is a critical yet underexplored task, offering the ability to quickly adapt to complex industrial scenarios. In few-shot settings, insufficient training samples often fail to cover the diverse patterns present in test samples. This challenge can be mi

Cited by 0SourcePDFScholar
2026

RSAgent: Learning to Reason and Act via Multi-Turn Tool Invocations for Text-Guided Segmentation

ICML 2026poster

Text-guided object segmentation requires both cross-modal reasoning and pixel grounding abilities. Most recent methods treat it as a single forward pass, where the model directly predicts pixel prompts to a segmentation model, which limits verification, refocusing and refinement when initial localiz…

Cited by 0SourceScholar
2025

General Compression Framework for Efficient Transformer Object Tracking

ICCV 2025poster

Previous works have attempted to improve tracking efficiency through lightweight architecture design or knowledge distillation from teacher models to compact student trackers. However, these solutions often sacrifice accuracy for speed to a great extent, and also have the problems of complex trainin…

2025

Scoring, Remember, and Reference: Catching Camouflaged Objects in Videos

ICCV 2025poster

Video Camouflaged Object Detection (VCOD) aims to segment objects whose appearances closely resemble their surroundings, posing a challenging and emerging task. Existing vision models often struggle in such scenarios due to the indistinguishable appearance of camouflaged objects and the insufficient…

Cited by 0SourcePDFScholar
2023

TINYCOD: Tiny and Effective Model for Camouflaged Object Detection

ICASSP 2023accepted

This paper introduces an effective and tiny model for real-time Camouflaged Object Detection (COD) named Tiny-COD. It achieves high performance with very low costs (Parameters < 5M, FLOPs < 1.5G), which can be applied on mobile devices. Specifically, we introduce a simple but effective Adjacent Scal…

Cited by 0SourceScholar
2022

FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos

CVPR 2022poster

Current benchmarks for facial expression recognition (FER) mainly focus on static images, while there are limited datasets for FER in videos. It is still ambiguous to evaluate whether performances of existing methods remain satisfactory in real-world application-oriented scenes. For example, the "Ha…

Cited by 110PDFcodeScholar
2022

Position-aware Joint Entity and Relation Extraction with Attention Mechanism

IJCAI 2022poster

Named entity recognition and relation extraction are two important core subtasks of information extraction, which aim to identify named entities and extract relations between them. In recent years, span representation methods have received a lot of attention and are widely used to extract entities a…

Cited by 6SourcePDFScholar
2022

Weakly-Supervised Salient Object Detection Using Point Supervision

AAAI 2022technical

Current state-of-the-art saliency detection models rely heavily on large datasets of accurate pixel-wise annotations, but manually labeling pixels is time-consuming and labor-intensive. There are some weakly supervised methods developed for alleviating the problem, such as image label, bounding box…

2021

Dual-Stream Network Based On Global Guidance for Salient Object Detection

ICASSP 2021accepted

High-level features can help low-level features eliminate semantic ambiguity, which is crucial for obtaining the precise salient object. Some methods use high-level features to provide global guidance for some layers of the network. However, there remain several problems: (1) the global guidance has…

Cited by 0SourceScholar