← Search

wenya guo

9 accepted papers

2026

Learning from Noisy Supervision: A Denoising-Debiasing Framework for Weakly Supervised Video Anomaly Detection

CVPR 2026

Weakly supervised video anomaly detection (WS-VAD) aims to localize frame-level anomalies using only video-level labels. This task is typically formulated within a multiple instance learning (MIL) paradigm, where each video is treated as a bag of snippets, achieving robust performance without requir

Cited by 0SourcecodeScholar
2025

BiDeV: Bilateral Defusing Verification for Complex Claim Fact-Checking

AAAI 2025technical

Complex claim fact-checking performs a crucial role in disinformation detection. However, existing fact-checking methods struggle with claim vagueness, specifically in effectively handling latent information and complex relations within claims. Moreover, evidence redundancy, where non-essential info…

2025

Zero-shot Document Retrieval with Hybrid Pseudo-document Retriever

ICASSP 2025accepted

The zero-shot retrieval task aims to retrieve the most relevant documents to a user’s query without relevance labels. Current approaches expand input queries by generating pseudo-documents with large language models (LLMs) and perform document retrieval based on the expanded queries. However, their…

Cited by 0SourceScholar
2024

Look before You Leap: Dual Logical Verification for Knowledge-based Visual Question Generation

COLING 2024main

Knowledge-based Visual Question Generation aims to generate visual questions with outside knowledge other than the image. Existing approaches are answer-aware, which incorporate answers into the question-generation process. However, these methods just focus on leveraging the semantics of inputs to p…

2023

AoM: Detecting Aspect-oriented Information for Multimodal Aspect-Based Sentiment Analysis

ACL 2023findings

Multimodal aspect-based sentiment analysis (MABSA) aims to extract aspects from text-image pairs and recognize their sentiments. Existing methods make great efforts to align the whole image to corresponding aspects. However, different regions of the image may relate to different aspects in the same…

2023

HyperPELT: Unified Parameter-Efficient Language Model Tuning for Both Language and Vision-and-Language Tasks

ACL 2023findings

With the scale and capacity of pretrained models growing rapidly, parameter-efficient language model tuning has emerged as a popular paradigm for solving various NLP and Vision-and-Language (V&L) tasks. In this paper, we design a unified parameter-efficient multitask learning framework that works ef…

Cited by 17SourcePDFScholar
2023

Licon: A Diverse, Controllable and Challenging Linguistic Concept Learning Benchmark

EMNLP 2023long findings

Concept Learning requires learning the definition of a general category from given training examples. Most of the existing methods focus on learning concepts from images. However, the visual information cannot present abstract concepts exactly, which struggles the introduction of novel concepts rela…

Cited by 0SourceScholar
2022

A Span-based Multimodal Variational Autoencoder for Semi-supervised Multimodal Named Entity Recognition

EMNLP 2022main

Multimodal named entity recognition (MNER) on social media is a challenging task which aims to extract named entities in free text and incorporate images to classify them into user-defined types. However, the annotation for named entities on social media demands a mount of human efforts. The existin…

2021

MTAAL: Multi-Task Adversarial Active Learning for Medical Named Entity Recognition and Normalization

AAAI 2021technical

Automated medical named entity recognition and normalization are fundamental for constructing knowledge graphs and building QA systems. When it comes to medical text, the annotation demands a foundation of expertise and professionalism. Existing methods utilize active learning to reduce costs in cor…

Cited by 20SourcePDFScholar