← Search

YuZhuo Fu

19 accepted papers

2026

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models

AAAI 2026technical

Recently, multimodal large language models (MLLMs) have achieved significant advancements across various domains, and corresponding evaluation benchmarks have been continuously refined and improved. In this process, benchmarks in the scientific domain have played an important role in assessing the r

Cited by 0SourcePDFScholar
2026

Unnoticed Yet Effective: A Hybrid Physical Camouflage Framework Against DNNs and Human Perception

AAAI 2026technical

While adversarial attacks can effectively deceive deep neural networks, their real-world applicability is often limited by complex and conspicuous patterns that reveal their attack intent to human observers. To overcome this limitation, we propose UYE, a novel camouflage framework designed to simult

Cited by 0SourcePDFScholar
2025

A Training-Free Correlation-Weighted Model for Zero-/Few-Shot Industrial Anomaly Detection with Retrieval Augmentation

ICASSP 2025accepted

Obtaining labeled data in the field of industrial anomaly detection is challenging, which necessitates the development of label-free frameworks. However, current methods mainly focus on the unsupervised paradigm, which uses a large number of normal samples of the same category to train the model, an…

Cited by 0SourceScholar
2025

Exploring Generalization Boundaries of Unsupervised Industrial Anomaly Detection Models through Attribute Perturbation

ICASSP 2025accepted

Industrial anomaly detection (IAD) plays a crucial role in large-scale industrial manufacturing. Recently, numerous unsupervised algorithms have been proposed and achieved remarkable performance on benchmark datasets. Given the high homogeneity of samples during training and testing, it appears that…

Cited by 0SourceScholar
2025

GPA: Enhancing Generalizable Physical Adversarial Attacks Across Multiple Vision Tasks

ICASSP 2025accepted

Adversarial attacks pose a significant challenge in deep learning, as carefully crafted perturbations can severely degrade even the most advanced models. In real-world scenarios, where the target models are often unknown, previous works often focus on creating adversarial patterns for specific known…

Cited by 0SourceScholar
2025

MM-CamObj: A Comprehensive Multimodal Dataset for Camouflaged Object Scenarios

AAAI 2025technical

Large visual-language models (LVLMs) have achieved great success in multiple applications. However, they still encounter challenges in complex scenes, especially those involving camouflaged objects. This is primarily due to the lack of samples related to camouflaged scenes in the training dataset. T…

2025

SmartRAG: Jointly Learn RAG-Related Tasks From the Environment Feedback

ICLR 2025poster

RAG systems consist of multiple modules to work together. However, these modules are usually separately trained. We argue that a system like RAG that incorporates multiple modules should be jointly optimized to achieve optimal performance. To demonstrate this, we design a specific pipeline called Sm…

Cited by 4SourcePDFScholar
2025

TTE: Two Tokens Are Enough to Improve Parameter-Efficient Tuning

AAAI 2025technical

Existing fine-tuning paradigms are predominantly characterized by Full Parameter Tuning (FPT) and Parameter-Efficient Tuning (PET). FPT fine-tunes all parameters of a pre-trained model on downstream tasks, whereas PET freezes the pre-trained model and employs only a minimal number of learnable param…

2025

VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models

ICCV 2025poster

Although large visual-language models (LVLMs) have demonstrated strong performance in multimodal tasks, errors may occasionally arise due to biases during the reasoning process. Recently, reward models (RMs) have become increasingly pivotal in the reasoning process. Specifically, process RMs evaluat…

2024

From Raw Video to Pedagogical Insights: A Unified Framework for Student Behavior Analysis

AAAI 2024technical

Understanding student behavior in educational settings is critical in improving both the quality of pedagogy and the level of student engagement. While various AI-based models exist for classroom analysis, they tend to specialize in limited tasks and lack generalizability across diverse educational…

Cited by 5SourcePDFScholar
2024

LAMM: Label Alignment for Multi-Modal Prompt Learning

AAAI 2024technical

With the success of pre-trained visual-language (VL) models such as CLIP in visual representation tasks, transferring pre-trained models to downstream tasks has become a crucial paradigm. Recently, the prompt tuning paradigm, which draws inspiration from natural language processing (NLP), has made s…

2023

AV-TAD: Audio-Visual Temporal Action Detection With Transformer

ICASSP 2023accepted

As an important and challenging task in video understanding, Temporal Action Detection (TAD) has been deeply studied in recent years. However, current works mainly tackle this task with visual information, while neglecting to explore the potential of the audio modality. To address this challenge, in…

Cited by 0SourceScholar
2023

CC-PoseNet: Towards Human Pose Estimation in Crowded Classrooms

ICASSP 2023accepted

Human pose estimation has long been motivated for its application in human behavior understanding and activity recognition. Despite recent advances in multi-person pose estimation, existing solutions remain challenging in crowded scenes, especially in classroom scenarios where students are extremely…

Cited by 0SourceScholar
2023

LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming

ACL 2023long

Open-domain dialogue systems have made promising progress in recent years. While the state-of-the-art dialogue agents are built upon large-scale social media data and large pre-trained models, there is no guarantee these agents could also perform well in fast-growing scenarios, such as live streamin…

2022

Development and Stiffness Optimization for a Flexible-Tail Robotic Fish

RA-L 2022

The integral flexible tail has the potential advantage of lifelike undulating motion. However, due to the complex manufacturing process and difficult modification of structural parameters, its application in robotic fish encounters many challenges. Combining rigid structure and flexible material, th

Cited by 17SourceScholar
2022

Synpose: A Large-Scale and Densely Annotated Synthetic Dataset for Human Pose Estimation in Classroom

ICASSP 2022accepted

Deep learning-based methods for human pose estimation require large volumes of training data to achieve superior performance. However, data acquisition in classroom environments raises privacy concerns, which will undoubtedly hinder the development of the latest deep learning techniques in education…

Cited by 0SourceScholar
2021

Improving Abstractive Dialogue Summarization with Hierarchical Pretraining and Topic Segment

EMNLP 2021finding

With the increasing abundance of meeting transcripts, meeting summary has attracted more and more attention from researchers. The unsupervised pre-training method based on transformer structure combined with fine-tuning of downstream tasks has achieved great success in the field of text summarizatio…

Cited by 22SourcePDFScholar
2021

Taking A Closer Look at Synthesis: Fine-Grained Attribute Analysis for Person Re-Identification

ICASSP 2021accepted

Person re-identification (re-ID) plays an important role in applications such as public security and video surveillance. Recently, learning from synthetic data, which benefits from the popularity of synthetic data engine, has achieved remarkable performance. However, in pursuit of high accuracy, res…

Cited by 0SourceScholar