← Search

Jiacheng Ruan

10 accepted papers

2026

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models

AAAI 2026technical

Recently, multimodal large language models (MLLMs) have achieved significant advancements across various domains, and corresponding evaluation benchmarks have been continuously refined and improved. In this process, benchmarks in the scientific domain have played an important role in assessing the r

Cited by 0SourcePDFScholar
2026

Unnoticed Yet Effective: A Hybrid Physical Camouflage Framework Against DNNs and Human Perception

AAAI 2026technical

While adversarial attacks can effectively deceive deep neural networks, their real-world applicability is often limited by complex and conspicuous patterns that reveal their attack intent to human observers. To overcome this limitation, we propose UYE, a novel camouflage framework designed to simult

Cited by 0SourcePDFScholar
2025

Dynamic Data Mixing Maximizes Instruction Tuning for Mixture-of-Experts

NAACL 2025long

Mixture-of-Experts (MoE) models have shown remarkable capability in instruction tuning, especially when the number of tasks scales. However, previous methods simply merge all training tasks (e.g. creative writing, coding, and mathematics) and apply fixed sampling weights, without considering the imp…

2025

GPA: Enhancing Generalizable Physical Adversarial Attacks Across Multiple Vision Tasks

ICASSP 2025accepted

Adversarial attacks pose a significant challenge in deep learning, as carefully crafted perturbations can severely degrade even the most advanced models. In real-world scenarios, where the target models are often unknown, previous works often focus on creating adversarial patterns for specific known…

Cited by 0SourceScholar
2025

MM-CamObj: A Comprehensive Multimodal Dataset for Camouflaged Object Scenarios

AAAI 2025technical

Large visual-language models (LVLMs) have achieved great success in multiple applications. However, they still encounter challenges in complex scenes, especially those involving camouflaged objects. This is primarily due to the lack of samples related to camouflaged scenes in the training dataset. T…

2025

TTE: Two Tokens Are Enough to Improve Parameter-Efficient Tuning

AAAI 2025technical

Existing fine-tuning paradigms are predominantly characterized by Full Parameter Tuning (FPT) and Parameter-Efficient Tuning (PET). FPT fine-tunes all parameters of a pre-trained model on downstream tasks, whereas PET freezes the pre-trained model and employs only a minimal number of learnable param…

2025

VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models

ICCV 2025poster

Although large visual-language models (LVLMs) have demonstrated strong performance in multimodal tasks, errors may occasionally arise due to biases during the reasoning process. Recently, reward models (RMs) have become increasingly pivotal in the reasoning process. Specifically, process RMs evaluat…

2024

LAMM: Label Alignment for Multi-Modal Prompt Learning

AAAI 2024technical

With the success of pre-trained visual-language (VL) models such as CLIP in visual representation tasks, transferring pre-trained models to downstream tasks has become a crucial paradigm. Recently, the prompt tuning paradigm, which draws inspiration from natural language processing (NLP), has made s…

2024

LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-Training

EMNLP 2024main

Mixture-of-Experts (MoE) has gained increasing popularity as a promising framework for scaling up large language models (LLMs). However, training MoE from scratch in a large-scale setting still suffers from data-hungry and instability problems. Motivated by this limit, we investigate building MoE mo…

2024

VT-ReID: Learning Discriminative Visual-Text Representation for Polyp Re-Identification

ICASSP 2024accepted

Colonoscopic Polyp Re-Identification (ReID) aims to match a specific polyp in a large gallery with different cameras and views, which plays a key role in the prevention and treatment of colorectal cancer in the computer-aided diagnosis. However, traditional methods mainly focus on the visual represe…

Cited by 0SourceScholar