← Search

Hengyuan Zhao

4 accepted papers

2026

CoPE: Continual Probe-guided Expansion for Large Vision-Language Models

ICML 2026poster

Mixture of Experts architectures have recently advanced the scalability and adaptability of Large Language Models for continual multimodal learning. However, extending these models to accommodate sequential tasks remains challenging. As new tasks arrive, naive model expansion leads to rapid paramete…

Cited by 0SourceScholar
2026

VaccineRAG: Boosting Multimodal Large Language Models’ Immunity to Harmful RAG Samples

AAAI 2026technical

Retrieval Augmented Generation enhances the response accuracy of Large Language Models (LLMs) by integrating retrieval and generation modules with external knowledge, demonstrating particular strength in real-time queries and Visual Question Answering tasks. However, the effectiveness of RAG is fre

Cited by 0SourcePDFScholar
2024

LOVA3: Learning to Visual Question Answering, Asking and Assessment

NeurIPS 2024poster

Question answering, asking, and assessment are three innate human traits crucial for understanding the world and acquiring knowledge. By enhancing these capabilities, humans can more effectively utilize data, leading to better comprehension and learning outcomes. However, current Multimodal Large La…

2021

ClassSR: A General Framework to Accelerate Super-Resolution Networks by Data Characteristic

CVPR 2021poster

We aim at accelerating super-resolution (SR) networks on large images (2K-8K). The large images are usually decomposed into small sub-images in practical usages. Based on this processing, we found that different image regions have different restoration difficulties and can be processed by networks w…

Cited by 217PDFcodeScholar