← Search

Chuanyi Zhang

13 accepted papers

2026

Knowledge Externalization: Reversible Unlearning and Modular Retrieval in Multimodal Large Language Models

ICLR 2026poster

Multimodal Large Language Models (MLLMs) achieve remarkable cross-modal understanding by training on vast web-scale datasets, but inadvertently internalize sensitive personal and proprietary information. Existing machine unlearning methods address this by irreversibly altering model parameters to pe…

Cited by 0SourceScholar
2026

RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow

AAAI 2026technical

Remote sensing imagery presents vast, inherently unstructured spatial data, necessitating sophisticated reasoning to interpret complex user intents and contextual relationships beyond simple recognition tasks. In this paper, we aim to construct an Earth observation workflow to handle complex queries

Cited by 0SourcePDFScholar
2025

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions

EMNLP 2025

While densely annotated image captions significantly facilitate the learning of robust vision-language alignment, methodologies for systematically optimizing human annotation efforts remain underexplored. We introduce Chain-of-Talkers (CoTalk), an AI-in-the-loop methodology designed to maximize the

Cited by 0SourcePDFScholar
2025

Forget the Token and Pixel: Rethinking Gradient Ascent for Concept Unlearning in Multimodal Generative Models

ACL 2025finding

Gradient Ascent (GA) has emerged as a promising approach for concept unlearning in Multimodal Generative Models (MGMs), such as Multimodal Large Language Models (MLLMs) and Stable Diffusion Models (SDMs). Despite its effectiveness in removing undesired knowledge, GA leads to severe utility degradati…

Cited by 0SourcePDFScholar
2025

Making Large Vision Language Models to Be Good Few-Shot Learners

AAAI 2025technical

Few-shot classification (FSC) is a fundamental yet challenging task in computer vision that involves recognizing novel classes from limited data. While previous methods have focused on enhancing visual features or incorporating additional modalities, Large Vision Language Models (LVLMs) offer a prom…

2025

Open-World Attribute Mining for E-Commerce Products with Multimodal Self-Correction Instruction Tuning

ACL 2025long

In e-commerce, effective product Attribute Mining (AM) is essential for improving product features and aiding consumer decisions. However, current AM methods often focus on extracting attributes from unimodal text, underutilizing multimodal data. In this paper, we propose a novel framework called Mu…

Cited by 0SourcePDFScholar
2025

Prompting DirectSAM for Semantic Contour Extraction in Remote Sensing Images

ICASSP 2025accepted

The Direct Segment Anything Model (DirectSAM) excels in class-agnostic contour extraction. In this paper, we explore its use by applying it to optical remote sensing imagery, where semantic contour extraction—such as identifying buildings, road networks, and coastlines-holds significant practical va…

Cited by 6SourceScholar
2025

RemoteTrimmer: Adaptive Structural Pruning for Remote Sensing Image Classification

ICASSP 2025accepted

Since high resolution remote sensing image classifi-cation often requires a relatively high computation complexity, lightweight models tend to be practical and efficient. Model pruning is an effective method for model compression. However, existing methods rarely take into account the specificity of…

Cited by 0SourceScholar
2024

MIKE: A New Benchmark for Fine-grained Multimodal Entity Knowledge Editing

ACL 2024findings

Multimodal knowledge editing represents a critical advancement in enhancing the capabilities of Multimodal Large Language Models (MLLMs). Despite its potential, current benchmarks predominantly focus on coarse-grained knowledge, leaving the intricacies of fine-grained (FG) multimodal entity knowledg…

Cited by 3SourcePDFScholar
2024

Single Image Unlearning: Efficient Machine Unlearning in Multimodal Large Language Models

NeurIPS 2024poster

Machine unlearning (MU) empowers individuals with the `right to be forgotten' by removing their private or sensitive information encoded in machine learning models. However, it remains uncertain whether MU can be effectively applied to Multimodal Large Language Models (MLLMs), particularly in scenar…

Cited by 8SourcePDFScholar
2023

Three Stream Based Multi-level Event Contrastive Learning for Text-Video Event Extraction

EMNLP 2023long main

Text-video based multimodal event extraction refers to identifying event information from the given text-video pairs. Existing methods predominantly utilize video appearance features (VAF) and text sequence features (TSF) as input information. Some of them employ contrastive learning to align VAF wi…

Cited by 0SourceScholar
2021

Jo-SRC: A Contrastive Approach for Combating Noisy Labels

CVPR 2021poster

Due to the memorization effect in Deep Neural Networks (DNNs), training with noisy labels usually results in inferior model performance. Existing state-of-the-art methods primarily adopt a sample selection strategy, which selects small-loss samples for subsequent training. However, prior literature…

Cited by 190PDFScholar
2021

Non-Salient Region Object Mining for Weakly Supervised Semantic Segmentation

CVPR 2021poster

Semantic segmentation aims to classify every pixel of an input image. Considering the difficulty of acquiring dense labels, researchers have recently been resorting to weak labels to alleviate the annotation burden of segmentation. However, existing works mainly concentrate on expanding the seed of…

Cited by 248PDFcodeScholar