← Search

Miaozeng Du

5 accepted papers

2026

Forget What Has Seen: Selective Concept Unlearning in Segmentation Foundation Models

AAAI 2026technical

Machine unlearning (MU) has emerged as a critical tool for removing sensitive or personal information from machine learning models, empowering individuals with the right to be forgotten. While MU has achieved success in classification and generative tasks, whether this technique can be effectively a

Cited by 0SourcePDFScholar
2025

Forget the Token and Pixel: Rethinking Gradient Ascent for Concept Unlearning in Multimodal Generative Models

ACL 2025finding

Gradient Ascent (GA) has emerged as a promising approach for concept unlearning in Multimodal Generative Models (MGMs), such as Multimodal Large Language Models (MLLMs) and Stable Diffusion Models (SDMs). Despite its effectiveness in removing undesired knowledge, GA leads to severe utility degradati…

Cited by 0SourcePDFScholar
2024

MIKE: A New Benchmark for Fine-grained Multimodal Entity Knowledge Editing

ACL 2024findings

Multimodal knowledge editing represents a critical advancement in enhancing the capabilities of Multimodal Large Language Models (MLLMs). Despite its potential, current benchmarks predominantly focus on coarse-grained knowledge, leaving the intricacies of fine-grained (FG) multimodal entity knowledg…

Cited by 3SourcePDFScholar
2024

Single Image Unlearning: Efficient Machine Unlearning in Multimodal Large Language Models

NeurIPS 2024poster

Machine unlearning (MU) empowers individuals with the `right to be forgotten' by removing their private or sensitive information encoded in machine learning models. However, it remains uncertain whether MU can be effectively applied to Multimodal Large Language Models (MLLMs), particularly in scenar…

Cited by 8SourcePDFScholar
2023

Three Stream Based Multi-level Event Contrastive Learning for Text-Video Event Extraction

EMNLP 2023long main

Text-video based multimodal event extraction refers to identifying event information from the given text-video pairs. Existing methods predominantly utilize video appearance features (VAF) and text sequence features (TSF) as input information. Some of them employ contrastive learning to align VAF wi…

Cited by 0SourceScholar