← Search

Zuopeng Yang

6 accepted papers

2026

Bi-LoRA: Efficient Sharpness-Aware Minimization for Fine-Tuning Large-Scale Models

ICLR 2026poster

Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning of large pre-trained models. Yet LoRA can face generalization challenges. One promising way to improve the generalization is Sharpness-Aware Minimization (SAM), which has proven effective for small-scale training scenarios. In this p…

Cited by 0SourceScholar
2025

Distraction is All You Need for Multimodal Large Language Model Jailbreaking

CVPR 2025highlight

Multimodal Large Language Models (MLLMs) bridge the gap between visual and textual data, enabling a range of advanced applications. However, complex internal interactions among visual elements and their alignment with text can introduce vulnerabilities, which may be exploited to bypass safety mechan…

Cited by 1SourcePDFScholar
2025

SGC-Net: Stratified Granular Comparison Network for Open-Vocabulary HOI Detection

CVPR 2025poster

Recent open-vocabulary human-object interaction (OV-HOI) detection methods primarily rely on large language model (LLM) for generating auxiliary descriptions and leverage knowledge distilled from CLIP to detect unseen interaction categories. Despite their effectiveness, these methods face two challe…

2024

TD²-Net: Toward Denoising and Debiasing for Video Scene Graph Generation

AAAI 2024technical

Dynamic scene graph generation (SGG) focuses on detecting objects in a video and determining their pairwise relationships. Existing dynamic SGG methods usually suffer from several issues, including 1) Contextual noise, as some frames might contain occluded and blurred objects. 2) Label bias, primari…

Cited by 4SourcePDFScholar
2023

Unified Discrete Diffusion for Simultaneous Vision-Language Generation

ICLR 2023poster

The recently developed discrete diffusion model performs extraordinarily well in generation tasks, especially in the text-to-image task, showing great potential for modeling multimodal signals. In this paper, we leverage these properties and present a unified multimodal generation model, which can p…

2022

Modeling Image Composition for Complex Scene Generation

CVPR 2022poster

We present a method that achieves state-of-the-art results on challenging (few-shot) layout-to-image generation tasks by accurately modeling textures, structures and relationships contained in a complex scene. After compressing RGB images into patch tokens, we propose the Transformer with Focal Atte…

Cited by 57PDFcodeScholar