← Search

Haopeng Zhang

14 accepted papers

2026

Breaking the Illusion: When Positive Meets Negative in Multimodal Decoding

CVPR 2026

Vision-Language Models (VLMs) are frequently undermined by object hallucination--generating content that contradicts visual reality--due to an over-reliance on linguistic priors. We introduce Positive-and-Negative Decoding (PND), a training-free inference framework that intervenes directly in the de

Cited by 0SourcecodeScholar
2026

Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots

ICML 2026poster

Large language model (LLM) answer engines are increasingly used for information seeking, shifting visibility from ranked lists to synthesized answers. This enables Generative Engine Optimization (GEO), which targets LLM answer engines' evidence pool and generation. We analyze the search engine optim…

Cited by 0SourceScholar
2025

Beyond Spatial Domain: Cross-domain Promoted Fourier Convolution Helps Single Image Dehazing

AAAI 2025technical

Vanilla convolution and window-based self-attention have shown significant success in image dehazing. However, they are constrained by limited receptive fields and ignore frequency gaps between dehazed and clear images. The former hampers the modeling of global dependencies, while the latter impedes…

Cited by 0SourcePDFScholar
2025

DomainSum: A Hierarchical Benchmark for Fine-Grained Domain Shift in Abstractive Text Summarization

NAACL 2025findings

Most research on abstractive summarization focuses on single-domain applications, often neglecting how domain shifts between documents affect performance and the generalization ability of summarization models. To address this issue, we introduce DomainSum, a hierarchical benchmark designed to captur…

2025

FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Models

EMNLP 2025

While large language models (LLMs) have demonstrated impressive performance across a wide range of natural language processing (NLP) tasks in high-resource languages, their capabilities in low-resource and minority languages remain significantly underexplored. Formosan languages—a subgroup of Austro

2024

Unveiling the Magic: Investigating Attention Distillation in Retrieval-Augmented Generation

NAACL 2024short

Retrieval-augmented generation framework addresses the limitations of large language models by enabling real-time knowledge updates for more accurate answers. An efficient way in the training phase of retrieval-augmented models is attention distillation, which uses attention scores as supervision si…

Cited by 3SourcePDFScholar
2024

XATU: A Fine-grained Instruction-based Benchmark for Explainable Text Updates

COLING 2024main

Text editing is a crucial task of modifying text to better align with user intents. However, existing text editing benchmark datasets contain only coarse-grained instructions and lack explainability, thus resulting in outputs that deviate from the intended changes outlined in the gold reference. To…

2023

DiffuSum: Generation Enhanced Extractive Summarization with Diffusion

ACL 2023findings

Extractive summarization aims to form a summary by directly extracting sentences from the source document. Existing works mostly formulate it as a sequence labeling problem by making individual sentence label predictions. This paper proposes DiffuSum, a novel paradigm for extractive summarization, b…

2022

An Accelerated Rank-(L, L, 1, 1) Block Term Decomposition Of Multi-Subject Fmri Data Under Spatial Orthonormality Constraint

ICASSP 2022accepted

The decomposition of multi-subject fMRI data using rank-(L,L,1,1) block term decomposition (BTD) can preserve higher-way data structure and is more robust to noise effects by decomposing shared spatial maps (SMs) into a product of two rank-L loading matrices. However, since the number of whole-brain…

Cited by 4SourceScholar
2022

Improving the Faithfulness of Abstractive Summarization via Entity Coverage Control

NAACL 2022findings

Abstractive summarization systems leveraging pre-training language models have achieved superior results on benchmark datasets. However, such models have been shown to be more prone to hallucinate facts that are unfaithful to the input context. In this paper, we propose a method to remedy entity-lev…

Cited by 37SourcePDFScholar
2022

Multi-Frame Super-Resolution With Raw Images Via Modified Deformable Convolution

ICASSP 2022accepted

In this paper we propose a novel model towards multi-frame super-resolution, which leverages multiple RAW images and yields a super-resolved RGB image. To facilitate the pixel misalignment in burst photography, we apply a refined Pyramid Cascading and Deformable Convolution (PCD) feature alignment m…

Cited by 0SourceScholar