← Search

Fengying Xie

6 accepted papers

2026

Breaking the Illusion: When Positive Meets Negative in Multimodal Decoding

CVPR 2026

Vision-Language Models (VLMs) are frequently undermined by object hallucination--generating content that contradicts visual reality--due to an over-reliance on linguistic priors. We introduce Positive-and-Negative Decoding (PND), a training-free inference framework that intervenes directly in the de

Cited by 0SourcecodeScholar
2026

Rectification Reimagined: A Unified Mamba Model for Image Correction and Rectangling with Prompts

AAAI 2026technical

Image correction and rectangling are valuable tasks in practical photography systems such as smartphones. Recent remarkable advancements in deep learning have undeniably brought about substantial performance improvements in these fields. Nevertheless, existing methods mainly rely on task-specific ar

Cited by 0SourcePDFScholar
2025

Beyond Spatial Domain: Cross-domain Promoted Fourier Convolution Helps Single Image Dehazing

AAAI 2025technical

Vanilla convolution and window-based self-attention have shown significant success in image dehazing. However, they are constrained by limited receptive fields and ignore frequency gaps between dehazed and clear images. The former hampers the modeling of global dependencies, while the latter impedes…

Cited by 0SourcePDFScholar
2025

Real-Time Scene-Adaptive Tone Mapping for High-Dynamic Range Object Detection

NeurIPS 2025poster

High dynamic range (HDR) images, with their rich tone and detail reproduction, hold significant potential to enhance computer vision systems, particularly in autonomous driving. However, most neural networks for embedded vision are trained on low dynamic range (LDR) inputs and suffer substantial per…

Cited by 0SourceScholar
2024

Prototypical Information Bottlenecking and Disentangling for Multimodal Cancer Survival Prediction

ICLR 2024spotlight

Multimodal learning significantly benefits cancer survival prediction, especially the integration of pathological images and genomic data. Despite advantages of multimodal learning for cancer survival prediction, massive redundancy in multimodal data prevents it from extracting discriminative and co…

2022

Multi-Frame Super-Resolution With Raw Images Via Modified Deformable Convolution

ICASSP 2022accepted

In this paper we propose a novel model towards multi-frame super-resolution, which leverages multiple RAW images and yields a super-resolved RGB image. To facilitate the pixel misalignment in burst photography, we apply a refined Pyramid Cascading and Deformable Convolution (PCD) feature alignment m…

Cited by 0SourceScholar