← Search

Yifei Fan

7 accepted papers

2026

Beyond Single Images: A Comprehensive Benchmark for Album-Level Vision-Language Understanding

CVPR 2026

Automatic album organization has been studied extensively over the past decades due to significant progress in digital photography. Recent vision-language models (VLMs) have shown strong performance on multi-image understanding, making them natural candidates for automating album organization workfl

Cited by 0SourcecodeScholar
2025

DiffTell: A High-Quality Dataset for Describing Image Manipulation Changes

ICCV 2025poster

The image difference captioning (IDC) task is to describe the distinctions between two images. However, existing datasets do not offer comprehensive coverage across all image-difference categories. In this work, we introduce a high-quality dataset, DiffTell with various types of image manipulations,…

Cited by 0SourcePDFScholar
2025

The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographers

CVPR 2025poster

Photographer, curator, and former director of photography at the Museum of Modern Art (MoMA), John Szarkowski remarked in *William Eggleston's Guide*, "While editing directly from life, photographers have found it too difficult to see simultaneously both the blue and the sky." Szarkowski insightfull…

Cited by 0SourcePDFScholar
2024

SPIN: Hierarchical Segmentation with Subpart Granularity in Natural Images

ECCV 2024poster

"Hierarchical segmentation entails creating segmentations at varying levels of granularity. We introduce the first hierarchical semantic segmentation dataset with subpart annotations for natural images, which we call SPIN (SubPartImageNet). We also introduce two novel evaluation metrics to evaluate…

Cited by 2SourcePDFScholar
2024

Uncertainty-aware Fine-tuning of Segmentation Foundation Models

NeurIPS 2024poster

The Segment Anything Model (SAM) is a large-scale foundation model that has revolutionized segmentation methodology. Despite its impressive generalization ability, the segmentation accuracy of SAM on images with intricate structures is often unsatisfactory. Recent works have proposed lightweight fin…

2024

VIXEN: Visual Text Comparison Network for Image Difference Captioning

AAAI 2024technical

We present VIXEN - a technique that succinctly summarizes in text the visual differences between a pair of images in order to highlight any content manipulation present. Our proposed network linearly maps image features in a pairwise manner, constructing a soft prompt for a pretrained large language…