← Search

Zhiyuan You

11 accepted papers

2026

PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models

ICML 2026oral

With the recent fast development of generative models, instruction-based image editing has shown great potential in generating high-quality images. However, the quality of editing highly depends on carefully designed instructions, placing the burden of task decomposition and sequencing entirely on t…

Cited by 0SourceScholar
2025

An Intelligent Agentic System for Complex Image Restoration Problems

ICLR 2025poster

Real-world image restoration (IR) is inherently complex and often requires combining multiple specialized models to address diverse degradations. Inspired by human problem-solving, we propose AgenticIR, an agentic system that mimics the human approach to image processing by following five key stages…

2025

RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts

NeurIPS 2025poster

Quality analysis of weather forecasts is an essential topic in meteorology. Although traditional score-based evaluation metrics can quantify certain forecast errors, they are still far from meteorological experts in terms of descriptive capability, interpretability, and understanding of dynamic evol…

Cited by 0SourceScholar
2025

ReinAD: Towards Real-world Industrial Anomaly Detection with a Comprehensive Contrastive Dataset

NeurIPS 2025poster

Recent years have witnessed significant advancements in industrial anomaly detection (IAD) thanks to existing anomaly detection datasets. However, the large performance gap between these benchmarks and real industrial practice reveals critical limitations in existing datasets. We argue that the mism…

Cited by 0SourcecodeScholar
2025

SAIL: Sample-Centric In-Context Learning for Document Information Extraction

AAAI 2025technical

Document Information Extraction (DIE) aims to extract structured information from Visually Rich Documents (VRDs). Previous full-training approaches have demonstrated strong performance but may struggle with generalization to unseen data. In contrast, training-free methods leverage powerful pre-train…

2025

Teaching Large Language Models to Regress Accurate Image Quality Scores Using Score Distribution

CVPR 2025poster

With the rapid advancement of Multi-modal Large Language Models (MLLMs), MLLM-based Image Quality Assessment (IQA) methods have shown promising performance in linguistic quality description. However, current methods still fall short in accurately scoring image quality. In this work, we aim to levera…

2025

UltraFusion: Ultra High Dynamic Imaging using Exposure Fusion

CVPR 2025highlight

Capturing high dynamic range (HDR) scenes is one of the most important issues in camera design. Majority of cameras use exposure fusion, which fuses images captured by different exposure levels, to increase dynamic range. However, this approach can only handle images with limited exposure difference…

2024

PhoCoLens: Photorealistic and Consistent Reconstruction in Lensless Imaging

NeurIPS 2024spotlight

Lensless cameras offer significant advantages in size, weight, and cost compared to traditional lens-based systems. Without a focusing lens, lensless cameras rely on computational algorithms to recover the scenes from multiplexed measurements. However, current algorithms struggle with inaccurate for…

Cited by 7SourcePDFScholar
2022

A Unified Model for Multi-class Anomaly Detection

NeurIPS 2022accept

Despite the rapid advance of unsupervised anomaly detection, existing methods require to train separate models for different objects. In this work, we present UniAD that accomplishes anomaly detection for multiple classes with a unified framework. Under such a challenging setting, popular reconstruc…