← Search

Huiyu Duan

17 accepted papers

2026

Agentic Retoucher for Text-To-Image Generation

CVPR 2026

Text-to-image (T2I) diffusion models such as SDXL and FLUX have achieved impressive photorealism, yet small-scale distortions remain pervasive in limbs, face, text and so on. Existing refinement approaches either perform costly iterative re-generation or rely on vision-language models (VLMs) with we

Cited by 3SourcecodeScholar
2026

EEmo-Logic: A Unified Dataset and Multi-Stage Framework for Comprehensive Image-Evoked Emotion Assessment

ICML 2026spotlight

Understanding the multi-dimensional attributes and intensity nuances of image-evoked emotions is pivotal for advancing machine empathy and empowering diverse human-computer interaction applications. However, existing models are still limited to coarse-grained emotion perception or deficient reasonin…

Cited by 0SourceScholar
2026

FVBench: Benchmarking Deepfake Video Detection Capability of Large Multimodal Models

CVPR 2026

As generative models rapidly evolve, the realism of AI-generated videos has reached new levels, posing significant challenges for detecting the authenticity of videos. Existing deepfake detection techniques generally rely on training datasets with limited generation methods and content diversity, wh

Cited by 0SourcecodeScholar
2026

GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal Models

AAAI 2026technical

Large multimodal models (LMMs) have demonstrated remarkable capabilities across a wide range of tasks, however their knowledge and abilities in the cross-view geo-localization and pose estimation domains remain unexplored, despite potential benefits for navigation, autonomous driving, outdoor roboti

Cited by 0SourcePDFScholar
2026

I2I-Bench: A Comprehensive Benchmark Suite for Image-to-Image Editing Models

CVPR 2026

Image editing models are advancing rapidly, yet comprehensive evaluation remains a significant challenge. Existing image editing benchmarks generally suffer from limited task scopes, insufficient evaluation dimensions, and heavy reliance on manual annotations, which significantly constrain their sca

Cited by 0SourceScholar
2026

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation

ICML 2026poster

Recent advancements in large multimodal models (LMMs) have driven substantial progress in both text-to-video (T2V) generation and video-to-text (V2T) interpretation tasks. However, current AI-generated videos (AIGVs) still exhibit limitations in terms of perceptual quality and text-video alignment. …

Cited by 0SourcecodeScholar
2026

Life-IQA: Boosting Blind Image Quality Assessment through GCN-enhanced Layer Interaction and MoE-based Feature Decoupling

CVPR 2026

Blind image quality assessment (BIQA) plays a crucial role in evaluating and optimizing visual experience. Most existing BIQA approaches fuse shallow and deep features extracted from backbone networks, while overlooking the unequal contributions to quality prediction. Moreover, while various vision

Cited by 0SourceScholar
2026

MIMIC-Bench: Exploring the User-Like Thinking and Mimicking Capabilities of Multimodal Large Language Models

ICLR 2026poster

The rapid advancement of multimodal large language models (MLLMs) has greatly prompted the video interpretation task, and numerous works have been proposed to explore and benchmark the cognition and basic visual reasoning capabilities of MLLMs. However, practical applications on social media platfo…

Cited by 0SourcecodeScholar
2026

ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?

ICLR 2026poster

Omnidirectional images (ODIs) provide full 360$^{\circ} \times$ 180$^{\circ}$ view which are widely adopted in VR, AR and embodied intelligence applications. While multi-modal large language models (MLLMs) have demonstrated remarkable performance on conventional 2D image and video understanding benc…

Cited by 0SourceScholar
2026

Steering and Rectifying Latent representation manifolds in Frozen Multi-modal LLMs for Video Anomaly Detection

ICLR 2026poster

Video anomaly detection (VAD) aims to identify abnormal events in videos. Traditional VAD methods generally suffer from the high costs of labeled data and full training, thus some recent works have explored leveraging frozen multi-modal large language models (MLLMs) in a tuning-free manner to perfor…

Cited by 0SourceScholar
2025

AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM

CVPR 2025poster

The rapid advancement of large multimodal models (LMMs) has led to the rapid expansion of artificial intelligence generated videos (AIGVs), which highlights the pressing need for effective video quality assessment (VQA) models designed specifically for AIGVs. Current VQA models generally fall short…

2025

F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and Restoration

ICCV 2025poster

Recent artificial intelligence (AI) generative models have demonstrated remarkable capabilities in image production, and have been widely applied to face image generation, customization, and restoration. However, many AI-generated faces (AIGFs) still suffer from issues such as unique distortions, un…

Cited by 0SourcePDFScholar
2025

FPEM: Face Prior Enhanced Facial Attractiveness Prediction for Live Videos with Face Retouching

ICCV 2025poster

Facial attractiveness prediction (FAP) has long been an important computer vision task, which could be widely applied in live videos with facial retouching. However, previous FAP datasets are either small or closed-source. Moreover, the corresponding FAP models exhibit limited generalization and ada…

2025

FineVQ: Fine-Grained User Generated Content Video Quality Assessment

CVPR 2025highlight

The rapid growth of user-generated content (UGC) videos has produced an urgent need for effective video quality assessment (VQA) algorithms to monitor video quality and guide optimization and recommendation procedures. However, current VQA models generally only give an overall rating for a UGC video…

2022

End-to-End Human-Gaze-Target Detection With Transformers

CVPR 2022poster

In this paper, we propose an effective and efficient method for Human-Gaze-Target (HGT) detection, i.e., gaze following. Current approaches decouple the HGT detection task into separate branches of salient object detection and human gaze prediction, employing a two-stage framework where human head l…

Cited by 65PDFScholar
2022

Iwin: Human-Object Interaction Detection via Transformer with Irregular Windows

ECCV 2022poster

"This paper presents a new vision Transformer, named Iwin Transformer, which is specifically designed for human-object interaction (HOI) detection, a detailed scene understanding task involving a sequential process of human/object detection and interaction recognition. Iwin Transformer is a hierarch…

Cited by 28SourcePDFScholar