← Search

Xiongkuo Min

49 accepted papers

2026

Adapter Shield: A Unified Framework with Built-in Authentication for Preventing Unauthorized Zero-Shot Image-to-Image Generation

CVPR 2026

With the rapid progress in diffusion models, image synthesis has advanced to the stage of zero-shot image-to-image generation, where high-fidelity replication of facial identities or artistic styles can be achieved using just one portrait or artwork, without modifying any model weights. Although the

Cited by 0SourceScholar
2026

Audio-Assisted Face Video Restoration with Temporal and Identity Complementary Learning

AAAI 2026technical

Face videos accompanied by audio have become integral to our daily lives, while they often suffer from complex degradations. Most face video restoration methods neglect the intrinsic correlations between visual and audio features, particularly in the mouth region. Several audio-aided face video rest

Cited by 0SourcePDFScholar
2026

EEmo-Logic: A Unified Dataset and Multi-Stage Framework for Comprehensive Image-Evoked Emotion Assessment

ICML 2026spotlight

Understanding the multi-dimensional attributes and intensity nuances of image-evoked emotions is pivotal for advancing machine empathy and empowering diverse human-computer interaction applications. However, existing models are still limited to coarse-grained emotion perception or deficient reasonin…

Cited by 0SourceScholar
2026

FVBench: Benchmarking Deepfake Video Detection Capability of Large Multimodal Models

CVPR 2026

As generative models rapidly evolve, the realism of AI-generated videos has reached new levels, posing significant challenges for detecting the authenticity of videos. Existing deepfake detection techniques generally rely on training datasets with limited generation methods and content diversity, wh

Cited by 0SourcecodeScholar
2026

Generalizable Video Quality Assessment via Weak-to-Strong Learning

CVPR 2026

Video quality assessment (VQA) seeks to predict the perceptual quality of a video in alignment with human visual perception, serving as a fundamental tool for quantifying quality degradation across video processing workflows. The dominant VQA paradigm relies on supervised training with human-labeled

Cited by 0SourcecodeScholar
2026

Grounding-IQA: Grounding Multimodal Language Model for Image Quality Assessment

ICLR 2026poster

The development of multimodal large language models (MLLMs) enables the evaluation of image quality through natural language descriptions. This advancement allows for more detailed assessments. However, these MLLM-based IQA methods primarily rely on general contextual descriptions, sometimes limitin…

Cited by 0SourcecodeScholar
2026

I2I-Bench: A Comprehensive Benchmark Suite for Image-to-Image Editing Models

CVPR 2026

Image editing models are advancing rapidly, yet comprehensive evaluation remains a significant challenge. Existing image editing benchmarks generally suffer from limited task scopes, insufficient evaluation dimensions, and heavy reliance on manual annotations, which significantly constrain their sca

Cited by 5SourceScholar
2026

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation

ICML 2026poster

Recent advancements in large multimodal models (LMMs) have driven substantial progress in both text-to-video (T2V) generation and video-to-text (V2T) interpretation tasks. However, current AI-generated videos (AIGVs) still exhibit limitations in terms of perceptual quality and text-video alignment. …

Cited by 0SourcecodeScholar
2026

LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks

CVPR 2026

The rapid progress of Multimodal Large Language Models (MLLMs) marks a significant step toward artificial general intelligence, offering great potential for augmenting human capabilities. However, their ability to provide effective assistance in dynamic, real-world environments remains largely under

Cited by 0SourceScholar
2026

MIMIC-Bench: Exploring the User-Like Thinking and Mimicking Capabilities of Multimodal Large Language Models

ICLR 2026poster

The rapid advancement of multimodal large language models (MLLMs) has greatly prompted the video interpretation task, and numerous works have been proposed to explore and benchmark the cognition and basic visual reasoning capabilities of MLLMs. However, practical applications on social media platfo…

Cited by 0SourcecodeScholar
2026

ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?

ICLR 2026poster

Omnidirectional images (ODIs) provide full 360$^{\circ} \times$ 180$^{\circ}$ view which are widely adopted in VR, AR and embodied intelligence applications. While multi-modal large language models (MLLMs) have demonstrated remarkable performance on conventional 2D image and video understanding benc…

Cited by 0SourceScholar
2026

Refine-IQA: Multi-Stage Reinforcement Finetuning for Perceptual Image Quality Assessment

AAAI 2026technical

Reinforcement fine-tuning (RFT) is a proliferating paradigm for LMM training. Analogous to high-level reasoning tasks, RFT is similarly applicable to low-level vision domains, including image quality assessment (IQA). Existing RFT-based IQA methods typically use rule-based output rewards to verify

Cited by 0SourcePDFScholar
2026

SalDiff-DTM: A Novel Dual-Temporal Modulated Diffusion Model for Omnidirectional Images Scanpath Prediction

AAAI 2026technical

Scanpath prediction in omnidirectional images (ODIs) serves as a critical component for optimizing foveated rendering efficiency and enhancing interactive quality in virtual reality systems. However, existing scanpath prediction methods for ODIs still suffer from fundamental limitations: (1) inadequ

Cited by 0SourcePDFScholar
2026

Scaling-up Perceptual Video Quality Assessment

AAAI 2026technical

The data scaling law has significantly enhanced large multi-modal models (LMMs) performance across various downstream tasks. However, in the domain of perceptual video quality assessment (VQA), the potential of data scaling remains unprecedented due to the scarcity of labeled resources and the insuf

Cited by 0SourcePDFScholar
2026

VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessment

CVPR 2026

Developing a robust visual quality assessment (VQualA) large multi-modal model (LMM) requires achieving versatility, powerfulness, and transferability. However, existing VQualA LMMs typically focus on a single task and rely on full-parameter fine-tuning, which makes them prone to overfitting on spec

Cited by 0SourcecodeScholar
2026

VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement Learning

AAAI 2026technical

Video quality assessment (VQA) aims to objectively quantify perceptual quality degradation in alignment with human visual perception. Despite recent advances, existing VQA models still suffer from two critical limitations: poor generalization to out-of-distribution (OOD) videos and limited explainab

Cited by 0SourcePDFScholar
2025

3DGCQA: A Quality Assessment Database for 3D AI-Generated Contents

ICASSP 2025accepted

Although 3D generated content (3DGC) offers advantages in reducing production costs and accelerating design timelines, its quality often falls short when compared to 3D professionally generated content. Common quality issues frequently affect 3DGC, highlighting the importance of timely and effective…

Cited by 0SourceScholar
2025

A-Bench: Are LMMs Masters at Evaluating AI-generated Images?

ICLR 2025poster

How to accurately and efficiently assess AI-generated images (AIGIs) remains a critical challenge for generative models. Given the high costs and extensive time commitments required for user studies, many researchers have turned towards employing large multi-modal models (LMMs) as AIGI evaluators, t…

2025

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment

ICML 2025poster

Many video-to-audio (VTA) methods have been proposed for dubbing silent AI-generated videos. An efficient quality assessment method for AI-generated audio-visual content (AGAV) is crucial for ensuring audio-visual quality. Existing audio-visual quality assessment methods struggle with unique distort…

2025

AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM

CVPR 2025poster

The rapid advancement of large multimodal models (LMMs) has led to the rapid expansion of artificial intelligence generated videos (AIGVs), which highlights the pressing need for effective video quality assessment (VQA) models designed specifically for AIGVs. Current VQA models generally fall short…

2025

Explore the Hallucination on Low-level Perception for MLLMs

ICASSP 2025accepted

The rapid development of Multi-modality Large Language Models (MLLMs) has significantly influenced various aspects of industry and daily life, showcasing impressive capabilities in visual perception and understanding. However, these models also exhibit hallucinations, which limit their reliability a…

Cited by 0SourceScholar
2025

FPEM: Face Prior Enhanced Facial Attractiveness Prediction for Live Videos with Face Retouching

ICCV 2025poster

Facial attractiveness prediction (FAP) has long been an important computer vision task, which could be widely applied in live videos with facial retouching. However, previous FAP datasets are either small or closed-source. Moreover, the corresponding FAP models exhibit limited generalization and ada…

2025

FineVQ: Fine-Grained User Generated Content Video Quality Assessment

CVPR 2025highlight

The rapid growth of user-generated content (UGC) videos has produced an urgent need for effective video quality assessment (VQA) algorithms to monitor video quality and guide optimization and recommendation procedures. However, current VQA models generally only give an overall rating for a UGC video…

2025

HazeCLIP: Towards Language Guided Real-World Image Dehazing

ICASSP 2025accepted

Existing methods have achieved remarkable performance in image dehazing, particularly on synthetic datasets. However, they often struggle with real-world hazy images due to domain shift, limiting their practical applicability. This paper introduces HazeCLIP, a language-guided adaptation framework de…

Cited by 0SourceScholar
2025

Image Quality Assessment: From Human to Machine Preference

CVPR 2025highlight

Image Quality Assessment (IQA) based on human subjective preferences has undergone extensive research in the past decades. However, with the development of communication protocols, the visual data consumption volume of machines has gradually surpassed that of humans. For machines, the preference dep…

2025

Information Density Principle for MLLM Benchmarks

ICCV 2025poster

With the emergence of Multimodal Large Language Models (MLLMs), hundreds of benchmarks have been developed to ensure the reliability of MLLMs in downstream tasks. However, the evaluation mechanism itself may not be reliable. For developers of MLLMs, questions remain about which benchmark to use and…

2025

Mesh Mamba: A Unified State Space Model for Saliency Prediction in Non-Textured and Textured Meshes

CVPR 2025poster

Mesh saliency enhances the adaptability of 3D vision by identifying and emphasizing regions that naturally attract visual attention. To investigate the interaction between geometric structure and texture in shaping visual attention, we establish a comprehensive mesh saliency dataset, which is the fi…

2025

Q-Bench-Video: Benchmark the Video Quality Understanding of LMMs

CVPR 2025poster

With the rising interest in research on Large Multi-modal Models (LMMs) for video understanding, many studies have emphasized general video comprehension capabilities, neglecting the systematic exploration into video quality understanding. To address this oversight, we introduce Q-Bench-Video in thi…

2025

Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content

CVPR 2025poster

Evaluating text-to-vision content hinges on two crucial aspects: **visual quality** and **alignment**. While significant progress has been made in developing objective models to assess these dimensions, the performance of such models heavily relies on the scale and quality of human annotations. Acco…

2025

Redundancy Principles for MLLMs Benchmarks

ACL 2025long

With the rapid iteration of Multi-modality Large Language Models (MLLMs) and the evolving demands of the field, the number of benchmarks produced annually has surged into the hundreds. The rapid growth has inevitably led to significant redundancy among benchmarks. Therefore, it is crucial to take a…

Cited by 0SourcePDFScholar
2025

Textured Mesh Saliency: Bridging Geometry and Texture for Human Perception in 3D Graphics

AAAI 2025technical

Textured meshes significantly enhance the realism and detail of objects by mapping intricate texture details onto the geometric structure of 3D models. This advancement is valuable across various applications, including entertainment, education, and industry. While traditional mesh saliency studies…

2025

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads

ICCV 2025poster

Speech-driven methods for portraits are figuratively known as "Talkers" because of their capability to synthesize speaking mouth shapes and facial movements. Especially with the rapid development of the Text-to-Image (T2I) models, AI-Generated Talking Heads (AGTHs) have gradually become an emerging…

2024

A Reduced-Reference Quality Assessment Metric for Textured Mesh Digital Humans

ICASSP 2024accepted

In an era where 3D Digital Humans (DHs) are becoming increasingly prevalent in fields like gaming, automotive, and the metaverse, the demand for high DH visual quality is rising. This paper presents the first-ever reduced-reference (RR) quality assessment metric tailored specifically for textured me…

Cited by 0SourceScholar
2024

GAIA: Rethinking Action Quality Assessment for AI-Generated Videos

NeurIPS 2024spotlight

Assessing action quality is both imperative and challenging due to its significant impact on the quality of AI-generated videos, further complicated by the inherently ambiguous nature of actions within AI-generated video (AIGV). Current action quality assessment (AQA) algorithms predominantly focus…

2024

GLARE: Low Light Image Enhancement via Generative Latent Feature based Codebook Retrieval

ECCV 2024poster

"Most existing Low-light Image Enhancement (LLIE) methods either directly map Low-Light (LL) to Normal-Light (NL) images or use semantic or illumination maps as guides. However, the ill-posed nature of LLIE and the difficulty of semantic retrieval from impaired inputs limit these methods, especially…

2024

Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

ICML 2024poster

The explosion of visual content available online underscores the requirement for an accurate machine assessor to robustly evaluate scores across diverse types of visual contents. While recent studies have demonstrated the exceptional potentials of large multi-modality models (LMMs) on a wide range o…

2024

UniProcessor: A Text-induced Unified Low-level Image Processor

ECCV 2024poster

"Image processing, including image restoration, image enhancement, etc., involves generating a high-quality clean image from a degraded input. Deep learning-based methods have shown superior performance for various image processing tasks in terms of single-task conditions. However, they require to t…

2023

MD-VQA: Multi-Dimensional Quality Assessment for UGC Live Videos

CVPR 2023poster

User-generated content (UGC) live videos are often bothered by various distortions during capture procedures and thus exhibit diverse visual qualities. Such source videos are further compressed and transcoded by media server providers before being distributed to end-users. Because of the flourishing…

2023

MM-PCQA: Multi-Modal Learning for No-reference Point Cloud Quality Assessment

IJCAI 2023poster

The visual quality of point clouds has been greatly emphasized since the ever-increasing 3D vision applications are expected to provide cost-effective and high-quality experiences for users. Looking back on the development of point cloud quality assessment (PCQA), the visual quality is usually eval…

2023

Perceptual Quality Assessment for Digital Human Heads

ICASSP 2023accepted

Digital humans are attracting more and more research interest during the last decade, the generation, representation, rendering, and animation of which have been put into large amounts of effort. However, the quality assessment of digital humans has fallen behind. Therefore, to tackle the challenge…

Cited by 0SourceScholar
2022

End-to-End Human-Gaze-Target Detection With Transformers

CVPR 2022poster

In this paper, we propose an effective and efficient method for Human-Gaze-Target (HGT) detection, i.e., gaze following. Current approaches decouple the HGT detection task into separate branches of salient object detection and human gaze prediction, employing a two-stage framework where human head l…

Cited by 65PDFScholar
2022

Iwin: Human-Object Interaction Detection via Transformer with Irregular Windows

ECCV 2022poster

"This paper presents a new vision Transformer, named Iwin Transformer, which is specifically designed for human-object interaction (HOI) detection, a detailed scene understanding task involving a sequential process of human/object detection and interaction recognition. Iwin Transformer is a hierarch…

Cited by 28SourcePDFScholar
2022

Learning Invisible Markers for Hidden Codes in Offline-to-Online Photography

CVPR 2022poster

QR (quick response) codes are widely used as an offline-to-online channel to convey information (e.g., links) from publicity materials (e.g., display and print) to mobile devices. However, QR Codes are not favorable for taking up valuable space of publicity materials. Recent works propose invisible…

Cited by 37PDFScholar
2022

Perceptual Attacks of No-Reference Image Quality Models with Human-in-the-Loop

NeurIPS 2022accept

No-reference image quality assessment (NR-IQA) aims to quantify how humans perceive visual distortions of digital images without access to their undistorted references. NR-IQA models are extensively studied in computational vision, and are widely used for performance evaluation and perceptual optimi…

2022

Video-based Human-Object Interaction Detection from Tubelet Tokens

NeurIPS 2022accept

We present a novel vision Transformer, named TUTOR, which is able to learn tubelet tokens, served as highly-abstracted spatial-temporal representations, for video-based human-object interaction (V-HOI) detection. The tubelet tokens structurize videos by agglomerating and linking semantically-related…

Cited by 18SourcePDFScholar
2021

Perceptual Quality Assessment for Recognizing True and Pseudo 4k Content

ICASSP 2021accepted

To meet the imperative demand for monitoring the quality of Ultra High-Definition (UHD) content in multimedia industries, we propose an efficient no-reference (NR) image quality assessment (IQA) metric to distinguish original and pseudo 4K contents and measure the quality of their quality in this pa…

Cited by 0SourceScholar
2021

Self-Conditioned Probabilistic Learning of Video Rescaling

ICCV 2021poster

Bicubic downscaling is a prevalent technique used to reduce the video storage burden or to accelerate the downstream processing speed. However, the inverse upscaling step is non-trivial, and the downscaled video may also deteriorate the performance of downstream tasks. In this paper, we propose a se…

Cited by 22PDFcodeScholar