← Search

Chunyi Li

23 accepted papers

2026

Can VLMs Diagnose and Recover from VLA Manipulation Faults?

ICML 2026poster

Existing VLA models frequently fail in robotic manipulation tasks, with poorly structured fault types that often require expert diagnosis.While VLMs offer strong explanatory capabilities, their effectiveness in assisting VLAs is limited by their unclear role in diagnostics and inadequate collaborati…

Cited by 0SourceScholar
2026

Exposing and Evaluating Hallucinations for GUI Grounding

CVPR 2026

Existing GUI benchmarks primarily focus on evaluating models' comprehensive capabilities but largely overlook hallucination phenomena in grounding tasks, which are crucial to the reliability of GUI understanding. In this work, we expose two major types of hallucinations in GUI grounding: 1) Confusio

Cited by 0SourceScholar
2026

GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal Models

AAAI 2026technical

Large multimodal models (LMMs) have demonstrated remarkable capabilities across a wide range of tasks, however their knowledge and abilities in the cross-view geo-localization and pose estimation domains remain unexplored, despite potential benefits for navigation, autonomous driving, outdoor roboti

Cited by 0SourcePDFScholar
2026

Image Quality Assessment for Embodied AI

ICLR 2026poster

Embodied AI has developed rapidly in recent years, but it is still mainly deployed in laboratories, with various distortions in the Real-world limiting its application. Traditionally, Image Quality Assessment (IQA) methods are applied to predict human preferences for distorted images; however, there…

Cited by 0SourcecodeScholar
2026

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond

ICML 2026poster

The success of large language models (LLMs) in scientific domains has heightened safety concerns, prompting numerous benchmarks to evaluate their scientific safety. Existing benchmarks often suffer from limited risk coverage and a reliance on subjective evaluation. To address thess problems, we intr…

Cited by 0SourceScholar
2026

Scaling-up Perceptual Video Quality Assessment

AAAI 2026technical

The data scaling law has significantly enhanced large multi-modal models (LMMs) performance across various downstream tasks. However, in the domain of perceptual video quality assessment (VQA), the potential of data scaling remains unprecedented due to the scarcity of labeled resources and the insuf

Cited by 0SourcePDFScholar
2025

A Multi-annotated and Multi-modal Dataset for Wide-angle Video Quality Assessment

ICASSP 2025accepted

Wide-angle video is favored for its wide viewing angle and ability to capture a large area of scenery, making it an ideal choice for sports and adventure recording. However, wide-angle video is prone to deformation, exposure and other distortions, resulting in poor video quality and affecting the pe…

Cited by 0SourceScholar
2025

A-Bench: Are LMMs Masters at Evaluating AI-generated Images?

ICLR 2025poster

How to accurately and efficiently assess AI-generated images (AIGIs) remains a critical challenge for generative models. Given the high costs and extensive time commitments required for user studies, many researchers have turned towards employing large multi-modal models (LMMs) as AIGI evaluators, t…

2025

Bidirectional Reference Image Quality Assessment via Content-Quality Correlation Modeling

ICASSP 2025accepted

The emphasis on no-reference image quality assessment has often overshadowed the significance of Full-Reference Image Quality Assessment (FR-IQA), which generally better reflects human contrastive perception mechanism. However, FRIQA presents challenges in obtaining content-aligned reference images.…

Cited by 0SourceScholar
2025

HazeCLIP: Towards Language Guided Real-World Image Dehazing

ICASSP 2025accepted

Existing methods have achieved remarkable performance in image dehazing, particularly on synthetic datasets. However, they often struggle with real-world hazy images due to domain shift, limiting their practical applicability. This paper introduces HazeCLIP, a language-guided adaptation framework de…

Cited by 0SourceScholar
2025

Image Quality Assessment: From Human to Machine Preference

CVPR 2025highlight

Image Quality Assessment (IQA) based on human subjective preferences has undergone extensive research in the past decades. However, with the development of communication protocols, the visual data consumption volume of machines has gradually surpassed that of humans. For machines, the preference dep…

2025

Information Density Principle for MLLM Benchmarks

ICCV 2025poster

With the emergence of Multimodal Large Language Models (MLLMs), hundreds of benchmarks have been developed to ensure the reliability of MLLMs in downstream tasks. However, the evaluation mechanism itself may not be reliable. For developers of MLLMs, questions remain about which benchmark to use and…

2025

Learning Hazing to Dehazing: Towards Realistic Haze Generation for Real-World Image Dehazing

CVPR 2025poster

Existing real-world image dehazing methods primarily attempt to fine-tune pre-trained models or adapt their inference procedures, thus heavily relying on the pre-trained models and associated training data. Moreover, restoring heavily distorted information under dense haze requires generative diffus…

2025

Q-Bench-Video: Benchmark the Video Quality Understanding of LMMs

CVPR 2025poster

With the rising interest in research on Large Multi-modal Models (LMMs) for video understanding, many studies have emphasized general video comprehension capabilities, neglecting the systematic exploration into video quality understanding. To address this oversight, we introduce Q-Bench-Video in thi…

2025

Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content

CVPR 2025poster

Evaluating text-to-vision content hinges on two crucial aspects: **visual quality** and **alignment**. While significant progress has been made in developing objective models to assess these dimensions, the performance of such models heavily relies on the scale and quality of human annotations. Acco…

2025

Redundancy Principles for MLLMs Benchmarks

ACL 2025long

With the rapid iteration of Multi-modality Large Language Models (MLLMs) and the evolving demands of the field, the number of benchmarks produced annually has surged into the hundreds. The rapid growth has inevitably led to significant redundancy among benchmarks. Therefore, it is crucial to take a…

Cited by 0SourcePDFScholar
2025

Towards All-in-One Medical Image Re-Identification

CVPR 2025poster

Medical image re-identification (MedReID) is under-explored so far, despite its critical applications in personalized healthcare and privacy protection.In this paper, we introduce a thorough benchmark and a unified model for this problem.First, to handle various medical modalities, we propose a nove…

2024

A Reduced-Reference Quality Assessment Metric for Textured Mesh Digital Humans

ICASSP 2024accepted

In an era where 3D Digital Humans (DHs) are becoming increasingly prevalent in fields like gaming, automotive, and the metaverse, the demand for high DH visual quality is rising. This paper presents the first-ever reduced-reference (RR) quality assessment metric tailored specifically for textured me…

Cited by 0SourceScholar
2024

Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

ICML 2024poster

The explosion of visual content available online underscores the requirement for an accurate machine assessor to robustly evaluate scores across diverse types of visual contents. While recent studies have demonstrated the exceptional potentials of large multi-modality models (LMMs) on a wide range o…

2024

Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

ICLR 2024spotlight

The rapid evolution of Multi-modality Large Language Models (MLLMs) has catalyzed a shift in computer vision from specialized models to general-purpose foundation models. Nevertheless, there is still an inadequacy in assessing the abilities of MLLMs on **low-level visual perception and understanding…

2024

Q-Instruct: Improving Low-level Visual Abilities for Multi-modality Foundation Models

CVPR 2024poster

Multi-modality large language models (MLLMs) as represented by GPT-4V have introduced a paradigm shift for visual perception and understanding tasks that a variety of abilities can be achieved within one foundation model. While current MLLMs demonstrate primary low-level visual abilities from the id…

2024

Towards Open-ended Visual Quality Comparison

ECCV 2024oral

"Comparative settings (pairwise choice, listwise ranking) have been adopted by a wide range of subjective studies for image quality assessment (IQA), as it inherently standardizes the evaluation criteria across different observers and offer more clear-cut responses. In this work, we extend the edge…