← Search

Liang Liao

20 accepted papers

2026

CHARM: Collaborative Harmonization Across Arbitrary Modalities for Modality-Agnostic Semantic Segmentation

AAAI 2026technical

Modality-agnostic Semantic Segmentation (MaSS) aims to achieve robust scene understanding across arbitrary combinations of input modality. Existing methods typically rely on explicit feature alignment to achieve modal homogenization, which dilutes the distinctive strengths of each modality and destr

Cited by 0SourcePDFScholar
2025

MDFG: Multi-Dimensional Fine-Grained Modeling for Fatigue Detection

AAAI 2025technical

Fatigue is a critical factor contributing to accidents in industries such as safety monitoring and engineering construction. Fatigue exhibits dynamic complexity and non-stationary characteristics, so there are many intermediate states of short-term variation between alert and fatigue. Capturing and…

2024

Boosting Image Quality Assessment through Efficient Transformer Adaptation with Local Feature Enhancement

CVPR 2024poster

Image Quality Assessment (IQA) constitutes a fundamental task within the field of computer vision yet it remains an unresolved challenge owing to the intricate distortion conditions diverse image contents and limited availability of data. Recently the community has witnessed the emergence of numerou…

2024

Enhancing Diffusion Models with Text-Encoder Reinforcement Learning

ECCV 2024poster

"Text-to-image diffusion models are typically trained to optimize the log-likelihood objective, which presents challenges in meeting specific requirements for downstream tasks, such as image aesthetics and image-text alignment. Recent research addresses this issue by refining the diffusion U-Net usi…

2024

Hidden Follower Detection: How Is the Gaze-Spacing Pattern Embodied in Frequency Domain?

AAAI 2024technical

Spatiotemporal social behavior analysis is a technique that studies the social behavior patterns of objects and estimates their risks based on their trajectories. In social public scenarios such as train stations, hidden following behavior has become one of the most challenging issues due to its pro…

Cited by 1SourcePDFScholar
2024

Iterative Token Evaluation and Refinement for Real-World Super-resolution

AAAI 2024technical

Real-world image super-resolution (RWSR) is a long-standing problem as low-quality (LQ) images often have complex and unidentified degradations. Existing methods such as Generative Adversarial Networks (GANs) or continuous diffusion models present their own issues including GANs being difficult to t…

2024

Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

ICML 2024poster

The explosion of visual content available online underscores the requirement for an accurate machine assessor to robustly evaluate scores across diverse types of visual contents. While recent studies have demonstrated the exceptional potentials of large multi-modality models (LMMs) on a wide range o…

2024

Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

ICLR 2024spotlight

The rapid evolution of Multi-modality Large Language Models (MLLMs) has catalyzed a shift in computer vision from specialized models to general-purpose foundation models. Nevertheless, there is still an inadequacy in assessing the abilities of MLLMs on **low-level visual perception and understanding…

2024

Q-Instruct: Improving Low-level Visual Abilities for Multi-modality Foundation Models

CVPR 2024poster

Multi-modality large language models (MLLMs) as represented by GPT-4V have introduced a paradigm shift for visual perception and understanding tasks that a variety of abilities can be achieved within one foundation model. While current MLLMs demonstrate primary low-level visual abilities from the id…

2024

Towards Open-ended Visual Quality Comparison

ECCV 2024oral

"Comparative settings (pairwise choice, listwise ranking) have been adopted by a wide range of subjective studies for image quality assessment (IQA), as it inherently standardizes the evaluation criteria across different observers and offer more clear-cut responses. In this work, we extend the edge…

2023

Bat: Bi-Alignment Based On Transformation in Multi-Target Domain Adaptation for Semantic Segmentation

ICASSP 2023accepted

While enlightening progress has been made recently in single-target domain adaptive semantic segmentation (ST-DASS), the multi-peak distributed multi-target domain cannot be directly aligned well with the single-peak distributed source domain. As a result, it is impossible for existing methods to ha…

Cited by 0SourceScholar
2023

Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives

ICCV 2023poster

The rapid increase in user-generated-content (UGC) videos calls for the development of effective video quality assessment (VQA) algorithms. However, the objective of the UGC-VQA problem is still ambiguous and can be viewed from two perspectives: the technical perspective, measuring the perception of…

Cited by 162PDFcodeScholar
2023

GCFAgg: Global and Cross-View Feature Aggregation for Multi-View Clustering

CVPR 2023poster

Multi-view clustering can partition data samples into their categories by learning a consensus representation in unsupervised way and has received more and more attention in recent years. However, most existing deep clustering methods learn consensus representation or view-specific representations f…

2023

Only a Few Classes Confusing: Pixel-Wise Candidate Labels Disambiguation for Foggy Scene Understanding

AAAI 2023technical

Not all semantics become confusing when deploying a semantic segmentation model for real-world scene understanding of adverse weather. The true semantics of most pixels have a high likelihood of appearing in the few top classes according to confidence ranking. In this paper, we replace the one-hot p…

Cited by 9SourcePDFScholar
2022

FAST-VQA: Efficient End-to-End Video Quality Assessment with Fragment Sampling

ECCV 2022poster

"Current deep video quality assessment (VQA) methods are usually with high computational costs when evaluating high-resolution videos. This cost hinders them from learning better video-quality-related representations via end-to-end training. Existing approaches typically consider naive sampling to r…

2022

Spatial-Temporal Space Hand-in-Hand: Spatial-Temporal Video Super-Resolution via Cycle-Projected Mutual Learning

CVPR 2022poster

Spatial-Temporal Video Super-Resolution (ST-VSR) aims to generate super-resolved videos with higher resolution (HR) and higher frame rate (HFR). Quite intuitively, pioneering two-stage based methods complete ST-VSR directly combining two sub-tasks: Spatial Video Super-Resolution (S-VSR) and Temporal…

Cited by 44PDFcodeScholar
2021

Image Inpainting Guided by Coherence Priors of Semantics and Textures

CVPR 2021poster

Existing inpainting methods have achieved promising performance in recovering defected images of specific scenes. However, filling holes involving multiple semantic categories remains challenging due to the obscure semantic boundaries and the mixture of different semantic textures. In this paper, we…

Cited by 120PDFScholar
2020

Guidance and Evaluation: Semantic-Aware Image Inpainting for Mixed Scenes

ECCV 2020poster

Completing a corrupted image with correct structures and reasonable textures for a mixed scene remains an elusive challenge. Since the missing hole in a mixed scene of a corrupted image often contains various semantic information, conventional two-stage approaches utilizing structural information of…

Cited by 154SourcePDFScholar