← Search

Shaohui Liu

33 accepted papers

2026

Bridging the Perceptual Gap: Residual-Enhanced Downscaling and Manifold-Aware Perception Alignment Adaptation for NR-IQA

ICML 2026poster

Leveraging Large Vision-Language Models like CLIP has recently set new benchmarks for No-Reference Image Quality Assessment (NR-IQA). However, the contrastive pretraining of CLIP inherently prioritizes semantic invariance, which often suppresses subtle perceptual signals, a phenomenon we term percep…

Cited by 0SourceScholar
2026

LLMInertia: Adaptive Counter-Inertial Reasoning to Improve Evidence Faithfulness in Large Language Models

ICML 2026poster

Large Language Models (LLMs) frequently generate output that contradicts explicit input evidence, limiting their reliability in real-world applications. We identify cognitive inertia in LLMs—a tendency to overly rely on co-occurrence associations learned during pretraining and to resist adaptation w…

Cited by 0SourceScholar
2026

PCLR: Progressively Compressed LoRA for Multimodal Continual Instruction Tuning

ICLR 2026poster

Continual Instruction Tuning (CIT) enables Large Multimodal Models (LMMs) to rapidly adapt to new tasks without retraining, but it suffers from the catastrophic forgetting problem. By adding new branches, model extension provides a great idea to accommodate novel knowledge while causing huge memory…

Cited by 0SourcecodeScholar
2026

Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation

ICML 2026poster

Accurate and efficient Video Quality Assessment (VQA) has long been a key research challenge. Current mainstream VQA methods typically improve performance by pretraining on large-scale classification datasets, followed by fine-tuning on VQA datasets. However, this strategy presents two significant c…

Cited by 0SourceScholar
2026

Unleashing Vision Transformer Potential in Image Quality Assessment via Global-Local Adaptive Interaction

ICASSP 2026poster

In the field of Blind Image Quality Assessment (BIQA), accurately predicting the perceptual quality of authentically distorted images remains highly challenging due to the diverse and complex distortions present in natural environments. Although existing methods have achieved notable accuracy, their…

Cited by 0SourcePDFScholar
2025

Benchmarking Egocentric Visual-Inertial SLAM at City Scale

ICCV 2025poster

Precise 6-DoF simultaneous localization and mapping (SLAM) from onboard sensors is critical for wearable devices capturing egocentric data, which exhibits specific challenges, such as a wider diversity of motions and viewpoints, prevalent dynamic visual content, or long sessions affected by time-var…

Cited by 0SourcePDFScholar
2025

Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMs

NeurIPS 2025poster

Large Language Models (LLMs) have emerged as powerful tools for diverse applications. However, their uniform token processing paradigm introduces critical vulnerabilities in instruction handling, particularly when exposed to adversarial scenarios. In this work, we identify and propose a novel class…

Cited by 0SourcecodeScholar
2025

Image Compressive Sensing With Adaptive Sampling by Median Filtering

ICASSP 2025accepted

Deep unfolding compressive sensing (CS) has experienced remarkable advancements. However, there still exist two challenges: (1) Many algorithms either use uniform block-based sampling, which ignore the fact that the content of different blocks is different, or allocate the sampling rate referring to…

Cited by 0SourceScholar
2025

MVQA: Mamba with Unified Sampling for Efficient Video Quality Assessment

ICCV 2025poster

The rapid growth of long-duration, high-definition videos has made efficient video quality assessment (VQA) a critical challenge. Existing research typically tackles this problem through two main strategies: reducing model parameters and resampling inputs. However, light-weight Convolution Neural Ne…

2025

PPTP: Performance-Guided Physiological Signal-Based Trust Prediction in Sequential Human-Robot Collaboration

RA-L 2025

Trust prediction is a key issue in human-robot collaboration, especially in construction scenarios where maintaining appropriate trust calibration is critical for safety and efficiency. This paper introduces the Performance-guided Physiological signal-based Trust Prediction (PPTP), a novel framework

Cited by 0SourceScholar
2025

Relative Pose Estimation through Affine Corrections of Monocular Depth Priors

CVPR 2025highlight

Monocular depth estimation (MDE) models have undergone significant advancements over recent years. Many MDE models aim to predict affine-invariant relative depth from monocular images, while recent developments in large-scale training and vision foundation models enable reasonable estimation of metr…

2025

Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization

CVPR 2025poster

Visual localization aims to determine the camera pose of a query image relative to a database of posed images. In recent years, deep neural networks that directly regress camera poses have gained popularity due to their fast inference capabilities. However, existing methods struggle to either genera…

2025

UAV-MaLO: Mamba-Augmented YOLO Hybrid Architecture for UAV Micro-Object Detection in Autonomous Robotics

IROS 2025

The rapid advancement of drone technology has led to the widespread application of micro-object detection in Unmanned Aerial Vehicle (UAV) systems. However, with the constraint of real-time computation, critical challenges remain in addressing extreme scale variations, low-resolution signatures and

Cited by 1SourceScholar
2024

AlphaTablets: A Generic Plane Representation for 3D Planar Reconstruction from Monocular Videos

NeurIPS 2024poster

We introduce AlphaTablets, a novel and generic representation of 3D planes that features continuous 3D surface and precise boundary delineation. By representing 3D planes as rectangles with alpha channels, AlphaTablets combine the advantages of current 2D and 3D plane representations, enabling accur…

Cited by 0SourcePDFScholar
2024

ZE-FESG: A Zero-Shot Feature Extraction Method Based on Semantic Guidance for No-Reference Video Quality Assessment

ICASSP 2024accepted

Although the current deep neural network based no-reference video quality assessment (NR-VQA) methods can effectively simulate the human visual system (HVS), their interpretability is getting worse. The current methods only extract the low-level features of space and time of the video and do not con…

Cited by 0SourceScholar
2023

Aprogressive Image Dehazing Framework with inter and Intra Contrastive Learning

ICASSP 2023accepted

Image dehazing, aims to estimate latent haze-free images from hazy images, suffering from a lot of lost information. Existing contrastive learning methods tend to utilize hazefree images as positive samples without consideration of negative samples. Even if negative samples are employed, the connect…

Cited by 0SourceScholar
2023

EI2SR: Learning an Enhanced Intra-Instance Semantic Relationship for Arbitrary-Shaped Scene Text Detection

ICASSP 2023accepted

Text detection in natural scenarios, has made significant progress with the deep learning architecture. Towards arbitrary-shaped text detection, fracture detection is the major concern due to the lack of semantic relationship within an instance in existing methods. To circumvent this dilemma, we pro…

Cited by 0SourceScholar
2023

LNPL-MIL: Learning from Noisy Pseudo Labels for Promoting Multiple Instance Learning in Whole Slide Image

ICCV 2023poster

Gigapixel Whole Slide Images (WSIs) aided patient diagnosis and prognosis analysis are promising directions in computational pathology. However, limited by expensive and time-consuming annotation costs, WSIs usually only have weak annotations, including 1) WSI-level Annotations (WA) and 2) Limited P…

Cited by 22PDFScholar
2023

Vanishing Point Estimation in Uncalibrated Images with Prior Gravity Direction

ICCV 2023poster

We tackle the problem of estimating a Manhattan frame, i.e. three orthogonal vanishing points, and the unknown focal length of the camera, leveraging a prior vertical direction. The direction can come from an Inertial Measurement Unit that is a standard component of recent consumer devices, e.g., sm…

Cited by 4PDFcodeScholar
2022

ParticleSfM: Exploiting Dense Point Trajectories for Localizing Moving Cameras in the Wild

ECCV 2022poster

"Estimating the pose of a moving camera from monocular video is a challenging problem, especially due to the presence of moving objects in dynamic environments, where the performance of existing camera pose estimation methods are susceptible to pixels that are not geometrically consistent. To tackle…

2021

A Confidence-Based Iterative Solver of Depths and Surface Normals for Deep Multi-View Stereo

ICCV 2021poster

In this paper, we introduce a deep multi-view stereo (MVS) system that jointly predicts depths, surface normals and per-view confidence maps. The key to our approach is a novel solver that iteratively solves for per-view depth map and normal map by optimizing an energy potential based upon the local…

Cited by 17PDFcodeScholar
2021

NerfingMVS: Guided Optimization of Neural Radiance Fields for Indoor Multi-View Stereo

ICCV 2021poster

In this work, we present a new multi-view depth estimation method that utilizes both conventional SfM reconstruction and learning-based priors over the recently proposed neural radiance fields (NeRF). Unlike existing neural network based optimization method that relies on estimated correspondences,…

Cited by 302PDFcodeScholar
2020

Classify and Explain: An Interpretable Convolutional Neural Network For Lung Cancer Diagnosis

ICASSP 2020accepted

The deep network-based computer-aided diagnosis systems have encountered many difficulties in practical applications because of its "black box" feature. The crux of the problem is that these models should be explainable - the model should provide doctors rationales that can explain the diagnosis. In…

Cited by 0SourceScholar
2020

DIST: Rendering Deep Implicit Signed Distance Function With Differentiable Sphere Tracing

CVPR 2020poster

We propose a differentiable sphere tracing algorithm to bridge the gap between inverse graphics methods and the recently proposed deep learning based implicit signed distance function. Due to the nature of the implicit function, the rendering process requires tremendous function queries, which is pa…

Cited by 350PDFcodeScholar
2020

Multi-Stage Residual Hiding for Image-Into-Audio Steganography

ICASSP 2020accepted

The widespread application of audio communication technologies has speeded up audio data flowing across the Internet, which made it a popular carrier for covert communication. In this paper, we present a cross-modal steganography method for hiding image content into audio carriers while preserving t…

Cited by 0SourceScholar
2020

Towards Better Generalization: Joint Depth-Pose Learning Without PoseNet

CVPR 2020poster

In this work, we tackle the essential problem of scale inconsistency for self supervised joint depth-pose learning. Most existing methods assume that a consistent scale of depth and pose can be learned across all input samples, which makes the learning problem harder, resulting in degraded performan…

Cited by 215PDFcodeScholar
2019

Image Inpainting With Learnable Bidirectional Attention Maps

ICCV 2019poster

Most convolutional network (CNN)-based inpainting methods adopt standard convolution to indistinguishably treat valid pixels and holes, making them limited in handling irregular holes and more likely to generate inpainting results with color discrepancy and blurriness. Partial convolution has been s…

Cited by 319PDFcodeScholar
2019

Scalable Convolutional Neural Network for Image Compressed Sensing

CVPR 2019poster

Recently, deep learning based image Compressed Sensing (CS) methods have been proposed and demonstrated superior reconstruction quality with low computational complexity. However, the existing deep learning based image CS methods need to train different models for different sampling ratios, which in…

Cited by 195PDFcodeScholar