← Search

Mai Xu

38 accepted papers

2026

Benchmarking and Enhancing VLM for Compressed Image Understanding

ICML 2026poster

With the rapid development of Vision-Language Models (VLMs) and the growing demand for their applications, efficient compression of the image inputs has become increasingly important. Existing VLMs predominantly digest and understand high-bitrate compressed images, while their ability to interpret l…

Cited by 0SourceScholar
2026

Burst Image Quality Assessment: A New Benchmark and Unified Framework for Multiple Downstream Tasks

AAAI 2026technical

In recent years, the development of burst imaging technology has improved the capture and processing capabilities of visual data, enabling a wide range of applications. However, the redundancy in burst images leads to the increased storage and transmission demands, as well as reduced efficiency of d

Cited by 0SourcePDFScholar
2026

Rethinking Diffusion Model-Based Video Super-Resolution: Leveraging Dense Guidance from Aligned Features

CVPR 2026

Diffusion model (DM) based Video Super-Resolution (VSR) approaches achieve impressive perceptual quality. Diffusion model (DM) based Video Super-Resolution (VSR) approaches achieve impressive perceptual quality. However, existing DM-based VSR methods over-prioritize perceptual synthesis while neglec

Cited by 0SourcecodeScholar
2026

Self-supervised Dynamic Heterogeneous Degradation Modeling for Unified Zero-Shot Image Restoration

CVPR 2026

Zero-shot image restoration provides a flexible way to handle diverse degradations without task-specific training. However, existing methods typically rely on stacked layers or pre-trained features to enhance degradation expression, while overlooking physically consistent priors. The insufficient de

Cited by 0SourcecodeScholar
2025

Spherical Manifold Guided Diffusion Model for Panoramic Image Generation

CVPR 2025poster

Panoramic image essentially acts as a pivotal role in emerging virtual reality and augmented reality scenarios; however, the generation of panoramic images are essentially challenging due to the intrinsic spherical geometry and spherical distortions caused by equirectangular projection (ERP). To add…

2025

Spherical-Nested Diffusion Model for Panoramic Image Outpainting

ICML 2025poster

Panoramic image outpainting acts as a pivotal role in immersive content generation, allowing for seamless restoration and completion of panoramic content. Given the fact that the majority of generative outpainting solutions operates on planar images, existing methods for panoramic images address the…

2025

Uncover Treasures in DCT: Advancing JPEG Quality Enhancement by Exploiting Latent Correlations

ICCV 2025poster

Joint Photographic Experts Group (JPEG) achieves data compression by quantizing Discrete Cosine Transform (DCT) coefficients, which inevitably introduces compression artifacts. Most existing JPEG quality enhancement methods operate in the pixel domain, suffering from the high computational costs of…

Cited by 0SourcePDFScholar
2024

Causal Context Adjustment Loss for Learned Image Compression

NeurIPS 2024poster

In recent years, learned image compression (LIC) technologies have surpassed conventional methods notably in terms of rate-distortion (RD) performance. Most present learned techniques are VAE-based with an autoregressive entropy model, which obviously promotes the RD performance by utilizing the dec…

2024

Enhancing Quality of Compressed Images by Mitigating Enhancement Bias Towards Compression Domain

CVPR 2024poster

Existing quality enhancement methods for compressed images focus on aligning the enhancement domain with the raw domain to yield realistic images. However these methods exhibit a pervasive enhancement bias towards the compression domain inadvertently regarding it as more realistic than the raw domai…

Cited by 3SourcePDFScholar
2024

Saliency Prediction of Sports Videos: A Large-Scale Database and a Self-Adaptive Approach

ICASSP 2024accepted

Predicting video saliency is crucial for improving sports video processing efficiency, thereby providing an enriched viewing experience for a wide-ranging audience. However, there is a long-term absence of well-established eye-tracking database and learning-based approach, particularly tailored for…

Cited by 0SourceScholar
2023

DINN360: Deformable Invertible Neural Network for Latitude-Aware 360deg Image Rescaling

CVPR 2023poster

With the rapid development of virtual reality, 360deg images have gained increasing popularity. Their wide field of view necessitates high resolution to ensure image quality. This, however, makes it harder to acquire, store and even process such 360deg images. To alleviate this issue, we propose the…

2023

Learning Noise-Induced Reward Functions for Surpassing Demonstrations in Imitation Learning

AAAI 2023technical

Imitation learning (IL) has recently shown impressive performance in training a reinforcement learning agent with human demonstrations, eliminating the difficulty of designing elaborate reward functions in complex environments. However, most IL methods work under the assumption of the optimality of…

Cited by 0SourcePDFScholar
2023

Neural Characteristic Function Learning for Conditional Image Generation

ICCV 2023poster

The emergence of conditional generative adversarial networks (cGANs) has revolutionised the way we approach and control the generation, by means of adversarially learning joint distributions of data and auxiliary information. Despite the success, cGANs have been consistently put under scrutiny due t…

Cited by 7PDFcodeScholar
2023

Uncertainty Guided Adaptive Warping for Robust and Efficient Stereo Matching

ICCV 2023poster

Correlation based stereo matching has achieved outstanding performance, which pursues cost volume between two feature maps. Unfortunately, current methods with a fixed trained model do not work uniformly well across various datasets, greatly limiting their real-world applicability. To tackle this is…

Cited by 24PDFScholar
2022

Does Text Attract Attention on E-Commerce Images: A Novel Saliency Prediction Dataset and Method

CVPR 2022poster

E-commerce images are playing a central role in attracting people's attention when retailing and shopping online, and an accurate attention prediction is of significant importance for both customers and retailers, where its research is yet to start. In this paper, we establish the first dataset of s…

Cited by 18PDFcodeScholar
2021

Deep Homography for Efficient Stereo Image Compression

CVPR 2021poster

In this paper, we propose HESIC, an end-to-end trainable deep network for stereo image compression (SIC). To fully explore the mutual information across two stereo images, we use a deep regression model to estimate the homography matrix, i.e., H matrix. Then, the left image is spatially transformed…

Cited by 54PDFcodeScholar
2021

Deep Multi-Task Learning for Diabetic Retinopathy Grading in Fundus Images

AAAI 2021technical

Recent years have witnessed the growing interest in disease severity grading, especially for ocular diseases based on fundus images. The existing grading methods are usually trained with high resolution (HR) images. However, the grading performance decreases a lot given low resolution (LR) images, w…

Cited by 42SourcePDFScholar
2021

LAU-Net: Latitude Adaptive Upscaling Network for Omnidirectional Image Super-Resolution

CVPR 2021poster

The omnidirectional images (ODIs) are usually at low-resolution, due to the constraints of collection, storage and transmission. The traditional two-dimensional (2D) image super-resolution methods are not effective for spherical ODIs, because ODIs tend to have non-uniformly distributed pixel density…

Cited by 63PDFcodeScholar
2020

Early Exit Or Not: Resource-Efficient Blind Quality Enhancement for Compressed Images

ECCV 2020poster

Lossy image compression is pervasively conducted to save communication bandwidth, resulting in undesirable compression artifacts. Recently, extensive approaches have been proposed to reduce image compression artifacts at the decoder side; however, they require a series of architecture-identical mode…

2020

Learning Diverse Sub-Policies via a Task-Agnostic Regularization on Action Distributions

ICASSP 2020accepted

Automatic sub-policy discovery has recently received much attention in hierarchical reinforcement learning (HRL). The conventional approaches to learning sub-policies suffer from collapsing into just one sub-policy dominating the whole task, lacking techniques to ensure the diversity of different su…

Cited by 0SourceScholar
2020

Learning to Predict Salient Faces: A Novel Visual-Audio Saliency Model

ECCV 2020poster

Recently, video streams have occupied a large proportion of Internet traffic, most of which contain human faces. Hence, it is necessary to predict saliency on multiple-face videos, which can provide attention cues for many content based applications. However, most of multiple-face prediction works o…

2020

Multi-level Wavelet-based Generative Adversarial Network for Perceptual Quality Enhancement of Compressed Video

ECCV 2020poster

The past few years have witnessed fast development in video quality enhancement via deep learning. Existing methods mainly focus on enhancing the objective quality of compressed videos while ignoring its perceptual quality. In this paper, we focus on enhancing the perceptual quality of compressed vi…

2019

Attention Based Glaucoma Detection: A Large-Scale Database and CNN Model

CVPR 2019poster

Recently, the attention mechanism has been successfully applied in convolutional neural networks (CNNs), significantly boosting the performance of many computer vision tasks. Unfortunately, few medical image recognition approaches incorporate the attention mechanism in the CNNs. In particular, there…

Cited by 302PDFcodeScholar
2019

Optimizing QoE of Multiple Users over DASH: A Meta-learning Approach

ICASSP 2019accepted

Dynamic adaptive video streaming over HTTP (DASH) plays a key role in video transmission over the Internet. The conventional DASH adaptation approaches concentrate on optimizing the overall quality of experience (QoE) for all client sides, neglecting the QoE diversity of different users. In this pap…

Cited by 0SourceScholar
2019

Wavelet Domain Style Transfer for an Effective Perception-Distortion Tradeoff in Single Image Super-Resolution

ICCV 2019oral

In single image super-resolution (SISR), given a low-resolution (LR) image, one wishes to find a high-resolution (HR) version of it which is both accurate and photorealistic. Recently, it has been shown that there exists a fundamental tradeoff between low distortion and high perceptual quality, and…

Cited by 97PDFcodeScholar
2018

DeepVS: A Deep Learning Based Video Saliency Prediction Approach

ECCV 2018poster

In this paper, we propose a novel deep learning based video saliency prediction method, named DeepVS. Specifically, we establish a large-scale eye-tracking database of videos (LEDOV), which includes 32 subjects' fixations on 538 videos. We find from LEDOV that human attention is more likely to be at…