← Search

Guangming Shi

38 accepted papers

2026

Scaling Dense Event-Stream Pretraining from Visual Foundation Models

CVPR 2026

Learning versatile, fine-grained representations from irregular event streams is pivotal yet nontrivial, primarily due to the heavy annotation that hinders scalability in dataset size, semantic richness, and application scope. To mitigate this dilemma, we launch a novel self-supervised pretraining m

Cited by 0SourcecodeScholar
2025

Asymmetric Hierarchical Difference-aware Interaction Network for Event-guided Motion Deblurring

AAAI 2025technical

Event cameras are bio-inspired sensors that are capable of capturing motion information with high temporal resolution, which show potential in aiding image motion deblurring recently. Most existing methods indiscriminately handle feature fusion of two modalities with symmetric unidirectional/bidirec…

2025

Feature Information Driven Position Gaussian Distribution Estimation for Tiny Object Detection

CVPR 2025poster

Tiny object detection remains challenging in spite of the success of generic detectors. The dramatic performance degradation of generic detectors on tiny objects is mainly due to the the weak representations of extremely limited pixels. To address this issue, we propose a plug-and-play architecture…

Cited by 0SourcePDFScholar
2025

Gain from Neighbors: Boosting Model Robustness in the Wild via Adversarial Perturbations Toward Neighboring Classes

CVPR 2025poster

Recent approaches, such as data augmentation, adversarial training, and transfer learning, have shown potential in addressing the issue of performance degradation caused by distributional shifts. However, they typically demand careful design in terms of data or models and lack awareness of the impac…

Cited by 0SourcePDFScholar
2025

Parameterized Blur Kernel Prior Learning for Local Motion Deblurring

CVPR 2025poster

Unlike global motion blur, Local Motion Deblurring (LMD) presents a more complex challenge, as it requires precise restoration of blurry regions while preserving the sharpness of the background. Existing LMD methods rely on manually annotated blur masks and often overlook the blur kernel's character…

Cited by 0SourcePDFScholar
2025

PatternCIR Benchmark and TisCIR: Advancing Zero-Shot Composed Image Retrieval in Remote Sensing

IJCAI 2025

Remote sensing composed image retrieval (RSCIR) is a new vision-language task that takes a composed query of an image and text, aiming to search for a target remote sensing image satisfying two conditions from intricate remote sensing imagery. However, the existing attribute-based benchmark Patternc

Cited by 0SourcePDFScholar
2025

Triad: Empowering LMM-based Anomaly Detection with Expert-guided Region-of-Interest Tokenizer and Manufacturing Process

ICCV 2025poster

Although recent methods have tried to introduce large multimodal models (LMMs) into industrial anomaly detection (IAD), their generalization in the IAD field is far inferior to that for general purposes. We summarize the main reasons for this gap into two aspects. On one hand, general-purpose LMMs l…

2024

BVT-IMA: Binary Vision Transformer with Information-Modified Attention

AAAI 2024technical

As a compression method that can significantly reduce the cost of calculations and memories, model binarization has been extensively studied in convolutional neural networks. However, the recently popular vision transformer models pose new challenges to such a technique, in which the binarized model…

Cited by 1SourcePDFScholar
2024

E-Motion: Future Motion Simulation via Event Sequence Diffusion

NeurIPS 2024poster

Forecasting a typical object's future motion is a critical task for interpreting and interacting with dynamic environments in computer vision. Event-based sensors, which could capture changes in the scene with exceptional temporal granularity, may potentially offer a unique opportunity to predict fu…

2024

Inverse Weight-Balancing for Deep Long-Tailed Learning

AAAI 2024technical

The performance of deep learning models often degrades rapidly when faced with imbalanced data characterized by a long-tailed distribution. Researchers have found that the fully connected layer trained by cross-entropy loss has large weight-norms for classes with many samples, but not for classes wi…

Cited by 3SourcePDFScholar
2024

Mosic: Multimodal Semantic Integrated Communication for Health Monitoring in Iot Scenarios

ICASSP 2024accepted

Monitoring multimodal signals provides a more comprehensive understanding of health conditions compared to singlemode monitoring. In the face of the significant volumes of multimodal signals, existing IoT health monitoring systems primarily focus on high-fidelity signal transmission by encoding mult…

Cited by 0SourceScholar
2024

Motion Deblurring via Spatial-Temporal Collaboration of Frames and Events

AAAI 2024technical

Motion deblurring can be advanced by exploiting informative features from supplementary sensors such as event cameras, which can capture rich motion information asynchronously with high temporal resolution. Existing event-based motion deblurring methods neither consider the modality redundancy in sp…

2024

SG2SC: A Generative Semantic Communication Framework for Scene Understanding-Oriented Image Transmission

ICASSP 2024accepted

In recent years, semantic communication based on deep learning for source-channel joint encoding has garnered significant attention. It utilizes network models trained end-to-end to represent signals as embedding vectors and has demonstrated superior performance compared to traditional methods. Howe…

Cited by 0SourceScholar
2024

Segment Any Event Streams via Weighted Adaptation of Pivotal Tokens

CVPR 2024poster

In this paper we delve into the nuanced challenge of tailoring the Segment Anything Models (SAMs) for integration with event data with the overarching objective of attaining robust and universal object segmentation within the event-centric domain. One pivotal issue at the heart of this endeavor is t…

2023

Gradient Corner Pooling for Keypoint-Based Object Detection

AAAI 2023technical

Detecting objects as multiple keypoints is an important approach in the anchor-free object detection methods while corner pooling is an effective feature encoding method for corner positioning. The corners of the bounding box are located by summing the feature maps which are max-pooled in the x and…

Cited by 1SourcePDFScholar
2023

Improving Robotic Tactile Localization Super-resolution via Spatiotemporal Continuity Learning and Overlapping Air Chambers

AAAI 2023technical

Human hand has amazing super-resolution ability in sensing the force and position of contact and this ability can be strengthened by practice. Inspired by this, we propose a method for robotic tactile super-resolution enhancement by learning spatiotemporal continuity of contact position and a tactil…

Cited by 5SourcePDFScholar
2023

Low-Light Image Enhancement with Multi-Stage Residue Quantization and Brightness-Aware Attention

ICCV 2023poster

Low-light image enhancement (LLIE) aims to recover illumination and improve the visibility of low-light images. Conventional LLIE methods often produce poor results because they neglect the effect of noise interference. Deep learning-based LLIE methods focus on learning a mapping function between lo…

Cited by 26PDFcodeScholar
2023

Residual Degradation Learning Unfolding Framework With Mixing Priors Across Spectral and Spatial for Compressive Spectral Imaging

CVPR 2023poster

To acquire a snapshot spectral image, coded aperture snapshot spectral imaging (CASSI) is proposed. A core problem of the CASSI system is to recover the reliable and fine underlying 3D spectral cube from the 2D measurement. By alternately solving a data subproblem and a prior subproblem, deep unfold…

2023

Self-Supervised Non-Uniform Kernel Estimation With Flow-Based Motion Prior for Blind Image Deblurring

CVPR 2023poster

Many deep learning-based solutions to blind image deblurring estimate the blur representation and reconstruct the target image from its blurry observation. However, these methods suffer from severe performance degradation in real-world scenarios because they ignore important prior information about…

2023

Vector Quantization With Self-Attention for Quality-Independent Representation Learning

CVPR 2023poster

Recently, the robustness of deep neural networks has drawn extensive attention due to the potential distribution shift between training and testing data (e.g., deep models trained on high-quality images are sensitive to corruption during testing). Many researchers attempt to make the model learn inv…

Cited by 9SourcePDFScholar
2022

Learning Degradation Uncertainty for Unsupervised Real-world Image Super-resolution

IJCAI 2022poster

Acquiring degraded images with paired high-resolution (HR) images is often challenging, impeding the advance of image super-resolution in real-world applications. By generating realistic low-resolution (LR) images with degradation similar to that in real-world scenarios, simulated paired LR-HR data…

Cited by 13SourcePDFScholar
2022

Robust Depth Completion with Uncertainty-Driven Loss Functions

AAAI 2022technical

Recovering a dense depth image from sparse LiDAR scans is a challenging task. Despite the popularity of color-guided methods for sparse-to-dense depth completion, they treated pixels equally during optimization, ignoring the uneven distribution characteristics in the sparse depth map and the accumul…

Cited by 51SourcePDFScholar
2022

Self-Feature Distillation with Uncertainty Modeling for Degraded Image Recognition

ECCV 2022poster

"Despite the remarkable performance on high-quality (HQ) data, the accuracy of deep image recognition models degrades rapidly in the presence of low-quality (LQ) images. Both feature de-drifting and quality agnostic models have been developed to make the features extracted from degraded images close…

Cited by 14SourcePDFScholar
2022

Uncertainty Learning in Kernel Estimation for Multi-stage Blind Image Super-Resolution

ECCV 2022poster

"Conventional wisdom in blind super-resolution (SR) first estimates the unknown degradation from the low-resolution image and then exploits the degradation information for image reconstruction. Such sequential approaches suffer from two fundamental weaknesses - i.e., the lack of robustness (the perf…

Cited by 19SourcePDFScholar
2021

Deep Gaussian Scale Mixture Prior for Spectral Compressive Imaging

CVPR 2021poster

In coded aperture snapshot spectral imaging (CASSI) system, the real-world hyperspectral image (HSI) can be reconstructed from the captured compressive image in a snapshot. Model-based HSI reconstruction methods employed hand-crafted priors to solve the reconstruction problem, but most of which achi…

Cited by 185PDFScholar
2021

EffiScene: Efficient Per-Pixel Rigidity Inference for Unsupervised Joint Learning of Optical Flow, Depth, Camera Pose and Motion Segmentation

CVPR 2021poster

This paper addresses the challenging unsupervised scene flow estimation problem by jointly learning four low-level vision sub-tasks: optical flow F, stereo-depth D, camera pose P and motion segmentation S. Our key insight is that the rigidity of the scene shares the same inherent geometrical structu…

Cited by 47PDFScholar
2021

PGNet: Real-time Arbitrarily-Shaped Text Spotting with Point Gathering Network

AAAI 2021technical

The reading of arbitrarily-shaped text has received increasing research attention. However, existing text spotters are mostly built on two-stage frameworks or character-based methods, which suffer from either Non-Maximum Suppression (NMS), Region-of-Interest (RoI) operations, or character-level anno…

2021

Uncertainty-Driven Loss for Single Image Super-Resolution

NeurIPS 2021poster

In low-level vision such as single image super-resolution (SISR), traditional MSE or L_1 loss function treats every pixel equally with the assumption that the importance of all pixels is the same. However, it has been long recognized that texture and edge areas carry more important visual informatio…

Cited by 75SourcePDFScholar
2021

Unsupervised Curriculum Domain Adaptation for No-Reference Video Quality Assessment

ICCV 2021poster

During the last years, convolutional neural networks (CNNs) have triumphed over video quality assessment (VQA) tasks. However, CNN-based approaches heavily rely on annotated data which are typically not available in VQA, leading to the difficulty of model generalization. Recent advances in domain ad…

Cited by 34PDFcodeScholar
2020

Beyond Network Pruning: a Joint Search-and-Training Approach

IJCAI 2020poster

Network pruning has been proposed as a remedy for alleviating the over-parameterization problem of deep neural networks. However, its value has been recently challenged especially from the perspective of neural architecture search (NAS). We challenge the conventional wisdom of pruning-after-training…

Cited by 0SourcePDFScholar
2020

Denoising of Event-Based Sensors with Spatial-Temporal Correlation

ICASSP 2020accepted

As a novel asynchronous-driven cameras, event-based sensors are with high sensitivity, fast speed, low power consumption and low data volume, but with abundant noise. Since the output of event-based sensors is in the form of address-event-representation (AER), the traditional frame-based denoising m…

Cited by 0SourceScholar
2020

MetaIQA: Deep Meta-Learning for No-Reference Image Quality Assessment

CVPR 2020poster

Recently, increasing interest has been drawn in exploiting deep convolutional neural networks (DCNNs) for no-reference image quality assessment (NR-IQA). Despite of the notable success achieved, there is a broad consensus that training DCNNs heavily relies on massive annotated data. Unfortunately, I…

Cited by 443PDFcodeScholar
2016

Enhanced just noticeable difference model with visual regularity consideration

ICASSP 2016accepted

Just noticeable difference (JND) reveals the visibility of our human visual system (HVS), below which changes cannot be perceived by the human. Though dozens of JND estimation models have been introduced during the past decade, how to accurately estimate the JND thresholds for different content regi…

Cited by 0SourceScholar
2016

Learning Parametric Sparse Models for Image Super-Resolution

NeurIPS 2016poster

Learning accurate prior knowledge of natural images is of great importance for single image super-resolution (SR). Existing SR methods either learn the prior from the low/high-resolution patch pairs or estimate the prior models from the input low-resolution (LR) image. Specifically, high-frequency d…

Cited by 9SourcePDFScholar
2015

High-Speed Hyperspectral Video Acquisition With a Dual-Camera Architecture

CVPR 2015poster

We propose a novel dual-camera design to acquire 4D high-speed hyperspectral (HSHS) videos with high spatial and spectral resolution. Our work has two key technical contributions. First, we build a dual-camera system that simultaneously captures a panchromatic video at a high frame rate and a hypers…

Cited by 116SourcePDFScholar
2015

Learning Parametric Distributions for Image Super-Resolution: Where Patch Matching Meets Sparse Coding

ICCV 2015poster

Existing approaches toward Image super-resolution (SR) is often either data-driven (e.g., based on internet-scale matching and web image retrieval) or model-based (e.g., formulated as an Maximizing a Posterior estimation problem). The former is conceptually simple yet heuristic; while the latter is…

Cited by 29PDFScholar
2015

Low-Rank Tensor Approximation With Laplacian Scale Mixture Modeling for Multiframe Image Denoising

ICCV 2015poster

Patch-based low-rank models have shown effective in exploiting spatial redundancy of natural images especially for the application of image denoising. However, two-dimensional low-rank model can not fully exploit the spatio-temporal correlation in larger data sets such as multispectral images and 3D…

Cited by 82PDFScholar