← Search

Yuming Fang

26 accepted papers

2026

Event-Guided Super-Resolving Blurry Image via Asymmetric Integral Driven Consistency

AAAI 2026technical

Super-Resolution from a Blurry low-resolution image (SRB) constitutes a severely ill-posed inverse problem. Current learning-based SRB approaches primarily rely on synthetic, well-labeled paired datasets to regularize solution spaces, yet they exhibit limited generalizability in practical applicatio

Cited by 0SourcePDFScholar
2026

Infrared-Privileged UAV Detection via Cross-Modal Vector-Quantization

AAAI 2026technical

RGB and infrared images has shown remarkable robustness for object detection based on unmanned aerial vehicles (UAV). However, the primitive RGB and infrared (IR) images are inevitably misaligned due to the device gap between RGB and infrared cameras. Most existing methods rely on manually filtered

Cited by 0SourcePDFScholar
2026

Learning and Aligning Click-Aware Shape Prior for Interactive Amodal Instance Segmentation

CVPR 2026

Amodal instance segmentation aims to segment both visible and occluded regions of object instance, which are challenging due to lacking inference support under occlusion. Most existing methods employ the prior knowledge about object mask (shape prior) to support the amodal estimation, but the shape

Cited by 0SourcecodeScholar
2026

PanFoMa: A Lightweight Foundation Model and Benchmark for Pan-Cancer

AAAI 2026technical

Single-cell RNA sequencing (scRNA-seq) is essential for decoding tumor heterogeneity. However, pan-cancer research still faces two key challenges: learning discriminative and efficient single-cell representations, and establishing a comprehensive evaluation benchmark. In this paper, we introduce \al

Cited by 0SourcePDFScholar
2026

TaylorMoDe-GS: Taylor-Driven Gaussian Splatting Motion Model for Multi-View Dynamic Scene Deblurring

IJCAI 2026

While 3D Gaussian Splatting (3DGS) has excelled in dynamic scene reconstruction, it struggles with multi-view object motion blur, where view-dependent non-uniform blur violates fundamental multi-view geometric constraints. Existing methods fail to balance complex motion fitting with physical consist

Cited by 0Scholar
2025

Enhancing Low-Light Images: A Synthetic Data Perspective on Practical and Generalizable Solutions

AAAI 2025technical

Recently, deep neural networks (DNNs) have emerged as the leading approach for low-light image enhancement (LLIE). However, training these models generally requires large-scale paired datasets, which are challenging to obtain due to the labor-intensive and time-consuming nature of real-world data co…

2025

PSReg: Prior-guided Sparse Mixture of Experts for Point Cloud Registration

AAAI 2025technical

The discriminative feature is crucial for point cloud registration. Recent methods improve the feature discriminative by distinguishing between non-overlapping and overlapping region points. However, they still face challenges in distinguishing the ambiguous structures in the overlapping regions. Th…

Cited by 1SourcePDFScholar
2025

PointGAC: Geometric-Aware Codebook for Masked Point Modeling

ICCV 2025poster

Most masked point cloud modeling (MPM) methods follow a regression paradigm to reconstruct the coordinate or feature of masked regions. However, they tend to over-constrain the model to learn the details of the masked region, resulting in failure to capture generalized features. To address this limi…

2025

Recurrent Feature Mining and Keypoint Mixup Padding for Category-Agnostic Pose Estimation

CVPR 2025poster

Category-agnostic pose estimation aims to locate keypoints on query images according to a few annotated support images for arbitrary novel classes. Existing methods generally extract support features via heatmap pooling, and obtain interacted features from support and query via cross-attention. Henc…

2025

Weak-shot Keypoint Estimation via Keyness and Correspondence Transfer

NeurIPS 2025poster

Keypoint estimation is a fundamental task in computer vision, but generally requires large-scale annotated data for training. Few-shot and unsupervised keypoint estimation are prevalent economical paradigms, but the former still requires annotations for extensive novel classes while the latter only…

Cited by 0SourceScholar
2024

Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare

NeurIPS 2024spotlight

While recent advancements in large multimodal models (LMMs) have significantly improved their abilities in image quality assessment (IQA) relying on absolute quality rating, how to transfer reliable relative quality comparison outputs to continuous perceptual quality scores remains largely unexplore…

2024

Arbitrary-Scale Video Super-Resolution with Structural and Textural Priors

ECCV 2024poster

"Arbitrary-scale video super-resolution (AVSR) aims to enhance the resolution of video frames, potentially at various scaling factors, which presents several challenges regarding spatial detail reproduction, temporal consistency, and computational complexity. In this paper, we first describe a stron…

2024

Comprehensive Visual Grounding for Video Description

AAAI 2024technical

The grounding accuracy of existing video captioners is still behind the expectation. The majority of existing methods perform grounded video captioning on sparse entity annotations, whereas the captioning accuracy often suffers from degenerated object appearances on the annotated area such as motion…

Cited by 2SourcePDFScholar
2024

Meta-Point Learning and Refining for Category-Agnostic Pose Estimation

CVPR 2024poster

Category-agnostic pose estimation (CAPE) aims to predict keypoints for arbitrary classes given a few support images annotated with keypoints. Existing methods only rely on the features extracted at support keypoints to predict or refine the keypoints on query image but a few support feature vectors…

2024

Multiscale Sliced Wasserstein Distances as Perceptual Color Difference Measures

ECCV 2024poster

"Contemporary color difference (CD) measures for photographic images typically operate by comparing co-located pixels, patches in a “perceptually uniform” color space, or features in a learned latent space. Consequently, these measures inadequately capture the human color perception of misaligned im…

2023

ScanDMM: A Deep Markov Model of Scanpath Prediction for 360deg Images

CVPR 2023poster

Scanpath prediction for 360deg images aims to produce dynamic gaze behaviors based on the human visual perception mechanism. Most existing scanpath prediction methods for 360deg images do not give a complete treatment of the time-dependency when predicting human scanpath, resulting in inferior perfo…

2022

GMF: General Multimodal Fusion Framework for Correspondence Outlier Rejection

RA-L 2022

Rejecting correspondence outliers enables to boost the correspondence quality, which is a critical step in achieving high point cloud registration accuracy. The current state-of-the-art correspondence outlier rejection methods only utilize the structure features of the correspondences. However, text

Cited by 15SourcecodeScholar
2022

IMFNet: Interpretable Multimodal Fusion for Point Cloud Registration

RA-L 2022

The existing state-of-the-art point descriptor relies on structure information only, which omits the texture information. However, texture information is crucial for our humans to distinguish a scene part. Moreover, the current learning-based point descriptors are all black boxes which are unclear h

Cited by 52SourcecodeScholar
2022

Perceptual Quality Assessment of Omnidirectional Images

AAAI 2022technical

Omnidirectional images, also called 360◦images, have attracted extensive attention in recent years, due to the rapid development of virtual reality (VR) technologies. During omnidirectional image processing including capture, transmission, consumption, and so on, measuring the perceptual quality of…

Cited by 167SourcePDFScholar
2022

Unsupervised Point Cloud Registration by Learning Unified Gaussian Mixture Models

RA-L 2022

Sampling noise and density variation widely exist in the point cloud acquisition process, leading to few accurate point-to-point correspondences. Since they rely on point-to-point correspondence search, existing state-of-the-art point cloud registration methods face difficulty in overcoming the samp

Cited by 32SourceScholar
2020

From Fidelity to Perceptual Quality: A Semi-Supervised Approach for Low-Light Image Enhancement

CVPR 2020poster

Under-exposure introduces a series of visual degradation, i.e. decreased visibility, intensive noise, and biased color, etc. To address these problems, we propose a novel semi-supervised learning approach for low-light image enhancement. A deep recursive band network (DRBN) is proposed to recover a…

Cited by 657PDFScholar
2016

Aspect Ratio Similarity (ARS) for image retargeting quality assessment

ICASSP 2016accepted

During the past few years, there have been various kinds of content-aware image retargeting methods proposed for image resizing. However, the lack of effective objective retargeting quality metric limits the further development of image retargeting. Different from the traditional image quality asses…

Cited by 0SourceScholar
2015

Multi-task rank learning for image quality assessment

ICASSP 2015accepted

In practice, multiple types of distortions are associated with an image quality degradation process. The existing machine learning (ML) based image quality assessment (IQA) approaches generally established a unified model for all distortion types, or each model is trained independently for each dist…

Cited by 0SourceScholar