← Search

Jufeng Yang

37 accepted papers

2026

Collaborative Feature Matching with Progressive Correspondence Learning

AAAI 2026technical

Accurate feature matching between image pairs is fundamental for various computer vision applications. In detector-base process, the feature matcher aims to find the optimal feature correspondences, and the match filter is used for further removing mismatches. However, their connection is rarely exp

Cited by 0SourcePDFScholar
2026

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation

CVPR 2026

Few-step generation has been a long-standing goal, with recent one-step generation methods exemplified by MeanFlow achieving remarkable results. Existing research on MeanFlow primarily focuses on class-to-image generation. However, an intuitive yet unexplored direction is to extend the condition fro

Cited by 0SourcecodeScholar
2026

It Takes Two: A Duet of Periodicity and Directionality for Burst Flicker Removal

CVPR 2026

Flicker artifacts, arising from unstable illumination and row-wise exposure inconsistencies, pose a significant challenge in short-exposure photography, severely degrading image quality. Unlike typical artifacts, e.g., noise and low-light, flicker is a structured degradation with specific spatial-te

Cited by 0SourcecodeScholar
2026

PHMRNet: Persistent Homology Based Mamba-RWKV Network for LiDAR Place Recognition

RA-L 2026

LiDAR-based place recognition (LPR) is a key component of visual localization and autonomous driving. Although LiDAR data are usually preprocessed by motion undistortion, which can greatly reduce scene distortion caused by sensor motion, 3-dimensional (3D) point clouds in complex scenes still show i

Cited by 0SourceScholar
2025

Boosting the Dual-Stream Architecture in Ultra-High Resolution Segmentation with Resolution-Biased Uncertainty Estimation

CVPR 2025poster

Over the last decade, significant efforts have been dedicated to designing efficient models for the challenge of ultra-high resolution (UHR) semantic segmentation. These models mainly follow the dual-stream architecture and generally fall into three subcategories according to the improvement objecti…

2025

BurstDeflicker: A Benchmark Dataset for Flicker Removal in Dynamic Scenes

NeurIPS 2025poster

Flicker artifacts in short-exposure images are caused by the interplay between the row-wise exposure mechanism of rolling shutter cameras and the temporal intensity variations of alternating current (AC)-powered lighting. These artifacts typically appear as uneven brightness distribution across the…

Cited by 0SourceScholar
2025

Devil is in the Uniformity: Exploring Diverse Learners within Transformer for Image Restoration

ICCV 2025poster

Transformer-based approaches have gained significant attention in image restoration, where the core component, i.e, Multi-Head Attention (MHA), plays a crucial role in capturing diverse features and recovering high-quality results. In MHA, heads perform attention calculation independently from unifo…

2025

FlareX: A Physics-Informed Dataset for Lens Flare Removal via 2D Synthesis and 3D Rendering

NeurIPS 2025poster

Lens flare occurs when shooting towards strong light sources, significantly degrading the visual quality of images. Due to the difficulty in capturing flare-corrupted and flare-free image pairs in the real world, existing datasets are typically synthesized in 2D by overlaying artificial flare templa…

Cited by 0SourceScholar
2025

Hybrid Re-matching for Continual Learning with Parameter-Efficient Tuning

NeurIPS 2025poster

Continual learning seeks to enable a model to assimilate knowledge from non-stationary data streams without catastrophic forgetting. Recently, methods based on Parameter-Efficient Tuning (PET) have achieved superior performance without even storing any historical exemplars, which train much fewer sp…

Cited by 0SourcecodeScholar
2025

MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding

ICML 2025spotlight

Multimodal large language models (MLLMs) recently showed strong capacity in integrating data among multiple modalities, empowered by generalizable attention architecture. Advanced methods predominantly focus on language-centric tuning while less exploring multimodal tokens mixed through attention, p…

Cited by 0SourcePDFScholar
2025

No Pains, More Gains: Recycling Sub-Salient Patches for Efficient High-Resolution Image Recognition

CVPR 2025highlight

Over the last decade, many notable methods have emerged to tackle the computational resource challenge of the high resolution image recognition (HRIR). They typically focus on identifying and aggregating a few salient regions for classification, discarding sub-salient areas for low training consumpt…

2025

PS-Diffusion: Photorealistic Subject-Driven Image Editing with Disentangled Control and Attention

CVPR 2025poster

Diffusion models pre-trained on large-scale paired image-text data achieve significant success in image editing. To convey more fine-grained visual details, subject-driven editing integrates subjects in user-provided reference images into existing scenes. However, it is challenging to obtain photore…

2025

Seek Common Ground While Reserving Differences: Semi-Supervised Image-Text Sentiment Recognition

CVPR 2025poster

Multimodal sentiment analysis has attracted extensive research attention as increasing users share images and texts to express their emotions and opinions on social media. Collecting large amounts of labeled sentiment data is an expensive and challenging task due to the high cost of labeling and una…

2025

VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models

NeurIPS 2025poster

Understanding and predicting emotions from videos has gathered significant attention in recent studies, driven by advancements in video large language models (VideoLLMs). While advanced methods have made progress in video emotion analysis, the intrinsic nature of emotions—characterized by their open…

Cited by 0SourceScholar
2024

Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image Restoration

CVPR 2024poster

Transformer-based approaches have achieved promising performance in image restoration tasks given their ability to model long-range dependencies which is crucial for recovering clear images. Though diverse efficient attention mechanism designs have addressed the intensive computations associated wit…

2024

ExtDM: Distribution Extrapolation Diffusion Model for Video Prediction

CVPR 2024poster

Video prediction is a challenging task due to its nature of uncertainty especially for forecasting a long period. To model the temporal dynamics advanced methods benefit from the recent success of diffusion models and repeatedly refine the predicted future frames with 3D spatiotemporal U-Net. Howeve…

Cited by 22SourcePDFScholar
2024

LAKE-RED: Camouflaged Images Generation by Latent Background Knowledge Retrieval-Augmented Diffusion

CVPR 2024poster

Camouflaged vision perception is an important vision task with numerous practical applications. Due to the expensive collection and labeling costs this community struggles with a major bottleneck that the species category of its datasets is limited to a small number of object species. However the ex…

2024

MART: Masked Affective RepresenTation Learning via Masked Temporal Distribution Distillation

CVPR 2024poster

Limited training data is a long-standing problem for video emotion analysis (VEA). Existing works leverage the power of large-scale image datasets for transferring while failing to extract the temporal correlation of affective cues in the video. Inspired by psychology research and empirical theory w…

Cited by 9SourcePDFScholar
2024

Seeing the Unseen: A Frequency Prompt Guided Transformer for Image Restoration

ECCV 2024poster

"How to explore useful features from images as prompts to guide the deep image restoration models is an effective way to solve image restoration. In contrast to mining spatial relations within images as prompt, which leads to characteristics of different frequencies being neglected and further remai…

2024

To Err Like Human: Affective Bias-Inspired Measures for Visual Emotion Recognition Evaluation

NeurIPS 2024poster

Accuracy is a commonly adopted performance metric in various classification tasks, which measures the proportion of correctly classified samples among all samples. It assumes equal importance for all classes, hence equal severity for misclassifications. However, in the task of emotional classificati…

2023

Probing Sentiment-Oriented Pre-Training Inspired by Human Sentiment Perception Mechanism

CVPR 2023poster

Pre-training of deep convolutional neural networks (DCNNs) plays a crucial role in the field of visual sentiment analysis (VSA). Most proposed methods employ the off-the-shelf backbones pre-trained on large-scale object classification datasets (i.e., ImageNet). While it boosts performance for a big…

2023

Weakly Supervised Video Emotion Detection and Prediction via Cross-Modal Temporal Erasing Network

CVPR 2023poster

Automatically predicting the emotions of user-generated videos (UGVs) receives increasing interest recently. However, existing methods mainly focus on a few key visual frames, which may limit their capacity to encode the context that depicts the intended emotions. To tackle that, in this paper, we p…

2020

BBS-Net: RGB-D Salient Object Detection with a Bifurcated Backbone Strategy Network

ECCV 2020poster

Multi-level feature fusion is a fundamental topic in computer vision for detecting, segmenting, and classifying objects at various scales. When multi-level features meet multi-modal cues, the optimal fusion problem becomes a hot potato. In this paper, we make the first attempt to leverage the inhere…

2019

Attention-Aware Polarity Sensitive Embedding for Affective Image Retrieval

ICCV 2019poster

Images play a crucial role for people to express their opinions online due to the increasing popularity of social networks. While an affective image retrieval system is useful for obtaining visual contents with desired emotions from a massive repository, the abstract and subjective characteristics m…

Cited by 45PDFScholar
2019

EGNet: Edge Guidance Network for Salient Object Detection

ICCV 2019poster

Fully convolutional neural networks (FCNs) have shown their advantages in the salient object detection task. However, most existing FCNs-based methods still suffer from coarse object boundaries. In this paper, to solve this problem, we focus on the complementarity between salient edge information an…

Cited by 1302PDFScholar
2019

IP102: A Large-Scale Benchmark Dataset for Insect Pest Recognition

CVPR 2019oral

Insect pests are one of the main factors affecting agricultural product yield. Accurate recognition of insect pests facilitates timely preventive measures to avoid economic losses. However, the existing datasets for the visual classification task mainly focus on common objects, e.g., flowers and do…

Cited by 519PDFcodeScholar
2019

Joint Acne Image Grading and Counting via Label Distribution Learning

ICCV 2019accepted

Accurate grading of skin disease severity plays a crucial role in precise treatment for patients. Acne vulgaris, the most common skin disease in adolescence, can be graded by evidence-based lesion counting as well as experience-based global estimation in the medical field. However, due to the appear…

2019

Zero-Shot Emotion Recognition via Affective Structural Embedding

ICCV 2019poster

Image emotion recognition attracts much attention in recent years due to its wide applications. It aims to classify the emotional response of humans, where candidate emotion categories are generally defined by specific psychological theories, such as Ekman's six basic emotions. However, with the dev…

Cited by 64PDFScholar
2018

Clinical Skin Lesion Diagnosis Using Representations Inspired by Dermatologist Criteria

CVPR 2018poster

The skin is the largest organ in human body. Around 30%-70% of individuals worldwide have skin related health problems, for whom effective and efficient diagnosis is necessary. Recently, computer aided diagnosis (CAD) systems have been successfully applied to the recognition of skin cancers in derma…

Cited by 134SourcePDFScholar
2018

Sub-GAN: An Unsupervised Generative Model via Subspaces

ECCV 2018poster

The recent years have witnessed significant growth in constructing robust generative models to capture informative distributions of natural data. However, it is difficult to fully exploit the distribution of complex data, like images and videos, due to the high dimensionality of ambient space. Seque…

Cited by 24SourcePDFScholar
2018

Weakly Supervised Coupled Networks for Visual Sentiment Analysis

CVPR 2018poster

Automatic assessment of sentiment from visual content has gained considerable attention with the increasing tendency of expressing opinions on-line. In this paper, we solve the problem of visual sentiment analysis using the high-level abstraction in the recognition process. Existing methods based on…

Cited by 161SourcePDFScholar