← Search

jiangtao wen

17 accepted papers

2026

FLOW INTELLIGENCE: ROBUST FEATURE MATCHING VIA TEMPORAL SIGNATURE CORRELATION

ICASSP 2026poster

Feature matching across video streams remains a cornerstone challenge in computer vision. Increasingly, robust multimodal matching has garnered interest in robotics, surveillance, remote sensing, and medical imaging. While traditional rely on detecting and matching spatial features, they break down…

Cited by 0SourcePDFScholar
2025

EP-SAM: An Edge-Detection Prompt SAM Based Efficient Framework for Ultra-Low Light Video Segmentation

ICASSP 2025accepted

The Segment Anything Model (SAM) excels at generating high-quality object masks with various prompts but struggles in ultra-low light. We developed EP-SAM (Edge-Detection Prompt SAM) with a Low Light Edge-Detection Network (LLEN), offering strong robustness and lightweight performance in ultra-low l…

Cited by 0SourceScholar
2023

DarkFeat: Noise-Robust Feature Detector and Descriptor for Extremely Low-Light RAW Images

AAAI 2023technical

Low-light visual perception, such as SLAM or SfM at night, has received increasing attention, in which keypoint detection and local feature description play an important role. Both handcraft designs and machine learning methods have been widely studied for local feature detection and description, ho…

2023

Efficient Semantic Segmentation by Altering Resolutions for Compressed Videos

CVPR 2023poster

Video semantic segmentation (VSS) is a computationally expensive task due to the per-frame prediction for videos of high frame rates. In recent work, compact models or adaptive network strategies have been proposed for efficient VSS. However, they did not consider a crucial factor that affects the c…

2022

Rate Control for Learned Video Compression

ICASSP 2022accepted

Rate control is a critical part for video compression, especially in bandwidth-limited tasks such as live and broadcast. The newly-rising learned video compression has shown advantageous rate-distortion (RD) performance in previous research, but lack of rate control heavily limits its usage in real…

Cited by 0SourceScholar
2021

Learning Model-Blind Temporal Denoisers without Ground Truths

ICASSP 2021accepted

Denoisers trained with synthetic noises often fail to cope with the diversity of real noises, giving way to methods that can adapt to unknown noise without noise modeling or ground truth. Previous image-based method leads to noise overfitting if directly applied to temporal denoising, and has inadeq…

Cited by 0SourceScholar
2021

Learning to Estimate Kernel Scale and Orientation of Defocus Blur with Asymmetric Coded Aperture

ICASSP 2021accepted

Consistent in-focus input imagery is an essential precondition for machine vision systems to perceive the dynamic environment. A de-focus blur severely degrades the performance of vision systems. To tackle this problem, we propose a deep-learning-based framework estimating the kernel scale and orien…

Cited by 0SourceScholar
2020

Deep Material Recognition in Light-Fields via Disentanglement of Spatial and Angular Information

ECCV 2020poster

Light-field cameras capture sub-views from multiple perspectives simultaneously, with possibly reflectance variations that can be used to augment material recognition in remote sensing, autonomous driving, etc. Existing approaches for light-field based material recognition suffer from the entangleme…

Cited by 10SourcePDFScholar
2020

Online Multi-modal Person Search in Videos

ECCV 2020poster

The task of searching certain people in videos has seen increasing potential in real-world applications, such as video organization and editing. Most existing approaches are devised to work in an offline manner, where identifies can only be inferred after an entire video is examined. This working ma…

Cited by 35SourcePDFScholar
2020

Social Data Assisted Multi-Modal Video Analysis For Saliency Detection

ICASSP 2020accepted

Video saliency should be taken into consideration to facilitate optimization of the end-to-end video production, delivery and consumption ecosystem to improve user experience at lowered cost. Although recent studies have significantly increased the accuracy of saliency prediction, the approaches are…

Cited by 0SourceScholar
2019

AGEM: Solving Linear Inverse Problems via Deep Priors and Sampling

NeurIPS 2019poster

In this paper we propose to use a denoising autoencoder (DAE) prior to simultaneously solve a linear inverse problem and estimate its noise parameter. Existing DAE-based methods estimate the noise parameter empirically or treat it as a tunable hyper-parameter. We instead propose autoencoder guided E…

2017

HEVC-based motion compensated joint temporal-spatial video denoising

ICASSP 2017accepted

A novel HEVC-based efficient video denoising algorithm is proposed in this paper. It uses a spatial Gaussian filter for the chrominance components and then utilizes the HEVC motion estimation process to find the best temporal correspondence for low-pass filtering. Other HEVC tools such as quantizati…

Cited by 0SourceScholar
2016

Novel 3D-WPP algorithms for parallel HEVC encoding

ICASSP 2016accepted

Although wavefront parallel processing (WPP) proposed in the HEVC standard and various inter frame WPP algorithms can achieve comparatively high parallelism, their scalability for its parallelism is still very limited due to various dependencies introduced in spatial and temporal prediction in HEVC.…

Cited by 0SourceScholar