← Search

Jinming Duan

11 accepted papers

2025

Exploring Temporal Constraints for Unsupervised Iris Motion Tracking in AS-OCT Videos

ICASSP 2025accepted

Iris motion tracking is critical for discriminating the iris stiffness and developmental stage of primary angle-closure disease (PACD). Anterior segment optical coherence tomography (AS-OCT) video is a highly efficient approach to observe the morphological determinant in iris motion. However, the ir…

Cited by 0SourceScholar
2025

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

CVPR 2025poster

Despite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks. However, real-world videos encompass omni-modal information (vision, audio, and speech) with a series of events forming a cohesive storyline. The lack of multi-modal vide…

2025

SACB-Net: Spatial-awareness Convolutions for Medical Image Registration

CVPR 2025highlight

Deep learning-based image registration methods have shown state-of-the-art performance and rapid inference speeds. Despite these advances, many existing approaches fall short in capturing spatially varying information in non-local regions of feature maps due to the reliance on spatially-shared convo…

2024

Optimizing ADMM and Over-Relaxed ADMM Parameters for Linear Quadratic Problems

AAAI 2024technical

The Alternating Direction Method of Multipliers (ADMM) has gained significant attention across a broad spectrum of machine learning applications. Incorporating the over-relaxation technique shows potential for enhancing the convergence rate of ADMM. However, determining optimal algorithmic parameter…

Cited by 3SourcePDFScholar
2023

Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and Baseline

CVPR 2023poster

Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them. However, this setting is unrealistic as natural videos often contain numerous audio-visual events with different categories. To better adapt to real-life applications, in this…

2023

Fourier-Net: Fast Image Registration with Band-Limited Deformation

AAAI 2023technical

Unsupervised image registration commonly adopts U-Net style networks to predict dense displacement fields in the full-resolution spatial domain. For high-resolution volumetric image data, this process is however resource-intensive and time-consuming. To tackle this problem, we propose the Fourier-Ne…

2023

UniFace: Unified Cross-Entropy Loss for Deep Face Recognition

ICCV 2023poster

As a widely used loss function in deep face recognition, the softmax loss cannot guarantee that the minimum positive sample-to-class similarity is larger than the maximum negative sample-to-class similarity. As a result, no unified threshold is available to separate positive sample-to-class pairs fr…

Cited by 29PDFcodeScholar
2023

UniTSFace: Unified Threshold Integrated Sample-to-Sample Loss for Face Recognition

NeurIPS 2023poster

Sample-to-class-based face recognition models can not fully explore the cross-sample relationship among large amounts of facial images, while sample-to-sample-based models require sophisticated pairing processes for training. Furthermore, neither method satisfies the requirements of real-world face…

2021

FS-Net: Fast Shape-Based Network for Category-Level 6D Object Pose Estimation With Decoupled Rotation Mechanism

CVPR 2021poster

In this paper, we focus on category-level 6D pose and size estimation from a monocular RGB-D image. Previous methods suffer from inefficient category-level pose feature extraction, which leads to low accuracy and inference speed. To tackle this problem, we propose a fast shape-based network (FS-Net)…

Cited by 197PDFcodeScholar
2020

G2L-Net: Global to Local Network for Real-Time 6D Pose Estimation With Embedding Vector Features

CVPR 2020poster

In this paper, we propose a novel real-time 6D object pose estimation framework, named G2L-Net. Our network operates on point clouds from RGB-D detection in a divide-and-conquer fashion. Specifically, our network consists of three steps. First, we extract the coarse object point cloud from the RGB-D…

Cited by 133PDFcodeScholar
2020

Geometry Constrained Weakly Supervised Object Localization

ECCV 2020poster

We propose a geometry constrained network, termed GCNet, for weakly supervised object localization (WSOL). GC-Net consists of three modules: a detector, a generator and a classifier. The detector predicts the object location defined by a set of coefficients describing a geometric shape (i.e. ellipse or…