← Search

Na Jiang

8 accepted papers

2026

Learning to Diversify and Focus: A Reinforcement Framework for Open-Vocabulary HOI Detection

CVPR 2026

Open-Vocabulary Human-Object Interaction (OV-HOI) detection aims to recognize novel HOI categories beyond the training set. Existing OV-HOI detection approaches typically leverage CLIP to extract global visual representations and perform cross-attention between learnable queries and global features

Cited by 0SourceScholar
2025

3D Lane Detection Based on Projection-Consistent Reference Points and Intra- & Inter-lane Context

ICRA 2025

3D lane detection aims to identify lane categories and trends in 3D space, which is a vital and challenging task in autonomous driving. Existing methods introduce various priors to guide 3D lane prediction, which generally consist of a series of reference points for context aggregation. However, due

Cited by 1SourceScholar
2025

ADC-GS: Pose-Free 3D Gaussian Splatting with Adaptive Depth Consistency

ICASSP 2025accepted

Recently proposed 3D Gaussian Splatting (3DGS) has achieved state-of-the-art results in the fields of novel view synthesis, but it heavily relies on pre-computed camera poses. Although recent methods mitigate by leveraging explicit representations achieve novel view synthesis without requiring camer…

Cited by 0SourceScholar
2024

AHRNET: Attention and Heatmap-Based Regressor for Hand Pose Estimation and Mesh Recovery

ICASSP 2024accepted

Estimating 3D hand pose and recovering the full hand surface mesh from a single RGB image is a challenging task due to self-occlusions, viewpoint changes, and the complexity of hand articulations. In this paper, we propose a novel framework that combines an attention mechanism with heatmap regressio…

Cited by 0SourceScholar
2023

Exploiting 3D Human Recovery for Action Recognition with Spatio-Temporal Bifurcation Fusion

ICASSP 2023accepted

Action recognition utilizes information in images or videos to analyze and classify human behaviors. The existing methods usually exploit 2D pose to improve classification features. Due to the lack of 3D cues, some approximate behaviors in 2D perspective cannot be recognized. In this paper, we propo…

Cited by 0SourceScholar
2023

Hankel Structured Low Rank and Sparse Representation Via L0-Norm Optimization for Compressed Ultrasound Plane Wave Signal Reconstruction

ICASSP 2023accepted

Ultrasound plane wave imaging is widely used in many applications thanks to its capability in reaching high frame rates. However, the amount of data acquisition and storage in a period of time can become a bottleneck in ultrasound system design for thousands frames per second. In our previous study,…

Cited by 0SourceScholar
2020

Co-Saliency Spatio-Temporal Interaction Network for Person Re-Identification in Videos

IJCAI 2020poster

Person re-identification aims at identifying a certain pedestrian across non-overlapping camera networks. Video-based person re-identification approaches have gained significant attention recently, expanding image-based approaches by learning features from multiple frames. In this work, we propose a…

Cited by 0SourcePDFScholar
2019

Multi-scale Vehicle Re-identification Using Self-adapting Label Smoothing Regularization

ICASSP 2019accepted

Vehicle re-identification (re-id) plays an important role in intelligent surveillance. Since difference vehicle models may have similar appearances, together with the problem of image scale variations, the vehicle re-id remains long-term challenging. We present a novel multi-scale vehicle re-id fram…

Cited by 0SourceScholar