← Search

Qilin Zhang

10 accepted papers

2026

VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models

CVPR 2026

Text-to-point-cloud (T2P) localization aims to infer precise spatial positions within 3D point cloud maps from natural language descriptions, reflecting how humans perceive and communicate spatial layouts through language. However, existing methods largely rely on shallow text-point cloud correspond

Cited by 0SourcecodeScholar
2025

Enhancing Tactile Sensing in Robotics Using Null-Space Diffusion Model with EIT-based Sensors

IROS 2025

Robotic tactile sensors based on Electrical Impedance Tomography (EIT) have gained great attention in robotic sensing applications due to their features such as no internal wiring, "all-in-one" structure, and continuous sensing capabilities. However, the effectiveness of EIT-based tactile sensors is

Cited by 0SourceScholar
2024

Stepwise Multi-grained Boundary Detector for Point-supervised Temporal Action Localization

ECCV 2024poster

"Point-supervised temporal action localization pursues high-accuracy action detection under low-cost data annotation. Despite recent advances, a significant challenge remains: sparse labeling of individual frames leads to semantic ambiguity in determining action boundaries due to the lack of continu…

Cited by 0SourcePDFScholar
2023

3D Point Cloud Completion Based on Multi-Scale Degradation

ICASSP 2023accepted

Recent advances in 3D point cloud completion adopt unsupervised deep learning-based methods, which does not rely on labeled data and improves generalization ability. However, existing methods tend to focus more on the generation overall shape rather than detailed structure. To explore unsupervised 3…

Cited by 0SourceScholar
2021

ACSNet: Action-Context Separation Network for Weakly Supervised Temporal Action Localization

AAAI 2021technical

The object of Weakly-supervised Temporal Action Localization (WS-TAL) is to localize all action instances in an untrimmed video with only video-level supervision. Due to the lack of frame-level annotations during training, current WS-TAL methods rely on attention mechanisms to localize the foregroun…

Cited by 86SourcePDFScholar
2021

Practical Relative Order Attack in Deep Ranking

ICCV 2021poster

Recent studies unveil the vulnerabilities of deep ranking models, where an imperceptible perturbation can trigger dramatic changes in the ranking result. While previous attempts focus on manipulating absolute ranks of certain candidates, the possibility of adjusting their relative order remains unde…

Cited by 21PDFcodeScholar
2020

Two-Stream Consensus Network for Weakly-Supervised Temporal Action Localization

ECCV 2020poster

Weakly-supervised Temporal Action Localization (W-TAL) aims to classify and localize all action instances in an untrimmed video under only video-level supervision. However, without frame-level annotations, it is challenging for W-TAL methods to identify false positive action proposals and generate a…

2019

Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation Networks

ICCV 2019poster

Weakly-supervised temporal action localization (WS-TAL) is a promising but challenging task with only video-level action categorical labels available during training. Without requiring temporal action boundary annotations in training data, WS-TAL could possibly exploit automatically retrieved video…

Cited by 135PDFcodeScholar