← Search

Jiaying Lin

17 accepted papers

2026

OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene Understanding

AAAI 2026technical

Open-vocabulary 3D scene understanding (OV-3D) aims to localize and classify novel objects beyond the closed set of object classes. However, existing approaches and benchmarks primarily focus on the open vocabulary problem within the context of object classes, which is insufficient in providing a ho

Cited by 0SourcePDFScholar
2025

Hierarchical Cross-Modal Alignment for Open-Vocabulary 3D Object Detection

AAAI 2025technical

Open-vocabulary 3D object detection (OV-3DOD) aims at localizing and classifying novel objects beyond closed sets. The recent success of vision-language models (VLMs) has demonstrated their remarkable capabilities to understand open vocabularies. Existing works that leverage VLMs for 3D object detec…

Cited by 0SourcePDFScholar
2025

HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented Generation

ACL 2025long

While Retrieval-Augmented Generation (RAG) has emerged as an effective approach for addressing the knowledge outdating problem in Large Language Models (LLMs), it still faces a critical challenge: the prevalence of outdated information in knowledge bases. Current research primarily focuses on incorp…

2025

Leveraging RGB-D Data with Cross-Modal Context Mining for Glass Surface Detection

AAAI 2025technical

Glass surfaces are becoming increasingly ubiquitous as modern buildings tend to use a lot of glass panels. This, however, poses substantial challenges to the operations of autonomous systems such as robots, self-driving cars, and drones, as the glass panels can become transparent obstacles to naviga…

Cited by 1SourcePDFScholar
2024

Boosting Weakly Supervised Referring Image Segmentation via Progressive Comprehension

NeurIPS 2024poster

This paper explores the weakly-supervised referring image segmentation (WRIS) problem, and focuses on a challenging setup where target localization is learned directly from image-text pairs. We note that the input text description typically already contains detailed information on how to localize t…

Cited by 2SourcePDFScholar
2024

Effective Video Mirror Detection with Inconsistent Motion Cues

CVPR 2024poster

Image-based mirror detection has recently undergone rapid research due to its significance in applications such as robotic navigation semantic segmentation and scene reconstruction. Recently VMD-Net was proposed as the first video mirror detection technique by modeling dual correspondences between t…

2024

Multi-View Dynamic Reflection Prior for Video Glass Surface Detection

AAAI 2024technical

Recent research has shown significant interest in image-based glass surface detection (GSD). However, detecting glass surfaces in dynamic scenes remains largely unexplored due to the lack of a high-quality dataset and an effective video glass surface detection (VGSD) method. In this paper, we propos…

2023

Efficient Mirror Detection via Multi-Level Heterogeneous Learning

AAAI 2023technical

We present HetNet (Multi-level Heterogeneous Network), a highly efficient mirror detection network. Current mirror detection methods focus more on performance than efficiency, limiting the real-time applications (such as drones). Their lack of efficiency is aroused by the common design of adopting h…

2023

Symmetry-Aware Transformer-Based Mirror Detection

AAAI 2023technical

Mirror detection aims to identify the mirror regions in the given input image. Existing works mainly focus on integrating the semantic features and structural features to mine specific relations between mirror and non-mirror regions, or introducing mirror properties like depth or chirality to help a…

2023

Weakly-Supervised Camouflaged Object Detection with Scribble Annotations

AAAI 2023technical

Existing camouflaged object detection (COD) methods rely heavily on large-scale datasets with pixel-wise annotations. However, due to the ambiguous boundary, annotating camouflage objects pixel-wisely is very time-consuming and labor-intensive, taking ~60mins to label one image. In this paper, we pr…