← Search

Shuquan Ye

5 accepted papers

2026

OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene Understanding

AAAI 2026technical

Open-vocabulary 3D scene understanding (OV-3D) aims to localize and classify novel objects beyond the closed set of object classes. However, existing approaches and benchmarks primarily focus on the open vocabulary problem within the context of object classes, which is insufficient in providing a ho

Cited by 0SourcePDFScholar
2025

Language-Guided Salient Object Ranking

CVPR 2025poster

Salient Object Ranking (SOR) aims to study human attention shifts across different objects in the scene. It is a challenging task, as it requires comprehension of the relations among the salient objects in the scene. However, existing works often overlook such relations or model them implicitly. In…

Cited by 0SourcePDFScholar
2025

Leveraging RGB-D Data with Cross-Modal Context Mining for Glass Surface Detection

AAAI 2025technical

Glass surfaces are becoming increasingly ubiquitous as modern buildings tend to use a lot of glass panels. This, however, poses substantial challenges to the operations of autonomous systems such as robots, self-driving cars, and drones, as the glass panels can become transparent obstacles to naviga…

Cited by 1SourcePDFScholar
2023

Improving Commonsense in Vision-Language Models via Knowledge Graph Riddles

CVPR 2023highlight

This paper focuses on analyzing and improving the commonsense ability of recent popular vision-language (VL) models. Despite the great success, we observe that existing VL-models still lack commonsense knowledge/reasoning ability (e.g., "Lemons are sour"), which is a vital component towards artifici…