← Search

Beining Xu

3 accepted papers

2026

IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding

CVPR 2026

Recent advances in vision-language models (VLMs) have significantly enhanced the visual grounding task, which involves locating objects in an image based on natural language queries. Despite these advancements, the security of VLM-based grounding systems has not been thoroughly investigated. This pa

Cited by 0SourcecodeScholar
2026

S$^2$-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance

CVPR 2026

3D Visual Grounding (3DVG) focuses on locating objects in 3D scenes based on natural language descriptions, serving as a fundamental task for embodied AI and robotics. Recent advances in Multi-modal Large Language Models (MLLMs) have motivated research into extending them to 3DVG. However, MLLMs pri

Cited by 0SourcecodeScholar
2025

SGLoc: Semantic Localization System for Camera Pose Estimation from 3D Gaussian Splatting Representation

IROS 2025

We propose SGLoc, a novel localization system that directly regresses camera poses from 3D Gaussian Splatting (3DGS) representation by leveraging semantic information. Our method utilizes the semantic relationship between 2D image and 3D scene representation to estimate the 6DoF pose without prior p

Cited by 1SourcecodeScholar