← Search

Yabing Wang

8 accepted papers

2026

HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses Through Reasoning MLLMs

AAAI 2026technical

While Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation frameworks for human-centered scenarios, encompassing both the understanding of complex human intentions and the provision of e

Cited by 0SourcePDFScholar
2026

Spatial Matters: Position-Guided 3D Referring Expression Segmentation

CVPR 2026

3D Referring Expression segmentation (3D-RES) is an emerging field that segments 3D objects in point cloud scenes based on given referring expressions. Although existing methods have achieved substantial progress, they primarily focus on semantic cues and often overlook spatial relations, which are

Cited by 0SourcecodeScholar
2025

Diversifying Query: Region-Guided Transformer for Temporal Sentence Grounding

AAAI 2025technical

Temporal sentence grounding is a challenging task that aims to localize the moment spans relevant to a language description. Although recent DETR-based models have achieved notable progress by leveraging multiple learnable moment queries, they suffer from overlapped and redundant proposals, leading…

2025

Moment Quantization for Video Temporal Grounding

ICCV 2025poster

Video temporal grounding is a critical video understanding task, which aims to localize moments relevant to a language description. The challenge of this task lies in distinguishing relevant and irrelevant moments. Previous methods focused on learning continuous features exhibit weak differentiation…

2025

RefDetector: A Simple Yet Effective Matching-based Method for Referring Expression Comprehension

AAAI 2025technical

Despite the rapid and substantial advancements in object detection, it continues to face limitations imposed by pre-defined category sets. Current methods for visual grounding primarily focus on how to better leverage the visual backbone to generate text-tailored visual features, which may require a…

Cited by 0SourcePDFScholar
2025

Towards Precise Embodied Dialogue Localization via Causality Guided Diffusion

CVPR 2025poster

Embodied localization based on vision and natural language dialogues presents a persistent challenge in embodied intelligence. Existing methods often approach this task as an image translation problem, leveraging encoder-decoder architectures to predict heatmaps. However, these methods frequently ex…

Cited by 0SourcePDFScholar
2024

CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge Transfer

AAAI 2024technical

Cross-lingual cross-modal retrieval has garnered increasing attention recently, which aims to achieve the alignment between vision and target language (V-T) without using any annotated V-T data pairs. Current methods employ machine translation (MT) to construct pseudo-parallel data pairs, which are…

Cited by 11SourcePDFScholar
2024

Referencing Where to Focus: Improving Visual Grounding with Referential Query

NeurIPS 2024poster

Visual Grounding aims to localize the referring object in an image given a natural language expression. Recent advancements in DETR-based visual grounding methods have attracted considerable attention, as they directly predict the coordinates of the target object without relying on additional effort…

Cited by 1SourcePDFScholar