← Search

Yihao Wu

5 accepted papers

2026

AR-Nav Benchmark: Augmented Reality Navigation with Vision and Language

AAAI 2026technical

Augmented Reality (AR) navigation has emerged as a transformative tool for spatial intelligence, enabling users to interactively explore complex environments through wearable and mobile AR devices. However, current AR navigation systems struggle with low indoor localization accuracy, weak semantic u

Cited by 0SourcePDFScholar
2026

EVALUATING BIAS IN SPOKEN DIALOGUE LLMS FOR REAL-WORLD DECISIONS AND RECOMMENDATIONS

ICASSP 2026poster

While biases in large language models (LLMs), such as stereotypes and cultural tendencies in outputs, have been examined and identified, their presence and characteristics in spoken dialogue models (SDMs) with audio input and output remain largely unexplored. Paralinguistic features, such as age, ge…

Cited by 0SourcePDFScholar
2025

MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

NeurIPS 2025poster

We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through ite…

Cited by 0SourcecodeScholar
2024

Autonomous Vision-Guided Two-Arm Collaborative Microassembly Using Learned Manipulation Model

RA-L 2024

This letter presents an integrated micromanipulation system suited for precise assembly of micro-parts in intricate environments. By embedding the Global Attention Mechanism (GAM) into YOLOv8, the system not only enhances its performance but also accurately identifies target keypoints, pinpointing t

Cited by 6SourceScholar