← Search

Liqi Yan

12 accepted papers

2026

AR-Nav Benchmark: Augmented Reality Navigation with Vision and Language

AAAI 2026technical

Augmented Reality (AR) navigation has emerged as a transformative tool for spatial intelligence, enabling users to interactively explore complex environments through wearable and mobile AR devices. However, current AR navigation systems struggle with low indoor localization accuracy, weak semantic u

Cited by 0SourcePDFScholar
2026

Chain-of-Search: Parameter-Efficient Reasoning for Zero-Shot Object Navigation

AAAI 2026technical

Zero-shot object navigation tasks agents with locating target objects in unseen environments—a core capability of embodied intelligence. While recent vision-language navigation methods leverage Large Language Models (LLMs) for multimodal reasoning, they suffer from two key limitations: (1) semantic

Cited by 0SourcePDFScholar
2026

NeuroMamba: A Universal Spatiotemporal Module for Robust Perception in Degraded Sensory Streams

ICML 2026poster

In open-world intelligent systems, processing continuous sensory streams disrupted by heterogeneous degradation sources presents a fundamental challenge: reconciling the inherent tension between observational completeness and reconstruction fidelity. Methods that prioritize completeness by bridging …

Cited by 0SourceScholar
2025

DroneSplat: 3D Gaussian Splatting for Robust 3D Reconstruction from In-the-Wild Drone Imagery

CVPR 2025highlight

Drones have become essential tools for reconstructing wild scenes due to their outstanding maneuverability. Recent advances in radiance field methods have achieved remarkable rendering quality, providing a new avenue for 3D reconstruction from drone imagery. However, dynamic distractors in wild env…

Cited by 2SourcePDFScholar
2025

Optimal Distributed Training With Co-Adaptive Data Parallelism in Heterogeneous Environments

IJCAI 2025

The computational power required for training deep learning models has been skyrocketing in the past decade as they scale with big data, and has become a very expensive and scarce resource. Therefore, distributed training, which can leverage distributed available computational power, is vital for ef

Cited by 0SourcePDFScholar
2024

Sparse Multi-Relational Graph Convolutional Network for Multi-type Object Trajectory Prediction

IJCAI 2024poster

Object trajectory prediction is a hot research issue with wide applications in video surveillance and autonomous driving. The previous studies consider the interaction sparsity mainly among the pedestrians instead of multi-type of objects, which brings new types of interactions and consequently supe…

Cited by 1SourcePDFScholar
2023

Prompt Learns Prompt: Exploring Knowledge-Aware Generative Prompt Collaboration For Video Captioning

IJCAI 2023poster

Fine-tuning large vision-language models is a challenging task. Prompt tuning approaches have been introduced to learn fixed textual or visual prompts while freezing the pre-trained model in downstream tasks. Despite the effectiveness of prompt tuning, what do those learnable prompts learn remains u…

Cited by 45SourcePDFScholar
2022

GL-RG: Global-Local Representation Granularity for Video Captioning

IJCAI 2022poster

Video captioning is a challenging task as it needs to accurately transform visual understanding into natural language description. To date, state-of-the-art methods inadequately model global-local representation across video frames for caption generation, leaving plenty of room for improve…

2021

DenserNet: Weakly Supervised Visual Localization Using Multi-Scale Feature Aggregation

AAAI 2021technical

In this work, we introduce a Denser Feature Network(DenserNet) for visual localization. Our work provides three principal contributions. First, we develop a convolutional neural network (CNN) architecture which aggregates feature maps at different semantic levels for image representations…

2020

Multimodal Aggregation Approach for Memory Vision-Voice Indoor Navigation with Meta-Learning

IROS 2020poster

Vision and voice are two vital keys for agents’ interaction and learning. In this paper, we present a novel indoor navigation model called Memory Vision-Voice Indoor Navigation (MVV-IN), which receives voice commands and analyzes multimodal information of visual observation in order to enhance robot…

Cited by 24SourceScholar