← Search

Qianli Xu

9 accepted papers

2025

DOTA: Distributional Test-time Adaptation of Vision-Language Models

NeurIPS 2025poster

Vision-language foundation models (VLMs), such as CLIP, exhibit remarkable performance across a wide range of tasks. However, deploying these models can be unreliable when significant distribution gaps exist between training and test data, while fine-tuning for diverse scenarios is often costly. Cac…

Cited by 0SourceScholar
2025

SPASCA: Social Presence and Support with Conversational Agent for Persons Living with Dementia

AAAI 2025technical

We present SPASCA - a conversational AI system that promotes psychological and cognitive well-being of persons living with dementia (PLWD). This system features an AI agent that provides social presence and support to PLWD through verbal communications, without physical presence of human caregivers.…

Cited by 0SourcePDFScholar
2025

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting

AAAI 2025technical

Large Language Model (LLM)-based agents have shown promise in procedural tasks, but the potential of multimodal instructions augmented by texts and videos to assist users remains under-explored. To address this gap, we propose the Visually Grounded Text-Video Prompting (VG-TVP) method which is a nov…

2024

VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation

NeurIPS 2024poster

A well-known dilemma in large vision-language models (e.g., GPT-4, LLaVA) is that while increasing the number of vision tokens generally enhances visual understanding, it also significantly raises memory and computational costs, especially in long-term, dense video frame streaming scenarios. Althoug…

2023

GazeVQA: A Video Question Answering Dataset for Multiview Eye-Gaze Task-Oriented Collaborations

EMNLP 2023long main

The usage of exocentric and egocentric videos in Video Question Answering (VQA) is a new endeavor in human-robot interaction and collaboration studies. Particularly for egocentric videos, one may leverage eye-gaze information to understand human intentions during the task. In this paper, we build a…

Cited by 0SourceScholar
2023

Visuo-Tactile Feedback-Based Robot Manipulation for Object Packing

RA-L 2023

Robots are increasingly expected to manipulate objects, of which properties have high perceptual uncertainty from any single sensory modality. This directly impacts successful object manipulation. Object packing is one of the challenging tasks in robot manipulation. In this work, a new visuo-tactile

Cited by 23SourceScholar
2021

Predicting Event Memorability from Contextual Visual Semantics

NeurIPS 2021poster

Episodic event memory is a key component of human cognition. Predicting event memorability,i.e., to what extent an event is recalled, is a tough challenge in memory research and has profound implications for artificial intelligence. In this study, we investigate factors that affect event memorabilit…

2021

Towards Efficient Multiview Object Detection with Adaptive Action Prediction

ICRA 2021poster

Active vision is a desirable perceptual feature for robots. Existing approaches usually make strong assumptions about the task and environment, thus are less robust and efficient. This study proposes an adaptive view planning approach to boost the efficiency and robustness of active object detection…

Cited by 9SourceScholar