← Search

Chongyang Li

3 accepted papers

2026

F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model

AAAI 2026technical

Traditional dialogue retrieval aims to select the most appropriate utterance or image from recent dialogue history. However, they often fail to meet users’ actual needs for revisiting semantically coherent content scattered across long-form conversations. To fill this gap, we define the Fine-grained

Cited by 0SourcePDFScholar
2026

Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants

ICASSP 2026oral

Approximately 283 million people worldwide live with visual impairments, motivating increasing research into leveraging Visual Language Models (VLMs) to develop effective walking assistance systems for blind and low vision individuals. However, existing VLMs in walking assistant task often have outp…

Cited by 0SourcePDFScholar
2021

A Model-Free Synchronous Control of Humanoid Robot Finger

ICRA 2021poster

For a multi-fingered robot hand, the individual control over single joints cannot guarantee their fine collaboration. For achieving a high-precision synchronization, a theory of synchronous control is introduced to multi-fingered robot hands. This paper introduced a new model-free and cross-coupling…

Cited by 2SourceScholar