← Search

Yik Lung Pang

5 accepted papers

2026

CHAIN-OF-CAPTION: TRAINING-FREE IMPROVEMENT OF MULTIMODAL LARGE LANGUAGE MODEL ON REFERRING EXPRESSION COMPREHENSION

ICASSP 2026poster

Given a textual description, the task of referring expression comprehension (REC) involves the localisation of the referred object in an image. Multimodal large language models (MLLMs) have achieved high accuracy on REC benchmarks through scaling up the model size and training data. Moreover, the pe…

Cited by 0SourcePDFScholar
2025

LaVA-Man: Learning Visual Action Representations for Robot Manipulation

CoRL 2025poster

Visual-textual understanding is essential for language-guided robot manipulation. Recent works leverage pre-trained vision-language models to measure the similarity between encoded visual observations and textual instructions, and then train a model to map this similarity to robot actions. However,…

Cited by 0SourceScholar
2022

Audio-Visual Object Classification for Human-Robot Collaboration

ICASSP 2022accepted

Human-robot collaboration requires the contactless estimation of the physical properties of containers manipulated by a person, for example while pouring content in a cup or moving a food box. Acoustic and visual signals can be used to estimate the physical properties of such objects, which may vary…

Cited by 0SourceScholar