← Search

Chaoran Zhu

3 accepted papers

2025

LaVA-Man: Learning Visual Action Representations for Robot Manipulation

CoRL 2025poster

Visual-textual understanding is essential for language-guided robot manipulation. Recent works leverage pre-trained vision-language models to measure the similarity between encoded visual observations and textual instructions, and then train a model to map this similarity to robot actions. However,…

Cited by 0SourceScholar
2022

Improving Generalization of Deep Networks for Estimating Physical Properties of Containers and Fillings

ICASSP 2022accepted

We present methods to estimate the physical properties of house-hold containers and their fillings manipulated by humans. We use a lightweight, pre-trained convolutional neural network with coordinate attention as a backbone model of the pipelines to accurately locate the object of interest and esti…

Cited by 0SourceScholar