← Search

Fengyu Zhou

5 accepted papers

2025

Semi-Supervised Language-Conditioned Grasping With Curriculum-Scheduled Augmentation and Geometric Consistency

RA-L 2025

Language-Conditioned Grasping (LCG) is an essential skill for robotic manipulation and has attracted increasing interest. Recent LCG models have made great progress, but need numerous paired image-text-pose annotations for fully supervised learning, which are tedious and expensive. Semi-supervised l

Cited by 1SourceScholar
2024

Hierarchical Multi-Modal Fusion for Language-Conditioned Robotic Grasping Detection in Clutter

RA-L 2024

This letter concentrates on the challenging task of language-conditioned grasping detection in clutter, where the grasping postures of objects should be generated for robots according to complicated human instructions. Existing methods typically employ well-trained object detectors and leverage lang

Cited by 4SourceScholar
2024

RP-SG: Relation Prediction in 3D Scene Graphs for Unobserved Objects Localization

RA-L 2024

The ability to search for objects is a fundamental prerequisite for mobile robots when addressing a wide range of automation tasks. However, how to effectively estimate the positions of unobserved objects in a continuously changing environment remains an open challenge. Previous works have utilized

Cited by 3SourceScholar
2023

InteMATs: Integrating Granularity-Specific Multilingual Adapters for Cross-Lingual Transfer

EMNLP 2023long findings

Multilingual language models (MLLMs) have achieved remarkable success in various cross-lingual transfer tasks. However, they suffer poor performance in zero-shot low-resource languages, particularly when dealing with longer contexts. Existing research mainly relies on full-model fine-tuning on large…

Cited by 0SourceScholar
2020

A Bottom-up Framework for Construction of Structured Semantic 3D Scene Graph

IROS 2020poster

For high-level human-robot interaction tasks, 3D scene understanding is important and non-trivial for autonomous robots. However, parsing and utilizing effective environment information of the 3D scene is not trivial due to the complexity of the 3D environment and the limited ability for reasoning a…

Cited by 7SourceScholar