← Search

Hamidreza Kasaei

22 accepted papers

2026

Co-NavGPT: Multi-Robot Cooperative Visual Semantic Navigation Using Vision Language Models

ICRA 2026poster

Visual target navigation is a critical capability for autonomous robots operating in unknown environments, particularly in human-robot interaction scenarios. While classical and learning-based methods have shown promise, most existing approaches lack common-sense reasoning and are typically designed…

2026

Co-NavGPT: Multirobot Cooperative Visual Semantic Navigation Using Vision Language Models

RA-L 2026

Visual target navigation is a critical capability for autonomous robots operating in unknown environments, particularly in human-robot interaction scenarios. While classical and learning-based methods have shown promise, most existing approaches lack common-sense reasoning and are typically designed

Cited by 4SourceScholar
2025

Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning

NeurIPS 2025poster

Learning semantically meaningful representations from unstructured 3D point clouds remains a central challenge in computer vision, especially in the absence of large-scale labeled datasets. While masked point modeling (MPM) is widely used in self-supervised 3D learning, its reconstruction-based obje…

Cited by 0SourcecodeScholar
2024

Harnessing the Synergy between Pushing, Grasping, and Throwing to Enhance Object Manipulation in Cluttered Scenarios

ICRA 2024poster

In this work, we delve into the intricate synergy among non-prehensile actions like pushing, and prehensile actions such as grasping and throwing, within the domain of robotic manipulation. We introduce an innovative approach to learning these synergies by leveraging model-free deep reinforcement le…

Cited by 1SourceScholar
2024

Lifelong Robot Library Learning: Bootstrapping Composable and Generalizable Skills for Embodied Control with Language Models

ICRA 2024poster

Large Language Models (LLMs) have emerged as a new paradigm for embodied reasoning and control, most recently by generating robot policy code that utilizes a custom library of vision and control primitive skills. However, prior arts fix their skills library and steer the LLM with carefully handcraft…

Cited by 8SourcecodeScholar
2024

Self-supervised Learning for Joint Pushing and Grasping Policies in Highly Cluttered Environments

ICRA 2024poster

Robotic systems often face challenges when attempting to grasp a target object due to interference from surrounding items. We propose a Deep Reinforcement Learning (DRL) method that develops joint policies for grasping and pushing, enabling effective manipulation of target objects within untrained,…

Cited by 13SourcecodeScholar
2024

SoftManiSim: A Fast Simulation Framework for Multi-Segment Continuum Manipulators Tailored for Robot Learning

CoRL 2024poster

This paper introduces SoftManiSim, a novel simulation framework for multi-segment continuum manipulators. Existing continuum robot simulators often rely on simplifying assumptions, such as constant curvature bending or ignoring contact forces, to meet real-time simulation and training demands. To br…

Cited by 1SourcecodeScholar
2024

TiV-ODE: A Neural ODE-based Approach for Controllable Video Generation From Text-Image Pairs

ICRA 2024poster

Videos capture the evolution of continuous dynamical systems over time in the form of discrete image sequences. Recently, video generation models have been widely used in robotic research. However, generating controllable videos from image-text pairs is an important yet underexplored research topic…

Cited by 0SourceScholar
2023

Early or Late Fusion Matters: Efficient RGB-D Fusion in Vision Transformers for 3D Object Recognition

IROS 2023poster

The Vision Transformer (ViT) architecture has established its place in computer vision literature, however, training ViTs for RGB-D object recognition remains an understudied topic, viewed in recent literature only through the lens of multi-task pretraining in multiple vision modalities. Such approa…

Cited by 16SourceScholar
2023

Enhancing Fine-Grained 3D Object Recognition Using Hybrid Multi-Modal Vision Transformer-CNN Models

IROS 2023poster

Robots operating in human-centered environments, such as retail stores, restaurants, and households, are often required to distinguish between similar objects in different contexts with a high degree of accuracy. However, fine-grained object recognition remains a challenge in robotics due to the hig…

Cited by 9SourcecodeScholar
2023

Language-guided Robot Grasping: CLIP-based Referring Grasp Synthesis in Clutter

CoRL 2023poster

Robots operating in human-centric environments require the integration of visual grounding and grasping capabilities to effectively manipulate objects based on user instructions. This work focuses on the task of referring grasp synthesis, which predicts a grasp pose for an object referred through na…

Cited by 27SourcecodeScholar