← Search

Mohammad Altillawi

8 accepted papers

2026

CoVAR: Co-Generation of Video and Action for Robotic Manipulation Via Multi-Modal Diffusion

ICRA 2026poster

We present a method to generate video–action pairs that follow text instructions, starting from an initial image observation and the robot’s joint states. Our approach automatically provides action labels for video diffusion mod- els, overcoming the common lack of action annotations and enabling the…

2026

DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos

ICRA 2026poster

Video diffusion models provide powerful real-world simulators for embodied AI but remain limited in controllability for robotic manipulation. Recent works on trajectory-conditioned video generation address this gap but often rely on 2D trajectories or single modality conditioning, which restricts th…

2026

VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents

CVPR 2026

Recent progress in video-to-video (V2V) translation has enabled realistic resimulation of embodied AI demonstrations, a capability that allows pretrained robot policies to be transferable to new environments without additional data collection. However, prior works can only operate on a single view a

Cited by 0SourceScholar
2025

RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation

IROS 2025

We address the problem of generating long-horizon videos for robotic manipulation tasks. Text-to-video diffusion models have made significant progress in photorealism, language understanding, and motion generation but struggle with long-horizon robotic tasks. Recent works use video diffusion models

Cited by 11SourceScholar
2025

RoboSwap: A GAN-driven Video Diffusion Framework For Unsupervised Robot Arm Swapping

IROS 2025

Recent advancements in generative models have revolutionized video synthesis and editing. However, the scarcity of diverse, high-quality datasets continues to hinder video-conditioned robotic learning, limiting cross-platform generalization. In this work, we address the challenge of swapping a robot

Cited by 1SourceScholar
2024

Implicit Learning of Scene Geometry From Poses for Global Localization

RA-L 2024

Global visual localization estimates the absolute pose of a camera using a single image, in a previously mapped area. Obtaining the pose from a single image enables many robotics and augmented/virtual reality applications. Inspired by latest advances in deep learning, many existing approaches direct

Cited by 3SourceScholar
2023

Global Localization: Utilizing Relative Spatio-Temporal Geometric Constraints from Adjacent and Distant Cameras

IROS 2023poster

Re-Iocalizing a camera from a single image in a previously mapped area is vital for many computer vision applications in robotics and augmented/virtual reality. In this work, we address the problem of estimating the 6 DoF camera pose relative to a global frame from a single image. We propose to leve…

Cited by 1SourceScholar