← Search

Mohammad Omama

6 accepted papers

2026

AsymLoc: Towards Asymmetric Feature Matching for Efficient Visual Localization

CVPR 2026

Precise and real-time visual localization is critical for applications like AR/VR and robotics, especially on resource-constrained edge devices such as smart glasses, where battery life and heat dissipation can be primary concerns. While many efficient models exist, further reducing compute without

Cited by 0SourceScholar
2025

Exploiting Distribution Constraints for Scalable and Efficient Image Retrieval

ICLR 2025poster

Image retrieval is crucial in robotics and computer vision, with downstream applications in robot place recognition and vision-based product recommendations. Modern retrieval systems face two key challenges: scalability and efficiency. State-of-the-art image retrieval systems train specific neural n…

Cited by 0SourcePDFScholar
2025

Neuro-Symbolic Evaluation of Text-to-Video Models using Formal Verification

CVPR 2025poster

Recent advancements in text-to-video models such as Sora, Gen-3, MovieGen, and CogVideoX are pushing the boundaries of synthetic video generation, with adoption seen in fields like robotics, autonomous driving, and entertainment. As these models become prevalent, various metrics and benchmarks have…

2025

R3DM: Enabling Role Discovery and Diversity Through Dynamics Models in Multi-agent Reinforcement Learning

ICML 2025poster

Multi-agent reinforcement learning (MARL) has achieved significant progress in large-scale traffic control, autonomous vehicles, and robotics. Drawing inspiration from biological systems where roles naturally emerge to enable coordination, role-based MARL methods have been proposed to enhance cooper…

2025

SparseLoc: Sparse Open-Set Landmark-based Global Localization for Autonomous Navigation

IROS 2025

Global localization is a critical problem in autonomous navigation, enabling precise positioning without reliance on GPS. Modern techniques often depend on dense LiDAR maps, which, while precise, require extensive storage and computational resources. Alternative approaches have explored sparse maps

Cited by 2SourceScholar
2024

Towards Neuro-Symbolic Video Understanding

ECCV 2024oral

"The unprecedented surge in video data production in recent years necessitates efficient tools to extract meaningful frames from videos for downstream tasks. Long-term temporal reasoning is a key desideratum for frame retrieval systems. While state-of-the-art foundation models, like VideoLLaMA and V…