← Search

Mohamad Alansari

3 accepted papers

2026

SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs

CVPR 2026

Multimodal large language models (MLLMs) have advanced from image-level reasoning to pixel-level grounding, but extending these capabilities to videos remains challenging as models must achieve spatial precision and temporally consistent reference tracking. Existing video MLLMs often rely on a stati

Cited by 0SourcecodeScholar
2025

STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection

CVPR 2025highlight

Advancements in Computer-Aided Screening (CAS) systems are essential for improving the detection of security threats in X-ray baggage scans. However, current datasets are limited in representing real-world, sophisticated threats and concealment tactics, and existing approaches are constrained by a c…

2024

PathFormer: A Transformer-Based Framework for Vision-Centric Autonomous Navigation in Off-Road Environments

IROS 2024poster

The efficient navigation of autonomous vehicles across rugged and unstructured terrains remains a significant challenge. Most existing research in this area emphasizes the need for complex mappings or intricate multi-step methodologies. However, these traditional approaches often struggle to adapt t…

Cited by 1SourceScholar