← Search

Hossein Nourkhiz Mahjoub

7 accepted papers

2026

MonoVLM: Monocular 3D Visual Grounding with Vision Language Models

CVPR 2026

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in instruction following and 2D visual understanding. However, state-of-the-art VLMs, including GPT-5, still struggle with 3D perception, particularly in tasks such as monocular 3D visual grounding. While specialized vision-only

Cited by 0SourcecodeScholar
2026

Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models

CVPR 2026

The recent success of reinforcement learning (RL) in large reasoning models has inspired the growing adoption of RL for post-training Multimodal Large Language Models (MLLMs) to enhance their visual reasoning capabilities. Although many studies have reported improved performance, it remains unclear

Cited by 0SourceScholar
2024

Active Learning with Dual Model Predictive Path-Integral Control for Interaction-Aware Autonomous Highway On-ramp Merging

ICRA 2024poster

Merging into dense highway traffic for an autonomous vehicle is a complex decision-making task, wherein the vehicle must identify a potential gap and coordinate with surrounding human drivers, each of whom may exhibit diverse driving behaviors. Many existing methods consider other drivers to be dyna…

Cited by 4SourceScholar
2024

Language Grounded Multi-agent Reinforcement Learning with Human-interpretable Communication

NeurIPS 2024poster

Multi-Agent Reinforcement Learning (MARL) methods have shown promise in enabling agents to learn a shared communication protocol from scratch and accomplish challenging team tasks. However, the learned language is usually not interpretable to humans or other agents not co-trained together, limiting…

Cited by 6SourcePDFScholar
2024

Multi-Robot Cooperative Navigation in Crowds: A Game-Theoretic Learning-Based Model Predictive Control Approach

ICRA 2024poster

In this paper, we develop a control framework for the coordination of multiple robots as they navigate through crowded environments. Our framework comprises of a local model predictive control (MPC) for each robot and a social long short-term memory model that forecasts pedestrians’ trajectories. We…

Cited by 8SourceScholar
2024

Social Navigation in Crowded Environments with Model Predictive Control and Deep Learning-Based Human Trajectory Prediction

IROS 2024poster

Navigating a robot among a crowd has received increasing attention from researchers over the last few decades, resulting in the emergence of numerous approaches aimed at addressing the problem of social navigation to date. Our proposed approach couples agent motion prediction and planning to avoid t…

Cited by 3SourceScholar
2023

Task-aware Distributed Source Coding under Dynamic Bandwidth

NeurIPS 2023poster

Efficient compression of correlated data is essential to minimize communication overload in multi-sensor networks. In such networks, each sensor independently compresses the data and transmits them to a central node. A decoder at the central node decompresses and passes the data to a pre-trained mac…