← Search

Mahmoud Ahmed

4 accepted papers

2025

InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows

EMNLP 2025

Understanding long-form videos, such as movies and TV episodes ranging from tens of minutes to two hours, remains a significant challenge for multi-modal models. Existing benchmarks often fail to test the full range of cognitive skills needed to process these temporally rich and narratively complex

Cited by 0SourcePDFScholar
2025

Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description

ICCV 2025poster

In this paper, we introduce Part-Aware Point Grounded Description (PaPGD), a challenging task aimed at advancing 3D multimodal learning for fine-grained, part-aware segmentation grounding and detailed explanation of 3D objects. Existing 3D datasets largely focus on either vision-only part segmentati…

Cited by 0SourcePDFScholar
2024

3DCoMPaT200: Language Grounded Large-Scale 3D Vision Dataset for Compositional Recognition

NeurIPS 2024poster

Understanding objects in 3D at the part level is essential for humans and robots to navigate and interact with the environment. Current datasets for part-level 3D object understanding encompass a limited range of categories. For instance, the ShapeNet-Part and PartNet datasets only include 16, and 2…

2024

CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding

ICLR 2024poster

3D visual grounding is the ability to localize objects in 3D scenes conditioned by utterances. Most existing methods devote the referring head to localize the referred object directly, causing failure in complex scenarios. In addition, it does not illustrate how and why the network reaches the final…