← Search

Eslam Mohamed Bakr

6 accepted papers

2025

InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows

EMNLP 2025

Understanding long-form videos, such as movies and TV episodes ranging from tens of minutes to two hours, remains a significant challenge for multi-modal models. Existing benchmarks often fail to test the full range of cognitive skills needed to process these temporally rich and narratively complex

Cited by 0SourcePDFScholar
2025

Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description

ICCV 2025poster

In this paper, we introduce Part-Aware Point Grounded Description (PaPGD), a challenging task aimed at advancing 3D multimodal learning for fine-grained, part-aware segmentation grounding and detailed explanation of 3D objects. Existing 3D datasets largely focus on either vision-only part segmentati…

Cited by 0SourcePDFScholar
2025

ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger Bridge

ICLR 2025poster

Diffusion models break down the challenging task of generating data from high-dimensional distributions into a series of easier denoising steps. Inspired by this paradigm, we propose a novel approach that extends the diffusion framework into modality space, decomposing the complex task of RGB image…

2024

CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding

ICLR 2024poster

3D visual grounding is the ability to localize objects in 3D scenes conditioned by utterances. Most existing methods devote the referring head to localize the referred object directly, causing failure in complex scenarios. In addition, it does not illustrate how and why the network reaches the final…

2023

HRS-Bench: Holistic, Reliable and Scalable Benchmark for Text-to-Image Models

ICCV 2023poster

Designing robust text-to-image (T2I) models have been extensively explored in recent years, especially with the emergence of diffusion models, which achieves state-of-the-art results on T2I synthesis tasks. Despite the significant effort and success in this direction, we observed that the existing m…

Cited by 73PDFcodeScholar
2022

Look Around and Refer: 2D Synthetic Semantics Knowledge Distillation for 3D Visual Grounding

NeurIPS 2022accept

3D visual grounding task has been explored with visual and language streams to comprehend referential language for identifying targeted objects in 3D scenes. However, most existing methods devote the visual stream to capture the 3D visual clues using off-the-shelf point clouds encoders. The main que…