← Search

Abdelrahman Eldesokey

8 accepted papers

2026

FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations

ICML 2026poster

We introduce FloorplanQA, a diagnostic benchmark for evaluating spatial reasoning in large-language models (LLMs). FloorplanQA is grounded in structured representations of indoor scenes (e.g., kitchens, living rooms, bedrooms, bathrooms, and others), encoded symbolically in JSON or XML layouts. The …

Cited by 0SourceScholar
2025

Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image Generation

ICLR 2025poster

We propose a diffusion-based approach for Text-to-Image (T2I) generation with interactive 3D layout control. Layout control has been widely studied to alleviate the shortcomings of T2I diffusion models in understanding objects' placement and relationships from text descriptions. Nevertheless, existi…

2025

EditCLIP: Representation Learning for Image Editing

ICCV 2025poster

We introduce EditCLIP, a novel representation-learning approach for image editing. Our method learns a unified representation of edits by jointly encoding an input image and its edited counterpart, effectively capturing their transformation. To evaluate its effectiveness, we employ EditCLIP to solve…

2025

Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation

NeurIPS 2025spotlight

We propose a novel approach for disentangling visual and semantic features from the backbones of pre-trained diffusion models, enabling visual correspondence in a manner analogous to the well-established semantic correspondence. While diffusion model backbones are known to encode semantically rich…

Cited by 0SourceScholar
2025

PlaceIt3D: Language-Guided Object Placement in Real 3D Scenes

ICCV 2025poster

We introduce the task of Language-Guided Object Placement in Real 3D Scenes. Given a 3D reconstructed point-cloud scene, a 3D asset, and a natural-language instruction, the goal is to place the asset so that the instruction is satisfied. The task demands tackling four intertwined challenges: (a) one…

Cited by 0SourcePDFScholar
2025

VidSeg: Training-free Video Semantic Segmentation based on Diffusion Models

CVPR 2025poster

We introduce the first training-free approach for Video Semantic Segmentation (VSS) based on pre-trained diffusion models. A growing research direction attempts to employ diffusion models to perform downstream vision tasks by exploiting their deep understanding of image semantics. Yet, the majority…

Cited by 0SourcePDFScholar
2025

ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models

ICCV 2025poster

We propose a novel zero-shot approach for keypoint detection on 3D shapes. Point-level reasoning on visual data is challenging as it requires precise localization capability, posing problems even for powerful models like DINO or CLIP. Traditional methods for 3D keypoint detection rely heavily on ann…

Cited by 0SourcePDFScholar
2020

Uncertainty-Aware CNNs for Depth Completion: Uncertainty from Beginning to End

CVPR 2020poster

The focus in deep learning research has been mostly to push the limits of prediction accuracy. However, this was often achieved at the cost of increased complexity, raising concerns about the interpretability and the reliability of deep networks. Recently, an increasing attention has been given to u…

Cited by 147PDFcodeScholar