← Search

Rim Assouel

4 accepted papers

2026

PGT: Procedurally Generated Tasks for improving fine-grained understanding in MLLMs

ICML 2026poster

Despite remarkable progress in Multimodal Large Language Models (MLLMs), these models still struggle with fine-grained understanding tasks. In this work, we propose **Procedurally Generated Tasks (PGT)** a simple data-driven framework that serves a dual purpose: inducing fine-grained visual understa…

Cited by 0SourceScholar
2026

Visual symbolic mechanisms: Emergent symbol processing in Vision Language Models

ICLR 2026oral

To accurately process a visual scene, observers must bind features together to represent individual objects. This capacity is necessary, for instance, to distinguish an image containing a red square and a blue circle from an image containing a blue square and a red circle. Recent work has found that…

Cited by 0SourceScholar
2025

Action abstractions for amortized sampling

ICLR 2025poster

As trajectories sampled by policies used by reinforcement learning (RL) and generative flow networks (GFlowNets) grow longer, credit assignment and exploration become more challenging, and the long planning horizon hinders mode discovery and generalization. The challenge is particularly pronounced i…

Cited by 0SourcePDFScholar
2025

Object-centric binding in Contrastive Language-Image Pretraining

NeurIPS 2025poster

Recent advances in vision language models (VLM) have been driven by contrastive models such as CLIP, which learn to associate visual information with their corresponding text descriptions. However, these models have limitations in understanding complex compositional scenes involving multiple objects…

Cited by 0SourceScholar