← Search

SYRIELLE MONTARIOL

13 accepted papers

2026

Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics

ICML 2026poster

Multimodal LLMs lack a systematic understanding of visual dynamics in complex human world activities, which requires the model to predict or simulate multiple levels of dynamic constituents, such as the general progression of actions and the associated changes of low-level details in the world. To a…

Cited by 0SourceScholar
2025

CAVE : Detecting and Explaining Commonsense Anomalies in Visual Environments

EMNLP 2025

Humans can naturally identify, reason about, and explain anomalies in their environment. In computer vision, this long-standing challenge remains limited to industrial defects or unrealistic, synthetically generated anomalies, failing to capture the richness and unpredictability of real-world anomal

Cited by 0SourcePDFScholar
2025

INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge

ICLR 2025spotlight

The performance differential of large language models (LLM) between languages hinders their effective deployment in many regions, inhibiting the potential economic and societal value of generative AI tools in many communities. However, the development of functional LLMs in many languages (i.e., mult…

Cited by 9SourcePDFScholar
2025

Intrinsic User-Centric Interpretability through Global Mixture of Experts

ICLR 2025poster

In human-centric settings like education or healthcare, model accuracy and model explainability are key factors for user adoption. Towards these two goals, intrinsically interpretable deep learning models have gained popularity, focusing on accurate predictions alongside faithful explanations. Howev…

2025

PICLe: Pseudo-annotations for In-Context Learning in Low-Resource Named Entity Detection

NAACL 2025long

In-context learning (ICL) enables Large Language Models (LLMs) to perform tasks using few demonstrations, facilitating task adaptation when labeled examples are hard to come by. However, ICL is sensitive to the choice of demonstrations, and it remains unclear which demonstration attributes enable in…

2025

VinaBench: Benchmark for Faithful and Consistent Visual Narratives

CVPR 2025poster

Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to the input text and self-consistent across generated images remains an open challenge, due to the lack of knowledge const…

Cited by 1SourcePDFScholar
2024

ConGeo: Robust Cross-view Geo-localization across Ground View Variations

ECCV 2024poster

"∗ Equal contribution Corresponding author (xuchangeis@whu.edu.cn) Cross-view geo-localization aims at localizing a ground-level query image by matching it to its corresponding geo-referenced aerial view. In real-world scenarios, the task requires accommodating diverse ground images captured by user…

Cited by 7SourcePDFScholar
2024

ConVQG: Contrastive Visual Question Generation with Multimodal Guidance

AAAI 2024technical

Asking questions about visual environments is a crucial way for intelligent agents to understand rich multi-faceted scenes, raising the importance of Visual Question Generation (VQG) systems. Apart from being grounded to the image, existing VQG systems can use textual constraints, such as expected a…

2024

“Flex Tape Can’t Fix That”: Bias and Misinformation in Edited Language Models

EMNLP 2024main

Weight-based model editing methods update the parametric knowledge of language models post-training. However, these methods can unintentionally alter unrelated parametric knowledge representations, potentially increasing the risk of harm. In this work, we investigate how weight editing methods unexp…

2023

CRAB: Assessing the Strength of Causal Relationships Between Real-world Events

EMNLP 2023long main

Understanding narratives requires reasoning about the cause-and-effect relationships between events mentioned in the text. While existing foundation models yield impressive results in many NLP tasks requiring reasoning, it is unclear whether they understand the complexity of the underlying network o…

Cited by 0SourcecodeScholar
2023

CRoW: Benchmarking Commonsense Reasoning in Real-World Tasks

EMNLP 2023long main

Recent efforts in natural language processing (NLP) commonsense reasoning research have yielded a considerable number of new datasets and benchmarks. However, most of these datasets formulate commonsense reasoning challenges in artificial scenarios that are not reflective of the tasks which real-wor…

Cited by 0SourcecodeScholar