← Search

Simon Jenni

22 accepted papers

2026

RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward

CVPR 2026

Recent advances in multimodal large language models (MLLMs) have shown great potential for extending vision-language reasoning to professional tool-based image editing, enabling intuitive and creative editing. A promising direction is to use reinforcement learning (RL) to enable MLLMs to reason abou

Cited by 0SourceScholar
2026

Seeing Through Words: Controlling Visual Retrieval Quality with Language

ICLR 2026poster

Text-to-image retrieval is a fundamental task in vision--language learning, yet in real-world scenarios it is often challenged by short and underspecified user queries. Such queries are typically only one or two words long, making them semantically ambiguous, prone to collisions across diverse visua…

Cited by 0SourcecodeScholar
2025

Improving Large Vision and Language Models by Learning from a Panel of Peers

ICCV 2025poster

Traditional alignment methods for Large Vision and Language Models (LVLMs) primarily rely on human-curated preference data. Human-generated preference data is costly; machine-generated preference data is limited in quality; and self-supervised preference data often introduces hallucinations. To over…

2025

MAGNET: Augmenting Generative Decoders with Representation Learning and Infilling Capabilities

ACL 2025long

While originally designed for unidirectional generative modeling, decoder-only large language models (LLMs) are increasingly being adapted for bidirectional modeling. However, unidirectional and bidirectional models are typically trained separately with distinct objectives (generation and representa…

Cited by 0SourcePDFScholar
2025

The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographers

CVPR 2025poster

Photographer, curator, and former director of photography at the Museum of Modern Art (MoMA), John Szarkowski remarked in *William Eggleston's Guide*, "While editing directly from life, photographers have found it too difficult to see simultaneously both the blue and the sky." Szarkowski insightfull…

Cited by 0SourcePDFScholar
2024

Building Vision-Language Models on Solid Foundations with Masked Distillation

CVPR 2024poster

Recent advancements in Vision-Language Models (VLMs) have marked a significant leap in bridging the gap between computer vision and natural language processing. However traditional VLMs trained through contrastive learning on limited and noisy image-text pairs often lack the spatial and linguistic u…

Cited by 8SourcePDFScholar
2024

Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image Models

CVPR 2024poster

While there has been significant progress in customizing text-to-image generation models generating images that combine multiple personalized concepts remains challenging. In this work we introduce Concept Weaver a method for composing customized text-to-image diffusion models at inference time. Spe…

Cited by 12SourcePDFScholar
2024

FineMatch: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction

ECCV 2024poster

"Recent progress in large-scale pre-training has led to the development of advanced vision-language models (VLMs) with remarkable proficiency in comprehending and generating multimodal content. Despite the impressive ability to perform complex reasoning for VLMs, current models often struggle to eff…

2024

No More Shortcuts: Realizing the Potential of Temporal Self-Supervision

AAAI 2024technical

Self-supervised approaches for video have shown impressive results in video understanding tasks. However, unlike early works that leverage temporal self-supervision, current state-of-the-art methods primarily rely on tasks from the image domain (e.g., contrastive learning) that do not explicitly pro…

2023

Meta-Personalizing Vision-Language Models To Find Named Instances in Video

CVPR 2023poster

Large-scale vision-language models (VLM) have shown impressive results for language-guided search applications. While these models allow category-level queries, they currently struggle with personalized searches for moments in a video where a specific object instance such as "My dog Biscuit" appears…

2023

Representation Learning by Detecting Incorrect Location Embeddings

AAAI 2023technical

In this paper, we introduce a novel self-supervised learning (SSL) loss for image representation learning. There is a growing belief that generalization in deep neural networks is linked to their ability to discriminate object shapes. Since object shape is related to the location of its parts, we pr…

2023

VADER: Video Alignment Differencing and Retrieval

ICCV 2023poster

We propose VADER, a spatio-temporal matching, alignment, and change summarization method to help fight misinformation spread via manipulated videos. VADER matches and coarsely aligns partial video fragments to candidate videos using a robust visual descriptor and scalable search over adaptively chun…

Cited by 5PDFcodeScholar
2018

Deep Bilevel Learning

ECCV 2018poster

We present a novel regularization approach to train neural networks that enjoys better generalization and test error than standard stochastic gradient descent. Our approach is based on the principles of cross-validation, where a validation set is used to limit the model overfitting. We formulate suc…

Cited by 150SourcePDFScholar