← Search

Leonid Sigal

67 accepted papers

2026

InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models

CVPR 2026

Despite the remarkable success, recent reconstruction-based anomaly detection (AD) methods via diffusion modeling still involve fine-grained noise-strength tuning and computationally expensive multi-step denoising, leading to a fundamental tension between fidelity and efficiency. In this paper, we p

Cited by 0SourcecodeScholar
2026

Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection

CVPR 2026

Instruction tuning has been central to the success of recent vision-language models (VLMs), but it remains expensive-requiring large-scale datasets, high-quality annotations, and large compute budgets. We propose PRioritized cOncept learninG via Relative Error-driven Sample Selection (PROGRESS), a d

Cited by 0SourceScholar
2026

SPIKE-RL: Video-LLMs meet Bayesian Surprise

ICLR 2026poster

Real-world videos often show routine activities punctuated by memorable, surprising events. However, most Video-LLMs process videos by sampling frames uniformly, likely missing critical moments that define a video's narrative. We introduce SPIKE, an inference-time framework that quantifies Bayesian…

Cited by 0SourcecodeScholar
2026

Segmentation From Attention: Training-Free Layer Selection and One-Shot Tuning for Segmentation in VLMs

ICML 2026poster

Large-scale vision-language models (VLMs), trained on extensive datasets of image-text pairs, exhibit strong multimodal understanding capabilities by implicitly learning associations between textual descriptions and image regions. This emergent ability enables zero-shot object detection and segmenta…

Cited by 0SourceScholar
2026

To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models

ICLR 2026poster

Large Vision Language Models (LVLMs) have recently emerged as powerful architectures capable of understanding and reasoning over both visual and textual information. These models typically rely on two key components: a Vision Transformer (ViT) and a Large Language Model (LLM). ViT encodes visual con…

Cited by 0SourceScholar
2025

Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events

CVPR 2025poster

The commonsense reasoning capabilities of vision-language models (VLMs), especially in abductive reasoning and defeasible reasoning, remain poorly understood. Most benchmarks focus on typical visual scenarios, making it difficult to discern whether model performance stems from keen perception and re…

Cited by 0SourcePDFScholar
2025

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?

NeurIPS 2025poster

Multi-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance, a key capability for future AI assistants. Effective guidance requires not only delivering instructions but also detecting their successful execution, a…

Cited by 0SourcecodeScholar
2025

ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement

EMNLP 2025

Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA), the task remains challenging, particularly when models attend to irrelevant regions of the chart. In this work, we prese

Cited by 0SourcePDFScholar
2025

LatentHOI: On the Generalizable Hand Object Motion Generation with Latent Hand Diffusion.

CVPR 2025poster

Current research on generating 3D hand-object interaction motion primarily focuses on in-domain objects. Generalization to unseen objects is essential for practical applications, yet it remains both challenging and largely unexplored.In this paper, we propose LatentHOI, a novel approach designed to…

Cited by 0SourcePDFScholar
2025

Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation

ICML 2025poster

Advances in Large Language Models (LLMs) have sparked interest in their ability to solve Olympiad-level math problems. However, the training and evaluation of these models are constrained by the limited size and quality of available datasets, as creating large-scale data for such advanced problems…

2025

Locality Sensitive Avatars From Video

ICLR 2025poster

We present locality-sensitive avatar, a neural radiance field (NeRF) based network to learn human motions from monocular videos. To this end, we estimate a canonical representation between different frames of a video with a non-linear mapping from observation to canonical space, which we decompose i…

2025

MM-R3: On (In-)Consistency of Vision-Language Models (VLMs)

ACL 2025finding

With the advent of LLMs and variants, a flurry of research has emerged, analyzing the performance of such models across an array of tasks. While most studies focus on evaluating the capabilities of state-of-the-art (SoTA) Vision Language Models (VLMs) through task accuracy (e.g., visual question ans…

Cited by 0SourcePDFScholar
2025

Mitigate One, Skew Another? Tackling Intersectional Biases in Text-to-Image Models

EMNLP 2025

The biases exhibited by text-to-image (TTI) models are often treated as independent, though in reality, they may be deeply interrelated. Addressing bias along one dimension—such as ethnicity or age—can inadvertently affect another, like gender, either mitigating or exacerbating existing disparities.

Cited by 0SourcePDFScholar
2025

Prompt2Perturb (P2P): Text-Guided Diffusion-Based Adversarial Attack on Breast Ultrasound Images

CVPR 2025poster

Deep neural networks (DNNs) offer significant promise for improving breast cancer diagnosis in medical imaging. However, these models are highly susceptible to adversarial attacks--small, imperceptible changes that can mislead classifiers--raising critical concerns about their reliability and secur…

Cited by 0SourcePDFScholar
2025

Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities

ACL 2025long

Vision-language Models (VLMs) have emerged as general-purpose tools for addressing a variety of complex computer vision problems. Such models have been shown to be highly capable, but, at the same time, lacking some basic visual understanding skills. In this paper, we set out to understand the limit…

Cited by 0SourcePDFScholar
2025

Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks

ICML 2025poster

Text watermarking aims to subtly embeds statistical signals into text by controlling the Large Language Model (LLM)'s sampling process, enabling watermark detectors to verify that the output was generated by the specified model. The robustness of these watermarking algorithms has become a key factor…

2024

Emergent Open-Vocabulary Semantic Segmentation from Off-the-shelf Vision-Language Models

CVPR 2024poster

From image-text pairs large-scale vision-language models (VLMs) learn to implicitly associate image regions with words which prove effective for tasks like visual question answering. However leveraging the learned association for open-vocabulary semantic segmentation remains a challenge. In this pap…

2024

Extending Video Masked Autoencoders to 128 frames

NeurIPS 2024poster

Video understanding has witnessed significant progress with recent video foundation models demonstrating strong performance owing to self-supervised pre-training objectives; Masked Autoencoders (MAE) being the design of choice. Nevertheless, the majority of prior works that leverage MAE pre-trainin…

Cited by 1SourcePDFScholar
2024

Preventing Catastrophic Forgetting through Memory Networks in Continuous Detection

ECCV 2024poster

"Modern pre-trained architectures struggle to retain previous information while undergoing continuous fine-tuning on new tasks. Despite notable progress in continual classification, systems designed for complex vision tasks such as detection or segmentation still struggle to attain satisfactory perf…

2024

Prompting Hard or Hardly Prompting: Prompt Inversion for Text-to-Image Diffusion Models

CVPR 2024poster

The quality of the prompts provided to text-to-image diffusion models determines how faithful the generated content is to the user's intent often requiring `prompt engineering'. To harness visual concepts from target images without prompt engineering current approaches largely rely on embedding inve…

Cited by 16SourcePDFScholar
2024

Saliency Prediction of Sports Videos: A Large-Scale Database and a Self-Adaptive Approach

ICASSP 2024accepted

Predicting video saliency is crucial for improving sports video processing efficiency, thereby providing an enriched viewing experience for a wide-ranging audience. However, there is a long-term absence of well-established eye-tracking database and learning-based approach, particularly tailored for…

Cited by 0SourceScholar
2024

TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models

ECCV 2024poster

"Text-to-Image (TTI) generative models have shown great progress in the past few years in terms of their ability to generate complex and high-quality imagery. At the same time, these models have been shown to suffer from harmful biases, including exaggerated societal biases (e.g., gender, ethnicity)…

2024

Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach

CVPR 2024poster

The emergence of attention-based transformer models has led to their extensive use in various tasks due to their superior generalization and transfer properties. Recent research has demonstrated that such models when prompted appropriately are excellent for few-shot inference. However such technique…

Cited by 9SourcePDFScholar
2023

DINN360: Deformable Invertible Neural Network for Latitude-Aware 360deg Image Rescaling

CVPR 2023poster

With the rapid development of virtual reality, 360deg images have gained increasing popularity. Their wide field of view necessitates high resolution to ensure image quality. This, however, makes it harder to acquire, store and even process such 360deg images. To alleviate this issue, we propose the…

2023

Make-a-Story: Visual Memory Conditioned Consistent Story Generation

CVPR 2023poster

There has been a recent explosion of impressive generative models that can produce high quality images (or videos) conditioned on text descriptions. However, all such approaches rely on conditional sentences that contain unambiguous descriptions of scenes and main actors in them. Therefore employing…

2023

Mitigating the Effect of Incidental Correlations on Part-based Learning

NeurIPS 2023poster

Intelligent systems possess a crucial characteristic of breaking complicated problems into smaller reusable components or parts and adjusting to new tasks using these part representations. However, current part-learners encounter difficulties in dealing with incidental correlations resulting from th…

2023

Omnimatte3D: Associating Objects and Their Effects in Unconstrained Monocular Video

CVPR 2023poster

We propose a method to decompose a video into a background and a set of foreground layers, where the background captures stationary elements while the foreground layers capture moving objects along with their associated effects (e.g. shadows and reflections). Our approach is designed for unconstrain…

Cited by 3SourcePDFScholar
2023

Self-supervision through Random Segments with Autoregressive Coding (RandSAC)

ICLR 2023poster

Inspired by the success of self-supervised autoregressive representation learning in natural language (GPT and its variants), and advances in recent visual architecture design with Vision Transformers (ViTs), in this paper, we explore the effects various design choices have on the success of applyin…

Cited by 15SourcePDFScholar
2023

Uncertainty Guided Adaptive Warping for Robust and Efficient Stereo Matching

ICCV 2023poster

Correlation based stereo matching has achieved outstanding performance, which pursues cost volume between two feature maps. Unfortunately, current methods with a fixed trained model do not work uniformly well across various datasets, greatly limiting their real-world applicability. To tackle this is…

Cited by 24PDFScholar
2021

Energy-Based Learning for Scene Graph Generation

CVPR 2021poster

Traditional scene graph generation methods are trained using cross-entropy losses that treat objects and relationships as independent entities. Such a formulation, however, ignores structure in the output space, in an inherently structured prediction problem. In this work, we introduce a novel energ…

Cited by 196PDFcodeScholar
2021

PROVIDE: a probabilistic framework for unsupervised video decomposition

UAI 2021poster

Unsupervised multi-object scene decomposition is a fast-emerging problem in representation learning. Despite significant progress in static scenes, such models are unable to leverage important dynamic cues present in videos. We propose PROVIDE, a novel unsupervised framework for PRObabilistic VIdeo…

2021

TriBERT: Human-centric Audio-visual Representation Learning

NeurIPS 2021poster

The recent success of transformer models in language, such as BERT, has motivated the use of such architectures for multi-modal feature learning and tasks. However, most multi-modal variants (e.g., ViLBERT) have limited themselves to visual-linguistic data. Relatively few have explored its use in au…

2021

UniT: Unified Knowledge Transfer for Any-Shot Object Detection and Segmentation

CVPR 2021poster

Methods for object detection and segmentation rely on large scale instance-level annotations for training, which are difficult and time-consuming to collect. Efforts to alleviate this look at varying degrees and quality of supervision. Weakly-supervised approaches draw on image-level labels to build…

Cited by 39PDFcodeScholar
2020

Front2Back: Single View 3D Shape Reconstruction via Front to Back Prediction

CVPR 2020poster

Reconstruction of a 3D shape from a single 2D image is a classical computer vision problem, whose difficulty stems from the inherent ambiguity of recovering occluded or only partially observed surfaces. Recent methods address this challenge through the use of largely unstructured neural networks tha…

Cited by 51PDFcodeScholar
2020

Generating Videos of Zero-Shot Compositions of Actions and Objects

ECCV 2020poster

Human activity videos involve rich, varied interactions between people and objects. In this paper we develop methods for generating such videos -- making progress toward addressing the important, open problem of video generation in complex scenes. In particular, we introduce the task of generating h…

Cited by 14SourcePDFScholar
2019

A Variational Auto-Encoder Model for Stochastic Point Processes

CVPR 2019poster

We propose a novel probabilistic generative model for action sequences. The model is termed the Action Point Process VAE (APP-VAE), a variational auto-encoder that can capture the distribution over the times and categories of action sequences. Modeling the variety of possible action sequences is a…

Cited by 70PDFScholar
2019

LayoutVAE: Stochastic Scene Layout Generation From a Label Set

ICCV 2019poster

Recently there is an increasing interest in scene generation within the research community. However, models used for generating scene layouts from textual description largely ignore plausible visual variations within the structure dictated by the text. We propose LayoutVAE, a variational autoencoder…

Cited by 185PDFScholar
2019

Watch, Listen and Tell: Multi-Modal Weakly Supervised Dense Event Captioning

ICCV 2019poster

Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from language grounding to dense event captioning. However, much of the research has been limited to approaches that either do no…

Cited by 115PDFcodeScholar
2018

A Neural Multi-Sequence Alignment TeCHnique (NeuMATCH)

CVPR 2018poster

The alignment of heterogeneous sequential data (video to text) is an important and challenging problem. Standard techniques for this task, including Dynamic Time Warping (DTW) and Conditional Random Fields (CRFs), suffer from inherent drawbacks. Mainly, the Markov assumption implies that, given the…

2018

Middle-Out Decoding

NeurIPS 2018poster

Despite being virtually ubiquitous, sequence-to-sequence models are challenged by their lack of diversity and inability to be externally controlled. In this paper, we speculate that a fundamental shortcoming of sequence generation models is that the decoding is done strictly from left-to-right, mean…

Cited by 24SourcePDFScholar
2018

Probabilistic Video Generation using Holistic Attribute Control

ECCV 2018poster

Videos express highly structured spatio-temporal patterns of visual data. A video can be thought of as being governed by two factors: (i) temporally invariant (e.g., person identity), or slowly varying (e.g., activity), attribute-induced appearance, encoding the persistent content of each frame, and…

Cited by 88SourcePDFScholar
2018

Show Me a Story: Towards Coherent Neural Story Illustration

CVPR 2018poster

We propose an end-to-end network for the visual illustration of a sequence of sentences forming a story. At the core of our model is the ability to model the inter-related nature of the sentences within a story, as well as the ability to learn coherence to support reference resolution. The framework…

2017

Visual Reference Resolution using Attention Memory for Visual Dialog

NeurIPS 2017poster

Visual dialog is a task of answering a series of inter-dependent questions given an input image, and often requires to resolve visual references among the questions. This problem is different from visual question answering (VQA), which relies on spatial attention ({\em a.k.a. visual grounding}) esti…

Cited by 143SourcePDFScholar
2016

Harnessing Object and Scene Semantics for Large-Scale Video Understanding

CVPR 2016spotlight

Large-scale action recognition and video categorization are important problems in computer vision. To address these problems, we propose a novel object- and scene-based semantic fusion network and representation. Our semantic fusion network combines three streams of information using a three-layer n…

Cited by 113PDFScholar
2015

Expanding Object Detector's Horizon: Incremental Learning Framework for Object Detection in Videos

CVPR 2015poster

Over the last several years it has been shown that image-based object detectors are sensitive to the training data and often fail to generalize to examples that fall outside the original training sample domain (e.g., videos). A number of domain adaptation (DA) techniques have been proposed to add…

Cited by 58SourcePDFScholar
2015

Ranking and Retrieval of Image Sequences From Multiple Paragraph Queries

CVPR 2015poster

We propose a method to rank and retrieve image sequences from a natural language text query, consisting of multiple sentences or paragraphs. One of the method's key applications is to visualize visitors' text-only reviews on TRIPADVISOR or YELP, by automatically retrieving the most illustrative imag…

Cited by 51SourcePDFScholar
2015

Storyline Representation of Egocentric Videos With an Applications to Story-Based Search

ICCV 2015poster

Egocentric videos are a valuable source of information as a daily log of our lives. However, large fraction of egocentric video content is typically irrelevant and boring to re-watch. It is an agonizing task, for example, to manually search for the moment when your daughter first met Mickey Mouse fr…

Cited by 63PDFScholar