← Search

Shiry Ginosar

10 accepted papers

2025

KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal Models

ICLR 2025poster

This paper investigates visual analogical reasoning in large multimodal models (LMMs) compared to human adults and children. A “visual analogy” is an abstract rule inferred from one image and applied to another. While benchmarks exist for testing visual reasoning in LMMs, they require advanced skill…

2025

Poly-Autoregressive Prediction for Modeling Interactions

CVPR 2025poster

We introduce a simple framework for predicting the behavior of an agent in multi-agent settings. In contrast to autoregressive (AR) tasks, such as language processing, our focus is on scenarios with multiple agents whose interactions are shaped by physical constraints and internal motivations. To th…

Cited by 0SourcePDFScholar
2025

Pose Priors from Language Models

CVPR 2025poster

Language is often used to describe physical interaction, yet most 3D human pose estimation methods overlook this rich source of information. We bridge this gap by leveraging large multimodal models (LMMs) as priors for reconstructing contact poses, offering a scalable alternative to traditional meth…

2024

Diffusion Models as Data Mining Tools

ECCV 2024poster

"This paper demonstrates how to use generative models trained for image synthesis as tools for visual data mining. Our insight is that since contemporary generative models learn an accurate representation of their training data, we can use them to summarize the data by mining for visual patterns. Co…

Cited by 3SourcePDFScholar
2023

Can Language Models Learn to Listen?

ICCV 2023poster

We present a framework for generating appropriate facial responses from a listener in dyadic social interactions based on the speaker's words. Given an input transcription of the speaker's words with their timestamps, our approach autoregressively predicts a response of a listener: a sequence of lis…

Cited by 24PDFScholar
2022

Learning To Listen: Modeling Non-Deterministic Dyadic Facial Motion

CVPR 2022poster

We present a framework for modeling interactional communication in dyadic conversations: given multimodal inputs of a speaker, we autoregressively output multiple possibilities of corresponding listener motion. We combine the motion and speech audio of the speaker using a motion-audio cross attentio…

Cited by 105PDFScholar
2021

Body2Hands: Learning To Infer 3D Hands From Conversational Gesture Body Dynamics

CVPR 2021poster

We propose a novel learned deep prior of body motion for 3D hand shape synthesis and estimation in the domain of conversational gestures. Our model builds upon the insight that body motion and hand gestures are strongly correlated in non-verbal communication settings. We formulate the learning of th…

Cited by 55PDFcodeScholar
2020

Learning to Factorize and Relight a City

ECCV 2020poster

We propose a learning-based framework for disentangling outdoor scenes into temporally-varying illumination and permanent scene factors. Inspired by the classic intrinsic image decomposition, our learning signal builds upon two insights: 1) combining the disentangled factors should reconstruct the o…

2019

Learning Individual Styles of Conversational Gesture

CVPR 2019poster

Human speech is often accompanied by hand and arm gestures. We present a method for cross-modal translation from "in-the-wild" monologue speech of a single speaker to their conversational gesture motion. We train on unlabeled videos for which we only have noisy pseudo ground truth from an automatic…

Cited by 403PDFcodeScholar