← Search

Prajwal Gatti

5 accepted papers

2025

HD-EPIC: A Highly-Detailed Egocentric Video Dataset

CVPR 2025poster

We present a validation dataset of newly-collected kitchen based egocentric videos, manually annotated with highly detailed and interconnected ground-truth labels covering: recipe steps, fine-grained actions, ingredients with nutritional values, moving objects, and audio annotations. Importantly, al…

Cited by 3SourcePDFScholar
2025

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

CVPR 2025poster

The goal of this work is to generate step-by-step visual instructions in the form of a sequence of images, given an input image that provides the scene context and the sequence of textual instructions. This is a challenging problem as it requires generating multi-step image sequences to achieve a co…

2024

Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions

AAAI 2024technical

Non-native speakers with limited vocabulary often struggle to name specific objects despite being able to visualize them, e.g., people outside Australia searching for ‘numbats.’ Further, users may want to search for such elusive objects with difficult-to-sketch interactions, e.g., “numbat digging in…

2023

SLBERT: A Novel Pre-Training Framework for Joint Speech and Language Modeling

ICASSP 2023accepted

We propose SLBERT (Speech and Language pre-training framework for BERT), an end-to-end trainable framework for learning joint representations of speech and language modalities. We enhance the well-known BERT architecture to provide a dual-stream multimodal architecture that processes both speech and…

Cited by 0SourceScholar