← Search

Kumar Ashutosh

12 accepted papers

2025

ExpertAF: Expert Actionable Feedback from Video

CVPR 2025poster

Feedback is essential for learning a new skill or improving one's current skill-level. However, current methods for skill-assessment from video only provide scores or compare demonstrations, leaving the burden of knowing what to do differently on the user. We introduce a novel method to generate act…

Cited by 3SourcePDFScholar
2025

LLMs can see and hear without any training

ICML 2025poster

We present MILS: Multimodal Iterative LLM Solver, a surprisingly simple, training-free approach, to imbue multimodal capabilities into your favorite LLM. Leveraging their innate ability to perform multi-step reasoning, MILS prompts the LLM to generate candidate outputs, each of which are scored and…

2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos

CVPR 2024poster

We propose a novel self-supervised embedding to learn how actions sound from narrated in-the-wild egocentric videos. Whereas existing methods rely on curated data with known audio-visual correspondence our multimodal contrastive-consensus coding (MC3) embedding reinforces the associations between au…

Cited by 8SourcePDFScholar
2023

HierVL: Learning Hierarchical Video-Language Embeddings

CVPR 2023highlight

Video-language embeddings are a promising avenue for injecting semantics into visual representations, but existing methods capture only short-term associations between seconds-long video clips and their accompanying text. We propose HierVL, a novel hierarchical video-language embedding that simultan…

Cited by 57SourcePDFScholar
2023

Video-Mined Task Graphs for Keystep Recognition in Instructional Videos

NeurIPS 2023poster

Procedural activity understanding requires perceiving human actions in terms of a broader task, where multiple keysteps are performed in sequence across a long video to reach a final goal state---such as the steps of a recipe or the steps of a DIY fix-it task. Prior work largely treats keystep reco…

Cited by 29SourcePDFScholar
2021

Bandit algorithms: Letting go of logarithmic regret for statistical robustness

AISTATS 2021poster

We study regret minimization in a stochastic multi-armed bandit setting, and establish a fundamental trade-off between the regret suffered under an algorithm, and its statistical robustness. Considering broad classes of underlying arms’ distributions, we show that bandit learning algorithms with log…

Cited by 18SourcePDFScholar