← Search

Jason J. Corso

27 accepted papers

2026

BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation

CVPR 2026

Text-guided dynamic 3D character generation has advanced rapidly, yet producing high-quality motion that faithfully reflects rich textual descriptions remains challenging. Existing methods tend to generate limited sub-actions or incoherent motion due to fixed-length temporal inputs and discrete fram

Cited by 0SourcecodeScholar
2026

Bridging Facial Understanding and Animation via Language Models

CVPR 2026

Text-guided human body animation has advanced rapidly, yet facial animation lags due to the scarcity of well-annotated, text-paired facial corpora. To close this gap, we leverage foundation generative models to synthesize a large, balanced corpus of facial behavior. We design prompts suite covering

Cited by 0SourceScholar
2026

Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos

CVPR 2026

We introduce Mistake Attribution (MATT), a new task for fine-grained understanding of human mistakes in egocentric videos. While prior work detects whether a mistake occurs, MATT attributes the mistake to what part of the instruction is violated (semantic role), when in the video the deviation becom

Cited by 0SourcecodeScholar
2026

R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space

CVPR 2026

Humans perceive and reason about their surroundings in four dimensions - three spatial and one temporal axis - by building persistent, structured internal representations that encode semantic meaning, spatial layout, and temporal dynamics. These multimodal memories enable them to recall past events,

Cited by 0SourceScholar
2026

When to Think and When to Look: Uncertainty-Guided Lookback

CVPR 2026

Test-time "thinking" (i.e., generating explicit intermediate reasoning chains) is known to boost performance in large language models and has recently shown strong gains for large vision-language models (LVLMs). However, despite these promising results, there is still no systematic analysis of how t

Cited by 0SourcecodeScholar
2025

Transparent and Coherent Procedural Mistake Detection

EMNLP 2025

Procedural mistake detection (PMD) is a challenging problem of classifying whether a human user (observed through egocentric video) has successfully executed a task (specified by a procedural text). Despite significant recent efforts, machine performance in the wild remains nonviable, and the reason

Cited by 0SourcePDFScholar
2025

VITRO: Vocabulary Inversion for Time-series Representation Optimization

ICASSP 2025accepted

Although LLMs have demonstrated remarkable capabilities in processing and generating textual data, their pretrained vocabularies are ill-suited for capturing the nuanced temporal dynamics and patterns inherent in time series. The discrete, symbolic nature of natural language tokens, which these voca…

Cited by 0SourceScholar
2020

Adversarial Background-Aware Loss for Weakly-supervised Temporal Activity Localization

ECCV 2020poster

Temporally localizing activities within untrimmed videos has been extensively studied in recent years. Despite recent advances, existing methods for weakly-supervised temporal activity localization struggle to recognize when an activity is not occurring. To address this issue, we propose a novel met…

2020

Novel Object Viewpoint Estimation Through Reconstruction Alignment

CVPR 2020poster

The goal of this paper is to estimate the viewpoint for a novel object. Standard viewpoint estimation approaches generally fail on this task due to their reliance on a 3D model for alignment or large amounts of class-specific training data and their corresponding canonical pose. We overcome those li…

Cited by 18PDFcodeScholar
2019

BubbleNets: Learning to Select the Guidance Frame in Video Object Segmentation by Deep Sorting Frames

CVPR 2019oral

Semi-supervised video object segmentation has made significant progress on real and challenging videos in recent years. The current paradigm for segmentation methods and benchmark datasets is to segment objects in video provided a single annotation in the first frame. However, we find that segmentat…

Cited by 55PDFcodeScholar
2019

Multi-Channel Attention Selection GAN With Cascaded Semantic Guidance for Cross-View Image Translation

CVPR 2019oral

Cross-view image translation is challenging because it involves images with drastically different views and severe deformation. In this paper, we propose a novel approach named Multi-Channel Attention SelectionGAN (SelectionGAN) that makes it possible to generate images of natural scenes in arbitrar…

Cited by 441PDFcodeScholar
2019

TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency Detection

ICCV 2019poster

TASED-Net is a 3D fully-convolutional network architecture for video saliency detection. It consists of two building blocks: first, the encoder network extracts low-resolution spatiotemporal features from an input clip of several consecutive frames, and then the following prediction network decodes…

Cited by 209PDFcodeScholar
2018

End-to-End Dense Video Captioning With Masked Transformer

CVPR 2018poster

Dense video captioning aims to generate text descriptions for all events in an untrimmed video. This involves both detecting and describing events. Therefore, all previous methods on dense video captioning tackle this problem by building two models, i.e. an event proposal and a captioning model, for…

Cited by 728SourcePDFScholar
2017

Weakly Supervised Actor-Action Segmentation via Robust Multi-Task Ranking

CVPR 2017poster

Fine-grained activity understanding in videos has attracted considerable recent attention with a shift from action classification to detailed actor and action understanding that provides compelling results for perceptual needs of cutting-edge autonomous systems. However, current methods for detailed…

Cited by 55PDFScholar
2016

A Continuous Occlusion Model for Road Scene Understanding

CVPR 2016poster

We present a physically interpretable, continuous 3D model for handling occlusions with applications to road scene understanding. We probabilistically assign each point in space to an object with a theoretical modeling of the reflection and transmission probabilities for the corresponding camera ray…

Cited by 37PDFScholar
2015

Can Humans Fly? Action Understanding With Multiple Classes of Actors

CVPR 2015poster

Can humans fly? Emphatically no. Can cars eat? Again, absolutely not. Yet, these absurd inferences result from the current disregard for particular types of actors in action understanding. There is no work we know of on simultaneously inferring actors and actions in the video, not to mention a datas…

Cited by 146SourcePDFScholar