← Search

Jonathan Huang

27 accepted papers

2025

Principles of Visual Tokens for Efficient Video Understanding

ICCV 2025poster

Video understanding has made huge strides in recent years, relying largely on the power of transformers. As this architecture is notoriously expensive and video data is highly redundant, research into improving efficiency has become particularly relevant. Some creative solutions include token select…

2025

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

ICLR 2025oral

Recent studies have shown that the denoising process in (generative) diffusion models can induce meaningful (discriminative) representations inside the model, though the quality of these representations still lags behind those learned through recent self-supervised learning methods. We argue that on…

2025

Visually Consistent Hierarchical Image Classification

ICLR 2025poster

Hierarchical classification predicts labels across multiple levels of a taxonomy, e.g., from coarse-level \textit{Bird} to mid-level \textit{Hummingbird} to fine-level \textit{Green hermit}, allowing flexible recognition under varying visual conditions. It is commonly framed as multiple single-leve…

Cited by 0SourcePDFScholar
2024

Optimizing Factorized Encoder Models: Time and Memory Reduction for Scalable and Efficient Action Recognition

ECCV 2024poster

"In this paper, we address the challenges posed by the substantial training time and memory consumption associated with video transformers, focusing on the ViViT (Video Vision Transformer) model, in particular the Factorised Encoder version, as our baseline for action recognition tasks. The factoris…

Cited by 1SourcePDFScholar
2024

Tree-D Fusion: Simulation-Ready Tree Dataset from Single Images with Diffusion Priors

ECCV 2024poster

"We introduce , featuring the first collection of 600,000 environmentally aware, 3D simulation-ready tree models generated through Diffusion priors. Each reconstructed 3D tree model corresponds to an image from Google’s Auto Arborist Dataset, comprising street view images and associated genus labels…

Cited by 4SourcePDFScholar
2024

VideoPoet: A Large Language Model for Zero-Shot Video Generation

ICML 2024oral

We present VideoPoet, a language model capable of synthesizing high-quality video from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs -- including images, videos, text, and audio. The training protocol follows that…

Cited by 257SourcePDFScholar
2023

DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation Model

NeurIPS 2023poster

Observing the close relationship among panoptic, semantic and instance segmentation tasks, we propose to train a universal multi-dataset multi-task segmentation model: DaTaSeg. We use a shared representation (mask proposals with class predictions) for all tasks. To tackle task discrepancy, we adopt…

2023

Learning to Detect Novel and Fine-Grained Acoustic Sequences Using Pretrained Audio Representations

ICASSP 2023accepted

This work investigates pretrained audio representations for few shot Sound Event Detection. We specifically address the task of few shot detection of novel acoustic sequences, or sound events with semantically meaningful temporal structure, without assuming access to non-target audio. We develop pro…

Cited by 0SourceScholar
2022

The Auto Arborist Dataset: A Large-Scale Benchmark for Multiview Urban Forest Monitoring Under Domain Shift

CVPR 2022poster

Generalization to novel domains is a fundamental challenge for computer vision. Near-perfect accuracy on benchmarks is common, but these models do not work as expected when deployed outside of the training distribution. To build computer vision systems that truly solve real-world problems at global…

Cited by 53PDFScholar
2021

The Surprising Impact of Mask-Head Architecture on Novel Class Segmentation

ICCV 2021poster

Instance segmentation models today are very accurate when trained on large annotated datasets, but collecting mask annotations at scale is prohibitively expensive. We address the partially supervised instance segmentation problem in which one can train on (significantly cheaper) bounding boxes for a…

Cited by 31PDFcodeScholar
2020

Context R-CNN: Long Term Temporal Context for Per-Camera Object Detection

CVPR 2020poster

In static monitoring cameras, useful contextual information can stretch far beyond the few seconds typical video understanding models might see: subjects may exhibit similar behavior over multiple days, and background objects remain static. Due to power and storage constraints, sampling frequencies…

Cited by 164PDFScholar
2020

RetinaTrack: Online Single Stage Joint Detection and Tracking

CVPR 2020poster

Traditionally multi-object tracking and object detection are performed using separate systems with most prior works focusing exclusively on one of these aspects over the other. Tracking systems clearly benefit from having access to accurate detections, however and there is ample evidence in literatu…

Cited by 279PDFcodeScholar
2020

Structural Sparsification for Far-Field Speaker Recognition with Intel® Gna

ICASSP 2020accepted

Recently, deep neural networks (DNN) have been widely used in speaker recognition area. In order to achieve fast response time and high accuracy, the requirements for hardware resources increase rapidly. However, as the speaker recognition application is often implemented on mobile devices, it is ne…

Cited by 0SourceScholar
2019

Uncertainty-Aware Audiovisual Activity Recognition Using Deep Bayesian Variational Inference

ICCV 2019oral

Deep neural networks (DNNs) provide state-of-the-art results for a multitude of applications, but the approaches using DNNs for multimodal audiovisual applications do not consider predictive uncertainty associated with individual modalities. Bayesian deep learning methods provide principled confiden…

Cited by 89PDFScholar
2018

Generative Models of Visually Grounded Imagination

ICLR 2018poster

It is easy for people to imagine what a man with pink hair looks like, even if they have never seen such a person before. We call the ability to create images of novel semantic concepts visually grounded imagination. In this paper, we show how we can modify variational auto-encoders to perform this…

Cited by 170SourcePDFScholar
2018

Progressive Neural Architecture Search

ECCV 2018poster

We propose a new method for learning the structure of convolutional neural networks (CNNs) that is more efficient than recent state-of-the-art methods based on reinforcement learning and evolutionary algorithms. Our approach uses a sequential model-based optimization (SMBO) strategy, in which we sea…

2018

Rethinking Spatiotemporal Feature Learning: Speed-Accuracy Trade-offs in Video Classification

ECCV 2018poster

Despite the steady progress in video analysis led by the adoption of convolutional neural networks (CNNs), the relative improvement has been less drastic as that in 2D static image classification. Three main challenges exist including spatial (image) feature representation, temporal information repr…

Cited by 1765SourcePDFScholar
2018

Sufficiency Quantification for Seamless Text-Independent Speaker Enrollment

ICASSP 2018accepted

Text-independent speaker recognition (TI-SR) requires a lengthy enrollment process that involves asking dedicated time from the user to create a reliable model of their voice. Seamless enrollment is a highly attractive feature which refers to the enrollment process that happens in the background and…

Cited by 0SourceScholar
2017

Spatially Adaptive Computation Time for Residual Networks

CVPR 2017poster

This paper proposes a deep learning architecture based on Residual Network that dynamically adjusts the number of executed layers for the regions of the image. This architecture is end-to-end trainable, deterministic and problem-agnostic. It is therefore applicable without any modifications to a wid…

Cited by 431PDFcodeScholar
2017

Speed/Accuracy Trade-Offs for Modern Convolutional Object Detectors

CVPR 2017spotlight

The goal of this paper is to serve as a guide for selecting a detection architecture that achieves the right speed/memory/accuracy balance for a given application and platform. To this end, we investigate various ways to trade accuracy for speed and memory usage in modern convolutional object detect…

Cited by 3693PDFcodeScholar
2016

Detecting Events and Key Actors in Multi-Person Videos

CVPR 2016oral

Multi-person event recognition is a challenging task, often with many people active in the scene but only a small subset contributing to an actual event. In this paper, we propose a model which learns to detect events in such videos while automatically "attending" to the people responsible for the e…

Cited by 284PDFScholar
2016

Generation and Comprehension of Unambiguous Object Descriptions

CVPR 2016oral

We propose a method that can generate an unambiguous description (known as a referring expression) of a specific object or region in an image, and which can also comprehend or interpret such an expression to infer which object is being described. We show that our method outperforms previous methods…

Cited by 1580PDFcodeScholar
2015

Deep Knowledge Tracing

NeurIPS 2015poster

Knowledge tracing, where a machine models the knowledge of a student as they interact with coursework, is an established and significantly unsolved problem in computer supported education.In this paper we explore the benefit of using recurrent neural networks to model student learning.This family of…

2015

Im2Calories: Towards an Automated Mobile Vision Food Diary

ICCV 2015poster

We present a system which can recognize the contents of your meal from a single image, and then predict its nutritional contents, such as calories. The simplest version assumes that the user is eating at a restaurant for which we know the menu. In this case, we can collect images offline to train a…

Cited by 602PDFScholar
2015

Learning Program Embeddings to Propagate Feedback on Student Code

ICML 2015poster

Providing feedback, both assessing final work and giving hints to stuck students, is difficult for open-ended assignments in massive online classes which can range from thousands to millions of students. We introduce a neural network method to encode programs as a linear mapping from an embedded pre…

Cited by 249SourcePDFScholar