← Search

Abhinav Shrivastava

86 accepted papers

2026

Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model

ICLR 2026poster

Diffusion-based image generation models excel at producing high-quality synthetic content, but suffer from slow and computationally expensive inference. Prior work has attempted to mitigate this by caching and reusing features within diffusion transformers across inference steps. These methods, howe…

Cited by 0SourcecodeScholar
2026

NeRV-Diffusion: Diffuse Implicit Neural Representation for Video Synthesis

ICLR 2026poster

We present NeRV-Diffusion, an implicit latent video diffusion model that synthesizes videos via generating neural network weights. The generated weights can be rearranged as the parameters of a convolutional neural network, which forms an implicit neural representation (INR), and decodes into videos…

Cited by 0SourceScholar
2026

UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders

CVPR 2026

The space of task-agnostic feature upsampling has emerged as a promising area of research to efficiently create denser features from pre-trained visual backbones. These methods act as a shortcut to achieve dense features for a fraction of the cost by learning to map low-resolution features to high-r

Cited by 0SourcecodeScholar
2026

VeriGraph: Scene Graphs for Execution Verifiable Robot Planning

ICRA 2026poster

Recent advancements in vision-language models (VLMs) offer potential for robot task planning, but challenges remain due to VLMs’ tendency to generate incorrect action sequences. To address these limitations, we propose VeriGraph, a novel framework that integrates VLMs for robotic planning while veri…

2025

CoLLM: A Large Language Model for Composed Image Retrieval

CVPR 2025poster

Composed Image Retrieval (CIR) is a complex task that aims to retrieve images based on a multimodal query. Typical training data consists of triplets containing a reference image, a textual description of desired modifications, and the target image, which are expensive and time-consuming to acquire.…

2025

Imagine, Verify, Execute: Memory-guided Agentic Exploration with Vision-Language Models

CoRL 2025poster

Exploration is key for general-purpose robotic learning, particularly in open-ended environments where explicit guidance or task-specific feedback is limited. Vision-language models (VLMs), which can reason about object semantics, spatial relations, and potential outcomes, offer a promising foundati…

Cited by 0SourceScholar
2025

LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior

ICLR 2025oral

We present LARP, a novel video tokenizer designed to overcome limitations in current video tokenization methods for autoregressive (AR) generative models. Unlike traditional patchwise tokenizers that directly encode local visual patches into discrete tokens, LARP introduces a holistic tokenization s…

2025

TREND: Tri-Teaching for Robust Preference-based Reinforcement Learning with Demonstrations

ICRA 2025

Preference feedback collected by human or VLM annotators is often noisy, presenting a significant challenge for preference-based reinforcement learning that relies on accurate preference labels. To address this challenge, we propose TREND, a novel framework that integrates few-shot expert demonstrat

Cited by 6SourceScholar
2025

Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition

ICCV 2025poster

Video understanding requires effective modeling of both motion and appearance information, particularly for few-shot action recognition. While recent advances in point tracking have been shown to improve few-shot action recognition, two fundamental challenges persist: selecting informative points to…

Cited by 0SourcePDFScholar
2024

ARDuP: Active Region Video Diffusion for Universal Policies

IROS 2024poster

Sequential decision-making can be formulated as a text-conditioned video generation problem, where a video planner, guided by a text-defined goal, generates future frames visualizing planned actions, from which control actions are subsequently derived. In this work, we introduce Active Region Video…

Cited by 3SourceScholar
2024

AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models

EMNLP 2024finding

Large vision-language models (LVLMs) are prone to hallucinations, where certain contextual cues in an image can trigger the language module to produce overconfident and incorrect reasoning about abnormal or hypothetical objects. While some benchmarks have been developed to investigate LVLM hallucina…

2024

Beyond Seen Primitive Concepts and Attribute-Object Compositional Learning

CVPR 2024poster

Learning from seen attribute-object pairs to generalize to unseen compositions has been studied extensively in Compositional Zero-Shot Learning (CZSL). However CZSL setup is still limited to seen attributes and objects and cannot generalize to unseen concepts and their compositions. To overcome this…

Cited by 0SourcePDFScholar
2024

Composing Object Relations and Attributes for Image-Text Matching

CVPR 2024poster

We study the visual semantic embedding problem for image-text matching. Most existing work utilizes a tailored cross-attention mechanism to perform local alignment across the two image and text modalities. This is computationally expensive even though it is more powerful than the unimodal dual-encod…

2024

Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models

ECCV 2024poster

"Image customization has been extensively studied in text-to-image (T2I) diffusion models, leading to impressive outcomes and applications. With the emergence of text-to-video (T2V) diffusion models, its temporal counterpart, motion customization, has not yet been well investigated. To address the c…

2024

Do text-free diffusion models learn discriminative visual representations?

ECCV 2024poster

"Diffusion models have proven to be state-of-the-art methods for generative tasks. These models involve training a U-Net to iteratively predict and remove noise, and the resulting model can synthesize high-fidelity, diverse, novel images. However, text-free diffusion models have typically not been e…

2024

EAGLES: Efficient Accelerated 3D Gaussians with Lightweight EncodingS

ECCV 2024poster

"Recently, 3D Gaussian splatting (3D-GS) has gained popularity in novel-view scene synthesis. It addresses the challenges of lengthy training times and slow rendering speeds associated with Neural Radiance Fields (NeRFs). Through rapid, differentiable rasterization of 3D Gaussians, 3D-GS achieves re…

2024

Explaining the Implicit Neural Canvas: Connecting Pixels to Neurons by Tracing their Contributions

CVPR 2024poster

The many variations of Implicit Neural Representations (INRs) where a neural network is trained as a continuous representation of a signal have tremendous practical utility for downstream tasks including novel view synthesis video compression and image super-resolution. Unfortunately the inner worki…

Cited by 1SourcePDFScholar
2024

Fast Encoding and Decoding for Implicit Video Representation

ECCV 2024poster

"Despite the abundant availability and content richness for video data, its high-dimensionality poses challenges for video research. Recent advancements have explored the implicit representation for videos using neural networks, demonstrating strong performance in applications such as video compress…

Cited by 1SourcePDFScholar
2024

Investigating Style Similarity in Diffusion Models

ECCV 2024poster

"Generative models are now widely used by graphic designers and artists. Prior works have shown that these models remember and often replicate content from their training data during generation. Hence as their proliferation increases, it has become important to perform a database search to determine…

Cited by 0SourcePDFScholar
2024

LEIA: Latent View-invariant Embeddings for Implicit 3D Articulation

ECCV 2024poster

"Neural Radiance Fields (NeRFs) have revolutionized the reconstruction of static scenes and objects in 3D, offering unprecedented quality. However, extending NeRFs to model dynamic objects or object articulations remains a challenging problem. Previous works have tackled this issue by focusing on pa…

Cited by 7SourcePDFScholar
2024

Latent-INR: A Flexible Framework for Implicit Representations of Videos with Discriminative Semantics

ECCV 2024poster

"Implicit Neural Networks (INRs) have emerged as powerful representations to encode all forms of data, including images, videos, audios, and scenes. With video, many INRs for video have been proposed for the compression task, and recent methods feature significant improvements with respect to encodi…

Cited by 2SourcePDFScholar
2024

LiFT: A Surprisingly Simple Lightweight Feature Transform for Dense ViT Descriptors

ECCV 2024poster

"We present a simple self-supervised method to enhance the performance of ViT features for dense downstream tasks. Our Lightweight Feature Transform (LiFT) is a straightforward and compact postprocessing network that can be applied to enhance the features of any pre-trained ViT backbone. LiFT is fas…

Cited by 5SourcePDFScholar
2024

MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

CVPR 2024poster

With the success of large language models (LLMs) integrating the vision model into LLMs to build vision-language foundation models has gained much more interest recently. However existing LLM-based large multimodal models (e.g. Video-LLaMA VideoChat) can only take in a limited number of frames for s…

2024

MaGGIe: Masked Guided Gradual Human Instance Matting

CVPR 2024poster

Human matting is a foundation task in image and video processing where human foreground pixels are extracted from the input. Prior works either improve the accuracy by additional guidance or improve the temporal consistency of a single instance across frames. We propose a new framework MaGGIe Masked…

2024

QUEEN: QUantized Efficient ENcoding of Dynamic Gaussians for Streaming Free-viewpoint Videos

NeurIPS 2024poster

Online free-viewpoint video (FVV) streaming is a challenging problem, which is relatively under-explored. It requires incremental on-the-fly updates to a volumetric representation, fast training and rendering to satisfy realtime constraints and a small memory footprint for efficient transmission. If…

Cited by 0SourcePDFScholar
2024

Trajectory-aligned Space-time Tokens for Few-shot Action Recognition

ECCV 2024poster

"We propose a simple yet effective approach for few-shot action recognition, emphasizing the disentanglement of motion and appearance representations. By harnessing recent progress in tracking, specifically point trajectories and self-supervised representation learning, we build trajectory-aligned t…

Cited by 2SourcePDFScholar
2024

Video Prediction by Modeling Videos as Continuous Multi-Dimensional Processes

CVPR 2024poster

Diffusion models have made significant strides in image generation mastering tasks such as unconditional image synthesis text-image translation and image-to-image conversions. However their capability falls short in the realm of video prediction mainly because they treat videos as a collection of in…

Cited by 13SourcePDFScholar
2023

ASIC: Aligning Sparse in-the-wild Image Collections

ICCV 2023oral

We present a method for joint alignment of sparse in-the-wild image collections of an object category. Most prior works assume either ground-truth keypoint annotations or a large dataset of images of a single object category. However, neither of the above assumptions hold true for the long-tail of t…

Cited by 21PDFcodeScholar
2023

Align and Attend: Multimodal Summarization With Dual Contrastive Losses

CVPR 2023poster

The goal of multimodal summarization is to extract the most important information from different modalities to form summaries. Unlike unimodal summarization, the multimodal summarization task explicitly leverages cross-modal information to help generate more reliable and high-quality summaries. Howe…

2023

BT^2: Backward-compatible Training with Basis Transformation

ICCV 2023poster

Modern retrieval system often requires recomputing the representation of every piece of data in the gallery when updating to a better representation model. This process is known as backfilling and can be especially costly in the real world where the gallery often contains billions of samples. Recent…

Cited by 6PDFcodeScholar
2023

Chop & Learn: Recognizing and Generating Object-State Compositions

ICCV 2023poster

Recognizing and generating object-state compositions has been a challenging task, especially when generalizing to unseen compositions. In this paper, we study the task of cutting objects in different styles and the resulting object state changes. We propose a new benchmark suite Chop & Learn, to acc…

Cited by 18PDFcodeScholar
2023

FlexNeRF: Photorealistic Free-Viewpoint Rendering of Moving Humans From Sparse Views

CVPR 2023poster

We present FlexNeRF, a method for photorealistic free-viewpoint rendering of humans in motion from monocular videos. Our approach works well with sparse views, which is a challenging scenario when the subject is exhibiting fast/complex motions. We propose a novel approach which jointly optimizes a c…

2023

HNeRV: A Hybrid Neural Representation for Videos

CVPR 2023poster

Implicit neural representations store videos as neural networks and have performed well for vision tasks such as video compression and denoising. With frame index and/or positional index as input, implicit representations (NeRV, E-NeRV, etc.) reconstruct video frames from fixed and content-agnostic…

2023

LilNetX: Lightweight Networks with EXtreme Model Compression and Structured Sparsification

ICLR 2023poster

We introduce LilNetX, an end-to-end trainable technique for neural networks that enables learning models with specified accuracy-rate-computation trade-off. Prior works approach these problems one at a time and often require post-processing or multistage training which become less practical and do n…

2023

MOST: Multiple Object Localization with Self-Supervised Transformers for Object Discovery

ICCV 2023oral

We tackle the challenging task of unsupervised object localization in this work. Recently, transformers trained with self-supervised learning have been shown to exhibit object localization properties without being trained for this task. In this work, we present Multiple Object localization with Self…

Cited by 12PDFcodeScholar
2023

NIRVANA: Neural Implicit Representations of Videos With Adaptive Networks and Autoregressive Patch-Wise Modeling

CVPR 2023poster

Implicit Neural Representations (INR) have recently shown to be powerful tool for high-quality video compression. However, existing works are limiting as they do not explicitly exploit the temporal redundancy in videos, leading to a long encoding time. Additionally, these methods have fixed architec…

2023

SHACIRA: Scalable HAsh-grid Compression for Implicit Neural Representations

ICCV 2023poster

Implicit Neural Representations (INR) or neural fields have emerged as a popular framework to encode multimedia signals such as images and radiance fields while retaining high-quality. Recently, learnable feature grids such as Instant-NGP have allowed significant speed-up in the training as well as…

Cited by 30PDFcodeScholar
2023

SimpSON: Simplifying Photo Cleanup With Single-Click Distracting Object Segmentation Network

CVPR 2023poster

In photo editing, it is common practice to remove visual distractions to improve the overall image quality and highlight the primary subject. However, manually selecting and removing these small and dense distracting regions can be a laborious and time-consuming task. In this paper, we propose an in…

2023

SparseDet: Improving Sparsely Annotated Object Detection with Pseudo-positive Mining

ICCV 2023poster

Training with sparse annotations is known to reduce the performance of object detectors. Previous methods have focused on proxies for missing ground truth annotations in the form of pseudo-labels for unlabeled boxes. We observe that existing methods suffer at higher levels of sparsity in the data du…

Cited by 13PDFScholar
2023

Teaching Matters: Investigating the Role of Supervision in Vision Transformers

CVPR 2023poster

Vision Transformers (ViTs) have gained significant popularity in recent years and have proliferated into many applications. However, their behavior under different learning paradigms is not well explored. We compare ViTs trained through different methods of supervision, and show that they learn a di…

2023

Towards Scalable Neural Representation for Diverse Videos

CVPR 2023poster

Implicit neural representations (INR) have gained increasing attention in representing 3D scenes and images, and have been recently applied to encode videos (e.g., NeRV, E-NeRV). While achieving promising results, existing INR-based methods are limited to encoding a handful of short videos (e.g., se…

Cited by 45SourcePDFScholar
2023

Video Dynamics Prior: An Internal Learning Approach for Robust Video Enhancements

NeurIPS 2023poster

In this paper, we present a novel robust framework for low-level vision tasks, including denoising, object removal, frame interpolation, and super-resolution, that does not require any external training data corpus. Our proposed approach directly learns the weights of neural modules by optimizing ov…

Cited by 12SourcePDFScholar
2022

ASM-Loc: Action-Aware Segment Modeling for Weakly-Supervised Temporal Action Localization

CVPR 2022poster

Weakly-supervised temporal action localization aims to recognize and localize action segments in untrimmed videos given only video-level action labels for training. Without the boundary information of action segments, existing methods mostly rely on multiple instance learning (MIL), where the predic…

Cited by 120PDFcodeScholar
2022

Beyond Supervised vs. Unsupervised: Representative Benchmarking and Analysis of Image Representation Learning

CVPR 2022poster

By leveraging contrastive learning, clustering, and other pretext tasks, unsupervised methods for learning image representations have reached impressive results on standard benchmarks. The result has been a crowded field -- many methods with substantially different implementations yield results that…

Cited by 25PDFcodeScholar
2022

Burn after Reading: Online Adaptation for Cross-Domain Streaming Data

ECCV 2022poster

"In the context of online privacy, many methods propose complex security preserving measures to protect sensitive data. In this paper, we note that: not storing any sensitive data is the best form of security. We propose an online framework called ""Burn After Reading"", i.e. each online sample is p…

Cited by 6SourcePDFScholar
2022

Dual-Key Multimodal Backdoors for Visual Question Answering

CVPR 2022poster

The success of deep learning has enabled advances in multimodal tasks that require non-trivial fusion of multiple input domains. Although multimodal models have shown potential in many problems, their increased complexity makes them more vulnerable to attacks. A Backdoor (or Trojan) attack is a clas…

Cited by 52PDFcodeScholar
2022

Improving Closed and Open-Vocabulary Attribute Prediction Using Transformers

ECCV 2022poster

"We study recognizing attributes for objects in visual scenes. We consider attributes to be any phrases that describe an object’s physical and semantic properties, and its relationships with other objects. Existing work studies attribute prediction in a closed setting with a fixed set of attributes,…

Cited by 24SourcePDFScholar
2022

Learning Semantic Correspondence with Sparse Annotations

ECCV 2022poster

"Finding dense semantic correspondence is a fundamental problem in computer vision, which remains challenging in complex scenes due to background clutter, extreme intra-class variation, and a severe lack of ground truth. In this paper, we aim to address the challenge of label sparsity in semantic co…

2022

ObjectFormer for Image Manipulation Detection and Localization

CVPR 2022poster

Recent advances in image editing techniques have posed serious challenges to the trustworthiness of multimedia data, which drives the research of image tampering detection. In this paper, we propose ObjectFormer to detect and localize image manipulations. To capture subtle manipulation traces that a…

Cited by 190PDFScholar
2022

Rethinking Pseudo Labels for Semi-supervised Object Detection

AAAI 2022technical

Recent advances in semi-supervised object detection (SSOD) are largely driven by consistency-based pseudo-labeling methods for image classification tasks, producing pseudo labels as supervisory signals. However, when using pseudo labels, there is a lack of consideration in localization precision and…

Cited by 97SourcePDFScholar
2021

2D or not 2D? Adaptive 3D Convolution Selection for Efficient Video Recognition

CVPR 2021poster

3D convolutional networks are prevalent for video recognition. While achieving excellent recognition performance on standard benchmarks, they operate on a sequence of frames with 3D convolutions and thus are computationally demanding. Exploiting large variations among different videos, we introduce…

Cited by 49PDFScholar
2021

Deep Co-Training With Task Decomposition for Semi-Supervised Domain Adaptation

ICCV 2021poster

Semi-supervised domain adaptation (SSDA) aims to adapt models trained from a labeled source domain to a different but related target domain, from which unlabeled data and a small set of labeled data are provided. Current methods that treat source and target supervision without distinction overlook t…

Cited by 118PDFcodeScholar
2021

Hierarchical Video Prediction Using Relational Layouts for Human-Object Interactions

CVPR 2021poster

Learning to model and predict how humans interact with objects while performing an action is challenging, and most of the existing video prediction models are ineffective in modeling complicated human-object interactions. Our work builds on hierarchical video prediction models, which disentangle the…

Cited by 28PDFScholar
2021

LayoutTransformer: Layout Generation and Completion With Self-Attention

ICCV 2021poster

We address the problem of scene layout generation for diverse domains such as images, mobile applications, documents, and 3D objects. Most complex scenes, natural or human-designed, can be expressed as a meaningful arrangement of simpler compositional graphical primitives. Generating a new layout or…

Cited by 184PDFcodeScholar
2021

Learned Spatial Representations for Few-Shot Talking-Head Synthesis

ICCV 2021poster

We propose a novel approach for few-shot talking-head synthesis. While recent works in neural talking heads have produced promising results, they can still produce images that do not preserve the identity of the subject in source images. We posit this is a result of the entangled representation of e…

Cited by 49PDFScholar
2021

Learning Graphs for Knowledge Transfer With Limited Labels

CVPR 2021poster

Fixed input graphs are a mainstay in approaches that utilize Graph Convolution Networks (GCNs) for knowledge transfer. The standard paradigm is to utilize relationships in the input graph to transfer information using GCNs from training to testing nodes in the graph; for example, the semi-supervised…

Cited by 12PDFScholar
2021

Learning To Predict Visual Attributes in the Wild

CVPR 2021poster

Visual attributes constitute a large portion of information contained in a scene. Objects can be described using a wide variety of attributes which portray their visual appearance (color, texture), geometry (shape, size, posture), and other intrinsic properties (state, action). Existing work is most…

Cited by 132PDFScholar
2021

NeRV: Neural Representations for Videos

NeurIPS 2021poster

We propose a novel neural representation for videos (NeRV) which encodes videos in neural networks. Unlike conventional representations that treat videos as frame sequences, we represent videos as neural networks taking frame index as input. Given a frame index, NeRV outputs the corresponding RGB i…

2021

PatchGame: Learning to Signal Mid-level Patches in Referential Games

NeurIPS 2021poster

We study a referential game (a type of signaling game) where two agents communicate with each other via a discrete bottleneck to achieve a common goal. In our referential game, the goal of the speaker is to compose a message or a symbolic representation of "important" image patches, while the task f…

2021

StEP: Style-Based Encoder Pre-Training for Multi-Modal Image Synthesis

CVPR 2021poster

We propose a novel approach for multi-modal Image-to-image (I2I) translation. To tackle the one-to-many relationship between input and output domains, previous works use complex training objectives to learn a latent embedding, jointly with the generator, that models the variability of the output dom…

Cited by 11PDFScholar
2021

The Lottery Ticket Hypothesis for Object Recognition

CVPR 2021poster

Recognition tasks, such as object recognition and keypoint estimation, have seen widespread adoption in recent years. Most state-of-the-art methods for these tasks use deep networks that are computationally expensive and have huge memory footprints. This makes it exceedingly difficult to deploy thes…

Cited by 76PDFcodeScholar
2021

The Pursuit of Knowledge: Discovering and Localizing Novel Categories Using Dual Memory

ICCV 2021poster

We tackle object category discovery, which is the problem of discovering and localizing novel objects in a large unlabeled dataset. While existing methods show results on datasets with less cluttered scenes and fewer object instances per image, we present our results on the challenging COCO dataset.…

Cited by 16PDFScholar
2021

Towards Discovery and Attribution of Open-World GAN Generated Images

ICCV 2021poster

With the recent progress in Generative Adversarial Networks (GANs), it is imperative for media and visual forensics to develop detectors which can identify and attribute images to the model generating them. Existing works have shown to attribute images to their corresponding GAN sources with high ac…

Cited by 71PDFcodeScholar
2020

A Generic Visualization Approach for Convolutional Neural Networks

ECCV 2020poster

Retrieval networks are essential for searching and indexing. Compared to classification networks, attention visualization for retrieval networks is hardly studied. We formulate attention visualization as a constrained optimization problem. We leverage the unit L2-Norm constraint as an attention filt…

2020

Curriculum Manager for Source Selection in Multi-Source Domain Adaptation

ECCV 2020poster

The performance of Multi-Source Unsupervised Domain Adaptation (MS-UDA) depends significantly on the effectiveness of transferring from labeled source domain samples. In this paper, we proposed an adversarial agent that learns a dynamic curriculum for source samples, called Curriculum Manager for So…

Cited by 151SourcePDFScholar
2020

Quantization Guided JPEG Artifact Correction

ECCV 2020poster

The JPEG image compression algorithm is the most popular method of image compression because of it’s ability for large compression ratios. However, to achieve such high compression, information is lost. For aggressive quantization settings, this leads to a noticeable reduction in image quality. Arti…

2020

Scalable Model Compression by Entropy Penalized Reparameterization

ICLR 2020poster

We describe a simple and general neural network weight compression approach, in which the network parameters (weights and biases) are represented in a “latent” space, amounting to a reparameterization. This space is equipped with a learned probability model, which is used to impose an entropy penalt…

Cited by 51SourceScholar
2019

Relational Action Forecasting

CVPR 2019oral

This paper focuses on multi-person action forecasting in videos. More precisely, given a history of H previous frames, the goal is to detect actors and to predict their future actions for the next T frames. Our approach jointly models temporal and spatial interactions among different actors by const…

Cited by 100PDFScholar
2018

Actor-centric Relation Network

ECCV 2018poster

Current state-of-the-art approaches for spatio-temporal action localization rely on detections at the frame level and model temporal context with 3D ConvNets. Here, we go one step further and model spatio-temporal relations to capture the interactions between human actors, relevant objects and scene…

Cited by 280SourcePDFScholar
2018

Tracking Emerges by Colorizing Videos

ECCV 2018poster

We use large amounts of unlabeled video to learn models for visual tracking without manual human supervision. We leverage the natural temporal coherency of color to create a model that learns to colorize gray-scale videos by copying colors from a reference frame. Quantitative and qualitative experim…

Cited by 497SourcePDFScholar
2017

A-Fast-RCNN: Hard Positive Generation via Adversary for Object Detection

CVPR 2017poster

How do we learn an object detector that is invariant to occlusions and deformations? Our current solution is to use a data-driven strategy -- collect large-scale datasets which have object instances under different conditions. The hope is that the final classifier can use these examples to learn inv…

Cited by 802PDFcodeScholar
2017

Revisiting Unreasonable Effectiveness of Data in Deep Learning Era

ICCV 2017spotlight

The success of deep learning in vision can be attributed to: (a) models with high capacity; (b) increased computational power; and (c) availability of large-scale labeled data. Since 2012, there have been significant advances in representation capabilities of the models and computational capabilitie…

Cited by 3433PDFScholar
2015

Watch and Learn: Semi-Supervised Learning for Object Detectors From Video

CVPR 2015poster

We present a semi-supervised approach that localizes multiple unknown object instances in long videos. We start with a handful of labeled boxes and iteratively learn and label hundreds of thousands of object instances. We propose criteria for reliable object detection and tracking for constraining t…

Cited by 157SourcePDFScholar