← Search

Saksham Suri

10 accepted papers

2026

UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders

CVPR 2026

The space of task-agnostic feature upsampling has emerged as a promising area of research to efficiently create denser features from pre-trained visual backbones. These methods act as a shortcut to achieve dense features for a fraction of the cost by learning to map low-resolution features to high-r

Cited by 0SourcecodeScholar
2026

VeriGraph: Scene Graphs for Execution Verifiable Robot Planning

ICRA 2026poster

Recent advancements in vision-language models (VLMs) offer potential for robot task planning, but challenges remain due to VLMs’ tendency to generate incorrect action sequences. To address these limitations, we propose VeriGraph, a novel framework that integrates VLMs for robotic planning while veri…

2026

VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice

CVPR 2026

Chain-of-thought (CoT) reasoning has emerged as a powerful tool for multimodal large language models on video understanding tasks. However, its necessity and advantages over direct answering remain underexplored. In this paper, we first demonstrate that for RL-trained video models, direct answering

Cited by 0SourceScholar
2025

EdgeTAM: On-Device Track Anything Model

CVPR 2025poster

On top of Segment Anything Model (SAM), SAM 2 further extends its capability from image to video inputs through a memory bank mechanism and obtains a remarkable performance compared with previous methods, making it a foundation model for video segmentation task. In this paper, we aim at making SAM 2…

2025

LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior

ICLR 2025oral

We present LARP, a novel video tokenizer designed to overcome limitations in current video tokenization methods for autoregressive (AR) generative models. Unlike traditional patchwise tokenizers that directly encode local visual patches into discrete tokens, LARP introduces a holistic tokenization s…

2023

SparseDet: Improving Sparsely Annotated Object Detection with Pseudo-positive Mining

ICCV 2023poster

Training with sparse annotations is known to reduce the performance of object detectors. Previous methods have focused on proxies for missing ground truth annotations in the form of pseudo-labels for unlabeled boxes. We observe that existing methods suffer at higher levels of sparsity in the data du…

Cited by 13PDFScholar
2023

Teaching Matters: Investigating the Role of Supervision in Vision Transformers

CVPR 2023poster

Vision Transformers (ViTs) have gained significant popularity in recent years and have proliferated into many applications. However, their behavior under different learning paradigms is not well explored. We compare ViTs trained through different methods of supervision, and show that they learn a di…

2021

Learned Spatial Representations for Few-Shot Talking-Head Synthesis

ICCV 2021poster

We propose a novel approach for few-shot talking-head synthesis. While recent works in neural talking heads have produced promising results, they can still produce images that do not preserve the identity of the subject in source images. We posit this is a result of the entangled representation of e…

Cited by 49PDFScholar
2021

Towards Discovery and Attribution of Open-World GAN Generated Images

ICCV 2021poster

With the recent progress in Generative Adversarial Networks (GANs), it is imperative for media and visual forensics to develop detectors which can identify and attribute images to the model generating them. Existing works have shown to attribute images to their corresponding GAN sources with high ac…

Cited by 71PDFcodeScholar