← Search

Sai Saketh Rambhatla

9 accepted papers

2025

Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition

ICCV 2025poster

Video understanding requires effective modeling of both motion and appearance information, particularly for few-shot action recognition. While recent advances in point tracking have been shown to improve few-shot action recognition, two fundamental challenges persist: selecting informative points to…

Cited by 0SourcePDFScholar
2024

Factorizing Text-to-Video Generation by Explicit Image Conditioning

ECCV 2024poster

"We present , a text-to-video generation model that factorizes the generation into two steps: first generating an image conditioned on the text, and then generating a video conditioned on the text and the generated image. We identify critical design decisions–adjusted noise schedules for diffusion,…

Cited by 84SourcePDFScholar
2024

InstanceDiffusion: Instance-level Control for Image Generation

CVPR 2024poster

Text-to-image diffusion models produce high quality images but do not offer control over individual instances in the image. We introduce InstanceDiffusion that adds precise instance-level control to text-to-image diffusion models. InstanceDiffusion supports free-form language conditions per instance…

2024

Trajectory-aligned Space-time Tokens for Few-shot Action Recognition

ECCV 2024poster

"We propose a simple yet effective approach for few-shot action recognition, emphasizing the disentanglement of motion and appearance representations. By harnessing recent progress in tracking, specifically point trajectories and self-supervised representation learning, we build trajectory-aligned t…

Cited by 2SourcePDFScholar
2023

MOST: Multiple Object Localization with Self-Supervised Transformers for Object Discovery

ICCV 2023oral

We tackle the challenging task of unsupervised object localization in this work. Recently, transformers trained with self-supervised learning have been shown to exhibit object localization properties without being trained for this task. In this work, we present Multiple Object localization with Self…

Cited by 12PDFcodeScholar
2021

The Pursuit of Knowledge: Discovering and Localizing Novel Categories Using Dual Memory

ICCV 2021poster

We tackle object category discovery, which is the problem of discovering and localizing novel objects in a large unlabeled dataset. While existing methods show results on datasets with less cluttered scenes and fewer object instances per image, we present our results on the challenging COCO dataset.…

Cited by 16PDFScholar
2021

Towards Discovery and Attribution of Open-World GAN Generated Images

ICCV 2021poster

With the recent progress in Generative Adversarial Networks (GANs), it is imperative for media and visual forensics to develop detectors which can identify and attribute images to the model generating them. Existing works have shown to attribute images to their corresponding GAN sources with high ac…

Cited by 71PDFcodeScholar
2019

A Dual-Path Model With Adaptive Attention for Vehicle Re-Identification

ICCV 2019oral

In recent years, attention models have been extensively used for person and vehicle re-identification. Most re-identification methods are designed to focus attention on key-point locations. However, depending on the orientation, the contribution of each key-point varies. In this paper, we present a…

Cited by 291PDFcodeScholar
2016

Camera based estimation of respiration rate by analyzing shape and size variation of structured light

ICASSP 2016accepted

Respiration rate is a key parameter that is monitored in intensive care units. The current solutions for respiration rate require that a sensor is placed in contact with the subject to derive it. However, this will cause discomfort to the subject and may damage the fragile skin if the subject were a…

Cited by 0SourceScholar