← Search

Himanshu Gaurav Singh

4 accepted papers

2026

Tracking by Predicting 3-D Gaussians Over Time

CVPR 2026

We propose Video Gaussian Masked Autoencoders (Video-GMAE), a self-supervised approach for representation learning that encodes a sequence of images into a set of Gaussian splats moving over time. Representing a video as a set of Gaussians enforces a reasonable inductive bias: that 2-D videos are of

Cited by 0SourcecodeScholar
2025

Hand-Object Interaction Pretraining from Videos

ICRA 2025

We present an approach to learn general robot manipulation priors from 3D hand-object interaction trajectories. We build a framework to use in-the-wild videos to generate sensorimotor robot trajectories. We do so by lifting both the human hand and the manipulated object in a shared 3D space and reta

Cited by 46SourcecodeScholar
2023

Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models

EMNLP 2023long main

The performance of large language models (LLMs) on existing reasoning benchmarks has significantly improved over the past years. In response, we present JEEBench, a considerably more challenging benchmark dataset for evaluating the problem solving abilities of LLMs. We curate 515 challenging pre-eng…

Cited by 0SourcecodeScholar