← Search

Suhas Lohit

16 accepted papers

2026

AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects

CVPR 2026

Assembling objects from parts requires understanding multimodal instructions, linking them to 3D components, and predicting physically plausible 6-DoF motions for each assembly step. Existing datasets focus on simplified scenarios, overlooking shape complexities and assembly trajectories in industri

Cited by 0SourceScholar
2026

Point4Cast: Streaming Dynamic Scene Reconstruction and Forecasting

CVPR 2026

Understanding how the 3D world evolves over time is a fundamental task in computer vision, essential for embodied settings, autonomous driving, etc. It requires not only the reconstruction of the observed scene but also the anticipation of how the scene dynamics will unfold in the future. While the

Cited by 0SourceScholar
2026

Understanding Dynamic Compute Allocation in Recurrent Transformers

ICML 2026poster

Token-level adaptive computation seeks to reduce inference cost by allocating more computation to harder tokens and less to easier ones. However, prior work is primarily evaluated on natural-language benchmarks using task-level metrics, where token-level difficulty is unobservable and confounded wit…

Cited by 0SourceScholar
2024

Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads

NeurIPS 2024poster

Recent years have seen a significant progress in the general-purpose problem solving abilities of large vision and language models (LVLMs), such as ChatGPT, Gemini, etc.; some of these breakthroughs even seem to enable AI models to outperform human abilities in varied tasks that demand higher-order…

Cited by 10SourcePDFScholar
2024

Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-aware Spatio-Temporal Sampling

CVPR 2024highlight

Extensions of Neural Radiance Fields (NeRFs) to model dynamic scenes have enabled their near photo-realistic free-viewpoint rendering. Although these methods have shown some potential in creating immersive experiences two drawbacks limit their ubiquity: (i) a significant reduction in reconstruction…

Cited by 4SourcePDFScholar
2024

TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models

CVPR 2024poster

Text-conditioned image-to-video generation (TI2V) aims to synthesize a realistic video starting from a given image (e.g. a woman's photo) and a text description (e.g. "a woman is drinking water."). Existing TI2V frameworks often require costly training on video-text datasets and specific model desig…

2023

Are Deep Neural Networks SMARTer Than Second Graders?

CVPR 2023poster

Recent times have witnessed an increasing number of applications of deep neural networks towards solving tasks that require superior cognitive abilities, e.g., playing Go, generating art, question answering (such as ChatGPT), etc. Such a dramatic progress raises the question: how generalizable are n…

2023

Robust Time Series Recovery and Classification Using Test-Time Noise Simulator Networks

ICASSP 2023accepted

Time-series are commonly susceptible to various types of corruption due to sensor-level changes and defects which can result in missing samples, sensor and quantization noise, unknown calibration, unknown phase shifts etc. These corruptions cannot be easily corrected as the noise model may be unknow…

Cited by 0SourceScholar
2023

Steered Diffusion: A Generalized Framework for Plug-and-Play Conditional Image Synthesis

ICCV 2023poster

Conditional generative models typically demand large annotated training sets to achieve high-quality synthesis. As a result, there has been significant interest in designing models that perform plug-and-play generation, i.e., to use a predefined or pretrained model, which is not explicitly trained o…

Cited by 13PDFcodeScholar
2022

Cross-Modal Knowledge Transfer without Task-Relevant Source Data

ECCV 2022poster

"Cost-effective depth and infrared sensors as alternatives to usual RGB sensors are now a reality, and have some advantages over RGB in domains like autonomous navigation and remote sensing. As such, building computer vision and deep learning systems for depth and infrared data are crucial. However,…

Cited by 19SourcePDFScholar
2022

What Makes a "Good" Data Augmentation in Knowledge Distillation - A Statistical Perspective

NeurIPS 2022accept

Knowledge distillation (KD) is a general neural network training approach that uses a teacher model to guide the student model. Existing works mainly study KD from the network output side (e.g., trying to design a better KD loss function), while few have attempted to understand it from the input sid…

2019

Temporal Transformer Networks: Joint Learning of Invariant and Discriminative Time Warping

CVPR 2019poster

Many time-series classification problems involve developing metrics that are invariant to temporal misalignment. In human activity analysis, temporal misalignment arises due to various reasons including differing initial phase, sensor sampling rates, and elastic time-warps due to subject-specific bi…

Cited by 86PDFScholar
2019

Unrolled Projected Gradient Descent for Multi-spectral Image Fusion

ICASSP 2019accepted

In this paper, we consider the problem of fusing low spatial resolution multi-spectral (MS) aerial images with their associated high spatial resolution panchromatic image. To solve this problem, various methods have been proposed, using either model-based or model-agnostic algorithms such as deep le…

Cited by 0SourceScholar
2016

ReconNet: Non-Iterative Reconstruction of Images From Compressively Sensed Measurements

CVPR 2016poster

The goal of this paper is to present a non-iterative and more importantly an extremely fast algorithm to reconstruct images from compressively sensed (CS) random measurements. To this end, we propose a novel convolutional neural network (CNN) architecture which takes in CS measurements of an image…

Cited by 854PDFScholar