← Search

Rohith Kukkala

3 accepted papers

2026

Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding

CVPR 2026

Recent advances in 3D vision-language models (VLMs) highlight a strong potential for 3D scene understanding and reasoning.However, effectively tokenizing 3D scenes into holistic scene tokens, and leveraging these tokens across diverse 3D understanding tasks, remain highly challenging. We present NDT

Cited by 0SourcecodeScholar
2026

Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animation

CVPR 2026

Generating dynamic 4D objects from sparse inputs is difficult because it demands joint preservation of appearance and motion coherence across views and time while suppressing artifacts and temporal drift. We hypothesize that the view discrepancy arises from supervision limited to pixel- or latent-sp

Cited by 0SourceScholar
2025

DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos

CVPR 2025poster

Long Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach taken by existing methods of dividing video into clips and processing each clip via a full-scale expert encoder is challen…