← Search

Neehar Peri

17 accepted papers

2026

RF-DETR: Neural Architecture Search for Real-Time Detection Transformers

ICLR 2026poster

Open-vocabulary detectors achieve impressive performance on COCO, but often fail to generalize to real-world datasets with out-of-distribution classes not typically found in their pre-training. Rather than simply fine-tuning a heavy-weight vision-language model (VLM) for new domains, we introduce RF…

Cited by 0SourcecodeScholar
2025

MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion

ICCV 2025poster

We address the problem of dynamic scene reconstruction from sparse-view videos. Prior work often requires dense multi-view captures with hundreds of calibrated cameras (e.g. Panoptic Studio) - such multi-view setups are prohibitively expensive to build and cannot capture diverse scenes in-the-wild.…

Cited by 0SourcePDFScholar
2025

Neural Eulerian Scene Flow Fields

ICLR 2025poster

We reframe scene flow as the task of estimating a continuous space-time ordinary differential equation (ODE) that describes motion for an entire observation sequence, represented with a neural prior. Our method, EulerFlow, optimizes this neural prior estimate against several multi-observation recons…

Cited by 1SourcePDFScholar
2025

Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models

NeurIPS 2025poster

Vision-language models (VLMs) trained on internet-scale data achieve remarkable zero-shot detection performance on common objects like car, truck, and pedestrian. However, state-of-the-art models still struggle to generalize to out-of-distribution classes, tasks and imaging modalities not typically…

Cited by 0SourcecodeScholar
2025

Towards Learning to Complete Anything in Lidar

ICML 2025poster

We propose CAL (Complete Anything in Lidar) for Lidar-based shape-completion in-the-wild. This is closely related to Lidar-based semantic/panoptic scene completion. However, contemporary methods can only complete and recognize objects from a closed vocabulary labeled in existing Lidar datasets. Diff…

Cited by 0SourcePDFScholar
2024

Better Call SAL: Towards Learning to Segment Anything in Lidar

ECCV 2024poster

"We propose the (Segment Anything in Lidar) method consisting of a text-promptable zero-shot model for segmenting and classifying any object in Lidar, and a pseudo-labeling engine that facilitates model training without manual supervision. While the established paradigm for (LPS) relies on manual su…

2024

Revisiting Few-Shot Object Detection with Vision-Language Models

NeurIPS 2024poster

The era of vision-language models (VLMs) trained on web-scale datasets challenges conventional formulations of “open-world" perception. In this work, we revisit the task of few-shot object detection (FSOD) in the context of recent foundational VLMs. First, we point out that zero-shot predictions fro…

2024

Shelf-Supervised Cross-Modal Pre-Training for 3D Object Detection

CoRL 2024poster

State-of-the-art 3D object detectors are often trained on massive labeled datasets. However, annotating 3D bounding boxes remains prohibitively expensive and time-consuming, particularly for LiDAR. Instead, recent works demonstrate that self-supervised pre-training with unlabeled data can improve de…

Cited by 0SourcecodeScholar
2024

ZeroFlow: Scalable Scene Flow via Distillation

ICLR 2024poster

Scene flow estimation is the task of describing the 3D motion field between temporally successive point clouds. State-of-the-art methods use strong priors and test-time optimization techniques, but require on the order of tens of seconds to process full-size point clouds, making them unusable as com…

2022

Forecasting From LiDAR via Future Object Detection

CVPR 2022poster

Object detection and forecasting are fundamental components of embodied perception. These two problems, however, are largely studied in isolation by the community. In this paper, we propose an end-to-end approach for motion forecasting based on raw sensor measurement as opposed to ground truth track…

Cited by 40PDFcodeScholar
2021

PreferenceNet: Encoding Human Preferences in Auction Design with Deep Learning

NeurIPS 2021poster

The design of optimal auctions is a problem of interest in economics, game theory and computer science. Despite decades of effort, strategyproof, revenue-maximizing auction designs are still not known outside of restricted settings. However, recent methods using deep learning have shown some success…

2020

The Devil is in the Details: Self-Supervised Attention for Vehicle Re-Identification

ECCV 2020poster

In recent years, the research community has approached the problem of vehicle re-identification (re-id) with attention-based models, specifically focusing on regions of a vehicle containing discriminative information. These re-id methods rely on expensive key-point labels, part annotations, and addi…

2019

A Dual-Path Model With Adaptive Attention for Vehicle Re-Identification

ICCV 2019oral

In recent years, attention models have been extensively used for person and vehicle re-identification. Most re-identification methods are designed to focus attention on key-point locations. However, depending on the orientation, the contribution of each key-point varies. In this paper, we present a…

Cited by 291PDFcodeScholar