← Search

Greg Shakhnarovich

20 accepted papers

2025

OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World

ICRA 2025

We would like to estimate the pose and full shape of an object from a single observation, without assuming known 3D model or category. In this work, we propose OmniShape, the first method of its kind to enable probabilistic pose and shape estimation. OmniShape is based on the key insight that shape

Cited by 1SourceScholar
2025

SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster Prediction

ACL 2025long

Sign language processing has traditionally relied on task-specific models, limiting the potential for transfer learning across tasks. Pre-training methods for sign language have typically focused on either supervised pre-training, which cannot take advantage of unlabeled data, or context-independent…

Cited by 0SourcePDFScholar
2025

SignMusketeers: An Efficient Multi-Stream Approach for Sign Language Translation at Scale

ACL 2025finding

A persistent challenge in sign language video processing, including the task of sign language to written language translation, is how we train efficient model given the nature of videos. Informed by the nature and linguistics of signed languages, our proposed method focuses on just the most relevant…

Cited by 0SourcePDFScholar
2025

SplArt: Articulation Estimation and Part-Level Reconstruction with 3D Gaussian Splatting

ICCV 2025poster

Reconstructing articulated objects prevalent in daily environments is crucial for applications in augmented/virtual reality and robotics. However, existing methods face scalability limitations (requiring 3D supervision or costly annotations), robustness issues (being susceptible to local optima), an…

2025

Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion

CVPR 2025poster

Current methods for 3D scene reconstruction from sparse posed images employ intermediate 3D representations such as neural fields, voxel grids, or 3D Gaussians, to achieve multi-view consistent scene appearance and geometry. In this paper we introduce MVGD, a diffusion-based architecture capable of…

Cited by 0SourcePDFScholar
2024

Alpha Invariance: On Inverse Scaling Between Distance and Volume Density in Neural Radiance Fields

CVPR 2024poster

Scale-ambiguity in 3D scene dimensions leads to magnitude-ambiguity of volumetric densities in neural radiance fields i.e. the densities double when scene size is halved and vice versa. We call this property alpha invariance. For NeRFs to better maintain alpha invariance we recommend 1) parameterizi…

Cited by 0SourcePDFScholar
2024

HyperFields: Towards Zero-Shot Generation of NeRFs from Text

ICML 2024poster

We introduce HyperFields, a method for generating text-conditioned Neural Radiance Fields (NeRFs) with a single forward pass and (optionally) some fine-tuning. Key to our approach are: (i) a dynamic hypernetwork, which learns a smooth mapping from text token embeddings to the space of NeRFs; (ii) Ne…

Cited by 10SourcePDFScholar
2024

Instant3D: Fast Text-to-3D with Sparse-view Generation and Large Reconstruction Model

ICLR 2024poster

Text-to-3D with diffusion models has achieved remarkable progress in recent years. However, existing methods either rely on score distillation-based optimization which suffer from slow inference, low diversity and Janus problems, or are feed-forward methods that generate low-quality results due to…

Cited by 250SourcePDFScholar
2023

Score Jacobian Chaining: Lifting Pretrained 2D Diffusion Models for 3D Generation

CVPR 2023poster

A diffusion model learns to predict a vector field of gradients. We propose to apply chain rule on the learned gradients, and back-propagate the score of a diffusion model through the Jacobian of a differentiable renderer, which we instantiate to be a voxel radiance field. This setup aggregates 2D s…

2022

Boosting Barely Robust Learners: A New Perspective on Adversarial Robustness

NeurIPS 2022accept

We present an oracle-efficient algorithm for boosting the adversarial robustness of barely robust learners. Barely robust learning algorithms learn predictors that are adversarially robust only on a small fraction $\beta \ll 1$ of the data distribution. Our proposed notion of barely robust learning…

Cited by 3SourcePDFScholar
2022

Depth Field Networks for Generalizable Multi-View Scene Representation

ECCV 2022poster

"Modern 3D computer vision leverages learning to boost geometric reasoning, mapping image data to classical structures such as cost volumes or epipolar constraints to improve matching. These architectures are specialized according to the particular problem, and thus require significant task-specific…

Cited by 16SourcePDFScholar
2022

Searching for fingerspelled content in American Sign Language

ACL 2022long

Natural language processing for sign language video—including tasks like recognition, translation, and search—is crucial for making artificial intelligence technologies accessible to deaf individuals, and is gaining research interest in recent years. In this paper, we address the problem of searchin…

Cited by 6SourcePDFScholar
2022

Self-Supervised Camera Self-Calibration from Video

ICRA 2022poster

Camera calibration is integral to robotics and computer vision algorithms that seek to infer geometric properties of the scene from visual input streams. In practice, calibration is a laborious procedure requiring specialized data collection and careful tuning. This process must be repeated whenever…

Cited by 31SourceScholar
2021

Information-Theoretic Segmentation by Inpainting Error Maximization

CVPR 2021poster

We study image segmentation from an information-theoretic perspective, proposing a novel adversarial method that performs unsupervised segmentation by partitioning images into maximally independent sets. More specifically, we group image pixels into foreground and background, with the goal of minimi…

Cited by 31PDFcodeScholar
2019

Fingerspelling Recognition in the Wild With Iterative Visual Attention

ICCV 2019poster

Sign language recognition is a challenging gesture sequence recognition problem, characterized by quick and highly coarticulated motion. In this paper we focus on recognition of fingerspelling sequences in American Sign Language (ASL) videos collected in the wild, mainly from YouTube and Deaf social…

Cited by 90PDFcodeScholar
2016

Depth from a Single Image by Harmonizing Overcomplete Local Network Predictions

NeurIPS 2016poster

A single color image can contain many cues informative towards different aspects of local geometric structure. We approach the problem of monocular depth estimation by using a neural network to produce a mid-level representation that summarizes these cues. This network is trained to characterize loc…

Cited by 161SourcePDFScholar