← Search

Amnon Geifman

9 accepted papers

2025

FFN Fusion: Rethinking Sequential Computation in Large Language Models

NeurIPS 2025spotlight

We introduce \textit{FFN Fusion}, an architectural optimization technique that reduces sequential computation in large language models by identifying and exploiting natural opportunities for parallelization. Our key insight is that sequences of Feed-Forward Network (FFN) layers, particularly those r…

Cited by 0SourceScholar
2025

Puzzle: Distillation-Based NAS for Inference-Optimized LLMs

ICML 2025poster

Large language models (LLMs) offer remarkable capabilities, yet their high inference costs restrict wider adoption. While increasing parameter counts improves accuracy, it also broadens the gap between state-of-the-art capabilities and practical deployability. We present **Puzzle**, a hardware-aware…

Cited by 2SourcePDFScholar
2023

A Kernel Perspective of Skip Connections in Convolutional Networks

ICLR 2023top-5%

Over-parameterized residual networks (ResNets) are amongst the most successful convolutional neural architectures for image processing. Here we study their properties through their Gaussian Process and Neural Tangent kernels. We derive explicit formulas for these kernels, analyze their spectra, and…

Cited by 13SourcePDFScholar
2022

On the Spectral Bias of Convolutional Neural Tangent and Gaussian Process Kernels

NeurIPS 2022accept

We study the properties of various over-parameterized convolutional neural architectures through their respective Gaussian Process and Neural Tangent kernels. We prove that, with normalized multi-channel input and ReLU activation, the eigenfunctions of these kernels with the uniform measure are form…

Cited by 16SourcePDFScholar
2020

Averaging Essential and Fundamental Matrices in Collinear Camera Settings

CVPR 2020poster

Global methods to Structure from Motion have gained popularity in recent years. A significant drawback of global methods is their sensitivity to collinear camera settings. In this paper, we introduce an analysis and algorithms for averaging bifocal tensors (essential or fundamental matrices) when ei…

Cited by 15PDFScholar
2020

Frequency Bias in Neural Networks for Input of Non-Uniform Density

ICML 2020poster

Recent works have partly attributed the generalization ability of over-parameterized neural networks to frequency bias – networks trained with gradient descent on data drawn from a uniform distribution find a low frequency fit before high frequency ones. As realistic training sets are not drawn from…

Cited by 220SourcePDFScholar
2020

On the Similarity between the Laplace and Neural Tangent Kernels

NeurIPS 2020poster

Recent theoretical work has shown that massively overparameterized neural networks are equivalent to kernel regressors that use Neural Tangent Kernels (NTKs). Experiments show that these kernel methods perform similarly to real neural networks. Here we show that NTK for fully connected networks wi…

Cited by 117SourcePDFScholar
2019

Algebraic Characterization of Essential Matrices and Their Averaging in Multiview Settings

ICCV 2019poster

Essential matrix averaging, i.e., the task of recovering camera locations and orientations in calibrated, multiview settings, is a first step in global approaches to Euclidean structure from motion. A common approach to essential matrix averaging is to separately solve for camera orientations and su…

Cited by 41PDFScholar
2019

GPSfM: Global Projective SFM Using Algebraic Constraints on Multi-View Fundamental Matrices

CVPR 2019poster

This paper addresses the problem of recovering projective camera matrices from collections of fundamental matrices in multiview settings. We make two main contributions. First, given n \choose 2 fundamental matrices computed for n images, we provide a complete algebraic characterization in the for…

Cited by 34PDFScholar