← Search

Amit H. Bermano

10 accepted papers

2026

CARLoS: Retrieval via Concise Assessment Representation of LoRAs at Scale

CVPR 2026

The rapid proliferation of generative components, such as LoRAs, has created a vast but unstructured ecosystem. Existing discovery methods depend on unreliable user descriptions or biased popularity metrics, hindering usability. We present CARLoS, a large-scale framework for characterizing LoRAs wit

Cited by 0SourcecodeScholar
2026

ImageRAG: Dynamic Image Retrieval for Reference-Guided Image Generation

ICLR 2026poster

While recent generative models synthesize high-quality visual content, they still struggle with generating rare or fine-grained concepts. To address this challenge, we explore the usage of Retrieval-Augmented Generation (RAG) for image generation, and introduce ImageRAG, a training-free method for r…

Cited by 0SourcecodeScholar
2025

Instant3dit: Multiview Inpainting for Fast Editing of 3D Objects

CVPR 2025poster

We propose a generative technique to edit 3D shapes, represented as meshes, NeRFs, or Gaussian Splats, in ~3 seconds, without the need for running an SDS type of optimization.Our key insight is to cast 3D editing as a multiview image inpainting problem, as this representation is generic and can be m…

Cited by 2SourcePDFScholar
2024

MAS: Multi-view Ancestral Sampling for 3D Motion Generation Using 2D Diffusion

CVPR 2024poster

We introduce Multi-view Ancestral Sampling (MAS) a method for 3D motion generation using 2D diffusion models that were trained on motions obtained from in-the-wild videos. As such MAS opens opportunities to exciting and diverse fields of motion previously under-explored as 3D data is scarce and hard…

2024

Performance Conditioning for Diffusion-Based Multi-Instrument Music Synthesis

ICASSP 2024accepted

Generating multi-instrument music from symbolic music representations is an important task in Music Information Retrieval (MIR). A central but still largely unsolved problem in this context is musically and acoustically informed control in the generation process. As the main contribution of this wor…

Cited by 0SourceScholar
2023

OReX: Object Reconstruction From Planar Cross-Sections Using Neural Fields

CVPR 2023poster

Reconstructing 3D shapes from planar cross-sections is a challenge inspired by downstream applications like medical imaging and geographic informatics. The input is an in/out indicator function fully defined on a sparse collection of planes in space, and the output is an interpolation of the indicat…

2022

MotionCLIP: Exposing Human Motion Generation to CLIP Space

ECCV 2022poster

"We introduce MotionCLIP, a 3D human motion auto-encoder featuring a latent embedding that is disentangled, well behaved, and supports highly semantic textual descriptions. MotionCLIP gains its unique power by aligning its latent space with that of the Contrastive Language-Image Pre-training (CLIP)…

2020

Unsupervised Multi-Modal Image Registration via Geometry Preserving Image-to-Image Translation

CVPR 2020poster

Many applications, such as autonomous driving, heavily rely on multi-modal data where spatial alignment between the modalities is required. Most multi-modal registration methods struggle computing the spatial correspondence between the images using prevalent cross-modality similarity measures. In th…

Cited by 180PDFScholar