← Search

Matheus Gadelha

22 accepted papers

2026

3D Space as a Scratchpad for Editable Text-to-Image Generation

CVPR 2026

Recent progress in large language models (LLMs) has shown that reasoning improves when intermediate thoughts are externalized into explicit workspaces, such as chain-of-thought traces or tool-augmented reasoning. Yet, visual language models (VLMs) lack an analogous mechanism for spatial reasoning, l

Cited by 0SourcecodeScholar
2026

Material Magic Wand: Material-Aware Grouping of 3D Parts in Untextured Meshes

CVPR 2026

We introduce the problem of material-aware part grouping in untextured meshes.Many real-world shapes, such as scales of pinecones or windows of buildings, contain repeated structures that share the same material but exhibit geometric variations.When assigning materials to such meshes, these repeated

Cited by 0SourceScholar
2026

MeshSplatting: Differentiable Rendering with Opaque Meshes

CVPR 2026

Primitive-based splatting methods like 3D Gaussian Splatting (3DGS) have revolutionized novel view synthesis with real-time rendering. However, their point-based representations remain incompatible with mesh-based pipelines that power AR/VR and game engines. We present Mesh Splatting, a mesh-based r

Cited by 0SourcecodeScholar
2026

Residual Primitive Fitting of 3D Shapes with SuperFrusta

CVPR 2026

We introduce a framework for converting 3D shapes into compact and editable assemblies of analytic primitives, directly addressing the persistent trade-off between reconstruction fidelity and parsimony. Our approach combines two key contributions: a novel primitive, termed SuperFrustum, and an itera

Cited by 0SourcecodeScholar
2026

SIGMA-GEN: STRUCTURE AND IDENTITY GUIDED MULTI-SUBJECT ASSEMBLY FOR IMAGE GENERATION

ICLR 2026poster

We present SIGMA-GEN, a unified framework for multi-identity preserving image generation. Unlike prior approaches, SIGMA-GEN is the first to enable single-pass multi-subject identity-preserved generation guided by both structural and spatial constraints. A key strength of our method is its ability t…

Cited by 0SourceScholar
2025

DMesh++: An Efficient Differentiable Mesh for Complex Shapes

ICCV 2025poster

Recent probabilistic methods for 3D triangular meshes capture diverse shapes by differentiable mesh connectivity, but face high computational costs with increased shape details. We introduce a new differentiable mesh processing method that addresses this challenge and efficiently handles meshes with…

2025

Frame In-N-Out: Unbounded Controllable Image-to-Video Generation

NeurIPS 2025poster

Controllability, temporal coherence, and detail synthesis remain the most critical challenges in video generation. In this paper, we focus on a commonly used yet underexplored cinematic technique known as Frame In and Frame Out. Specifically, starting from image-to-video generation, users can contro…

Cited by 0SourceScholar
2025

Instant3dit: Multiview Inpainting for Fast Editing of 3D Objects

CVPR 2025poster

We propose a generative technique to edit 3D shapes, represented as meshes, NeRFs, or Gaussian Splats, in ~3 seconds, without the need for running an SDS type of optimization.Our key insight is to cast 3D editing as a multiview image inpainting problem, as this representation is generic and can be m…

Cited by 2SourcePDFScholar
2025

Motion Modes: What Could Happen Next?

CVPR 2025poster

Predicting diverse object motions from a single static image remains challenging, as current video generation models often entangle object movement with camera motion and other scene changes. While recent methods can predict specific motions from motion arrow input, they rely on synthetic data and p…

Cited by 1SourcePDFScholar
2025

PreciseCam: Precise Camera Control for Text-to-Image Generation

CVPR 2025poster

Images as an artistic medium often rely on specific camera angles and lens distortions to convey ideas or emotions; however, such precise control is missing in current text-to-image models. We propose an efficient and general solution that allows precise control over the camera when generating both…

2025

Reusing Computation in Text-to-Image Diffusion for Efficient Generation of Image Sets

ICCV 2025poster

Text-to-image diffusion models enable high-quality image generation but are computationally expensive, especially when producing large image collections. While prior work optimizes per-inference efficiency, we explore an orthogonal approach: reducing redundancy across multiple correlated prompts. Ou…

Cited by 0SourcePDFScholar
2024

DMesh: A Differentiable Mesh Representation

NeurIPS 2024poster

We present a differentiable representation, DMesh, for general 3D triangular meshes. DMesh considers both the geometry and connectivity information of a mesh. In our design, we first get a set of convex tetrahedra that compactly tessellates the domain based on Weighted Delaunay Triangulation (WDT),…

2024

Diffusion Handles Enabling 3D Edits for Diffusion Models by Lifting Activations to 3D

CVPR 2024highlight

Diffusion handles is a novel approach to enable 3D object edits on diffusion images requiring only existing pre-trained diffusion models depth estimation without any fine-tuning or 3D object retrieval. The edited results remain plausible photo-real and preserve object identity. Diffusion handles add…

Cited by 20SourcePDFScholar
2024

Generative Rendering: Controllable 4D-Guided Video Generation with 2D Diffusion Models

CVPR 2024poster

Traditional 3D content creation tools empower users to bring their imagination to life by giving them direct control over a scene's geometry appearance motion and camera path. Creating computer-generated videos however is a tedious manual process which can be automated by emerging text-to-video diff…

Cited by 15SourcePDFScholar
2024

Learning Continuous 3D Words for Text-to-Image Generation

CVPR 2024poster

Current controls over diffusion models (e.g. through text or ControlNet) for image generation fall short in recognizing abstract continuous attributes like illumination direction or non-rigid shape change. In this paper we present an approach for allowing users of text-to-image models to have fine-g…

2023

3DMiner: Discovering Shapes from Large-Scale Unannotated Image Datasets

ICCV 2023poster

We present 3DMiner -- a pipeline for mining 3D shapes from challenging large-scale unannotated image datasets. Unlike other unsupervised 3D reconstruction methods, we assume that, within a large-enough dataset, there must exist images of objects with similar shapes but varying backgrounds, textures,…

Cited by 0PDFcodeScholar
2022

PlanarRecon: Real-Time 3D Plane Detection and Reconstruction From Posed Monocular Videos

CVPR 2022poster

We present PlanarRecon -- a novel framework for globally coherent detection and reconstruction of 3D planes from a posed monocular video. Unlike previous works that detect planes in 2D from a single image, PlanarRecon incrementally detects planes in 3D for each video fragment, which consists of a se…

Cited by 30PDFcodeScholar
2020

Label-Efficient Learning on Point Clouds using Approximate Convex Decompositions

ECCV 2020poster

The problems of shape classification and part segmentation from 3D point clouds have garnered increasing attention in the last few years. Both of these problems, however, suffer from relatively small training sets, creating the need for statistically efficient methods to learn 3D shape representatio…

2020

Learning Generative Models of Shape Handles

CVPR 2020poster

We present a generative model to synthesize 3D shapes as sets of handles -- lightweight proxies that approximate the original 3D shape -- for applications in interactive editing, shape parsing, and building compact 3D representations. Our model can generate handle sets with varying cardinality and d…

Cited by 34PDFScholar