← Search

Jan Eric Lenssen

28 accepted papers

2026

AnyUp: Universal Feature Upsampling

ICLR 2026oral

We introduce AnyUp, a method for feature upsampling that can be applied to any vision feature at any resolution, without encoder-specific training. Existing learning-based upsamplers for features like DINO or CLIP need to be re-trained for every feature extractor and thus do not generalize to differ…

Cited by 0SourcecodeScholar
2026

MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation

CVPR 2026

We introduce MoLingo, a text-to-motion (T2M) model that generates realistic, lifelike human motion by denoising in a continuous latent space. Recent works perform latent space diffusion, either on the whole latent at once or auto-regressively over multiple latents. In this paper, we study how to mak

Cited by 0SourceScholar
2026

Rewis3d: Reconstruction Improves Weakly-Supervised Semantic Segmentation

CVPR 2026

We present Rewis3d, a framework that leverages recent advances in feed-forward 3D reconstruction to significantly improve weakly supervised semantic segmentation on 2D images. Obtaining dense, pixel-level annotations remains a costly bottleneck for training segmentation models. Alleviating this issu

Cited by 0SourcecodeScholar
2026

SemanticNVS: Improving Semantic Scene Understanding in Generative Novel View Synthesis

ICML 2026poster

We present SemanticNVS, a camera-conditioned multi-view diffusion model for novel view synthesis (NVS), which improves generation quality and consistency by integrating pre-trained semantic feature extractors. Existing NVS methods perform well for views near the input view, however, they tend to gen…

Cited by 0SourceScholar
2025

ContextGNN: Beyond Two-Tower Recommendation Systems

ICLR 2025poster

Recommendation systems predominantly utilize two-tower architectures, which evaluate user-item rankings through the inner product of their respective embeddings. However, one key limitation of two-tower models is that they learn a pair-agnostic representation of users and items. In contrast, pair-wi…

2025

MET3R: Measuring Multi-View Consistency in Generated Images

CVPR 2025poster

We introduce MEt3R, a metric for multi-view consistency in generated images. Large-scale generative models for multi-view image generation are rapidly advancing the field of 3D inference from sparse observations. However, due to the nature of generative modeling, traditional reconstruction metrics a…

Cited by 1SourcePDFScholar
2025

PersonaHOI: Effortlessly Improving Face Personalization in Human-Object Interaction Generation

CVPR 2025poster

We introduce PersonaHOI, a training- and tuning-free framework that fuses a general StableDiffusion model with a personalized face diffusion (PFD) model to generate identity-consistent human-object interaction (HOI) images. While existing PFD models have advanced significantly, they often overemphas…

2025

Solving Inverse Problems with FLAIR

NeurIPS 2025poster

Flow-based latent generative models such as Stable Diffusion 3 are able to generate images with remarkable quality, even enabling photorealistic text-to-image generation. Their impressive performance suggests that these models should also constitute powerful priors for inverse imaging problems, but…

Cited by 0SourcecodeScholar
2025

TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters

ICLR 2025spotlight

Transformers have become the predominant architecture in foundation models due to their excellent performance across various domains. However, the substantial cost of scaling these models remains a significant concern. This problem arises primarily from their dependence on a fixed number of paramete…

2024

From Similarity to Superiority: Channel Clustering for Time Series Forecasting

NeurIPS 2024poster

Time series forecasting has attracted significant attention in recent decades. Previous studies have demonstrated that the Channel-Independent (CI) strategy improves forecasting performance by treating different channels individually, while it leads to poor generalization on unseen instances and…

2024

GEARS: Local Geometry-aware Hand-object Interaction Synthesis

CVPR 2024poster

Generating realistic hand motion sequences in interaction with objects has gained increasing attention with the growing interest in digital humans. Prior work has illustrated the effectiveness of employing occupancy-based or distance-based virtual sensors to extract hand-object interaction features.…

Cited by 10SourcePDFScholar
2024

Improving 2D Feature Representations by 3D-Aware Fine-Tuning

ECCV 2024poster

"Current visual foundation models are trained purely on unstructured 2D data, limiting their understanding of 3D structure of objects and scenes. In this work, we show that fine-tuning on 3D-aware data improves the quality of emerging semantic features. We design a method to lift semantic 2D feature…

2024

NRDF: Neural Riemannian Distance Fields for Learning Articulated Pose Priors

CVPR 2024highlight

Faithfully modeling the space of articulations is a crucial task that allows recovery and generation of realistic poses and remains a notorious challenge. To this end we introduce Neural Riemannian Distance Fields (NRDFs) data-driven priors modeling the space of plausible articulations represented a…

Cited by 11SourcePDFScholar
2024

Neural Parametric Gaussians for Monocular Non-Rigid Object Reconstruction

CVPR 2024poster

Reconstructing dynamic objects from monocular videos is a severely underconstrained and challenging problem and recent work has approached it in various directions. However owing to the ill-posed nature of this problem there has been no solution that can provide consistent high-quality novel views f…

Cited by 22SourcePDFScholar
2024

Neural Point Cloud Diffusion for Disentangled 3D Shape and Appearance Generation

CVPR 2024poster

Controllable generation of 3D assets is important for many practical applications like content creation in movies games and engineering as well as in AR/VR. Recently diffusion models have shown remarkable results in generation quality of 3D objects. However none of the existing models enable disenta…

2024

Position: Relational Deep Learning - Graph Representation Learning on Relational Databases

ICML 2024poster

Much of the world's most valued data is stored in relational databases and data warehouses, where the data is organized into tables connected by primary-foreign key relations. However, building machine learning models using this data is both challenging and time consuming because no ML algorithm can…

Cited by 12SourcePDFScholar
2024

RelBench: A Benchmark for Deep Learning on Relational Databases

NeurIPS 2024poster

We present RelBench, a public benchmark for solving predictive tasks in relational databases with deep learning. RelBench provides databases and tasks spanning diverse domains, scales, and database dimensions, and is intended to be a foundational infrastructure for future research in this direction…

Cited by 11SourcePDFScholar
2024

Scribbles for All: Benchmarking Scribble Supervised Segmentation Across Datasets

NeurIPS 2024spotlight

In this work, we introduce *Scribbles for All*, a label and training data generation algorithm for semantic segmentation trained on scribble labels. Training or fine-tuning semantic segmentation models with weak supervision has become an important topic recently and was subject to significant advanc…

2024

Template Free Reconstruction of Human-object Interaction with Procedural Interaction Generation

CVPR 2024highlight

Reconstructing human-object interaction in 3D from a single RGB image is a challenging task and existing data driven methods do not generalize beyond the objects present in the carefully curated 3D interaction datasets. Capturing large-scale real data to learn strong interaction and 3D shape priors…

Cited by 14SourcePDFScholar
2022

Pose-NDF: Modeling Human Pose Manifolds with Neural Distance Fields

ECCV 2022poster

"We present Pose-NDF, a continuous model for plausible human poses based on neural distance fields (NDFs). Pose or motion priors are important for generating realistic new poses and for reconstructing accurate poses from noisy or partial observations. Pose-NDF learns a manifold of plausible poses as…

2022

TOCH: Spatio-Temporal Object-to-Hand Correspondence for Motion Refinement

ECCV 2022poster

"We present TOCH, a method for refining incorrect 3D hand-object interaction sequences using a data prior. Existing hand trackers, especially those that rely on very few cameras, often produce visually unrealistic results with hand-object intersection or missing contacts. Although correcting such er…

Cited by 56SourcePDFScholar
2020

Quaternion Equivariant Capsule Networks for 3D Point Clouds

ECCV 2020poster

We present a 3D capsule module for processing point clouds that is equivariant to 3D rotations and translations, as well as invariant to permutations of the input points. The operator receives a sparse set of local reference frames, computed from an input point cloud and establishes end-to-end trans…

Cited by 112SourcePDFScholar
2018

SplineCNN: Fast Geometric Deep Learning With Continuous B-Spline Kernels

CVPR 2018poster

We present Spline-based Convolutional Neural Networks (SplineCNNs), a variant of deep neural networks for irregular structured and geometric input, e.g., graphs or meshes. Our main contribution is a novel convolution operator based on B-splines, that makes the computation time independent from the k…