← Search

Lars Petersson

28 accepted papers

2026

DTO-KD: Dynamic Trade-off Optimization for Effective Knowledge Distillation

ICLR 2026oral

Knowledge Distillation (KD) is a widely adopted framework for compressing large models into compact student models by transferring knowledge from a high-capacity teacher. Despite its success, KD presents two persistent challenges: (1) the trade-off between optimizing for the primary task loss and mi…

Cited by 0SourceScholar
2025

GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction

CVPR 2025poster

3D modeling of highly reflective objects remains challenging due to strong view-dependent appearances. While previous SDF-based methods can recover high-quality meshes, they are often time-consuming and tend to produce over-smoothed surfaces. In contrast, 3D Gaussian Splatting (3DGS) offers the adva…

Cited by 0SourcePDFScholar
2025

Open Set Label Shift with Test Time Out-of-Distribution Reference

CVPR 2025poster

Open set label shift (OSLS) occurs when label distributions change from a source to a target distribution, and the target distribution has an additional out-of-distribution (OOD) class.In this work, we build estimators for both source and target open set label distributions using a source domain in-…

2024

Backpropagation-free Network for 3D Test-time Adaptation

CVPR 2024poster

Real-world systems often encounter new data over time which leads to experiencing target domain shifts. Existing Test-Time Adaptation (TTA) methods tend to apply computationally heavy and memory-intensive backpropagation-based approaches to handle this. Here we propose a novel method that uses a bac…

2024

Canonical Shape Projection is All You Need for 3D Few-shot Class Incremental Learning

ECCV 2024poster

"In recent years, robust pre-trained foundation models have been successfully used in many downstream tasks. Here, we would like to use such powerful models to address the problem of few-shot class incremental learning (FSCIL) tasks on 3D point cloud objects. Our approach is to reprogram the well-kn…

2023

Hyperbolic Audio-visual Zero-shot Learning

ICCV 2023poster

Audio-visual zero-shot learning aims to classify samples consisting of a pair of corresponding audio and video sequences from classes that are not present during training. An analysis of the audio-visual data reveals a large degree of hyperbolicity, indicating the potential benefit of using a hyperb…

Cited by 22PDFScholar
2022

Blind Image Decomposition

ECCV 2022poster

"We propose and study a novel task named Blind Image Decomposition (BID), which requires separating a superimposed image into constituent underlying images in a blind setting, that is, both the source components involved in mixing as well as the mixing mechanism are unknown. For example, rain may co…

2022

Declarative nets that are equilibrium models

ICLR 2022poster

Implicit layers are computational modules that output the solution to some problem depending on the input and the layer parameters. Deep equilibrium models (DEQs) output a solution to a fixed point equation. Deep declarative networks (DDNs) solve an optimisation problem in their forward pass, an arg…

Cited by 7SourcePDFScholar
2022

You Only Cut Once: Boosting Data Augmentation with a Single Cut

ICML 2022spotlight

We present You Only Cut Once (YOCO) for performing data augmentations. YOCO cuts one image into two pieces and performs data augmentations individually within each piece. Applying YOCO improves the diversity of the augmentation per sample and encourages neural networks to recognize objects from part…

2021

Contextually Plausible and Diverse 3D Human Motion Prediction

ICCV 2021poster

We tackle the task of diverse 3D human motion prediction, that is, forecasting multiple plausible future 3D poses given a sequence of observed 3D poses. In this context, a popular approach consists of using a Conditional Variational Autoencoder (CVAE). However, existing approaches that do so either…

Cited by 52PDFcodeScholar
2021

Reinforced Attention for Few-Shot Learning and Beyond

CVPR 2021poster

Few-shot learning aims to correctly recognize query samples from unseen classes given a limited number of support samples, often by relying on global embeddings of images. In this paper, we propose to equip the backbone network with an attention agent, which is trained by reinforcement learning. The…

Cited by 53PDFScholar
2021

Semantic-Aware Knowledge Distillation for Few-Shot Class-Incremental Learning

CVPR 2021poster

Few-shot class incremental learning (FSCIL) portrays the problem of learning new concepts gradually, where only a few examples per concept are available to the learner. Due to the limited number of examples for training, the techniques developed for standard incremental learning cannot be applied ve…

Cited by 241PDFScholar
2021

Synthesized Feature Based Few-Shot Class-Incremental Learning on a Mixture of Subspaces

ICCV 2021poster

Few-shot class incremental learning (FSCIL) aims to incrementally add sets of novel classes to a well-trained base model in multiple training sessions with the restriction that only a few novel instances are available per class. While learning novel classes, FSCIL methods gradually forget base (old)…

Cited by 86PDFScholar
2020

A Stochastic Conditioning Scheme for Diverse Human Motion Prediction

CVPR 2020poster

Human motion prediction, the task of predicting future 3D human poses given a sequence of observed ones, has been mostly treated as a deterministic problem. However, human motion is a stochastic process: Given an observed sequence of poses, multiple future motions are plausible. Existing approaches…

Cited by 148PDFcodeScholar
2020

Transferring Cross-Domain Knowledge for Video Sign Language Recognition

CVPR 2020oral

Word-level sign language recognition (WSLR) is a fundamental task in sign language interpretation. It requires models to recognize isolated sign words from videos. However, annotating WSLR data needs expert knowledge, thus limiting WSLR dataset acquisition. On the contrary, there are abundant subtit…

Cited by 164PDFScholar
2019

Bilinear Attention Networks for Person Retrieval

ICCV 2019poster

This paper investigates a novel Bilinear attention (Bi-attention) block, which discovers and uses second order statistical information in an input feature map, for the purpose of person retrieval. The Bi-attention block uses bilinear pooling to model the local pairwise feature interactions along eac…

Cited by 184PDFScholar
2019

The Alignment of the Spheres: Globally-Optimal Spherical Mixture Alignment for Camera Pose Estimation

CVPR 2019poster

Determining the position and orientation of a calibrated camera from a single image with respect to a 3D model is an essential task for many applications. When 2D-3D correspondences can be obtained reliably, perspective-n-point solvers can be used to recover the camera pose. However, without the pos…

Cited by 42PDFScholar
2018

Effective Use of Synthetic Data for Urban Scene Semantic Segmentation

ECCV 2018poster

Training a deep network to perform semantic segmentation requires large amounts of labeled data. To alleviate the manual effort of annotating real images, researchers have investigated the use of synthetic data, which can be labeled automatically. Unfortunately, a network trained on synthetic data p…

2018

Improving Object Localization With Fitness NMS and Bounded IoU Loss

CVPR 2018poster

We demonstrate that many detection methods are designed to identify only a sufficently accurate bounding box, rather than the best available one. To address this issue we propose a simple and fast modification to the existing methods called Fitness NMS. This method is tested with the DeNet model and…

2017

Bringing Background Into the Foreground: Making All Classes Equal in Weakly-Supervised Video Semantic Segmentation

ICCV 2017poster

Pixel-level annotations are expensive and time-consuming to obtain. Hence, weak supervision using only image tags could have a significant impact in semantic segmentation. Recent years have seen great progress in weakly-supervised semantic segmentation, whether from a single image or from videos. Ho…

Cited by 47PDFScholar
2017

Encouraging LSTMs to Anticipate Actions Very Early

ICCV 2017poster

In contrast to the widely studied problem of recognizing an action given a complete sequence, action anticipation aims to identify the action from only partially available videos. As such, it is therefore key to the success of computer vision applications requiring to react as early as possible, suc…

Cited by 212PDFScholar
2017

Globally-Optimal Inlier Set Maximisation for Simultaneous Camera Pose and Feature Correspondence

ICCV 2017oral

Estimating the 6-DoF pose of a camera from a single image relative to a pre-computed 3D point-set is an important task for many computer vision applications. Perspective-n-Point (PnP) solvers are routinely used for camera pose estimation, provided that a good quality set of 2D-3D feature corresponde…

Cited by 73PDFScholar
2016

Sample and Filter: Nonparametric Scene Parsing via Efficient Filtering

CVPR 2016poster

Scene parsing has attracted a lot of attention in computer vision. While parametric models have proven effective for this task, they cannot easily incorporate new training data. By contrast, nonparametric approaches, which bypass any learning phase and directly transfer the labels from the training…

Cited by 15PDFScholar
2015

Cutting Edge: Soft Correspondences in Multimodal Scene Parsing

ICCV 2015poster

Exploiting multiple modalities for semantic scene parsing has been shown to improve accuracy over the single modality scenario. Existing methods, however, assume that corresponding regions in two modalities have the same label. In this paper, we address the problem of data misalignment and label inc…

Cited by 11PDFScholar