← Search

Oswald Lanz

9 accepted papers

2026

Action-Guided Attention for Video Action Anticipation

ICLR 2026poster

Anticipating future actions in videos is challenging, as the observed frames provide only evidence of past activities, requiring the inference of latent intentions to predict upcoming actions. Existing transformer-based approaches, which rely on dot-product attention over pixel representations, ofte…

Cited by 0SourcecodeScholar
2025

L-SWAG: Layer-Sample Wise Activation with Gradients Information for Zero-Shot NAS on Vision Transformers

CVPR 2025poster

Training-free Neural Architecture Search (NAS) efficiently identifies high-performing neural networks using zero-cost (ZC) proxies. Unlike multi-shot and one-shot NAS approaches, ZC-NAS is both (i) time-efficient, eliminating the need for model training, and (ii) interpretable, with proxy designs of…

Cited by 0SourcePDFScholar
2024

Your Image is My Video: Reshaping the Receptive Field via Image-To-Video Differentiable AutoAugmentation and Fusion

CVPR 2024poster

The landscape of deep learning research is moving towards innovative strategies to harness the true potential of data. Traditionally emphasis has been on scaling model architectures resulting in large and complex neural networks which can be difficult to train with limited computational resources. H…

Cited by 0SourcePDFScholar
2019

Accurate Target Annotation in 3D from Multimodal Streams

ICASSP 2019accepted

Accurate annotation is fundamental to quantify the performance of multi-sensor and multi-modal object detectors and trackers. However, invasive or expensive instrumentation is needed to automatically generate these annotations. To mitigate this problem, we present a multi-modal approach that leverag…

Cited by 0SourceScholar
2019

LSTA: Long Short-Term Attention for Egocentric Action Recognition

CVPR 2019poster

Egocentric activity recognition is one of the most challenging tasks in video analysis. It requires a fine-grained discrimination of small objects and their manipulation. While some methods base on strong supervision and attention mechanisms, they are either annotation consuming or do not take spati…

Cited by 208PDFcodeScholar
2018

3D Mouth Tracking from a Compact Microphone Array Co-Located with a camera

ICASSP 2018accepted

We address the 3D audio-visual mouth tracking problem when using a compact platform with co-located audio-visual sensors, without a depth camera. In particular, we propose a multi-modal particle filter that combines a face detector and 3D hypothesis mapping to the image plane. The audio likelihood c…

Cited by 0SourceScholar
2015

Uncovering Interactions and Interactors: Joint Estimation of Head, Body Orientation and F-Formations From Surveillance Videos

ICCV 2015poster

We present a novel approach for jointly estimating tar- gets' head, body orientations and conversational groups called F-formations from a distant social scene (e.g., a cocktail party captured by surveillance cameras). Differing from related works that have (i) coupled head and body pose learning by…

Cited by 85PDFScholar