← Search

Francois Bremond

12 accepted papers

2026

3D Gaussian Splatting at Arbitrary Resolutions with Compact Proxy Anchors

CVPR 2026

Despite achieving high-quality rendering, 3D Gaussian Splatting suffers from aliasing when the rendering resolution changes, as it is typically trained at a fixed resolution. To address this limitation, we introduce a method that enables the model to generate resolution-adaptive 3D Gaussians under a

Cited by 0SourcecodeScholar
2025

Just Dance with pi! A Poly-modal Inductor for Weakly-supervised Video Anomaly Detection

CVPR 2025highlight

Weakly-supervised methods for video anomaly detection (VAD) are conventionally based merely on RGB spatio-temporal features, which continues to limit their reliability in real-world scenarios. This is due to the fact that RGB-features are not sufficiently distinctive in setting apart categories such…

2025

LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living

CVPR 2025poster

Current Large Language Vision Models (LLVMs) trained on web videos perform well in general video understanding but struggle with fine-grained details, complex human-object interactions (HOI), and view-invariant representation learning essential for Activities of Daily Living (ADL). This limitation s…

2025

Mixture of Experts Guided by Gaussian Splatters Matters: A new Approach to Weakly-Supervised Video Anomaly Detection

ICCV 2025poster

Video Anomaly Detection (VAD) is a challenging task due to the variability of anomalous events and the limited availability of labeled data. Under the Weakly-Supervised VAD (WSVAD) paradigm, only video-level labels are provided during training, while predictions are made at the frame level. Although…

2025

SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living

AAAI 2025technical

The introduction of vision-language models like CLIP has enabled the development of foundational video models capable of generalizing to unseen videos and human actions. However, these models are typically trained on web videos, which often fail to capture the challenges present in Activities of Dai…

2025

Scaling Action Detection: AdaTAD++ with Transformer-Enhanced Temporal-Spatial Adaptation

ICCV 2025poster

Temporal Action Detection (TAD) is essential for analyzing long-form videos by identifying and segmenting actions within untrimmed sequences. While recent innovations like Temporal Informative Adapters (TIA) have improved resolution, memory constraints still limit large video processing. To address…

Cited by 0SourcePDFScholar
2023

LAC - Latent Action Composition for Skeleton-based Action Segmentation

ICCV 2023poster

Skeleton-based action segmentation requires recognizing composable actions in untrimmed videos. Current approaches decouple this problem by first extracting local visual features from skeleton sequences and then processing them by a temporal model to classify frame-wise actions. However, their perfo…

Cited by 14PDFScholar
2023

StressID: a Multimodal Dataset for Stress Identification

NeurIPS 2023poster

StressID is a new dataset specifically designed for stress identification from unimodal and multimodal data. It contains videos of facial expressions, audio recordings, and physiological signals. The video and audio recordings are acquired using an RGB camera with an integrated microphone. The physi…

2022

Latent Image Animator: Learning to Animate Images via Latent Space Navigation

ICLR 2022poster

Due to the remarkable progress of deep generative models, animating images has become increasingly efficient, whereas associated results have become increasingly realistic. Current animation-approaches commonly exploit structure representation extracted from driving videos. Such structure representa…

Cited by 178SourcePDFScholar
2021

Joint Generative and Contrastive Learning for Unsupervised Person Re-Identification

CVPR 2021poster

Recent self-supervised contrastive learning provides an effective approach for unsupervised person re-identification (ReID) by learning invariance from different views (transformed versions) of an input. In this paper, we incorporate a Generative Adversarial Network (GAN) and a contrastive learning…

Cited by 220PDFcodeScholar
2020

G3AN: Disentangling Appearance and Motion for Video Generation

CVPR 2020poster

Creating realistic human videos entails the challenge of being able to simultaneously generate both appearance, as well as motion. To tackle this challenge, we introduce G3AN, a novel spatio-temporal generative model, which seeks to capture the distribution of high dimensional video data and to mode…

Cited by 112PDFcodeScholar
2019

Toyota Smarthome: Real-World Activities of Daily Living

ICCV 2019poster

The performance of deep neural networks is strongly influenced by the quantity and quality of annotated data. Most of the large activity recognition datasets consist of data sourced from the web, which does not reflect challenges that exist in activities of daily living. In this paper, we introduce…

Cited by 204PDFScholar