← Search

Ricard Marxer

8 accepted papers

2025

Aligning Multimodal Representations through an Information Bottleneck

ICML 2025poster

Contrastive losses have been extensively used as a tool for multimodal representation learning. However, it has been empirically observed that their use is not effective to learn an aligned representation space. In this paper, we argue that this phenomenon is caused by the presence of modality-spec…

Cited by 0SourcePDFScholar
2025

Optimizing Underwater Robot Navigation: A Study of DRL Algorithms and Multi-Modal Sensor Fusion

ICRA 2025

Autonomous underwater navigation faces significant challenges due to the complexity of the environment, limited localization methods, and poor visibility. This paper investigates the performance of various reinforcement learning (RL) algorithms-Proximal Policy Optimization (PPO), Trust Region Policy

Cited by 0SourcecodeScholar
2024

Speech Foundation Models on Intelligibility Prediction for Hearing-Impaired Listeners

ICASSP 2024accepted

Speech foundation models (SFMs) have been benchmarked on many speech processing tasks, often achieving state-of-the-art performance with minimal adaptation. However, the SFM paradigm has been significantly less explored for applications of interest to the speech perception community. In this paper w…

Cited by 0SourceScholar
2022

Contrastive Prediction Strategies for Unsupervised Segmentation and Categorization of Phonemes and Words

ICASSP 2022accepted

We identify a performance trade-off between the tasks of phoneme categorization and phoneme and word segmentation in several self-supervised learning algorithms based on Contrastive Predictive Coding (CPC). Our experiments suggest that context building networks, albeit necessary for high performance…

Cited by 0SourceScholar
2022

Homography-Based Loss Function for Camera Pose Regression

RA-L 2022

Some recent visual-based relocalization algorithms rely on deep learning methods to perform camera pose regression from image data. This letter focuses on the loss functions that embed the error between two poses to perform deep learning based camera pose regression. Existing loss functions are eith

Cited by 6SourcecodeScholar
2022

Variable-rate hierarchical CPC leads to acoustic unit discovery in speech

NeurIPS 2022accept

The success of deep learning comes from its ability to capture the hierarchical structure of data by learning high-level representations defined in terms of low-level ones. In this paper we explore self-supervised learning of hierarchical representations of speech by applying multiple levels of Cont…

2019

Real-time Passive Acoustic 3D Tracking of Deep Diving Cetacean by Small Non-uniform Mobile Surface Antenna

ICASSP 2019accepted

Detecting and localizing the echolocation clicks of sperm whales provides insight into their diving behavior, but existing methods are limited in range, imprecise, or costly. In this work, we demonstrate that we can obtain a high definition 3D track of deep diving cetaceans from a five-channel, smal…

Cited by 0SourceScholar