← Search

Lucas Goncalves

6 accepted papers

2026

Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent

CVPR 2026

Unifying image clustering across different clustering scenarios remains challenging due to fundamental gaps among tasks. We introduce a Guideline-Driven Image Clustering Agent, the first universal framework that bridges these gaps through textual guidelines. To incorporate complex guidelines without

Cited by 0SourceScholar
2025

Efficient Fusion of Computationally Diverse Modalities Using Chunking and Cross-Attention

ICASSP 2025accepted

Emotion recognition is inherently a multimodal problem. Humans use both audible and visual cues to determine a person’s emotions. There has been extensive improvement in the methods we use to fuse audio and visual representations between two unimodal deep-learning models. However, there is a lack of…

Cited by 0SourceScholar
2025

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation

ICASSP 2025accepted

Audio-Visual Speech-to-Speech Translation (AVS2S) typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony—ensuring that the movements of the lips match the spoken content—essential for maintaining realism in du…

Cited by 0SourceScholar
2024

Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers’ Opinion Scores

ECCV 2024poster

"Recent advancements in audio-visual generative modeling have been propelled by progress in deep learning and the availability of data-rich benchmarks. However, the growth is not attributed solely to models and benchmarks. Universally accepted evaluation metrics also play an important role in advanc…

Cited by 1SourcePDFScholar
2023

Learning Cross-Modal Audiovisual Representations with Ladder Networks for Emotion Recognition

ICASSP 2023accepted

Representation learning is a challenging, but essential task in audiovisual learning. A key challenge is to generate strong cross-modal representations while still capturing discriminative information contained in unimodal features. Properly capturing this information is important to increase accura…

Cited by 0SourceScholar