← Search

Dimitrios Hatzinakos

11 accepted papers

2026

Hear What You See: Video-to-Audio Generation with Diffusion Transformer and Semantic-Temporal Alignment-Ranked Direct Preference Optimization

CVPR 2026

Generating high-fidelity audio that is both semantically meaningful and temporally synchronized with silent videos remains a challenging problem in video-to-audio generation. Existing approaches often fail to capture fine-grained temporal correspondence between visual events and audio dynamics, lead

Cited by 0SourcecodeScholar
2026

JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation

ICLR 2026poster

Recent AIGC advances have rapidly expanded from text-to-image generation toward high-quality multimodal synthesis across video and audio. Within this context, joint audio-video generation (JAVG) has emerged as a fundamental task that produces synchronized and semantically aligned sound and vision fr…

Cited by 0SourcecodeScholar
2025

OSR: Toward Developing Efficient Federated Learning-based Human Activity Recognition using Optimal Server Representations

ICASSP 2025accepted

Federated Learning (FL) is a privacy-preserving algorithm that enables multiple clients to collaboratively train a global model without sharing their local data. This learning algorithm is particularly valuable in privacy-sensitive applications such as Human Activity Recognition (HAR), where users a…

Cited by 0SourceScholar
2024

MOMA: Mixture-of-Modality-Adaptations for Transferring Knowledge from Image Models Towards Efficient Audio-Visual Action Recognition

ICASSP 2024accepted

In this work, we investigate how to transfer learned knowledge from pre-trained image models for the audio-visual domain without relying on a full finetuning paradigm. To achieve this objective, we propose a novel parameter-efficient scheme called Mixture-of-Modality-Adaptations (MoMA) for audio-vis…

Cited by 0SourceScholar
2022

Hierarchical Deep Learning Model with Inertial and Physiological Sensors Fusion for Wearable-Based Human Activity Recognition

ICASSP 2022accepted

This paper presents a human activity recognition (HAR) system with wearable devices. While various approaches have been suggested for HAR, most of them focus on either 1) the inertial sensors to capture the physical movement or 2) subject-dependent evaluations that are less practical to real world c…

Cited by 0SourceScholar
2021

Detection of Post-Traumatic Stress Disorder Using Learned Time-Frequency Representations from Pupillometry

ICASSP 2021accepted

Post-traumatic stress disorder is a major public health concern with a lifetime prevalence rate of 6.1-9.2% in North America. PTSD is known to alter the autonomic nervous system leading to chronic sympathetic arousal including heightened anxiety and hypervigilance. Pupillometry offers a quick and ac…

Cited by 0SourceScholar
2019

Improving Eye Movement Biometrics Using Remote Registration of Eye Blinking Patterns

ICASSP 2019accepted

In this paper, the biometric potential of eye movement and eye blinking for human recognition task is investigated. These modalities might be useful for specific biometric applications like driver authentication for law enforcement. For this purpose, a database of 22 subjects was build where eye mov…

Cited by 0SourceScholar
2016

An adaptive multi-level wavelet denoising method for 40-Hz ASSR

ICASSP 2016accepted

This paper presents a novel method for extracting auditory steady state response (ASSR) signals from background electroencephalogram. 40-Hz ASSR signals are sensitive to subject's state of consciousness and can be used as a monitor for the depth of anaesthesia. The suggested method is a multilevel a…

Cited by 0SourceScholar