← Search

PETROS MARAGOS

53 accepted papers

2026

Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic Correspondence

CVPR 2026

Recent frameworks like ToFu and TEMPEH provide an automated alternative to classical registration pipelines by predicting 3D meshes in dense semantic correspondence directly from calibrated multi-view images. However, these learning-based methods rely on the slow, manual registration pipelines they

Cited by 0SourcecodeScholar
2026

SHARED-WEIGHTS EXTENDER AND GRADIENT VOTING FOR NEURAL NETWORK EXPANSION

ICASSP 2026poster

Expanding neural networks during training is a promising way to augment capacity without retraining larger models from scratch. However, newly added neurons often fail to adjust to a trained network and become inactive, providing no contribution to capacity growth. We propose the Shared-Weights Exte…

Cited by 0SourcePDFScholar
2025

Category-Level 6D Object Pose Estimation in Agricultural Settings Using a Lattice-Deformation Framework and Diffusion-Augmented Synthetic Data

IROS 2025

Accurate 6D object pose estimation is essential for robotic grasping and manipulation, particularly in agriculture, where fruits and vegetables exhibit high intra-class variability in shape, size, and texture. The vast majority of existing methods rely on instance-specific CAD models or require dept

Cited by 1SourcecodeScholar
2025

Instance-Level Composed Image Retrieval

NeurIPS 2025poster

The progress of composed image retrieval (CIR), a popular research direction in image retrieval, where a combined visual and textual query is used, is held back by the absence of high-quality training and evaluation data. We introduce a new evaluation dataset, i-CIR, which, unlike existing datasets,…

Cited by 0SourceScholar
2025

Power in Unity: Combining in-Domain and out-of-Domain Pre-Training Strategies for EEG-Based Person Identification

ICASSP 2025accepted

We present the NTUA-IRAL team’s solution for the Person Identification track of the Signal Processing EEG-Music Emotion Recognition Grand Challenge, hosted at ICASSP. Our approach employs an ensemble of three CNNs, each pretrained using a distinct strategy: contrastive pre-training, traditional Imag…

Cited by 0SourceScholar
2025

Proactive Tactile Exploration for Object-Agnostic Shape Reconstruction from Minimal Visual Priors

ICRA 2025

The perception of an object's surface is important for robotic applications enabling robust object manipulation. The level of accuracy in such a representation affects the outcome of the action planning, especially during tasks that require physical contact, e.g. grasping. In this paper, we propose

Cited by 2SourceScholar
2025

Towards Open-Ended Robotic Exploration Using Vision-Inspired Similarity and Foundation Models

ICRA 2025

In the domain of robotics, achieving Lifelong Open-ended Learning Autonomy (LOLA) represents a significant milestone, especially in contexts where autonomous agents must adapt to unforeseen environmental variations and evolving objectives. This paper introduces VISOR (VisionSimilarity for Open-ended

Cited by 0SourceScholar
2024

3D Facial Expressions through Analysis-by-Neural-Synthesis

CVPR 2024poster

While existing methods for 3D face reconstruction from in-the-wild images excel at recovering the overall face shape they commonly miss subtle extreme asymmetric or rarely observed expressions. We improve upon these methods with SMIRK (Spatial Modeling for Image-based Reconstruction of Kinesics) whi…

2024

Augmenting Transformer Autoencoders with Phenotype Classification for Robust Detection of Psychotic Relapses

ICASSP 2024accepted

Recently, deep autoencoder architectures have received attention for the problem of unsupervised anomaly detection. Detecting psychotic relapses in mental health patients is a crucial challenge, often framed as anomaly detection, given the limited availability of data during relapsing states. In thi…

Cited by 0SourceScholar
2024

Matrix Factorization in Tropical and Mixed Tropical-Linear Algebras

ICASSP 2024accepted

Matrix Factorization (MF) has found numerous applications in Machine Learning and Data Mining, including collaborative filtering recommendation systems, dimensionality reduction, data visualization, and community detection. Motivated by the recent successes of tropical algebra and geometry in machin…

Cited by 0SourceScholar
2024

SDPL-SLAM: Introducing Lines in Dynamic Visual SLAM and Multi-Object Tracking

IROS 2024poster

The need for a robust visual SLAM system operating in real human environments has led to the gradual abandonment of the static world assumption and to the creation of many dynamic SLAM algorithms. Even though there have been many dynamic SLAM proposals, the vast majority of them relied on point feat…

Cited by 0SourceScholar
2023

Convolutional Recurrent Neural Networks for the Classification of Cetacean Bioacoustic Patterns

ICASSP 2023accepted

In this paper we focus on the development of a convolutional recurrent neural network (CRNN) to categorize biosignals collected in the Hellenic Trench, generated by two cetacean species, sperm whales (Physeter macrocephalus) and striped dolphins (Stenella coeruleoalba). We convert audio signals into…

Cited by 0SourceScholar
2023

E-Prevention: The ICASSP-2023 Challenge on Person Identification and Relapse Detection from Continuous Recordings of Biosignals

ICASSP 2023accepted

The e-Prevention challenge concerns the analysis and processing of long-term continuous recordings of biosignals recorded from wearable sensors, i.e., accelerometers, gyroscopes and heart rate monitors embedded in smartwatches, as well as sleep information and daily step count, in order to extract h…

Cited by 0SourceScholar
2023

Newton-Based Trainable Learning Rate

ICASSP 2023accepted

Selecting an appropriate learning rate for efficiently training deep neural networks is a difficult process that can be affected by numerous parameters, such as the dataset, the model architecture or even the batch size. In this work, we propose an algorithm for automatically adjusting the learning…

Cited by 0SourceScholar
2023

Relapse Prediction from Long-Term Wearable Data Using Self-Supervised Learning and Survival Analysis

ICASSP 2023accepted

The introduction of biometric signal analysis in psychiatry could potentially reshape the field by making it more accurate, proactive and personalized. Such biosignals usually acquired from wearables encompass the quantification of human behavior and traits. In this study, we use long-term data acqu…

Cited by 0SourceScholar
2022

A Few-Sample Strategy for Guitar Tablature Transcription Based on Inharmonicity Analysis and Playability Constraints

ICASSP 2022accepted

The prominent strategical approaches regarding the problem of guitar tablature transcription rely either on fingering patterns encoding or on the extraction of string-related audio features. The current work combines the two aforementioned strategies in an explicit manner by employing two discrete c…

Cited by 0SourceScholar
2022

Child Engagement Estimation in Heterogeneous Child-Robot Interactions Using Spatiotemporal Visual Cues

IROS 2022poster

Robots are increasingly introduced in various Child-Robot Interactions with educational, entertainment or even therapeutic goals. In order to achieve qualitative inter-actions, robots need to adjust their behavior according to children's response. A robot's ability to successfully estimate partner's…

Cited by 3SourceScholar
2022

Enhancing Affective Representations Of Music-Induced Eeg Through Multimodal Supervision And Latent Domain Adaptation

ICASSP 2022accepted

The study of Music Cognition and neural responses to music has been invaluable in understanding human emotions. Brain signals, though, manifest a highly complex structure that makes processing and retrieving meaningful features challenging, particularly of abstract constructs like affect. Moreover,…

Cited by 0SourceScholar
2022

Neural Emotion Director: Speech-Preserving Semantic Control of Facial Expressions in "In-the-Wild" Videos

CVPR 2022oral

In this paper, we introduce a novel deep learning method for photo-realistic manipulation of the emotional state of actors in "in-the-wild" videos. The proposed method is based on a parametric 3D face representation of the actor in the input scene that offers a reliable disentanglement of the facial…

Cited by 41PDFcodeScholar
2022

Neural Network Approximation based on Hausdorff distance of Tropical Zonotopes

ICLR 2022poster

In this work we theoretically contribute to neural network approximation by providing a novel tropical geometrical viewpoint to structured neural network compression. In particular, we show that the approximation error between two neural networks with ReLU activations and one hidden layer depends on…

Cited by 10SourcePDFScholar
2022

Spatio-Temporal Graph Convolutional Networks for Continuous Sign Language Recognition

ICASSP 2022accepted

We address the challenging problem of continuous sign language recognition (CSLR) from RGB videos, proposing a novel deep-learning framework that employs spatio-temporal graph convolutional networks (ST-GCNs), which operate on multiple, appropriately fused feature streams, capturing the signer’s pos…

Cited by 0SourceScholar
2021

Advances in Morphological Neural Networks: Training, Pruning and Enforcing Shape Constraints

ICASSP 2021accepted

In this paper, we study an emerging class of neural networks, the Morphological Neural networks, from some modern perspectives. Our approach utilizes ideas from tropical geometry and mathematical morphology. First, we state the training of a binary morphological classifier as a Difference-of-Convex…

Cited by 0SourceScholar
2021

Deep Convolutional and Recurrent Networks for Polyphonic Instrument Classification from Monophonic Raw Audio Waveforms

ICASSP 2021accepted

Sound Event Detection and Audio Classification tasks are traditionally addressed through time-frequency representations of audio signals such as spectrograms. However, the emergence of deep neural networks as efficient feature extractors has enabled the direct use of audio signals for classification…

Cited by 0SourceScholar
2021

Engagement Estimation During Child Robot Interaction Using Deep Convolutional Networks Focusing on ASD Children

ICRA 2021poster

Estimating the engagement of children is an essential prerequisite for constructing natural Child-Robot Interaction. Especially in the case of children with Autism Spectrum Disorder, monitoring the engagement of the other party allows robots to adjust their actions according to the educational and t…

Cited by 16SourceScholar
2021

Grounding Consistency: Distilling Spatial Common Sense for Precise Visual Relationship Detection

ICCV 2021poster

Scene Graph Generators (SGGs) are models that, given an image, build a directed graph where each edge represents a predicted subject predicate object triplet. Most SGGs silently exploit datasets' bias on relationships' context, i.e. its subject and object, to improve recall and neglect spatial and v…

Cited by 13PDFcodeScholar
2021

Independent Sign Language Recognition with 3d Body, Hands, and Face Reconstruction

ICASSP 2021accepted

Independent Sign Language Recognition is a complex visual recognition problem that combines several challenging tasks of Computer Vision due to the necessity to exploit and fuse information from hand gestures, body features and facial expressions. While many state-of-the-art works have managed to de…

Cited by 0SourceScholar
2021

Sparsity in Max-Plus Algebra and Applications in Multivariate Convex Regression

ICASSP 2021accepted

In this paper, we study concepts of sparsity in the max-plus algebra and apply them to the problem of multivariate convex regression. We show how to efficiently find sparse (containing many −∞ elements) approximate solutions to max-plus equations by leveraging notions from submodular optimization. S…

Cited by 0SourceScholar
2020

An LSTM-Based Dynamic Chord Progression Generation System for Interactive Music Performance

ICASSP 2020accepted

In this paper, we describe an interactive generative music system, designed to handle polyphonic guitar music. We formulate the problem of chord progression generation as a prediction problem. Thus, we propose utilization of an LSTM-based network architecture incorporating neural attention that is a…

Cited by 0SourceScholar
2020

Maxpolynomial Division with Application To Neural Network Simplification

ICASSP 2020accepted

In this work, we further the link between neural networks with piecewise linear activations and tropical algebra. To that end, we introduce the process of Maxpolynomial Division, a geometric method which simulates division of polynomials in the max-plus semiring, while highlighting its key propertie…

Cited by 0SourceScholar
2020

Multiclass Neural Network Minimization via Tropical Newton Polytope Approximation

ICML 2020poster

The field of tropical algebra is closely linked with the domain of neural networks with piecewise linear activations, since their output can be described via tropical polynomials in the max-plus semiring. In this work, we attempt to make use of methods stemming from a form of approximate division of…

2020

Person Identification Using Deep Convolutional Neural Networks on Short-Term Signals from Wearable Sensors

ICASSP 2020accepted

In this work, we explore the discriminating ability of short-term signal patterns (e.g. few minutes long) with respect to the person identification task. We focus on signals recorded by simple wearable devices, such as smart watches, which can measure movements (accelerometer and gyroscope sensors)…

Cited by 0SourceScholar
2019

A Deep Learning Approach for Multi-View Engagement Estimation of Children in a Child-Robot Joint Attention Task

IROS 2019poster

In this work, we tackle the problem of child engagement estimation while children freely interact with a robot in a friendly, room-like environment. We propose a deep learning-based multi-view solution that takes advantage of recent developments in human pose detection. We extract the child's pose f…

Cited by 35SourceScholar
2019

Fusing Body Posture With Facial Expressions for Joint Recognition of Affect in Child-Robot Interaction

RA-L 2019

In this letter, we address the problem of multi-cue affect recognition in challenging scenarios such as child–robot interaction. Toward this goal we propose a method for automatic recognition of affect that leverages body expressions alongside facial ones, as opposed to traditional methods that typi

Cited by 62SourceScholar
2019

LSTM-based Network for Human Gait Stability Prediction in an Intelligent Robotic Rollator

ICRA 2019poster

In this work, we present a novel framework for on-line human gait stability prediction of the elderly users of an intelligent robotic rollator using Long Short Term Memory (LSTM) networks, fusing multimodal RGB-D and Laser Range Finder (LRF) data from non-wearable sensors. A Deep Learning (DL) based…

Cited by 37SourceScholar
2019

Learn to Adapt to Human Walking: A Model-Based Reinforcement Learning Approach for a Robotic Assistant Rollator

RA-L 2019

In this letter, we tackle the problem of adapting the motion of a robotic assistant rollator to patients with different mobility status. The goal is to achieve a coupled human–robot motion in a front-following setting as if the patient was pushing the rollator himself/herself. To this end, we propos

Cited by 22SourceScholar
2018

Augmented Human State Estimation Using Interacting Multiple Model Particle Filters With Probabilistic Data Association

RA-L 2018

The accurate human gait tracking is an important factor for various robotic applications, such as robotic walkers aiming to provide assistance to patients with different mobility impairment, social robot companions, etc. A context-aware robot control architecture needs constant knowledge of the user

Cited by 32SourceScholar
2018

Far-Field Audio-Visual Scene Perception of Multi-Party Human-Robot Interaction for Children and Adults

ICASSP 2018accepted

Human-robot interaction (HRI) is a research area of growing interest with a multitude of applications for both children and adult user groups, as, for example, in edutainment and social robotics. Crucial, however, to its wider adoption remains the robust perception of HRI scenes in natural, untether…

Cited by 0SourceScholar
2018

Multi-View Audio-Articulatory Features for Phonetic Recognition on RTMRI-TIMIT Database

ICASSP 2018accepted

In this paper, we investigate the use of articulatory information, and more specifically real time Magnetic Resonance Imaging (rtMRI) data of the vocal tract, to improve speech recognition performance. For the purpose of our experiments, we use data from the rtMRI-TIMIT database. Firstly, Scale Inva…

Cited by 0SourceScholar
2018

Multi3: Multi-Sensory Perception System for Multi-Modal Child Interaction with Multiple Robots

ICRA 2018poster

Child-robot interaction is an interdisciplinary research area that has been attracting growing interest, primarily focusing on edutainment applications. A crucial factor to the successful deployment and wide adoption of such applications remains the robust perception of the child's multi-modal actio…

Cited by 0SourceScholar
2018

Multimodal Signal Processing and Learning Aspects of Human-Robot Interaction for an Assistive Bathing Robot

ICASSP 2018accepted

We explore new aspects of assistive living on smart human-robot interaction (HRI) that involve automatic recognition and online validation of speech and gestures in a natural interface, providing social features for HRI. We introduce a whole framework and resources of a real-life scenario for elderl…

Cited by 0SourceScholar
2018

Multimodal Visual Concept Learning With Weakly Supervised Techniques

CVPR 2018poster

Despite the availability of a huge amount of video data accompanied by descriptive texts, it is not always easy to exploit the information contained in natural language in order to automatically recognize video concepts. Towards this goal, in this paper we use textual cues as means of supervision, i…

2018

Object Assembly Guidance in Child-Robot Interaction using RGB-D based 3D Tracking

IROS 2018poster

This work examines how and to what benefit an autonomous humanoid robot can supervise a child in an object assembly task. In order to understand the child's actions, a novel 3D object tracking algorithm for RGB-D data is employed. The tracker consists of two stages: the first performs a tracking-by-…

Cited by 8SourceScholar
2018

User-Adaptive Human-Robot Formation Control for an Intelligent Robotic Walker Using Augmented Human State Estimation and Pathological Gait Characterization

IROS 2018poster

In this paper we describe a control strategy for a user-adaptive human-robot system for an intelligent robotic Mobility Assistive Device (MAD)using raw data from a single laser-range-finder (LRF)mounted on the MAD and scanning the walking area. The proposed control architecture consists of three mod…

Cited by 25SourceScholar
2017

Comparative experimental validation of human gait tracking algorithms for an intelligent robotic rollator

ICRA 2017poster

Tracking human gait accurately and robustly constitutes a key factor for a smart robotic walker, aiming to provide assistance to patients with different mobility impairment. A context-aware assistive robot needs constant knowledge of the user's kinematic state to assess the gait status and adjust it…

Cited by 14SourceScholar
2017

Real-time end-effector motion behavior planning approach using on-line point-cloud data towards a user adaptive assistive bath robot

IROS 2017poster

Elderly people have particular needs in performing bathing activities, since these tasks require body flexibility. Our aim is to build an assistive robotic bath system, in order to increase the independence and safety of this procedure. Towards this end, the expertise of professional carers for bath…

Cited by 15SourceScholar
2016

Multimodal human action recognition in assistive human-robot interaction

ICASSP 2016accepted

Within the context of assistive robotics we develop an intelligent interface that provides multimodal sensory processing capabilities for human action recognition. Human action is considered in multimodal terms, containing inputs such as audio from microphone arrays, and visual inputs from high defi…

Cited by 0SourceScholar
2016

Towards a behaviorally-validated computational audiovisual saliency model

ICASSP 2016accepted

Computational saliency models aim at predicting, in a bottom-up fashion, where human attention is drawn in the presented (visual, auditory or audiovisual) scene and have been proven useful in applications like robotic navigation, image compression and movie summarization. Despite the fact that well-…

Cited by 0SourceScholar
2015

Hidden markov modeling of human pathological gait using laser range finder for an assisted living intelligent robotic walker

IROS 2015poster

The precise analysis of a patient's or an elderly person's walking pattern is very important for an effective intelligent active mobility assistance robot. This walking pattern can be described by a cyclic motion, which can be modeled using the consecutive gait phases. In this paper, we present a co…

Cited by 26SourceScholar
2015

Max-product dynamical systems and applications to audio-visual salient event detection in videos

ICASSP 2015accepted

This paper introduces a theory for max-product systems by analyzing them as discrete-time nonlinear dynamical systems that obey a superposition of a weighted maximum type and evolve on nonlinear spaces which we call complete weighted lattices. Special cases of such systems have found applications in…

Cited by 0SourceScholar
2015

Multichannel speech enhancement using MEMS microphones

ICASSP 2015accepted

In this work, we investigate the efficacy of Micro Electro-Mechanical System (MEMS) microphones, a newly developed technology of very compact sensors, for multichannel speech enhancement. Experiments are conducted on real speech data collected using a MEMS microphone array. First, the effectiveness…

Cited by 0SourceScholar