← Search

Gerasimos Potamianos

12 accepted papers

2025

Resource-Efficient and Noise-Robust Modality Fusion for Audio-Visual Speech Recognition

ICASSP 2025accepted

Resource-efficient audio-visual fusion techniques often struggle to maintain robust performance across varying acoustic noise conditions in speech recognition tasks. This paper introduces a dynamic routing approach for noise-robust audio-visual fusion, which adaptively directs features to noise-spec…

Cited by 0SourceScholar
2023

Sign Language Recognition via Deformable 3D Convolutions and Modulated Graph Convolutional Networks

ICASSP 2023accepted

Automatic sign language recognition (SLR) remains challenging, especially when employing RGB video alone (i.e., with no depth or special glove-based input) and under a signer-independent (SI) framework, due to inter-personal signing variation. In this paper, we address SI isolated SLR from RGB video…

Cited by 0SourceScholar
2022

Accurate and Resource-Efficient Lipreading with Efficientnetv2 and Transformers

ICASSP 2022accepted

We present a novel resource-efficient end-to-end architecture for lipreading that achieves state-of-the-art results on a popular and challenging benchmark. In particular, we make the following contributions: First, inspired by the recent success of the EfficientNet architecture in image classificati…

Cited by 0SourceScholar
2022

Spatio-Temporal Graph Convolutional Networks for Continuous Sign Language Recognition

ICASSP 2022accepted

We address the challenging problem of continuous sign language recognition (CSLR) from RGB videos, proposing a novel deep-learning framework that employs spatio-temporal graph convolutional networks (ST-GCNs), which operate on multiple, appropriately fused feature streams, capturing the signer’s pos…

Cited by 0SourceScholar
2020

Audio-Assisted Image Inpainting for Talking Faces

ICASSP 2020accepted

The goal of our work is to complete missing areas of images of talking faces, exploiting information from both the visual and audio modalities. Existing image inpainting methods rely solely on visual content that doesn't always provide sufficient information for the task. To counter this, we propose…

Cited by 0SourceScholar
2019

Fusing Body Posture With Facial Expressions for Joint Recognition of Affect in Child-Robot Interaction

RA-L 2019

In this letter, we address the problem of multi-cue affect recognition in challenging scenarios such as child–robot interaction. Toward this goal we propose a method for automatic recognition of affect that leverages body expressions alongside facial ones, as opposed to traditional methods that typi

Cited by 62SourceScholar
2018

Far-Field Audio-Visual Scene Perception of Multi-Party Human-Robot Interaction for Children and Adults

ICASSP 2018accepted

Human-robot interaction (HRI) is a research area of growing interest with a multitude of applications for both children and adult user groups, as, for example, in edutainment and social robotics. Crucial, however, to its wider adoption remains the robust perception of HRI scenes in natural, untether…

Cited by 0SourceScholar
2018

Multi3: Multi-Sensory Perception System for Multi-Modal Child Interaction with Multiple Robots

ICRA 2018poster

Child-robot interaction is an interdisciplinary research area that has been attracting growing interest, primarily focusing on edutainment applications. A crucial factor to the successful deployment and wide adoption of such applications remains the robust perception of the child's multi-modal actio…

Cited by 0SourceScholar
2018

Object Assembly Guidance in Child-Robot Interaction using RGB-D based 3D Tracking

IROS 2018poster

This work examines how and to what benefit an autonomous humanoid robot can supervise a child in an object assembly task. In order to understand the child's actions, a novel 3D object tracking algorithm for RGB-D data is employed. The tracker consists of two stages: the first performs a tracking-by-…

Cited by 8SourceScholar
2017

Deep Affordance-Grounded Sensorimotor Object Recognition

CVPR 2017spotlight

It is well-established by cognitive neuroscience that human perception of objects constitutes a complex process, where object appearance information is combined with evidence about the so-called object "affordances", namely the types of actions that humans typically perform when interacting with the…

Cited by 44PDFScholar
2015

Multichannel speech enhancement using MEMS microphones

ICASSP 2015accepted

In this work, we investigate the efficacy of Micro Electro-Mechanical System (MEMS) microphones, a newly developed technology of very compact sensors, for multichannel speech enhancement. Experiments are conducted on real speech data collected using a MEMS microphone array. First, the effectiveness…

Cited by 0SourceScholar