← Search

Ali Etemad

34 accepted papers

2026

Personalized Feature Translation for Expression Recognition: An Efficient Source-Free Domain Adaptation Method

ICLR 2026poster

Facial expression recognition (FER) models are employed in many video-based affective computing applications, such as human-computer interaction and healthcare monitoring. However, deep FER models often struggle with subtle expressions and high inter-subject variability, limiting their performance…

Cited by 0SourceScholar
2025

Dynamic Prototype Rehearsal for Continual ECG Arrhythmia Detection

ICASSP 2025accepted

Continual Learning (CL) methods aim to learn from a sequence of tasks while avoiding the challenge of forgetting previous knowledge. We present DREAM-CL, a novel CL method for ECG arrhythmia detection that introduces dynamic prototype rehearsal memory. DREAM-CL selects representative prototypes by c…

Cited by 0SourceScholar
2025

Federated Domain Generalization with Label Smoothing and Balanced Decentralized Training

ICASSP 2025accepted

In this paper, we propose a novel approach, Federated Domain Generalization with Label Smoothing and Balanced Decentralized Training (FedSB), to address the challenges of data heterogeneity within a federated learning framework. FedSB utilizes label smoothing at the client level to prevent overfitti…

Cited by 0SourceScholar
2025

Federated Unsupervised Domain Generalization Using Global and Local Alignment of Gradients

AAAI 2025technical

We address the problem of federated domain generalization in an unsupervised setting for the first time. We first theoretically establish a connection between domain shift and alignment of gradients in unsupervised federated learning and show that aligning the gradients at both client and server lev…

2025

Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment

ICLR 2025poster

Despite their significant advancements, Multimodal Large Language Models (MLLMs) often generate factually inaccurate information, referred to as hallucination. In this work, we address object hallucinations in MLLMs, where information is generated about an object not present in the input image. We i…

Cited by 0SourcePDFScholar
2025

PRISM: Reducing Spurious Implicit Biases in Vision-Language Models with LLM-Guided Embedding Projection

ICCV 2025poster

We introduce Projection-based Reduction of Implicit Spurious bias in vision-language Models (PRISM), a new data-free and task-agnostic solution for bias mitigation in VLMs like CLIP. VLMs often inherit and amplify biases in their training data, leading to skewed predictions.PRISM is designed to debi…

2025

Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM Inference

AAAI 2025technical

Large language models (LLMs) have triggered a new stream of research focusing on compressing the context length to reduce the computational cost while ensuring the retention of helpful information for LLMs to answer the given question. Token-based removal methods are one of the most prominent approa…

2025

Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization

NeurIPS 2025poster

Despite recent advances in Large Video Language Models (LVLMs), they still struggle with fine-grained temporal understanding, hallucinate, and often make simple mistakes on even simple video question-answering tasks, all of which pose significant challenges to their safe and reliable deployment in r…

Cited by 0SourceScholar
2025

Some Optimizers are More Equal: Understanding the Role of Optimizers in Group Fairness

NeurIPS 2025spotlight

We study whether and how the choice of optimization algorithm can impact group fairness in deep neural networks. Through stochastic differential equation analysis of optimization dynamics in an analytically tractable setup, we demonstrate that the choice of optimization algorithm indeed influences f…

Cited by 0SourcecodeScholar
2025

Subject Representation Learning from EEG using Graph Convolutional Variational Autoencoders

ICASSP 2025accepted

We propose GC-VASE, a graph convolutional-based variational autoencoder that leverages contrastive learning for subject representation learning from EEG data. Our method successfully learns robust subject-specific latent representations using the split-latent space architecture tailored for subject…

Cited by 0SourceScholar
2024

Region-Disentangled Diffusion Model for High-Fidelity PPG-to-ECG Translation

AAAI 2024technical

The high prevalence of cardiovascular diseases (CVDs) calls for accessible and cost-effective continuous cardiac monitoring tools. Despite Electrocardiography (ECG) being the gold standard, continuous monitoring remains a challenge, leading to the exploration of Photoplethysmography (PPG), a promisi…

2024

Segment, Shuffle, and Stitch: A Simple Layer for Improving Time-Series Representations

NeurIPS 2024poster

Existing approaches for learning representations of time-series keep the temporal arrangement of the time-steps intact with the presumption that the original order is the most optimal for learning. However, non-adjacent sections of real-world time-series may have strong dependencies. Accordingly, we…

2024

Speech Emotion Recognition with Distilled Prosodic and Linguistic Affect Representations

ICASSP 2024accepted

We propose EmoDistill, a novel speech emotion recognition (SER) framework that leverages cross-modal knowledge distillation during training to learn strong linguistic and prosodic representations of emotion from speech. During inference, our method only uses a stream of speech signals to perform uni…

Cited by 0SourceScholar
2024

UPose3D: Uncertainty-Aware 3D Human Pose Estimation with Cross-View and Temporal Cues

ECCV 2024poster

"We introduce UPose3D, a novel approach for multi-view 3D human pose estimation, addressing challenges in accuracy and scalability. Our method advances existing pose estimation frameworks by improving robustness and flexibility without requiring direct 3D annotations. At the core of our method, a po…

2024

XKD: Cross-Modal Knowledge Distillation with Domain Alignment for Video Representation Learning

AAAI 2024technical

We present XKD, a novel self-supervised framework to learn meaningful representations from unlabelled videos. XKD is trained with two pseudo objectives. First, masked data reconstruction is performed to learn modality-specific representations from audio and visual streams. Next, self-supervised cros…

2023

AVCAffe: A Large Scale Audio-Visual Dataset of Cognitive Load and Affect for Remote Work

AAAI 2023technical

We introduce AVCAffe, the first Audio-Visual dataset consisting of Cognitive load and Affect attributes. We record AVCAffe by simulating remote work scenarios over a video-conferencing platform, where subjects collaborate to complete a number of cognitively engaging tasks. AVCAffe is the largest ori…

2023

Human Pose Estimation from Ambiguous Pressure Recordings with Spatio-Temporal Masked Transformers

ICASSP 2023accepted

Despite the impressive performance of vision-based pose estimators, they generally fail to perform well under adverse vision conditions and often don’t satisfy the privacy demands of customers. As a result, researchers have begun to study tactile sensing systems as an alternative. However, these sys…

Cited by 0SourceScholar
2023

Self-Supervised Audio-Visual Representation Learning with Relaxed Cross-Modal Synchronicity

AAAI 2023technical

We present CrissCross, a self-supervised framework for learning audio-visual representations. A novel notion is introduced in our framework whereby in addition to learning the intra-modal and standard 'synchronous' cross-modal relations, CrissCross also learns 'asynchronous' cross-modal relationship…

2023

Uncovering the Hidden Dynamics of Video Self-supervised Learning under Distribution Shifts

NeurIPS 2023spotlight

Video self-supervised learning (VSSL) has made significant progress in recent years. However, the exact behavior and dynamics of these models under different forms of distribution shift are not yet known. In this paper, we comprehensively study the behavior of six popular self-supervised methods (v-…

2022

Multiscale Crowd Counting and Localization By Multitask Point Supervision

ICASSP 2022accepted

We propose a multitask approach for crowd counting and person localization in a unified framework. As the detection and localization tasks are well-correlated and can be jointly tackled, our model benefits from a multitask solution by learning multiscale representations of encoded crowd images, and…

Cited by 0SourceScholar
2022

ObjectBox: From Centers to Boxes for Anchor-Free Object Detection

ECCV 2022poster

"We present ObjectBox, a novel single-stage anchor-free and highly generalizable object detection approach. As opposed to both existing anchor-based and anchor-free detectors, which are more biased toward specific object scales in their label assignments, we use only object center locations as posit…

2022

Vote from the Center: 6 DoF Pose Estimation in RGB-D Images by Radial Keypoint Voting

ECCV 2022poster

"We propose a novel keypoint voting scheme based on intersecting spheres, that is more accurate than existing schemes and allows for fewer, more disperse keypoints. The scheme is based upon the distance between points, which as a 1D quantity can be regressed more accurately than the 2D and 3D vector…

2021

CardioGAN: Attentive Generative Adversarial Network with Dual Discriminators for Synthesis of ECG from PPG

AAAI 2021technical

Electrocardiogram (ECG) is the electrical measurement of cardiac activity, whereas Photoplethysmogram (PPG) is the optical measurement of volumetric changes in blood circulation. While both signals are used for heart rate monitoring, from a medical perspective, ECG is more useful as it carries addit…

2021

In-Bed Pressure-Based Pose Estimation Using Image Space Representation Learning

ICASSP 2021accepted

Recent advances in deep pose estimation models have proven to be effective in a wide range of applications such as health monitoring, sports, animations, and robotics. However, pose estimation models fail to generalize when facing images acquired from in-bed pressure sensing systems. In this paper,…

Cited by 0SourceScholar
2021

Multi-Perspective LSTM for Joint Visual Representation Learning

CVPR 2021poster

We present a novel LSTM cell architecture capable of learning both intra- and inter-perspective relationships available in visual sequences captured from multiple perspectives. Our architecture adopts a novel recurrent joint learning strategy that uses additional gates and memories at the cell level…

Cited by 12PDFcodeScholar
2021

Teacher-Student Adversarial Depth Hallucination To Improve Face Recognition

ICCV 2021poster

We present the Teacher-Student Generative Adversarial Network (TS-GAN) to generate depth images from single RGB images in order to boost the performance of face recognition systems. For our method to generalize well across unseen datasets, we design two components in the architecture, a teacher and…

Cited by 11PDFcodeScholar
2020

Detecting Multiple Speech Disfluencies Using a Deep Residual Network with Bidirectional Long Short-Term Memory

ICASSP 2020accepted

Stuttering is a speech impediment affecting tens of millions of people on an everyday basis. Even with its commonality, there is minimal data and research on the identification and classification of stuttered speech. This paper tackles the problem of detection and classification of different forms o…

Cited by 0SourceScholar