← Search

Naveen Kumar

8 accepted papers

2023

Transformer-Based Neural Augmentation of Robot Simulation Representations

RA-L 2023

Simulation representations of robots have advanced in recent years. Yet, there remain significant sim-to-real gaps because of modeling assumptions and hard-to-model behaviors such as friction. In this letter, we propose to augment common simulation representations with a transformer-inspired archite

Cited by 6SourceScholar
2019

Hierarchy-aware Loss Function on a Tree Structured Label Space for Audio Event Detection

ICASSP 2019accepted

The paper introduces a hierarchy-aware loss function in a Deep Neural Network for an audio event detection task that has a bi-level tree structured label space. The goal is not only to improve audio event detection performance at all levels in the label hierarchy, but also to produce better audio em…

Cited by 0SourceScholar
2018

A Deep Reinforcement Learning Framework for Identifying Funny Scenes in Movies

ICASSP 2018accepted

This paper presents a novel deep Reinforcement Learning (RL) framework for classifying movie scenes based on affect using the face images detected in the video stream as input. Extracting affective information from the video is a challenging task modulating complex visual and temporal representation…

Cited by 0SourceScholar
2017

Device Placement Optimization with Reinforcement Learning

ICML 2017poster

The past few years have witnessed a growth in size and computational requirements for training and inference with neural networks. Currently, a common approach to address these requirements is to use a heterogeneous distributed environment with a mixture of hardware devices such as CPUs and GPUs. Im…

Cited by 556SourcePDFScholar
2016

A multimodal mixture-of-experts model for dynamic emotion prediction in movies

ICASSP 2016accepted

This paper addresses the problem of continuous emotion prediction in movies from multimodal cues. The rich emotion content in movies is inherently multimodal, where emotion is evoked through both audio (music, speech) and video modalities. To capture such affective information, we put forth a set of…

Cited by 0SourceScholar
2016

Opening big in box office? Trailer content can help

ICASSP 2016accepted

Computational prediction of a movie's financial success usually relies only on metadata such as - genre, budget, actors, Motion Picture Association of America (MPAA) rating and critics' reviews. We argue that movie trailers, created to invoke viewers' interest and curiosity about a movie, carry comp…

Cited by 0SourceScholar
2016

Pathological speech processing: State-of-the-art, current challenges, and future directions

ICASSP 2016accepted

The study of speech pathology involves evaluation and treatment of speech production related disorders affecting phonation, fluency, intonation and aeromechanical components of respiration. Recently, speech pathology has garnered special interest amongst machine learning and signal processing (ML-SP…

Cited by 0SourceScholar
2015

Computationally deconstructing movie narratives: An informatics approach

ICASSP 2015accepted

In general, popular films and screenplays follow a well defined storytelling paradigm that comprises three essential segments or acts: exposition (act I), conflict (act II) and resolution (act III). Deconstructing a movie into its narrative units can enrich semantic understanding of movies, and help…

Cited by 0SourceScholar