← Search

Viktor Rozgic

11 accepted papers

2025

Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge

NeurIPS 2025poster

Agentic search such as Deep Research systems-where agents autonomously browse the web, synthesize information, and return comprehensive citation-backed answers-represents a major shift in how users interact with web-scale information. While promising greater efficiency and cognitive offloading, the…

Cited by 0SourceScholar
2023

FedRPO: Federated Relaxed Pareto Optimization for Acoustic Event Classification

ICASSP 2023accepted

Performance and robustness of real-world Acoustic Event Classification (AEC) solutions depend on ability to train on diverse data from wide range of end-point devices and acoustic environments. Federated Learning (FL) provides a framework to leverage annotated and non-annotated AEC data from servers…

Cited by 0SourceScholar
2023

Weight-Sharing Supernet for Searching Specialized Acoustic Event Classification Networks Across Device Constraints

ICASSP 2023accepted

Acoustic Event Classification (AEC) has been widely used in devices such as smart speakers and mobile phones for home safety or accessibility support [1]. As AEC models run on more and more devices with diverse computation resource constraints, it became increasingly expensive to develop models that…

Cited by 1SourceScholar
2022

Confidence Estimation for Speech Emotion Recognition Based on the Relationship Between Emotion Categories and Primitives

ICASSP 2022accepted

Confidence estimation for Speech Emotion Recognition (SER) is instrumental in improving the reliability in the behavior of downstream applications. In this work we propose (1) a novel confidence metric for SER based on the relationship between emotion primitives: arousal, valence, and dominance (AVD…

Cited by 0SourceScholar
2022

Federated Self-Supervised Learning for Acoustic Event Classification

ICASSP 2022accepted

Standard acoustic event classification (AEC) solutions require large-scale collection of data from client devices for model optimization. Federated learning (FL) is a compelling frame- work that decouples data collection and model training to enhance customer privacy. In this work, we investigate th…

Cited by 14SourceScholar
2022

Improved Representation Learning For Acoustic Event Classification Using Tree-Structured Ontology

ICASSP 2022accepted

Acoustic events have a hierarchical structure analogous to a tree (or a directed acyclic graph). In this work, we propose a structure-aware semi-supervised learning framework for acoustic event classification (AEC). Our hypothesis is that the audio label structure contains useful information that is…

Cited by 0SourceScholar
2022

Sentiment-Aware Automatic Speech Recognition Pre-Training for Enhanced Speech Emotion Recognition

ICASSP 2022accepted

We propose a novel multi-task pre-training method for Speech Emotion Recognition (SER). We pre-train SER model simultaneously on Automatic Speech Recognition (ASR) and sentiment classification tasks to make the acoustic ASR model more "emotion aware". We generate targets for the sentiment classifica…

Cited by 0SourceScholar
2021

Contrastive Unsupervised Learning for Speech Emotion Recognition

ICASSP 2021accepted

Speech emotion recognition (SER) is a key technology to enable more natural human-machine communication. However, SER has long suffered from a lack of public large-scale labeled datasets. To circumvent this problem, we investigate how unsupervised representation learning on unlabeled datasets can be…

Cited by 0SourceScholar
2019

Hierarchical Residual-pyramidal Model for Large Context Based Media Presence Detection

ICASSP 2019accepted

We study media presence detection, that is, learning to recognize if a sound segment (typically lasting for a few seconds) of a long recorded stream contains media (TV) sound. This problem is difficult because non-media sound sources can be quite diverse (e.g. human voicing, non-vocal sounds and non…

Cited by 0SourceScholar
2019

Improving Emotion Classification through Variational Inference of Latent Variables

ICASSP 2019accepted

Conventional models for emotion recognition from speech signal are trained in supervised fashion using speech utterances with emotion labels. In this study we hypothesize that speech signal depends on multiple latent variables including the emotional state, age, gender, and speech content. We propos…

Cited by 0SourceScholar
2019

Semi-supervised Acoustic Event Detection Based on Tri-training

ICASSP 2019accepted

This paper presents our work of training acoustic event detection (AED) models using unlabeled dataset. Recent acoustic event detectors are based on large-scale neural networks, which are typically trained with huge amounts of labeled data. Labels for acoustic events are expensive to obtain, and rel…

Cited by 0SourceScholar