← Search

Bhiksha Raj

93 accepted papers

2026

$\phi$-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal Models

CVPR 2026

Fairness in Continual Learning for Large Multimodal Models (LMMs) is an emerging yet underexplored challenge, particularly in the presence of imbalanced data distributions that can lead to biased model updates and suboptimal performance across tasks. While recent continual learning studies have made

Cited by 0SourceScholar
2026

VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding

ICLR 2026poster

Precisely evaluating video understanding models remains challenging: commonly used metrics such as BLEU, ROUGE, and BERTScore fail to capture the fineness of human judgment, while obtaining such judgments through manual evaluation is costly. Recent work has explored using large language models (LLMs…

Cited by 0SourceScholar
2025

ADIFF: Explaining audio difference using natural language

ICLR 2025spotlight

Understanding and explaining differences between audio recordings is crucial for fields like audio forensics, quality assessment, and audio generation. This involves identifying and describing audio events, acoustic scenes, signal characteristics, and their emotional impact on listeners. This paper…

2025

Audio Entailment: Assessing Deductive Reasoning for Audio Understanding

AAAI 2025technical

Recent literature uses language to build foundation models for audio. These Audio-Language Models (ALMs) are trained on a vast number of audio-text pairs and show remarkable performance in tasks including Text-to-Audio Retrieval, Captioning, and Question Answering. However, their ability to engage i…

2025

CAARMA: Class Augmentation with Adversarial Mixup Regularization

EMNLP 2025

Speaker verification is a typical zero-shot learning task, where inference of unseen classes is performed by comparing embeddings of test instances to known examples. The models performing inference must hence naturally generate embeddings that cluster same-class instances compactly, while maintaini

2025

Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models

NeurIPS 2025poster

Large multimodal models (LMMs) have gained impressive performance due to their outstanding capability in various understanding tasks. However, these models still suffer from some fundamental limitations related to robustness and generalization due to the alignment and correlation between visual and…

Cited by 0SourceScholar
2025

FALCON: Fairness Learning via Contrastive Attention Approach to Continual Semantic Scene Understanding

CVPR 2025poster

Continual Learning in semantic scene segmentation aims to continually learn new unseen classes in dynamic environments while maintaining previously learned knowledge. Prior studies focused on modeling the catastrophic forgetting and background shift challenges in continual learning. However, fairnes…

Cited by 0SourcePDFScholar
2025

ImageFolder: Autoregressive Image Generation with Folded Tokens

ICLR 2025poster

Image tokenizers are crucial for visual generative models, \eg, diffusion models (DMs) and autoregressive (AR) models, as they construct the latent representation for modeling. Increasing token length is a common approach to improve image reconstruction quality. However, tokenizers with longer token…

2025

Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models

ACL 2025finding

Speech foundation models trained at a massive scale, both in terms of model and data size, result in robust systems capable of performing multiple speech tasks, including automatic speech recognition (ASR). These models transcend language and domain barriers, yet effectively measuring their performa…

2025

Masked Autoencoders Are Effective Tokenizers for Diffusion Models

ICML 2025spotlight

Recent advances in latent diffusion models have demonstrated their effectiveness for high-resolution image synthesis. However, the properties of the latent space from tokenizer for better learning and generation of diffusion models remain under-explored. Theoretically and empirically, we find that i…

Cited by 8SourcePDFScholar
2025

On Fairness of Unified Multimodal Large Language Model for Image Generation

NeurIPS 2025poster

Unified multimodal large language models (U-MLLMs) have demonstrated impressive performance in end-to-end visual understanding and generation tasks. However, compared to generation-only systems (e.g., Stable Diffusion), the unified architecture of U-MLLMs introduces new risks of propagating demograp…

Cited by 0SourceScholar
2025

PhoniTale: Phonologically Grounded Mnemonic Generation for Typologically Distant Language Pairs

EMNLP 2025

Vocabulary acquisition poses a significant challenge for second-language (L2) learners, especially when learning typologically distant languages such as English and Korean, where phonological and structural mismatches complicate vocabulary learning. Recently, large language models (LLMs) have been u

2025

SVeritas: Benchmark for Robust Speaker Verification under Diverse Conditions

EMNLP 2025

Speaker verification (SV) models are increasingly integrated into security, personalization, and access control systems, yet their robustness to many real-world challenges remains inadequately benchmarked. Real-world systems can face diverse conditions, some naturally occurring, and others that may

2025

Scalable Benchmarking and Robust Learning for Noise-Free Ego-Motion and 3D Reconstruction from Noisy Video

ICLR 2025poster

We aim to redefine robust ego-motion estimation and photorealistic 3D reconstruction by addressing a critical limitation: the reliance on noise-free data in existing models. While such sanitized conditions simplify evaluation, they fail to capture the unpredictable, noisy complexities of real-world…

2025

SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer

CVPR 2025poster

Efficient image tokenization with high compression ratios remains a critical challenge for training generative models.We present SoftVQ-VAE, a continuous image tokenizer that leverages soft categorical posteriors to aggregate multiple codewords into each latent token, substantially increasing the re…

2025

Speech Robust Bench: A Robustness Benchmark For Speech Recognition

ICLR 2025poster

As Automatic Speech Recognition (ASR) models become ever more pervasive, it is important to ensure that they make reliable predictions under corruptions present in the physical and digital world. We propose Speech Robust Bench (SRB), a comprehensive benchmark for evaluating the robustness of ASR mo…

2025

Toward Material-Agnostic System Identification from Videos

ICCV 2025poster

System identification from videos aims to recover object geometry and governing physical laws. Existing methods integrate differentiable rendering with simulation but rely on predefined material priors, limiting their ability to handle unknown ones. We introduce MASIV, the first vision-based framewo…

2025

Unsupervised Disentanglement of Content and Style via Variance-Invariance Constraints

ICLR 2025poster

We contribute an unsupervised method that effectively learns disentangled content and style representations from sequences of observations. Unlike most disentanglement algorithms that rely on domain-specific labels or knowledge, our method is based on the insight of domain-general statistical differ…

Cited by 0SourcePDFScholar
2025

uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes

NAACL 2025long

Recent work on distilling Whisper’s knowledge into small models using pseudo-labels shows promising performance while reducing the size by up to 50%. This results in small, efficient, and dedicated models. However, a critical step of distillation using pseudo-labels involves filtering high-quality p…

2024

A General Framework for Learning from Weak Supervision

ICML 2024poster

Weakly supervised learning generally faces challenges in applicability to various scenarios with diverse weak supervision and in scalability due to the complexity of existing algorithms, thereby hindering the practical deployment. This paper introduces a general framework for learning from weak supe…

2024

AugSumm: Towards Generalizable Speech Summarization Using Synthetic Labels from Large Language Models

ICASSP 2024accepted

Abstractive speech summarization (SSUM) aims to generate humanlike summaries from speech. Given variations in information captured and phrasing, recordings can be summarized in multiple ways. Therefore, it is more reasonable to consider a probabilistic distribution of all potential summaries rather…

Cited by 0SourceScholar
2024

AutoPRM: Automating Procedural Supervision for Multi-Step Reasoning via Controllable Question Decomposition

NAACL 2024long

Recent advancements in large language models (LLMs) have shown promise in multi-step reasoning tasks, yet their reliance on extensive manual labeling to provide procedural feedback remains a significant impediment. To address this challenge, in this paper, we propose a novel self-supervised framewor…

Cited by 25SourcePDFScholar
2024

Completing Visual Objects via Bridging Generation and Segmentation

ICML 2024poster

This paper presents a novel approach to object completion, with the primary goal of reconstructing a complete object from its partially visible components. Our method, named MaskComp, delineates the completion process through iterative stages of generation and segmentation. In each iteration, the ob…

Cited by 6SourcePDFScholar
2024

Continual Contrastive Spoken Language Understanding

ACL 2024findings

Recently, neural networks have shown impressive progress across diverse fields, with speech processing being no exception. However, recent breakthroughs in this area require extensive offline training using large datasets and tremendous computing resources. Unfortunately, these models struggle to re…

Cited by 2SourcePDFScholar
2024

EAGLE: Efficient Adaptive Geometry-based Learning in Cross-view Understanding

NeurIPS 2024poster

Unsupervised Domain Adaptation has been an efficient approach to transferring the semantic segmentation model across data distributions. Meanwhile, the recent Open-vocabulary Semantic Scene understanding based on large-scale vision language models is effective in open-set settings because it can lea…

Cited by 1SourcePDFScholar
2024

Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label Configurations

NeurIPS 2024poster

Learning with reduced labeling standards, such as noisy label, partial label, and supplementary unlabeled data, which we generically refer to as imprecise label, is a commonplace challenge in machine learning tasks. Previous methods tend to propose specific designs for every emerging imprecise label…

2024

Improving Continual Learning of Acoustic Scene Classification via Mutual Information Optimization

ICASSP 2024accepted

Continual learning, which aims to incrementally accumulate knowledge, has been an increasingly significant but challenging research topic for deep models that are prone to catastrophic forgetting. In this paper, we propose a novel replay-based continual learning approach in the context of class-incr…

Cited by 0SourceScholar
2024

Metric from Human: Zero-shot Monocular Metric Depth Estimation via Test-time Adaptation

NeurIPS 2024poster

Monocular depth estimation (MDE) is fundamental for deriving 3D scene structures from 2D images. While state-of-the-art monocular relative depth estimation (MRDE) excels in estimating relative depths for in-the-wild images, current monocular metric depth estimation (MMDE) approaches still face chall…

Cited by 4SourcePDFScholar
2024

Prompting Audios Using Acoustic Properties for Emotion Representation

ICASSP 2024accepted

Emotions lie on a continuum, but current models treat emotions as a finite valued discrete variable. This representation does not capture the diversity in the expression of emotion. To better represent emotions we propose the use of natural language descriptions (or prompts). In this work, we addres…

Cited by 0SourceScholar
2024

QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic Decomposition

CVPR 2024poster

Audiovisual segmentation (AVS) is a challenging task that aims to segment visual objects in videos according to their associated acoustic cues. With multiple sound sources and background disturbances involved establishing robust correspondences between audio and visual contents poses unique challeng…

2024

R-BASS : Relevance-aided Block-wise Adaptation for Speech Summarization

NAACL 2024findings

End-to-end speech summarization on long recordings is challenging because of the high computational cost. Block-wise Adaptation for Speech Summarization (BASS) summarizes arbitrarily long sequences by sequentially processing abutting chunks of audio. Despite the benefits of BASS, it has higher compu…

Cited by 0SourcePDFScholar
2024

R^2-Bench: Benchmarking the Robustness of Referring Perception Models under Perturbations

ECCV 2024poster

"Referring perception, which aims at grounding visual objects with multimodal referring guidance, is essential for bridging the gap between humans, who provide instructions, and the environment where intelligent systems perceive. Despite progress in this field, the robustness of referring perception…

Cited by 3SourcePDFScholar
2024

Slight Corruption in Pre-training Data Makes Better Diffusion Models

NeurIPS 2024spotlight

Diffusion models (DMs) have shown remarkable capabilities in generating realistic high-quality images, audios, and videos. They benefit significantly from extensive pre-training on large-scale datasets, including web-crawled data with paired data and conditions, such as image-text and image-class p…

Cited by 6SourcePDFScholar
2024

Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?

ACL 2024long

Reference summaries for abstractive speech summarization require human annotation, which can be performed by listening to an audio recording or by reading textual transcripts of the recording. In this paper, we examine whether summaries based on annotators listening to the recordings differ from tho…

2024

Synergistic Global-space Camera and Human Reconstruction from Videos

CVPR 2024poster

Remarkable strides have been made in reconstructing static scenes or human bodies from monocular videos. Yet the two problems have largely been approached independently without much synergy. Most visual SLAM methods can only reconstruct camera trajectories and scene structures up to scale while most…

Cited by 3SourcePDFScholar
2024

Training Audio Captioning Models without Audio

ICASSP 2024accepted

Automated Audio Captioning (AAC) is the task of generating natural language descriptions given an audio stream. A typical AAC system requires manually curated training data of audio segments and corresponding text caption annotations. The creation of these audio-caption pairs is costly, resulting in…

Cited by 0SourceScholar
2024

Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks

ICLR 2024spotlight

Pre-training on large-scale datasets and then fine-tuning on downstream tasks have become a standard practice in deep learning. However, pre-training data often contain label noise that may adversely affect the generalization of the model. This paper aims to understand the nature of noise in pre-tra…

2024

uSee: Unified Speech Enhancement And Editing with Conditional Diffusion Models

ICASSP 2024accepted

Speech enhancement aims to improve the quality of speech signals in terms of quality and intelligibility, and speech editing refers to the process of editing the speech according to specific user needs. In this paper, we propose a Unified Speech Enhancement and Editing (uSee) model with conditional…

Cited by 15SourceScholar
2023

An Approach to Ontological Learning from Weak Labels

ICASSP 2023accepted

Ontologies encompass a formal representation of knowledge through the definition of concepts or properties of a domain, and the relationships between those concepts. In this work, we seek to investigate whether using this ontological information will improve learning from weakly labeled data, which…

Cited by 0SourceScholar
2023

FREDOM: Fairness Domain Adaptation Approach to Semantic Scene Understanding

CVPR 2023poster

Although Domain Adaptation in Semantic Scene Segmentation has shown impressive improvement in recent years, the fairness concerns in the domain adaptation have yet to be well defined and addressed. In addition, fairness is one of the most critical aspects when deploying the segmentation models into…

2023

Fairness Continual Learning Approach to Semantic Scene Understanding in Open-World Environments

NeurIPS 2023poster

Continual semantic segmentation aims to learn new classes while maintaining the information from the previous classes. Although prior studies have shown impressive progress in recent years, the fairness concern in the continual semantic segmentation needs to be better addressed. Meanwhile, fairness…

Cited by 14SourcePDFScholar
2023

FreeMatch: Self-adaptive Thresholding for Semi-supervised Learning

ICLR 2023poster

Semi-supervised Learning (SSL) has witnessed great success owing to the impressive performances brought by various methods based on pseudo labeling and consistency regularization. However, we argue that existing methods might fail to utilize the unlabeled data more effectively since they either use…

2023

How Many Perturbations Break This Model? Evaluating Robustness Beyond Adversarial Accuracy

ICML 2023poster

Robustness to adversarial attacks is typically evaluated with adversarial accuracy. While essential, this metric does not capture all aspects of robustness and in particular leaves out the question of how many perturbations can be found for each point. In this work, we introduce an alternative appro…

2023

Paaploss: A Phonetic-Aligned Acoustic Parameter Loss for Speech Enhancement

ICASSP 2023accepted

Despite rapid advancement in recent years, current speech enhancement models often produce speech that differs in perceptual quality from real clean speech. We propose a learning objective that formalizes differences in perceptual quality, by using domain knowledge of acoustic-phonetics. We identify…

Cited by 0SourceScholar
2023

PaintSeg: Painting Pixels for Training-free Segmentation

NeurIPS 2023poster

The paper introduces PaintSeg, a new unsupervised method for segmenting objects without any training. We propose an adversarial masked contrastive painting (AMCP) process, which creates a contrast between the original image and a painted image in which a masked area is painted using off-the-shelf ge…

2023

Pairwise Similarity Learning is SimPLE

ICCV 2023poster

In this paper, we focus on a general yet important learning problem, pairwise similarity learning (PSL). PSL subsumes a wide range of important applications, such as open-set face recognition, speaker verification, image retrieval and person re-identification. The goal of PSL is to learn a pairwise…

Cited by 10PDFcodeScholar
2023

Panoramic Video Salient Object Detection with Ambisonic Audio Guidance

AAAI 2023technical

Video salient object detection (VSOD), as a fundamental computer vision problem, has been extensively discussed in the last decade. However, all existing works focus on addressing the VSOD problem in 2D scenarios. With the rapid development of VR devices, panoramic videos have been a promising alter…

Cited by 16SourcePDFScholar
2023

Robust Referring Video Object Segmentation with Cyclic Structural Consensus

ICCV 2023poster

Referring Video Object Segmentation (R-VOS) is a challenging task that aims to segment an object in a video based on a linguistic expression. Most existing R-VOS methods have a critical assumption: the object referred to must appear in the video. This assumption, which we refer to as "semantic conse…

Cited by 35PDFScholar
2023

SoftMatch: Addressing the Quantity-Quality Tradeoff in Semi-supervised Learning

ICLR 2023poster

The critical challenge of Semi-Supervised Learning (SSL) is how to effectively leverage the limited labeled data and massive unlabeled data to improve the model's generalization performance. In this paper, we first revisit the popular pseudo-labeling methods via a unified sample weighting formulatio…

2023

TAPLoss: A Temporal Acoustic Parameter Loss for Speech Enhancement

ICASSP 2023accepted

Speech enhancement models have greatly progressed in recent years, but still show limits in perceptual quality of their speech outputs. We propose an objective for perceptual quality based on temporal acoustic parameters. These are fundamental speech features that play an essential role in various a…

Cited by 0SourceScholar
2023

Token Prediction as Implicit Classification to Identify LLM-Generated Text

EMNLP 2023short main

This paper introduces a novel approach for identifying the possible large language models (LLMs) involved in text generation. Instead of adding an additional classification layer to a base LM, we reframe the classification task as a next-token prediction task and directly fine-tune the base LM to pe…

Cited by 0SourcecodeScholar
2023

Towards Noise-Tolerant Speech-Referring Video Object Segmentation: Bridging Speech and Text

EMNLP 2023long main

Linguistic communication is prevalent in Human-Computer Interaction (HCI). Speech (spoken language) serves as a convenient yet potentially ambiguous form due to noise and accents, exposing a gap compared to text. In this study, we investigate the prominent HCI task, Referring Video Object Segmentati…

Cited by 0SourceScholar
2023

Training on Foveated Images Improves Robustness to Adversarial Attacks

NeurIPS 2023poster

Deep neural networks (DNNs) have been shown to be vulnerable to adversarial attacks -- subtle, perceptually indistinguishable perturbations of inputs that change the response of the model. In the context of vision, we hypothesize that an important contributor to the robustness of human visual perce…

Cited by 3SourcePDFScholar
2023

VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning

AAAI 2023technical

Video Paragraph Captioning aims to generate a multi-sentence description of an untrimmed video with multiple temporal event locations in a coherent storytelling. Following the human perception process, where the scene is effectively understood by decomposing it into visual (e.g. human, animal) and…

2022

SphereFace2: Binary Classification is All You Need for Deep Face Recognition

ICLR 2022spotlight

State-of-the-art deep face recognition methods are mostly trained with a softmax-based multi-class classification framework. Despite being popular and effective, these methods still have a few shortcomings that limit empirical performance. In this paper, we start by identifying the discrepancy betwe…

Cited by 63SourcePDFScholar
2022

USB: A Unified Semi-supervised Learning Benchmark for Classification

NeurIPS 2022accept

Semi-supervised learning (SSL) improves model generalization by leveraging massive unlabeled data to augment limited labeled samples. However, currently, popular SSL evaluation protocols are often constrained to computer vision (CV) tasks. In addition, previous work typically trains deep neural netw…

2021

Contrast and Order Representations for Video Self-Supervised Learning

ICCV 2021poster

This paper studies the problem of learning self-supervised representations on videos. In contrast to image modality that only requires appearance information on objects or scenes, video needs to further explore the relations between multiple frames/clips along the temporal dimension. However, the re…

Cited by 77PDFcodeScholar
2021

FoolHD: Fooling Speaker Identification by Highly Imperceptible Adversarial Disturbances

ICASSP 2021accepted

Speaker identification models are vulnerable to carefully designed adversarial perturbations of their input signals that induce misclassification. In this work, we propose a white-box steganography-inspired adversarial attack that generates imperceptible adversarial perturbations against a speaker i…

Cited by 0SourceScholar
2021

The Right To Talk: An Audio-Visual Transformer Approach

ICCV 2021poster

Turn-taking has played an essential role in structuring the regulation of a conversation. The task of identifying the main speaker (who is properly taking his/her turn of speaking) and the interrupters (who are interrupting or reacting to the main speaker's utterances) remains a challenging task. Al…

Cited by 46PDFcodeScholar
2021

The in-the-Wild Speech Medical Corpus

ICASSP 2021accepted

Automatic detection of speech affecting (SA) diseases has received significant attention, particularly in clinical scenarios. However, the same task in in-the-wild conditions is often neglected, in part, due to the lack of appropriate datasets.In this work, we present the in-the-Wild Speech Medical…

Cited by 0SourceScholar
2020

Is normalization indispensable for training deep neural network?

NeurIPS 2020oral

Normalization operations are widely used to train deep neural networks, and they can improve both convergence and generalization in most tasks. The theories for normalization's effectiveness and new forms of normalization have always been hot topics in research. To better understand normalization, o…

2019

Cross Modal Audio Search and Retrieval with Joint Embeddings Based on Text and Audio

ICASSP 2019accepted

Existing audio search engines use one of two approaches: matching text-text or audio-audio pairs. In the former, text queries are matched to semantically similar words in an index of audio metadata to retrieve corresponding audio clips or segments, while in the latter, audio signals are directly use…

Cited by 65SourceScholar
2019

Disjoint Mapping Network for Cross-modal Matching of Voices and Faces

ICLR 2019poster

We propose a novel framework, called Disjoint Mapping Network (DIMNet), for cross-modal biometric matching, in particular of voices and faces. Different from the existing methods, DIMNet does not explicitly learn the joint relationship between the modalities. Instead, DIMNet learns a shared represen…

Cited by 92SourcePDFScholar
2019

Face Reconstruction from Voice using Generative Adversarial Networks

NeurIPS 2019poster

Voice profiling aims at inferring various human parameters from their speech, e.g. gender, age, etc. In this paper, we address the challenge posed by a subtask of voice profiling - reconstructing someone's face from their voice. The task is designed to answer the question: given an audio clip spoken…

2018

A Corrective Learning Approach for Text-Independent Speaker Verification

ICASSP 2018accepted

We present a conceptually plausible approach for text-independent speaker verification (TISV) which treats speech recordings as a collection of segments providing incremental evidence. This approach, called corrective learning, gradually improves an initial prediction of speaker identity based on in…

Cited by 0SourceScholar
2018

Acoustic Scene Classification Using Discrete Random Hashing for Laplacian Kernel Machines

ICASSP 2018accepted

State of the art acoustic scene classification techniques often employ features of large dimensionality, which are then used to train and perform inferences with kernel machines such as Support Vector Machines. However, the complexity of computing the non-linear kernel matrix for these methods incre…

Cited by 0SourceScholar
2018

Content-Based Representations of Audio Using Siamese Neural Networks

ICASSP 2018accepted

In this paper, we focus on the problem of content-based retrieval for audio, which aims to retrieve all semantically similar audio recordings for a given audio clip query. This problem is similar to the problem of query by example of audio, which aims to retrieve media samples from a database, which…

Cited by 0SourceScholar
2018

Framework for Evaluation of Sound Event Detection in Web Videos

ICASSP 2018accepted

The largest source of sound events is web videos. Most videos lack sound event labels at segment level, however, a significant number of them do respond to text queries, from a match found using metadata by search engines. In this paper we explore the extent to which a search query can be used as th…

Cited by 0SourceScholar
2017

SphereFace: Deep Hypersphere Embedding for Face Recognition

CVPR 2017poster

This paper addresses deep face recognition (FR) problem under open-set protocol, where ideal face features are expected to have smaller maximal intra-class distance than minimal inter-class distance under a suitably chosen metric space. However, few existing algorithms can effectively achieve this c…

Cited by 3736PDFcodeScholar
2016

The relationship of voice onset time and Voice Offset Time to physical age

ICASSP 2016accepted

In a speech signal, Voice Onset Time (VOT) is the period between the release of a plosive and the onset of vocal cord vibrations in the production of the following sound. Voice Offset Time (VOFT), on the other hand, is the period between the end of a voiced sound and the release of the following plo…

Cited by 0SourceScholar
2015

Beyond Gaussian Pyramid: Multi-Skip Feature Stacking for Action Recognition

CVPR 2015poster

Most state-of-the-art action feature extractors involve differential operators, which act as highpass filters and tend to attenuate low frequency action information. This attenuation introduces bias to the resulting features and generates ill-conditioned feature matrices. The Gaussian Pyramid has be…

Cited by 365SourcePDFScholar
2015

Reducing communication overhead in distributed learning by an order of magnitude (almost)

ICASSP 2015accepted

Large-scale distributed learning plays an ever-more increasing role in modern computing. However, whether using a compute cluster with thousands of nodes, or a single multi-GPU machine, the most significant bottleneck is that of communication. In this work, we explore the effects of applying quantiz…

Cited by 0SourceScholar