ICASSP 2021 Accepted Papers
The full list of 1,712 papers accepted at ICASSP 2021 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- MRI Image Recovery using Damped Denoising Vector AMP
- MSR-GAN: Multi-Segment Reconstruction via Adversarial Learning
- Machine Translation Verbosity Control for Automatic Dubbing
- Makf-Sr: Multi-Agent Adaptive Kalman Filtering-Based Successor Representations
- Making Punctuation Restoration Robust and Fast with Multi-Task Learning and Knowledge Distillation
- MarbleNet: Deep 1D Time-Channel Separable Convolutional Neural Network for Voice Activity Detection
- Mask4D: 4D Convolution Network for Light Field Occlusion Removal
- Maskcyclegan-VC: Learning Non-Parallel Voice Conversion with Filling in Frames
- Matching as Color Images: Thermal Image Local Feature Detection and Description
- Maximum a Posteriori Estimator for Convolutive Sound Source Separation with Sub-Source Based NTF Model and the Localization Probabilistic Prior on the Mixing Matrix
- Measure-Transformed Covariance Test for Robust Spectrum Sensing
- Measurement Coding Framework with Adjacent Pixels Based Measurement Matrix for Compressively Sensed Images
- Melody Harmonization Using Orderless Nade, Chord Balancing, and Blocked Gibbs Sampling
- Melon Playlist Dataset: A Public Dataset for Audio-Based Playlist Generation and Music Tagging
- Memory Layers with Multi-Head Attention Mechanisms for Text-Dependent Speaker Verification
- Memory-Efficient Speech Recognition on Smart Devices
- Message Transmission Over Rapidly Time-Varying Channels
- Meta Ordinal Weighting Net For Improving Lung Nodule Classification
- Meta-Adapter: Efficient Cross-Lingual Adaptation With Meta-Learning
- Meta-Cognition-Based Simple And Effective Approach To Object Detection
- Meta-Learning for 6G Communication Networks with Reconfigurable Intelligent Surfaces
- Meta-Learning for Cross-Channel Speaker Verification
- Meta-Learning for Improving Rare Word Recognition in End-to-End ASR
- Meta-Learning for Low-Resource Speech Emotion Recognition
- Meta-Learning with Attention for Improved Few-Shot Learning
- Micaugment: One-Shot Microphone Style Transfer
- Microsoft Speaker Diarization System for the Voxceleb Speaker Recognition Challenge 2020
- Millimeter Wave MIMO Channel Estimation with 1-bit Spatial Sigma-Delta Analog-to-Digital Converters
- Mind the Beat: Detecting Audio Onsets from EEG Recordings of Music Listening
- Minimizing Weighted Concave Impurity Partition Under Constraints
- Minimum Bayes Risk Training for End-to-End Speaker-Attributed ASR
- Misalignment Recognition in Acoustic Sensor Networks Using a Semi-Supervised Source Estimation Method and Markov Random Fields
- Mispronunciation Detection in Non-Native (L2) English with Uncertainty Modeling
- Mitigating Clipping Distortion in OFDM Using Deep Residual Learning
- Mitigating Inter-Subject Brain Signal Variability FOR EEG-Based Driver Fatigue State Classification
- MixSpeech: Data Augmentation for Low-Resource Automatic Speech Recognition
- Mixed Precision Quantization of Transformer Language Models for Speech Recognition
- Mixture of Informed Experts for Multilingual Speech Recognition
- Mixup Regularized Adversarial Networks for Multi-Domain Text Classification
- Model-Inspired Deep Learning for Light-Field Microscopy with Application to Neuron Localization
- Modeling Homophone Noise for Robust Neural Machine Translation
- Modelling Paralinguistic Properties in Conversational Speech to Detect Bipolar Disorder and Borderline Personality Disorder
- Modified Arcsine Law for One-Bit Sampled Stationary Signals with Time-Varying Thresholds
- Modular Binary Tree Architecture for Distributed Large Intelligent Surface
- Modurec: Recommender Systems with Feature and Time Modulation
- Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses
- More: A Metric Learning Based Framework for Open-Domain Relation Extraction
- Movement Detection Using A Reciprocal Received Signal Strength Model
- Moving Object Classification with a Sub-6 GHz Massive MIMO Array Using Real Data
- MuG: A Multipath-Exploited and Grid-Free Localisation Method
- Multi Path Training Framework for Data-Driven Open-Domain Conversation System
- Multi-Branch Tomlinson-Harashima Precoding for Rate Splitting Based Systems with Multiple Antennas
- Multi-Channel Speech Enhancement Using Graph Neural Networks
- Multi-Channel Target Speech Extraction with Channel Decorrelation and Target Speaker Adaptation
- Multi-Decoder Dprnn: Source Separation for Variable Number of Speakers
- Multi-Dialect Speech Recognition in English Using Attention on Ensemble of Experts
- Multi-Directional Convolution Networks with Spatial-Temporal Feature Pyramid Module for Action Recognition
- Multi-Entity Collaborative Relation Extraction
- Multi-Granularity Feature Interaction and Relation Reasoning for 3D Dense Alignment and Face Reconstruction
- Multi-Granularity Heterogeneous Graph for Document-Level Relation Extraction
- Multi-Initialization Meta-Learning with Domain Adaptation
- Multi-Level Adaptive Region of Interest and Graph Learning for Facial Action Unit Recognition
- Multi-Level Group Testing with Application to One-Shot Pooled COVID-19 Tests
- Multi-Level Reversible Encryption for ECG Signals Using Compressive Sensing
- Multi-Modal Label Dequantized Gaussian Process Latent Variable Model for Ordinal Label Estimation
- Multi-Models Fusion for Light Field Angular Super-Resolution
- Multi-Object Tracking Using Poisson Multi-Bernoulli Mixture Filtering For Autonomous Vehicles
- Multi-Order Adversarial Representation Learning for Composed Query Image Retrieval
- Multi-Rate Attention Architecture for Fast Streamable Text-to-Speech Spectrum Modeling
- Multi-Sample Online Learning for Spiking Neural Networks Based on Generalized Expectation Maximization
- Multi-Scale Cascade Disparity Refinement Stereo Network
- Multi-Scale Feature-Guided Stereoscopic Video Quality Assessment Based on 3d Convolutional Neural Network
- Multi-Scale Residual Network for Covid-19 Diagnosis Using Ct-Scans
- Multi-Scale Speaker Diarization with Neural Affinity Score Fusion
- Multi-Scale and Multi-Region Facial Discriminative Representation for Automatic Depression Level Prediction
- Multi-Speaker Emotional Speech Synthesis with Fine-Grained Prosody Modeling
- Multi-Stage Speaker Extraction with Utterance and Frame-Level Reference Signals
- Multi-Step Spoken Language Understanding System Based on Adversarial Learning
- Multi-Target DoA Estimation with an Audio-Visual Fusion Mechanism
- Multi-Task Estimation of Age and Cognitive Decline from Speech
- Multi-Task Learning Via Sharing Inexact Low-Rank Subspace
- Multi-Task Self-Supervised Pre-Training for Music Classification
- Multi-Task Transformer with Input Feature Reconstruction for Dysarthric Speech Recognition
- Multi-Tier Federated Learning for Vertically Partitioned Data
- Multi-Vehicle Velocity Estimation Using IEEE 802.11ad Waveform
- Multi-View Audio And Music Classification
- Multi-View Contrastive Learning for Online Knowledge Distillation
- Multichannel Overlapping Speaker Segmentation Using Multiple Hypothesis Tracking Of Acoustic And Spatial Features
- Multichannel-based Learning for Audio Object Extraction
- Multilabel 12-Lead Electrocardiogram Classification Using Beat to Sequence Autoencoders
- Multilingual Phonetic Dataset for Low Resource Speech Recognition
- Multimodal Cross- and Self-Attention Network for Speech Emotion Recognition
- Multimodal Emotion Recognition with Capsule Graph Convolutional Based Representation Fusion
- Multimodal Metric Learning for Tag-Based Music Retrieval
- Multimodal Punctuation Prediction with Contextual Dropout
- Multiphish: Multi-Modal Features Fusion Networks for Phishing Detection
- Multiple Auxiliary Networks for Single Blind Image Deblurring
- Multiple Human Tracking in Non-Specific Coverage with Wearable Cameras
- Multiple-Hypothesis CTC-Based Semi-Supervised Adaptation of End-to-End Speech Recognition
- Multiple-Input Multiple-Output Fusion Network for Generalized Zero-Shot Learning
- Multistream CNN for Robust Acoustic Modeling
- Multitask Learning and Joint Optimization for Transformer-RNN-Transducer Speech Recognition
- Multivariate Non-Negative Matrix Factorization with Application to Energy Disaggregation
- Multiview Sensing with Unknown Permutations: an Optimal Transport Approach
- Multiview Variational Graph Autoencoders for Canonical Correlation Analysis
- Muse: Multi-Modal Target Speaker Extraction with Visual Cues
- Mutual Information Flows in a Bivariate Point Process
- Mutually-Constrained Monotonic Multihead Attention for Online ASR
- NASA: A Noise-Adaptive and Structure-Aware Learning Framework for Image Deblurring
- NISP: A Multi-lingual Multi-accent Dataset for Speaker Profiling
- NMF-SAE: An Interpretable Sparse Autoencoder for Hyperspectral Unmixing
- NN-KOG2P: A Novel Grapheme-to-Phoneme Model for Korean Language
- NNAKF: A Neural Network Adapted Kalman Filter for Target Tracking
- Near-Optimal Algorithms for Piecewise-Stationary Cascading Bandits
- Near-Optimal Resampling in Particle Filters Using the Ising Energy Model
- Nested Error Map Generation Network for No-Reference Image Quality Assessment
- Nested Learning for Multi-Level Classification
- Network Classifiers Based on Social Learning
- Network Pruning Using Linear Dependency Analysis on Feature Maps
- Network Topology Change-Point Detection from Graph Signals with Prior Spectral Signatures
- Network Topology Inference with Graphon Spectral Penalties
- Network and Content-Dependent Bitrate Ladder Estimation for Adaptive Bitrate Video Streaming
- Network-Aware Optimal Microphone Channel Selection in Wireless Acoustic Sensor Networks
- Neural Architecture Search for LF-MMI Trained Time Delay Neural Networks
- Neural Audio Fingerprint for High-Specific Audio Retrieval Based on Contrastive Learning
- Neural Inverse Text Normalization
- Neural Kalman Filtering for Speech Enhancement
- Neural Layered Min-Sum Decoding for Protograph LDPC Codes
- Neural Network-Based Virtual Microphone Estimator
- Neural Noise Embedding for End-To-End Speech Enhancement with Conditional Layer Normalization
- Neural Utterance Confidence Measure for RNN-Transducers and Two Pass Models
- Neuro-Steered Music Source Separation With EEG-Based Auditory Attention Decoding And Contrastive-NMF
- New Variants of DFA Based on Loess and Lowess Methods: Generalization of the Detrending Moving Average
- Nlkd: Using Coarse Annotations For Semantic Segmentation Based on Knowledge Distillation
- No Relaxation: Guaranteed Recovery of Finite-Valued Signals from Undersampled Measurements
- No-Reference Stereoscopic Image Quality Assessment Based on the Human Visual System
- Node Attribute Completion in Knowledge Graphs with Multi-Relational Propagation
- Noise Level Limited Sub-Modeling for Diffusion Probabilistic Vocoders
- Noise-Assisted Multivariate Variational Mode Decomposition
- Noise-Robust Adaptation Control for Supervised Acoustic System Identification Exploiting a Noise Dictionary
- Non-Autoregressive Sequence-To-Sequence Voice Conversion
- Non-Autoregressive Transformer ASR with CTC-Enhanced Decoder Input
- Non-Coherent DOA Estimation of Off-Grid Signals With Uniform Circular Arrays
- Non-Intrusive Binaural Prediction of Speech Intelligibility Based on Phoneme Classification
- Non-Iterative Blind Calibration of Nested Arrays with Asymptotically Optimal Weighting
- Non-Local Single Image DE-Raining Without Decomposition
- Non-Parallel Many-To-Many Voice Conversion Using Local Linguistic Tokens
- Non-Parallel Many-To-Many Voice Conversion by Knowledge Transfer from a Text-To-Speech Model
- Non-Recursive Graph Convolutional Networks
- Non-Singular Adversarial Robustness of Neural Networks
- Noncontact Heartbeat Detection by Viterbi Algorithm with Fusion of Beat-Beat Interval and Deep Learning-Driven Branch Metrics
- Nonlinear State-Space Generalizations of Graph Convolutional Neural Networks
- Nonnegative Unimodal Matrix Factorization
- Nonstationary Portfolios: Diversification in the Spectral Domain
- Numerical Solution of Stochastic Differential Equations in Stiefel Manifolds via Tangent Space Parametrization
- OAS-Net: Occlusion Aware Sampling Network for Accurate Optical Flow
- ORTHROS: non-autoregressive end-to-end speech translation With dual-decoder
- Object-Oriented Relational Distillation for Object Detection
- On Distributed Composite Tests with Dependent Observations in WSN
- On Information Asymmetry in Online Reinforcement Learning
- On Loss Functions for Deep-Learning Based T60 Estimation
- On Overfitting in Discrete Super-Resolution Recovery
- On Permutation Invariant Training For Speech Source Separation
- On Scaling Contrastive Representations for Low-Resource Speech Recognition
- On Strategic Jamming in Distributed Detection Networks
- On The Accuracy Limit of Joint Time-Delay/Doppler/Acceleration Estimation with a Band-Limited Signal
- On The Adversarial Robustness of Principal Component Analysis
- On The Asymptotic Performance of One-Bit Co-Array-Based Music
- On The Camera Position Dithering In Visual 3d Reconstruction
- On The Effect of Spatial Correlation on Distributed Energy Detection of a Stochastic Process
- On The Power of Deep But Naive Partial Label Learning
- On The Relationship Between Speech-Based Breathing Signal Prediction Evaluation Measures and Breathing Parameters Estimation
- On The Role of Visual Cues in Audiovisual Speech Enhancement
- On The Stability of Graph Convolutional Neural Networks Under Edge Rewiring
- On a Guided Nonnegative Matrix Factorization
- On the Convergence of Randomized Bregman Coordinate Descent for Non-Lipschitz Composite Problems
- On the Design of Square Differential Microphone Arrays with a Multistage Structure
- On the Detection of Pitch-Shifted Voice: Machines and Human Listeners
- On the Marginal Benefit of Active Learning: Does Self-Supervision Eat its Cake?
- On the Optimality of Backward Regression: Sparse Recovery and Subset Selection
- On the Performance-Complexity Tradeoff in Stochastic Greedy Weak Submodular Optimization
- On the Predictability of Hrtfs from Ear Shapes Using Deep Networks
- On the Preparation and Validation of a Large-Scale Dataset of Singing Transcription
- One Shot Learning for Speech Separation
- One-Bit Autocorrelation Estimation With Non-Zero Thresholds
- One-Bit Compressed Sensing Using Untrained Network Prior
- One-Shot Conditional Audio Filtering of Arbitrary Sounds
- One-Shot Voice Conversion Based on Speaker Aware Module
- Online Antenna Selection for Enhanced DOA Estimation
- Online Classification of Dynamic Multilayer-Network Time Series in Riemannian Manifolds
- Online Dynamic Window (ODW) Assisted 2-Stage LSTM Indoor Localization for Smart Phones
- Online Hyper-Parameter Tuning for the Contextual Bandit
- Online Learning of Time-Varying Signals and Graphs
- Online Multi-Hop Information Based Kernel Learning Over Graphs
- Online Time-Varying Topology Identification Via Prediction-Correction Algorithms
- Online Unsupervised Learning Using Ensemble Gaussian Processes with Random Features
- Optimal Attacking Strategy Against Online Reputation Systems with Consideration of the Message-Based Persuasion Phenomenon
- Optimal Detection in the Presence of Non-Gaussian Jamming
- Optimal Importance Sampling for Federated Learning
- Optimal Questionnaires for Screening of Strategic Agents
- Optimal Selection of Matrix Shape and Decomposition Scheme for Neural Network Compression
- Optimal TOA Localization for Moving Sensor in Asymmetric Network
- Optimize What Matters: Training DNN-Hmm Keyword Spotting Model Using End Metric
- Optimizing Coverage and Capacity in Cellular Networks using Machine Learning
- Optimizing Short-Time Fourier Transform Parameters via Gradient Descent
- Optimum Feature Ordering for Dynamic Instance-Wise Joint Feature Selection and Classification
- Ordered Reliability Bits Guessing Random Additive Noise Decoding
- Orthogonality and Zero DC Tradeoffs in Biorthogonal Graph Filterbanks
- Outlier-Robust Kernel Hierarchical-Optimization RLS on a Budget with Affine Constraints
- Overcoming Measurement Inconsistency In Deep Learning For Linear Inverse Problems: Applications In Medical Imaging
- PD-GAN: Perceptual-Details GAN for Extremely Noisy Low Light Image Enhancement
- POLA: Online Time Series Prediction by Adaptive Learning Rates
- PPG-Based Singing Voice Conversion with Adversarial Representation Learning
- Paragraph Level Multi-Perspective Context Modeling for Question Generation
- Parallel Iterated Extended and Sigma-Point Kalman Smoothers
- Parallel Tacotron: Non-Autoregressive and Controllable TTS
- Parallel Waveform Synthesis Based on Generative Adversarial Networks with Voicing-Aware Conditional Discriminators
- Parameter Estimation for Coherent Passive MIMO Radar with Unknown Signals under Direct Path Influence
- Parameter Estimation for Student's t VAR Model with Missing Data
- Parameter Identifiability Of Spatial-Smoothing-Based Bistatic Mimo Radar
- Parametric Spectral Filters for Fast Converging, Scalable Convolutional Neural Networks
- Part-Aligned Network with Background for Misaligned Person Search
- Partial Feature Aggregation Network for Real-Time Object Counting
- Partially Overlapped Inference for Long-Form Speech Recognition
- Particle Gibbs Sampling for Regime-Switching State-Space Models
- Patch Decoder-Side Depth Estimation In Mpeg Immersive Video
- Patnet : A Phoneme-Level Autoregressive Transformer Network for Speech Synthesis
- Pause-Encoded Language Models for Recognition of Alzheimer's Disease and Emotion
- Perceptual Loss Based Speech Denoising with an Ensemble of Audio Pattern Recognition and Self-Supervised Models
- Perceptual Quality Assessment for Recognizing True and Pseudo 4k Content
- Performance Analysis of Spatial and Frequency Domain Index-Modulated Reconfigurable Intelligent Metasurfaces
- Periodic Signal Denoising: An Analysis-Synthesis Framework Based on Ramanujan Filter Banks and Dictionaries
- Periodnet: A Non-Autoregressive Waveform Generation Model with a Structure Separating Periodic and Aperiodic Components
- Personalization Strategies for End-to-End Speech Recognition Systems
- Personalized HRTF Modeling Using DNN-Augmented BEM
- Phase Recovery with Bregman Divergences for Audio Source Separation
- Phase Transitions for One-Vs-One and One-Vs-All Linear Separability in Multiclass Gaussian Mixtures
- Phone Distribution Estimation for Low Resource Languages
- Phoneme Based Neural Transducer for Large Vocabulary Speech Recognition
- Phoneme-Based Distribution Regularization for Speech Enhancement
- Physical-Layer Security via Distributed Beamforming in the Presence of Adversaries with Unknown Locations
- Pipeline Safety Early Warning Method for Distributed Signal using Bilinear CNN and LightGBM
- Pitch-Timbre Disentanglement Of Musical Instrument Sounds Based On Vae-Based Metric Learning
- Planar Array Geometry Optimization for Region Sound Acquisition
- Playing a Part: Speaker Verification at the movies
- Plug-And-Play Learned Gaussian-mixture Approximate Message Passing
- Point of Care Image Analysis for COVID-19
- Pointer Networks for Arbitrary-Shaped Text Spotting
- Policy Augmentation: An Exploration Strategy For Faster Convergence of Deep Reinforcement Learning Algorithms
- Polynomial Matrix Eigenvalue Decomposition of Spherical Harmonics for Speech Enhancement
- Portable Photoglottography for Monitoring Vocal Fold Vibrations in Speech Production
- Positnn: Training Deep Neural Networks with Mixed Low-Precision Posit
- Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Text Data
- Prediction of Egfr Mutation Status in Lung Adenocarcinoma Using Multi-Source Feature Representations
- Prediction of Object Geometry from Acoustic Scattering Using Convolutional Neural Networks
- Predictive Coding for Lossless Dataset Compression
- Preventing Early Endpointing for Online Automatic Speech Recognition
- Privacy-Accuracy Trade-Off of Inference as Service
- Privacy-Preserving Cloud-Based DNN Inference
- Privacy-Preserving Optimal Insulin Dosing Decision
- Privacy-Preserving near Neighbor Search via Sparse Coding with Ambiguation
- Private Wireless Federated Learning with Anonymous Over-the-Air Computation
- Probabilistic Graph Neural Networks for Traffic Signal Control
- Probability of Resolution of G-MUSIC: An Asymptotic Approach
- Probing Acoustic Representations for Phonetic Properties
- Processing Pipelines for Efficient, Physically-Accurate Simulation of Microphone Array Signals in Dynamic Sound Scenes
- Progressive Co-Teaching for Ambiguous Speech Emotion Recognition
- Progressive Multi-Stage Feature Mix for Person Re-Identification
- Progressive Spatio-Temporal Graph Convolutional Network for Skeleton-Based Human Action Recognition
- Progressive Voice Trigger Detection: Accuracy vs Latency
- Prosodic Clustering for Phoneme-Level Prosody Control in End-to-End Speech Synthesis
- Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech
- Prosody and Voice Factorization for Few-Shot Speaker Adaptation in the Challenge M2voc 2021
- Prototype-Based Personalized Pruning
- Prototypical Networks for Domain Adaptation in Acoustic Scene Classification
- Provably Fast Asynchronous And Distributed Algorithms For Pagerank Centrality Computation
- Pruning of Convolutional Neural Networks using ising Energy Model
- Pushing The Limit of Type I Codebook For Fdd Massive Mimo Beamforming: A Channel Covariance Reconstruction Approach
- Pushing the Limit of Phase Offset for Contactless Sensing Using Commodity Wifi
- Pyramid U-Net for Retinal Vessel Segmentation
- QUERYD: A Video Dataset with High-Quality Text and Audio Narrations
- QoE-Driven and Tile-Based Adaptive Streaming for Point Clouds
- Query-By-Example Keyword Spotting System Using Multi-Head Attention and Soft-triple Loss
- Quickest Change Detection With Time Inconsistent Anticipatory Agents In Cyber-Physical Systems
- Quickest Joint Detection and Classification of Faults in Statistically Periodic Processes
- REDAT: Accent-Invariant Representation for End-To-End ASR by Domain Adversarial Training with Relabeling
- REPAC: Reliable Estimation of Phase-Amplitude Coupling in Brain Networks
- REST: Robust lEarned Shrinkage-Thresholding Network Taming Inverse Problems with Model Mismatch
- RGLN: Robust Residual Graph Learning Networks via Similarity-Preserving Mapping on Graphs
- RIS-Aided Joint Localization and Synchronization with a Single-Antenna Mmwave Receiver
- RNN Transducer Models for Spoken Language Understanding
- RNN-T Based Open-Vocabulary Keyword Spotting in Mandarin with Multi-Level Detection
- Radar Clutter Classification Using Expectation-Maximization Method
- Radio Frequency Based Heart Rate Variability Monitoring
- Random Projection Streams for (Weighted) Nonnegative Matrix Factorization
- Range Guided Depth Refinement and Uncertainty-Aware Aggregation for View Synthesis
- Rank-Revealing Block-Term Decomposition for Tensor Completion
- Rate 1 Quasi Orthogonal Universal Transmission and Combining for MIMO Systems Achieving Full Diversity
- Rate-Distortion Optimized Motion Estimation for on-the-Sphere Compression of 360 Videos
- Raw Data Processing for Practical Time-of-Flight Super-Resolution
- Real Image Super-Resolution Using Token Based Contextual Attention
- Real Number Signal Processing can Detect Denial-of-Service Attacks
- Real Versus Fake 4k - Authentic Resolution Assessment
- Real-Time Denoising and Dereverberation wtih Tiny Recurrent U-Net
- Real-Time Interaural Time Delay Estimation via Onset Detection
- Real-Time Radio Modulation Classification With An LSTM Auto-Encoder
- Real-Time Speech Enhancement for Mobile Communication Based on Dual-Channel Complex Spectral Mapping
- Real-Time Speech Frequency Bandwidth Extension
- Real-Time Synchronization in Neural Networks for Multivariate Time Series Anomaly Detection
- Recent Advances in Arabic Syntactic Diacritics Restoration
- Recent Developments on Espnet Toolkit Boosted By Conformer
- Recognition of Dynamic Hand Gesture Based on Mm-Wave Fmcw Radar Micro-Doppler Signatures
- Recurrent Phase Reconstruction Using Estimated Phase Derivatives from Deep Neural Networks
- Recursive Input and State Estimation: a General Framework for Learning from Time Series With Missing Data
- Reduced-Complexity Channel Estimation by Hierarchical Interpolation Exploiting Sparsity for Massive MIMO Systems with Uniform Rectangular Array
- Reduced-Complexity Modular Polynomial Multiplication for R-LWE Cryptosystems
- Reducing Modal Error Propagation through Correcting Mismatched Microphone Gains Using Rapid
- Reducing Spelling Inconsistencies in Code-Switching ASR Using Contextualized CTC Loss
- Refinement of Direction of Arrival Estimators by Majorization-Minimization Optimization on the Array Manifold
- Refining Automatic Speech Recognition System for Older Adults
- Reflectance-Oriented Probabilistic Equalization for Image Enhancement
- Regression or classification? New methods to evaluate no-reference picture and video quality models
- Regularized Recovery by Multi-Order Partial Hypergraph Total Variation
- Reinforcement Stacked Learning with Semantic-Associated Attention for Visual Question Answering
- Relaxed Wasserstein with Applications to GANs
- Reliability Assessment of Singing Voice F0-Estimates Using Multiple Algorithms
- Relying on a Rate Constraint to Reduce Motion Estimation Complexity
- Replacing Human Audio with Synthetic Audio for on-Device Unspoken Punctuation Prediction
- Replay and Synthetic Speech Detection with Res2Net Architecture
- Replay-Attack Detection Using Features With Adaptive Spectro-Temporal Resolution
- Representation Learning for Speech Recognition Using Feedback Based Relevance Weighting
- Representation Learning with Spectro-Temporal-Channel Attention for Speech Emotion Recognition
- Representative Local Feature Mining for Few-Shot Learning
- Resolution Limits of 20 Questions Search Strategies for Moving Targets
- Respipe: Resilient Model-Distributed DNN Training at Edge Networks
- Rethinking The Separation Layers In Speech Separation Networks
- Reverb Conversion Of Mixed Vocal Tracks Using An End-To-End Convolutional Deep Neural Network
- Reversible Data Hiding in Jpeg Images for Privacy Protection
- Reweighted Dynamic Group Convolution
- Riemannian Geometric Optimization Methods for Joint Design of Transmit Sequence and Receive Filter of MIMO Radar
- Riemannian Geometry on Connectivity for Clinical BCI
- Riemannian Geometry-Based Decoding of the Directional Focus of Auditory Attention Using EEG
- Robust Binary Loss for Multi-Category Classification with Label Noise
- Robust Deep Reinforcement Learning for Underwater Navigation with Unknown Disturbances
- Robust Device-Free Proximity Detection Using Wifi
- Robust Domain-Free Domain Generalization with Class-Aware Alignment
- Robust Graph Autoencoder for Hyperspectral Anomaly Detection
- Robust Graph-Filter Identification with Graph Denoising Regularization
- Robust Latent Representations Via Cross-Modal Translation and Alignment
- Robust Maml: Prioritization Task Buffer with Adaptive Learning Process for Model-Agnostic Meta-Learning
- Robust PCA Through Maximum Correntropy Power Iterations
- Robust Recursive Least M-Estimate Adaptive Filter for the Identification of Low-Rank Acoustic Systems
- Robust STFT Domain Multi-Channel Acoustic Echo Cancellation with Adaptive Decorrelation of the Reference Signals
- Robust Spatial-Temporal Correlation Model for Background Initialization in Severe Scene
- Robust Steerable Differential Beamformers with Null Constraints for Concentric Circular Microphone Arrays
- Robust Voice Activity Detection Using a Masked Auditory Encoder Based Convolutional Neural Network
- Robust estimation of high-order phase dynamics using Variational Bayes inference
- Robustness and Diversity Seeking Data-Free Knowledge Distillation
- Role Aware Multi-Party Dialogue Question Answering
- Room Adaptive Conditioning Method for Sound Event Classification in Reverberant Environments
- Room Impulse Response Interpolation from a Sparse Set of Measurements Using a Modal Architecture
- Rotation Invariance Analysis of Local Convolutional Features in Image Retrieval
- Rotation-Robust Beamforming Based on Sound Field Interpolation with Regularly Circular Microphone Array
- Routinggan: Routing Age Progression and Regression with Disentangled Learning
- Rule-Embedded Network for Audio-Visual Voice Activity Detection in Live Musical Video Streams
- SA-Net: Shuffle Attention for Deep Convolutional Neural Networks
- SANet++: Enhanced Scale Aggregation with Densely Connected Feature Fusion for Crowd Counting
- SEP-28k: A Dataset for Stuttering Event Detection from Podcasts with People Who Stutter
- SEQ-CPC : Sequential Contrastive Predictive Coding for Automatic Speech Recognition
- SERN: Stance Extraction and Reasoning Network for Fake News Detection
- SESQA: Semi-Supervised Learning for Speech Quality Assessment
- SIML: Sieved Maximum Likelihood for Array Signal Processing
- SLAP: a Split Latency Adaptive VLIW Pipeline Architecture Which Enables on-The-Fly Variable SIMD Vector-Length
- SM+: Refined Scale Match for Tiny Person Detection
- SNR-Adaptive Deep Joint Source-Channel Coding for Wireless Image Transmission
- SQWA: Stochastic Quantized Weight Averaging For Improving The Generalization Capability Of Low-Precision Deep Neural Networks
- SRF-Net: Selective Receptive Field Network for Anchor-Free Temporal Action Detection
- SSFENet: Spatial and Semantic Feature Enhancement Network for Object Detection
- SSLIDE: Sound Source Localization for Indoors Based on Deep Learning
- STEP-GAN: A One-Class Anomaly Detection Model with Applications to Power System Security
- Safe Screening for Sparse Regression with the Kullback-Leibler Divergence
- Saga: Sparse Adversarial Attack on EEG-Based Brain Computer Interface
- Saliency-Driven Versatile Video Coding for Neural Object Detection
- Sample Efficient Subspace-Based Representations for Nonlinear Meta-Learning
- Sandglasset: A Light Multi-Granularity Self-Attentive Network for Time-Domain Speech Separation
- SapAugment: Learning A Sample Adaptive Policy for Data Augmentation
- Sar Image Autofocusing Using Wirtinger Calculus and Cauchy Regularization
- Scalable Discriminative Discrete Hashing For Large-Scale Cross-Modal Retrieval
- Scalable Multilevel Quantization for Distributed Detection
- Scalable Privacy-Preserving Distributed Extremely Randomized Trees for Structured Data With Multiple Colluding Parties
- Scalable Reinforcement Learning For Routing In Ad-Hoc Networks Based On Physical-Layer Attributes
- Scalable and Distributed MMSE Algorithms for Uplink Receive Combining in Cell-Free Massive MIMO Systems
- Scaled Fast Nested Key Equation Solver for Generalized Integrated Interleaved BCH Decoders
- Scene Completeness-Aware Lidar Depth Completion for Driving Scenario
- Score-Based Change Detection For Gradient-Based Learning Machines
- Searching for Anomalies with Multiple Plays under Delay and Switching Costs
- Secret Key Generation Over Wireless Channels using short Blocklength Multilevel Source Polar Coding
- Secure UAV Communications Under Uncertain Eavesdroppers Locations
- SeeHear: Signer Diarisation and a New Dataset
- Seen and Unseen Emotional Style Transfer for Voice Conversion with A New Emotional Speech Dataset
- Segmental Dtw: A Parallelizable Alternative to Dynamic Time Warping
- Segregation in Social Networks: MARKOV Bridge Models and Estimation
- Seizure Detection Using Power Spectral Density via Hyperdimensional Computing
- Selection Based on Statistical Characteristics for Object Detection
- Self-Attention Generative Adversarial Network for Speech Enhancement
- Self-Attentive VAD: Context-Aware Detection of Voice from Noise
- Self-Augmented Multi-Modal Feature Embedding
- Self-Convolution: A Highly-Efficient Operator for Non-Local Image Restoration
- Self-Inference Of Others' Policies For Homogeneous Agents In Cooperative Multi-Agent Reinforcement Learning
- Self-Supervised Depth Estimation Via Implicit Cues from Videos
- Self-Supervised Learning Based Domain Adaptation for Robust Speaker Verification
- Self-Supervised Learning for Few-Shot Image Classification
- Self-Supervised Learning for Sleep Stage Classification with Predictive and Discriminative Contrastive Coding
- Self-Supervised Text-Independent Speaker Verification Using Prototypical Momentum Contrastive Learning
- Self-Supervised VQ-VAE for One-Shot Music Style Transfer
- Self-Training and Pre-Training are Complementary for Speech Recognition
- Self-Training for Sound Event Detection in Audio Mixtures
- Selfgait: A Spatiotemporal Representation Learning Method for Self-Supervised Gait Recognition
- Semantic Image Synthesis from Inaccurate and Coarse Masks
- Semantic-Aware Context Aggregation for Image Inpainting
- Semantic-Aware Unpaired Image-to-Image Translation for Urban Scene Images
- Semi-Supervised Batch Active Learning Via Bilevel Optimization
- Semi-Supervised Feature Embedding for Data Sanitization in Real-World Events
- Semi-Supervised Learning for Singing Synthesis Timbre
- Semi-Supervised Multimodal Image Translation for Missing Modality Imputation
- Semi-Supervised Singing Voice Separation With Noisy Self-Training
- Semi-Supervised Skin Lesion Segmentation with Learning Model Confidence
- Semi-Supervised Speech Recognition Via Graph-Based Temporal Classification
- Semi-Supervised Spoken Language Understanding via Self-Supervised Speech and Language Model Pretraining
- Semi-Supervised Time Series Classification by Temporal Relation Prediction
- Senone-Aware Adversarial Multi-Task Training for Unsupervised Child to Adult Speech Adaptation
- Sensor Networks TDOA Self-Calibration: 2D Complexity Analysis and Solutions
- Sentence Boundary Augmentation for Neural Machine Translation Robustness
- Sentiment Injected Iteratively Co-Interactive Network for Spoken Language Understanding
- SepNet: A Deep Separation Matrix Prediction Network for Multichannel Audio Source Separation
- Sequence-Level Self-Teaching Regularization
- Sequence-To-Sequence Singing Voice Synthesis With Perceptual Entropy Loss
- Sequential Adversarial Anomaly Detection with Deep Fourier Kernel
- Shapelet Based Visual Assessment of Cluster Tendency in Analyzing Complex Upper Limb Motion
- Short-Time Spectral Aggregation for Speaker Embedding
- Show and Speak: Directly Synthesize Spoken Description of Images
- Siamese Capsule Network for End-to-End Speaker Recognition in the Wild
- Sig2Sig: Signal Translation Networks to Take the Remains of the Past
- Sign Language Segmentation with Temporal Convolutional Networks
- Signature Feature Marking Enhanced IRM Framework for Drone Image Analysis in Precision Agriculture
- Similarity Analysis of Self-Supervised Speech Representations
- Simpleflat: A Simple Whole-Network Pre-Training Approach for RNN Transducer-Based End-to-End Speech Recognition
- Singer Identification Using Deep Timbre Feature Learning with KNN-NET
- Singing Language Identification Using a Deep Phonotactic Approach
- Singing Melody Extraction from Polyphonic Music based on Spectral Correlation Modeling
- Single Channel Voice Separation for Unknown Number of Speakers Under Reverberant and Noisy Settings
- Single-Point Array Response Control with Minimum Pattern Deviation
- Skip Attention GAN for Remote Sensing Image Synthesis
- Sliding-Capon Based Convolutional Beamspace for Linear Arrays
- Slow-Fast Auditory Streams for Audio Recognition
- Small Footprint Text-Independent Speaker Verification For Embedded Systems
- Social Learning Under Inferential Attacks
- Solving a Class of Non-Convex Min-Max Games Using Adaptive Momentum Methods
- Sound Event Detection Based on Curriculum Learning Considering Learning Difficulty of Events
- Sound Event Detection and Separation: A Benchmark on Desed Synthetic Soundscapes
- Sound Event Detection by Consistency Training and Pseudo-Labeling With Feature-Pyramid Convolutional Recurrent Neural Networks
- Sound Event Detection in Urban Audio with Single and Multi-Rate Pcen
- Sound Recovery From Radio Signals
- Source-Aware Neural Speech Coding for Noisy Speech Compression
- Sparse Array Transceiver Design for Enhanced Adaptive Beamforming in MIMO Radar
- Sparse Bayesian Learning for Acoustic Source Localization
- Sparse Factorization-Based Detection of Off-the-Grid Moving Targets Using FMCW Radars
- Sparse Flow Adversarial Model For Robust Image Compression
- Sparse Graph Based Sketching for Fast Numerical Linear Algebra
- Sparse High-Order Portfolios Via Proximal Dca And Sca
- Sparse Parameter Estimation for PMCW MIMO Radar Using Few-Bit ADCs
- Sparse Recovery Beamforming and Upscaling in the Ray Space
- Sparse Representation of Complex-Valued fMRI Data Based on Hard Thresholding of Spatial Source Phase
- Sparse Time-Frequency Representation Via Atomic Norm Minimization
- Sparse-Coded Dynamic Mode Decomposition on Graph for Prediction of River Water Level Distribution
- Sparsification via Compressed Sensing for Automatic Speech Recognition
- Sparsity And Nonnegativity Constrained Krylov Approach For Direction Of Arrival Estimation
- Sparsity Driven Latent Space Sampling for Generative Prior Based Compressive Sensing
- Sparsity in Max-Plus Algebra and Applications in Multivariate Convex Regression
- Spatial Equalization Before Reception: Reconfigurable Intelligent Surfaces for Multi-Path Mitigation
- Spatiotemporal Attention for Multivariate Time Series Prediction and Interpretation
- Speaker Activity Driven Neural Speech Extraction
- Speaker Embeddings for Diarization of Broadcast Data In The Allies Challenge
- Speaker and Direction Inferred Dual-Channel Speech Separation
- Speaker-Independent Brain Enhanced Speech Denoising
- Speaking Rate and Tonal Realization in Mandarin Chinese: What Can We Learn From Large Speech Corpora?
- Specialized Embedding Approximation for Edge Intelligence: A Case Study in Urban Sound Classification
- Spectral Domain Convolutional Neural Network
- Spectral Folding And Two-Channel Filter-Banks On Arbitrary Graphs
- Speech Acoustic Modelling from Raw Phase Spectrum
- Speech Bert Embedding for Improving Prosody in Neural TTS
- Speech Dereverberation Using Variational Autoencoders
- Speech Emotion Recognition Based on Listener Adaptive Models
- Speech Emotion Recognition Using Quaternion Convolutional Neural Networks
- Speech Emotion Recognition Using Semantic Information
- Speech Emotion Recognition with Multiscale Area Attention and Data Augmentation
- Speech Enhancement Aided End-To-End Multi-Task Learning for Voice Activity Detection
- Speech Enhancement Autoencoder with Hierarchical Latent Structure
- Speech Enhancement with Mixture of Deep Experts with Clean Clustering Pre-Training
- Speech Prediction in Silent Videos Using Variational Autoencoders
- Speech Recognition by Simply Fine-Tuning Bert
- Speech-Based Depression Prediction Using Encoder-Weight-Only Transfer Learning and a Large Corpus
- Speech-Language Pre-Training for End-to-End Spoken Language Understanding
- Speeding Up of Kernel-Based Learning for High-Order Tensors
- Spherical Harmonic Representation for Dynamic Sound-Field Measurements
- Spoken Language Identification in Unseen Target Domain Using Within-Sample Similarity Loss
- Squeezing Value of Cross-Domain Labels: A Decoupled Scoring Approach for Speaker Verification
- St-Bert: Cross-Modal Language Model Pre-Training for End-to-End Spoken Language Understanding
- Stability Analysis of the RC-PLMS Adaptive Beamformer Using a Simple Transfer Function Approximation
- Stability of Algebraic Neural Networks to Small Perturbations
- Stable Checkpoint Selection and Evaluation in Sequence to Sequence Speech Synthesis
- Stable and Effective One-Step Method for Person Search
- Statistical Correction of Transcribed Melody Notes Based on Probabilistic Integration of a Music Language Model and a Transcription Error Model
- Statistical Distance Metric Learning for Image Set Retrieval
- Statistical Properties of a Modified Welch Method That Uses Sample Percentiles
- Stereo Rectification Based on Epipolar Constrained Neural Network
- Stochastic Deep Unfolding for Imaging Inverse Problems
- Stochastic Successive Weighted Sum-Rate Maximization for Multiuser MIMO Systems with Finite-Alphabet Inputs
- Stock Movement Prediction and Portfolio Management via Multimodal Learning with Transformer
- Streaming End-to-End Speech Recognition with Jointly Trained Neural Feature Enhancement
- Streaming Multi-Speaker ASR with RNN-T
- Streaming Simultaneous Speech Translation with Augmented Memory Transformer
- Strong Data Augmentation Sanitizes Poisoning and Backdoor Attacks Without an Accuracy Tradeoff
- Structure-Aware Audio-to-Score Alignment Using Progressively Dilated Convolutional Neural Networks
- Structure-Enhanced Attentive Learning For Spine Segmentation From Ultrasound Volume Projection Images
- Structured Support Exploration for Multilayer Sparse Matrix Factorization
- StyleMelGAN: An Efficient High-Fidelity Adversarial Vocoder with Temporal Adaptive Normalization
- Sub-Band Grouping Spectral Feature-Attention Block for Hyperspectral Image Classification
- Sub-NYQUIST Multichannel Blind Deconvolution
- Subject-Invariant Eeg Representation Learning For Emotion Recognition
- Subjective and Objective Evaluation of Deepfake Videos
- Subspace Oddity - Optimization on Product of Stiefel Manifolds for EEG Data
- Subspectral Normalization for Neural Audio Data Processing
- Super-Resolution Of Periodic Signals From Short Sequences Of Samples
- Super-Resolution and Infection Edge Detection Co-Guided Learning for Covid-19 Ct Segmentation
- Supervised Chorus Detection for Popular Music Using Convolutional Neural Network and Multi-Task Learning
- Supervised Direct-Path Relative Transfer Function Learning for Binaural Sound Source Localization
- Suremap: Predicting Uncertainty in Cnn-Based Image Reconstructions Using Stein's Unbiased Risk Estimate
- Surrogate Source Model Learning for Determined Source Separation
- Switched Hawkes Processes
- Switching Variational Auto-Encoders for Noise-Agnostic Audio-Visual Speech Enhancement
- Symmetric Sub-graph Spatio-Temporal Graph Convolution and its application in Complex Activity Recognition
- SynAug: Synthesis-Based Data Augmentation for Text-Dependent Speaker Verification
- Synchronous Multi-Bit Audio Watermarking Based on Phase Shifting
- Synergic Feature Attention for Image Restoration
- Syntactic Representation Learning For Neural Network Based TTS with Syntactic Parse Tree Traversal
- Synthesis of New Words for Improved Dysarthric Speech Recognition on an Expanded Vocabulary
- Synthetic Aperture Acoustic Imaging with Deep Generative Model Based Source Distribution Prior
- Synthetic Data For Dnn-Based Doa Estimation of Indoor Speech
- TCLA Array: A New Sparse Array Design with Less Mutual Coupling
- TSTNN: Two-Stage Transformer Based Neural Network for Speech Enhancement in the Time Domain
- TTS-by-TTS: TTS-Driven Data Augmentation for Fast and High-Quality Speech Synthesis
- Tabular Transformers for Modeling Multivariate Time Series
- Taking A Closer Look at Synthesis: Fine-Grained Attribute Analysis for Person Re-Identification
- Taming Voting Algorithms on Gpus for an Efficient Connected Component Analysis Algorithm
- Target Detection from Distributed Passive Sensors: Semi-Labeled Data Quantization
- Target Detection in Frequency Hopping MIMO Dual-Function Radar-Communication Systems
- Task Aware Multi-Task Learning for Speech to Text Tasks
- Task-Aware Neural Architecture Search
- Task-Related Self-Supervised Learning For Remote Sensing Image Change Detection
- Teacher-Assisted Mini-Batch Sampling for Blind Distillation Using Metric Learning
- Teacher-Student Learning With Multi-Granularity Constraint Towards Compact Facial Feature Representation
- Teacher-Student Learning for Low-Latency Online Speech Enhancement Using Wave-U-Net
- Temporal Exemplar Channels In High-Multipath Environments
- Temporal Link Prediction Via Reinforcement Learning
- Temporal Rain Decomposition with Spatial Structure Guidance for Video Deraining
- Tensor Decomposition Via Core Tensor Networks
- Tensor Reordering for CNN Compression
- Text-to-Audio Grounding: Building Correspondence Between Captions and Sound Events
- The Accented English Speech Recognition Challenge 2020: Open Datasets, Tracks, Baselines, Results and Methods
- The Benefit of Temporally-Strong Labels in Audio Event Classification
- The Far-Field Equatorial Array for Binaural Rendering
- The Huya Multi-Speaker and Multi-Style Speech Synthesis System for M2voc Challenge 2020
- The Idlab Voxsrc-20 Submission: Large Margin Fine-Tuning and Quality-Aware Score Calibration in DNN Based Speaker Verification
- The Multi-Speaker Multi-Style Voice Cloning Challenge 2021
- The Role of Task and Acoustic Similarity in Audio Transfer Learning: Insights from the Speech Emotion Recognition Case
- The Thinkit System for Icassp2021 M2voc Challenge
- The in-the-Wild Speech Medical Corpus
- The ins and outs of speaker recognition: lessons from VoxSRC 2020
- The use of Voice Source Features for Sung Speech Recognition
- Time-Domain Concentration and Approximation of Computable Bandlimited Signals
- Time-Domain Loss Modulation Based on Overlap Ratio for Monaural Conversational Speaker Separation
- Time-Domain Speaker Verification Using Temporal Convolutional Networks
- Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning Mechanism
- Time-Varying Graph Signal Inpainting Via Unrolling Networks
- Tiny Transducer: A Highly-Efficient Speech Recognition Model on Edge Devices
- Top-Down Attention in End-to-End Spoken Language Understanding
- Topic Sequence Embedding for User Identity Linkage from Heterogeneous Behavior Data
- Topic-Aware Dialogue Generation with Two-Hop Based Graph Attention
- Topological Volterra Filters
- Toward Skills Dialog Orchestration with Online Learning
- Towards Adversarial Robustness Via Compact Feature Representations
- Towards An ASR Approach Using Acoustic and Language Models for Speech Enhancement
- Towards Data Selection on TTS Data for Children's Speech Recognition
- Towards Efficient Age Estimation by Embedding Potential Gender Features
- Towards Efficient Models for Real-Time Deep Noise Suppression
- Towards Efficiently Diversifying Dialogue Generation Via Embedding Augmentation
- Towards Explaining Expressive Qualities in Piano Recordings: Transfer of Explanatory Features Via Acoustic Domain Adaptation
- Towards Immediate Backchannel Generation Using Attention-Based Early Prediction Model
- Towards Listening to 10 People Simultaneously: An Efficient Permutation Invariant Training of Audio Source Separation Using Sinkhorn's Algorithm
- Towards Low-Resource Stargan Voice Conversion Using Weight Adaptive Instance Normalization
- Towards Natural and Controllable Cross-Lingual Voice Conversion Based on Neural TTS Model and Phonetic Posteriorgram
- Towards Parkinson's Disease Prognosis Using Self-Supervised Learning and Anomaly Detection
- Towards Practical Lipreading with Distilled and Efficient Models
- Towards Practical Near-Maximum-Likelihood Decoding of Error-Correcting Codes: An Overview
- Towards Robust Speaker Verification with Target Speaker Enhancement
- Towards Robust Training of Multi-Sensor Data Fusion Network Against Adversarial Examples in Semantic Segmentation
- Towards The Development of Subject-Independent Inverse Metabolic Models
- Towards an Intrinsic Definition of Robustness for a Classifier
- Traffic Speed Forecasting Via Spatio-Temporal Attentive Graph Isomorphism Network
- Train Your Classifier First: Cascade Neural Networks Training from Upper Layers to Lower Layers
- Training Logical Neural Networks by Primal-Dual Methods for Neuro-Symbolic Reasoning
- Training Neural Networks with Domain Pattern-Aware Auxiliary Task for Sleep Staging
- Training Noisy Single-Channel Speech Separation with Noisy Oracle Sources: A Large Gap and a Small Step
- Training Real-Time Panoramic Object Detectors with Virtual Dataset
- Training Speech Recognition Models with Federated Learning: A Quality/Cost Framework
- Training a Bank of Wiener Models with a Novel Quadratic Mutual Information Cost Function
- TransMask: A Compact and Fast Speech Separation Model Based on Transformer
- Transcription Is All You Need: Learning To Separate Musical Mixtures With Score As Supervision
- Transfer Learning for Input Estimation of Vehicle Systems
- Transformer Based Unsupervised Pre-Training for Acoustic Representation Learning
- Transformer Language Models with LSTM-Based Cross-Utterance Information Representation
- Transformer in Action: A Comparative Study of Transformer-Based Acoustic Models for Large Scale Speech Recognition Applications
- Transformer-Based End-to-End Speech Recognition with Local Dense Synthesizer Attention
- Transformer-Transducers for Code-Switched Speech Recognition
- Transitive Transfer Sparse Coding for Distant Domain
- Transmittance Regularizer for Binary coded Aperture Design in a Computational Imaging end-to-end Approach
- Treatment Effect Estimation Using Invariant Risk Minimization
- Triple Sequence Generative Adversarial Nets for Unsupervised Image Captioning
- Tucker Decomposition for Extracting Shared and Individual Spatial Maps from Multi-Subject Resting-State fMRI Data
- Two-Stage Adaptive Pooling with RT-QPCR for Covid-19 Screening
- Two-Stage Framework for Seasonal Time Series Forecasting
- Two-Stage Graph-Constrained Group Testing: Theory and Application
- Two-Stage Textual Knowledge Distillation for End-to-End Spoken Language Understanding
- Typingwristband: A Human Slight Motion Sensing System Based on Vibration Detection
- U-Convolution Based Residual Echo Suppression with Multiple Encoders
- UTDN: An Unsupervised Two-Stream Dirichlet-Net for Hyperspectral Unmixing
- Ultra-Lightweight Speech Separation Via Group Communication
- Ultra-Low Bitrate Video Conferencing Using Deep Image Animation
- Ultrasound Elasticity Imaging Using Physics-Based Models and Learning-Based Plug-and-Play Priors
- Uncertainty-Based Biological Age Estimation of Brain MRI Scans
- Unfolding Neural Networks for Compressive Multichannel Blind Deconvolution
- Unidirectional Memory-Self-Attention Transducer for Online Speech Recognition
- Unified Clustering and Outlier Detection on Specialized Hardware
- Unified Gradient Reweighting for Model Biasing with Applications to Source Separation
- Unit Selection Synthesis Based Data Augmentation for Fixed Phrase Speaker Verification
- Universal Neural Vocoding with Parallel Wavenet
- Unrolling of Deep Graph Total Variation for Image Denoising
- Unsupervised Audio-Visual Subspace Alignment for High-Stakes Deception Detection
- Unsupervised Clustering of Time Series Signals Using Neuromorphic Energy-Efficient Temporal Neural Networks
- Unsupervised Common Particular Object Discovery and Localization by Analyzing a Match Graph
- Unsupervised Contrastive Learning of Sound Event Representations
- Unsupervised Discriminative Learning of Sounds for Audio Event Classification
- Unsupervised Domain Adaptation for Speech Recognition via Uncertainty Driven Self-Training
- Unsupervised Heart Abnormality Detection Based on Phonocardiogram Analysis with Beta Variational Auto-Encoders
- Unsupervised Learning for Asynchronous Resource Allocation In Ad-Hoc Wireless Networks
- Unsupervised Learning for Multi-Style Speech Synthesis with Limited Data
- Unsupervised Motion Representation Enhanced Network for Action Recognition
- Unsupervised Multimodal Image Registration with Adaptative Gradient Guidance
- Unsupervised Musical Timbre Transfer for Notification Sounds
- Unsupervised Neural Adaptation Model Based on Optimal Transport for Spoken Language Identification
- Unsupervised Reconstruction of Sea Surface Currents from AIS Maritime Traffic Data Using Learnable Variational Models
- Unsupervised Stacked Capsule Autoencoder for Hyperspectral Image Classification
- Unsupervised and Semi-Supervised Few-Shot Acoustic Event Classification
- Unveiling Anomalous Nodes Via Random Sampling and Consensus on Graphs
- Upsampling Artifacts in Neural Audio Synthesis
- UserReg: A Simple but Strong Model for Rating Prediction
- Using Deep Image Priors to Generate Counterfactual Explanations
- Using Synthetic Audio to Improve the Recognition of Out-of-Vocabulary Words in End-to-End Asr Systems
- VGAI: End-to-End Learning of Vision-Based Decentralized Controllers for Robot Swarms
- VK-Net: Category-Level Point Cloud Registration with Unsupervised Rotation Invariant Keypoints
- Validating the Inspired Sinewave Technique to Measure Lung Heterogeneity Compared to Atelectasis & Over-Distended Volume in Computed Tomography Images
- Variance-Constrained Learning for Stochastic Graph Neural Networks
- Variation-Stable Fusion for PPG-Based Biometric System
- Variational Autoencoder for Speech Enhancement with a Noise-Aware Encoder
- Variational Autoencoders for Hyperspectral Unmixing with Endmember Variability
- Variational Dialogue Generation with Normalizing Flows
- Variational Parameter Learning in Sequential State-Space Model Via Particle Filtering
- Vehicle 3d Localization in Road Scenes VIA a Monocular Moving Camera
- Video Quality Prediction Using Voxel-Wise fMRI Models of the Visual Cortex
- Violence Detection in Videos Based on Fusing Visual and Audio Information
- Visual Privacy Protection via Mapping Distortion
- Visualizing Association in Exemplar-Based Classification
- Voting-Based Ensemble Model for Network Anomaly Detection
- Vowel Non-Vowel Based Spectral Warping and Time Scale Modification for Improvement in Children's ASR
- Vset: A Multimodal Transformer for Visual Speech Enhancement
- Wake Word Detection with Streaming Transformers
- Warp-Q: Quality Prediction for Generative Neural Speech Codecs
- Wase: Learning When to Attend for Speaker Extraction in Cocktail Party Environments
- Wasserstein Barycenter Transport for Acoustic Adaptation
- Wave-Tacotron: Spectrogram-Free End-to-End Text-to-Speech Synthesis
- Waveform Design for the Joint MIMO Radar and Communications with Low Integrated Sidelobe Levels and Accurate Information Embedding
- Weakly Supervised Patch Label Inference Network with Image Pyramid for Pavement Diseases Recognition in the Wild
- Wearing A Mask: Compressed Representations of Variable-Length Sequences Using Recurrent Neural Tangent Kernels
- Webly Supervised Deep Attentive Quantization
- Weight Identification Through Global Optimization in a New Hysteretic Neural Network Model
- Weighted Magnitude-Phase Loss for Speech Dereverberation
- Weighted Recursive Least Square Filter and Neural Network Based Residual ECHO Suppression for the AEC-Challenge
- What And Where To Focus In Person Search
- What's all the Fuss about Free Universal Sound Separation Data?
- When Face Recognition Meets Occlusion: A New Benchmark
- Wide and Deep Graph Neural Networks with Distributed Online Learning
- Wiener Filter on Meet/Join Lattices
- Wifi-Based Device-Free Gesture Recognition Through-the-Wall
- Window Beamformer for Sparse Concentric Circular Array
- Word-Level ASL Recognition and Trigger Sign Detection with RF Sensors
- Yapa: Accelerated Proximal Algorithm for Convex Composite Problems
- Zero-Gradient Constraints for Destriping of Remote-Sensing Data
- Zero-Shot Audio Classification with Factored Linear and Nonlinear Acoustic-Semantic Projections
- Zero-Shot Voice Conversion with Adjusted Speaker Embeddings and Simple Acoustic Features
- m-Activity: Accurate and Real-Time Human Activity Recognition Via Millimeter Wave Radar
- t-k-means: A ROBUST AND STABLE k-means VARIANT
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.