← All conferences

ICASSP 2021 Accepted Papers

The full list of 1,712 papers accepted at ICASSP 2021 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. MRI Image Recovery using Damped Denoising Vector AMP
  2. MSR-GAN: Multi-Segment Reconstruction via Adversarial Learning
  3. Machine Translation Verbosity Control for Automatic Dubbing
  4. Makf-Sr: Multi-Agent Adaptive Kalman Filtering-Based Successor Representations
  5. Making Punctuation Restoration Robust and Fast with Multi-Task Learning and Knowledge Distillation
  6. MarbleNet: Deep 1D Time-Channel Separable Convolutional Neural Network for Voice Activity Detection
  7. Mask4D: 4D Convolution Network for Light Field Occlusion Removal
  8. Maskcyclegan-VC: Learning Non-Parallel Voice Conversion with Filling in Frames
  9. Matching as Color Images: Thermal Image Local Feature Detection and Description
  10. Maximum a Posteriori Estimator for Convolutive Sound Source Separation with Sub-Source Based NTF Model and the Localization Probabilistic Prior on the Mixing Matrix
  11. Measure-Transformed Covariance Test for Robust Spectrum Sensing
  12. Measurement Coding Framework with Adjacent Pixels Based Measurement Matrix for Compressively Sensed Images
  13. Melody Harmonization Using Orderless Nade, Chord Balancing, and Blocked Gibbs Sampling
  14. Melon Playlist Dataset: A Public Dataset for Audio-Based Playlist Generation and Music Tagging
  15. Memory Layers with Multi-Head Attention Mechanisms for Text-Dependent Speaker Verification
  16. Memory-Efficient Speech Recognition on Smart Devices
  17. Message Transmission Over Rapidly Time-Varying Channels
  18. Meta Ordinal Weighting Net For Improving Lung Nodule Classification
  19. Meta-Adapter: Efficient Cross-Lingual Adaptation With Meta-Learning
  20. Meta-Cognition-Based Simple And Effective Approach To Object Detection
  21. Meta-Learning for 6G Communication Networks with Reconfigurable Intelligent Surfaces
  22. Meta-Learning for Cross-Channel Speaker Verification
  23. Meta-Learning for Improving Rare Word Recognition in End-to-End ASR
  24. Meta-Learning for Low-Resource Speech Emotion Recognition
  25. Meta-Learning with Attention for Improved Few-Shot Learning
  26. Micaugment: One-Shot Microphone Style Transfer
  27. Microsoft Speaker Diarization System for the Voxceleb Speaker Recognition Challenge 2020
  28. Millimeter Wave MIMO Channel Estimation with 1-bit Spatial Sigma-Delta Analog-to-Digital Converters
  29. Mind the Beat: Detecting Audio Onsets from EEG Recordings of Music Listening
  30. Minimizing Weighted Concave Impurity Partition Under Constraints
  31. Minimum Bayes Risk Training for End-to-End Speaker-Attributed ASR
  32. Misalignment Recognition in Acoustic Sensor Networks Using a Semi-Supervised Source Estimation Method and Markov Random Fields
  33. Mispronunciation Detection in Non-Native (L2) English with Uncertainty Modeling
  34. Mitigating Clipping Distortion in OFDM Using Deep Residual Learning
  35. Mitigating Inter-Subject Brain Signal Variability FOR EEG-Based Driver Fatigue State Classification
  36. MixSpeech: Data Augmentation for Low-Resource Automatic Speech Recognition
  37. Mixed Precision Quantization of Transformer Language Models for Speech Recognition
  38. Mixture of Informed Experts for Multilingual Speech Recognition
  39. Mixup Regularized Adversarial Networks for Multi-Domain Text Classification
  40. Model-Inspired Deep Learning for Light-Field Microscopy with Application to Neuron Localization
  41. Modeling Homophone Noise for Robust Neural Machine Translation
  42. Modelling Paralinguistic Properties in Conversational Speech to Detect Bipolar Disorder and Borderline Personality Disorder
  43. Modified Arcsine Law for One-Bit Sampled Stationary Signals with Time-Varying Thresholds
  44. Modular Binary Tree Architecture for Distributed Large Intelligent Surface
  45. Modurec: Recommender Systems with Feature and Time Modulation
  46. Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses
  47. More: A Metric Learning Based Framework for Open-Domain Relation Extraction
  48. Movement Detection Using A Reciprocal Received Signal Strength Model
  49. Moving Object Classification with a Sub-6 GHz Massive MIMO Array Using Real Data
  50. MuG: A Multipath-Exploited and Grid-Free Localisation Method
  51. Multi Path Training Framework for Data-Driven Open-Domain Conversation System
  52. Multi-Branch Tomlinson-Harashima Precoding for Rate Splitting Based Systems with Multiple Antennas
  53. Multi-Channel Speech Enhancement Using Graph Neural Networks
  54. Multi-Channel Target Speech Extraction with Channel Decorrelation and Target Speaker Adaptation
  55. Multi-Decoder Dprnn: Source Separation for Variable Number of Speakers
  56. Multi-Dialect Speech Recognition in English Using Attention on Ensemble of Experts
  57. Multi-Directional Convolution Networks with Spatial-Temporal Feature Pyramid Module for Action Recognition
  58. Multi-Entity Collaborative Relation Extraction
  59. Multi-Granularity Feature Interaction and Relation Reasoning for 3D Dense Alignment and Face Reconstruction
  60. Multi-Granularity Heterogeneous Graph for Document-Level Relation Extraction
  61. Multi-Initialization Meta-Learning with Domain Adaptation
  62. Multi-Level Adaptive Region of Interest and Graph Learning for Facial Action Unit Recognition
  63. Multi-Level Group Testing with Application to One-Shot Pooled COVID-19 Tests
  64. Multi-Level Reversible Encryption for ECG Signals Using Compressive Sensing
  65. Multi-Modal Label Dequantized Gaussian Process Latent Variable Model for Ordinal Label Estimation
  66. Multi-Models Fusion for Light Field Angular Super-Resolution
  67. Multi-Object Tracking Using Poisson Multi-Bernoulli Mixture Filtering For Autonomous Vehicles
  68. Multi-Order Adversarial Representation Learning for Composed Query Image Retrieval
  69. Multi-Rate Attention Architecture for Fast Streamable Text-to-Speech Spectrum Modeling
  70. Multi-Sample Online Learning for Spiking Neural Networks Based on Generalized Expectation Maximization
  71. Multi-Scale Cascade Disparity Refinement Stereo Network
  72. Multi-Scale Feature-Guided Stereoscopic Video Quality Assessment Based on 3d Convolutional Neural Network
  73. Multi-Scale Residual Network for Covid-19 Diagnosis Using Ct-Scans
  74. Multi-Scale Speaker Diarization with Neural Affinity Score Fusion
  75. Multi-Scale and Multi-Region Facial Discriminative Representation for Automatic Depression Level Prediction
  76. Multi-Speaker Emotional Speech Synthesis with Fine-Grained Prosody Modeling
  77. Multi-Stage Speaker Extraction with Utterance and Frame-Level Reference Signals
  78. Multi-Step Spoken Language Understanding System Based on Adversarial Learning
  79. Multi-Target DoA Estimation with an Audio-Visual Fusion Mechanism
  80. Multi-Task Estimation of Age and Cognitive Decline from Speech
  81. Multi-Task Learning Via Sharing Inexact Low-Rank Subspace
  82. Multi-Task Self-Supervised Pre-Training for Music Classification
  83. Multi-Task Transformer with Input Feature Reconstruction for Dysarthric Speech Recognition
  84. Multi-Tier Federated Learning for Vertically Partitioned Data
  85. Multi-Vehicle Velocity Estimation Using IEEE 802.11ad Waveform
  86. Multi-View Audio And Music Classification
  87. Multi-View Contrastive Learning for Online Knowledge Distillation
  88. Multichannel Overlapping Speaker Segmentation Using Multiple Hypothesis Tracking Of Acoustic And Spatial Features
  89. Multichannel-based Learning for Audio Object Extraction
  90. Multilabel 12-Lead Electrocardiogram Classification Using Beat to Sequence Autoencoders
  91. Multilingual Phonetic Dataset for Low Resource Speech Recognition
  92. Multimodal Cross- and Self-Attention Network for Speech Emotion Recognition
  93. Multimodal Emotion Recognition with Capsule Graph Convolutional Based Representation Fusion
  94. Multimodal Metric Learning for Tag-Based Music Retrieval
  95. Multimodal Punctuation Prediction with Contextual Dropout
  96. Multiphish: Multi-Modal Features Fusion Networks for Phishing Detection
  97. Multiple Auxiliary Networks for Single Blind Image Deblurring
  98. Multiple Human Tracking in Non-Specific Coverage with Wearable Cameras
  99. Multiple-Hypothesis CTC-Based Semi-Supervised Adaptation of End-to-End Speech Recognition
  100. Multiple-Input Multiple-Output Fusion Network for Generalized Zero-Shot Learning
  101. Multistream CNN for Robust Acoustic Modeling
  102. Multitask Learning and Joint Optimization for Transformer-RNN-Transducer Speech Recognition
  103. Multivariate Non-Negative Matrix Factorization with Application to Energy Disaggregation
  104. Multiview Sensing with Unknown Permutations: an Optimal Transport Approach
  105. Multiview Variational Graph Autoencoders for Canonical Correlation Analysis
  106. Muse: Multi-Modal Target Speaker Extraction with Visual Cues
  107. Mutual Information Flows in a Bivariate Point Process
  108. Mutually-Constrained Monotonic Multihead Attention for Online ASR
  109. NASA: A Noise-Adaptive and Structure-Aware Learning Framework for Image Deblurring
  110. NISP: A Multi-lingual Multi-accent Dataset for Speaker Profiling
  111. NMF-SAE: An Interpretable Sparse Autoencoder for Hyperspectral Unmixing
  112. NN-KOG2P: A Novel Grapheme-to-Phoneme Model for Korean Language
  113. NNAKF: A Neural Network Adapted Kalman Filter for Target Tracking
  114. Near-Optimal Algorithms for Piecewise-Stationary Cascading Bandits
  115. Near-Optimal Resampling in Particle Filters Using the Ising Energy Model
  116. Nested Error Map Generation Network for No-Reference Image Quality Assessment
  117. Nested Learning for Multi-Level Classification
  118. Network Classifiers Based on Social Learning
  119. Network Pruning Using Linear Dependency Analysis on Feature Maps
  120. Network Topology Change-Point Detection from Graph Signals with Prior Spectral Signatures
  121. Network Topology Inference with Graphon Spectral Penalties
  122. Network and Content-Dependent Bitrate Ladder Estimation for Adaptive Bitrate Video Streaming
  123. Network-Aware Optimal Microphone Channel Selection in Wireless Acoustic Sensor Networks
  124. Neural Architecture Search for LF-MMI Trained Time Delay Neural Networks
  125. Neural Audio Fingerprint for High-Specific Audio Retrieval Based on Contrastive Learning
  126. Neural Inverse Text Normalization
  127. Neural Kalman Filtering for Speech Enhancement
  128. Neural Layered Min-Sum Decoding for Protograph LDPC Codes
  129. Neural Network-Based Virtual Microphone Estimator
  130. Neural Noise Embedding for End-To-End Speech Enhancement with Conditional Layer Normalization
  131. Neural Utterance Confidence Measure for RNN-Transducers and Two Pass Models
  132. Neuro-Steered Music Source Separation With EEG-Based Auditory Attention Decoding And Contrastive-NMF
  133. New Variants of DFA Based on Loess and Lowess Methods: Generalization of the Detrending Moving Average
  134. Nlkd: Using Coarse Annotations For Semantic Segmentation Based on Knowledge Distillation
  135. No Relaxation: Guaranteed Recovery of Finite-Valued Signals from Undersampled Measurements
  136. No-Reference Stereoscopic Image Quality Assessment Based on the Human Visual System
  137. Node Attribute Completion in Knowledge Graphs with Multi-Relational Propagation
  138. Noise Level Limited Sub-Modeling for Diffusion Probabilistic Vocoders
  139. Noise-Assisted Multivariate Variational Mode Decomposition
  140. Noise-Robust Adaptation Control for Supervised Acoustic System Identification Exploiting a Noise Dictionary
  141. Non-Autoregressive Sequence-To-Sequence Voice Conversion
  142. Non-Autoregressive Transformer ASR with CTC-Enhanced Decoder Input
  143. Non-Coherent DOA Estimation of Off-Grid Signals With Uniform Circular Arrays
  144. Non-Intrusive Binaural Prediction of Speech Intelligibility Based on Phoneme Classification
  145. Non-Iterative Blind Calibration of Nested Arrays with Asymptotically Optimal Weighting
  146. Non-Local Single Image DE-Raining Without Decomposition
  147. Non-Parallel Many-To-Many Voice Conversion Using Local Linguistic Tokens
  148. Non-Parallel Many-To-Many Voice Conversion by Knowledge Transfer from a Text-To-Speech Model
  149. Non-Recursive Graph Convolutional Networks
  150. Non-Singular Adversarial Robustness of Neural Networks
  151. Noncontact Heartbeat Detection by Viterbi Algorithm with Fusion of Beat-Beat Interval and Deep Learning-Driven Branch Metrics
  152. Nonlinear State-Space Generalizations of Graph Convolutional Neural Networks
  153. Nonnegative Unimodal Matrix Factorization
  154. Nonstationary Portfolios: Diversification in the Spectral Domain
  155. Numerical Solution of Stochastic Differential Equations in Stiefel Manifolds via Tangent Space Parametrization
  156. OAS-Net: Occlusion Aware Sampling Network for Accurate Optical Flow
  157. ORTHROS: non-autoregressive end-to-end speech translation With dual-decoder
  158. Object-Oriented Relational Distillation for Object Detection
  159. On Distributed Composite Tests with Dependent Observations in WSN
  160. On Information Asymmetry in Online Reinforcement Learning
  161. On Loss Functions for Deep-Learning Based T60 Estimation
  162. On Overfitting in Discrete Super-Resolution Recovery
  163. On Permutation Invariant Training For Speech Source Separation
  164. On Scaling Contrastive Representations for Low-Resource Speech Recognition
  165. On Strategic Jamming in Distributed Detection Networks
  166. On The Accuracy Limit of Joint Time-Delay/Doppler/Acceleration Estimation with a Band-Limited Signal
  167. On The Adversarial Robustness of Principal Component Analysis
  168. On The Asymptotic Performance of One-Bit Co-Array-Based Music
  169. On The Camera Position Dithering In Visual 3d Reconstruction
  170. On The Effect of Spatial Correlation on Distributed Energy Detection of a Stochastic Process
  171. On The Power of Deep But Naive Partial Label Learning
  172. On The Relationship Between Speech-Based Breathing Signal Prediction Evaluation Measures and Breathing Parameters Estimation
  173. On The Role of Visual Cues in Audiovisual Speech Enhancement
  174. On The Stability of Graph Convolutional Neural Networks Under Edge Rewiring
  175. On a Guided Nonnegative Matrix Factorization
  176. On the Convergence of Randomized Bregman Coordinate Descent for Non-Lipschitz Composite Problems
  177. On the Design of Square Differential Microphone Arrays with a Multistage Structure
  178. On the Detection of Pitch-Shifted Voice: Machines and Human Listeners
  179. On the Marginal Benefit of Active Learning: Does Self-Supervision Eat its Cake?
  180. On the Optimality of Backward Regression: Sparse Recovery and Subset Selection
  181. On the Performance-Complexity Tradeoff in Stochastic Greedy Weak Submodular Optimization
  182. On the Predictability of Hrtfs from Ear Shapes Using Deep Networks
  183. On the Preparation and Validation of a Large-Scale Dataset of Singing Transcription
  184. One Shot Learning for Speech Separation
  185. One-Bit Autocorrelation Estimation With Non-Zero Thresholds
  186. One-Bit Compressed Sensing Using Untrained Network Prior
  187. One-Shot Conditional Audio Filtering of Arbitrary Sounds
  188. One-Shot Voice Conversion Based on Speaker Aware Module
  189. Online Antenna Selection for Enhanced DOA Estimation
  190. Online Classification of Dynamic Multilayer-Network Time Series in Riemannian Manifolds
  191. Online Dynamic Window (ODW) Assisted 2-Stage LSTM Indoor Localization for Smart Phones
  192. Online Hyper-Parameter Tuning for the Contextual Bandit
  193. Online Learning of Time-Varying Signals and Graphs
  194. Online Multi-Hop Information Based Kernel Learning Over Graphs
  195. Online Time-Varying Topology Identification Via Prediction-Correction Algorithms
  196. Online Unsupervised Learning Using Ensemble Gaussian Processes with Random Features
  197. Optimal Attacking Strategy Against Online Reputation Systems with Consideration of the Message-Based Persuasion Phenomenon
  198. Optimal Detection in the Presence of Non-Gaussian Jamming
  199. Optimal Importance Sampling for Federated Learning
  200. Optimal Questionnaires for Screening of Strategic Agents
  201. Optimal Selection of Matrix Shape and Decomposition Scheme for Neural Network Compression
  202. Optimal TOA Localization for Moving Sensor in Asymmetric Network
  203. Optimize What Matters: Training DNN-Hmm Keyword Spotting Model Using End Metric
  204. Optimizing Coverage and Capacity in Cellular Networks using Machine Learning
  205. Optimizing Short-Time Fourier Transform Parameters via Gradient Descent
  206. Optimum Feature Ordering for Dynamic Instance-Wise Joint Feature Selection and Classification
  207. Ordered Reliability Bits Guessing Random Additive Noise Decoding
  208. Orthogonality and Zero DC Tradeoffs in Biorthogonal Graph Filterbanks
  209. Outlier-Robust Kernel Hierarchical-Optimization RLS on a Budget with Affine Constraints
  210. Overcoming Measurement Inconsistency In Deep Learning For Linear Inverse Problems: Applications In Medical Imaging
  211. PD-GAN: Perceptual-Details GAN for Extremely Noisy Low Light Image Enhancement
  212. POLA: Online Time Series Prediction by Adaptive Learning Rates
  213. PPG-Based Singing Voice Conversion with Adversarial Representation Learning
  214. Paragraph Level Multi-Perspective Context Modeling for Question Generation
  215. Parallel Iterated Extended and Sigma-Point Kalman Smoothers
  216. Parallel Tacotron: Non-Autoregressive and Controllable TTS
  217. Parallel Waveform Synthesis Based on Generative Adversarial Networks with Voicing-Aware Conditional Discriminators
  218. Parameter Estimation for Coherent Passive MIMO Radar with Unknown Signals under Direct Path Influence
  219. Parameter Estimation for Student's t VAR Model with Missing Data
  220. Parameter Identifiability Of Spatial-Smoothing-Based Bistatic Mimo Radar
  221. Parametric Spectral Filters for Fast Converging, Scalable Convolutional Neural Networks
  222. Part-Aligned Network with Background for Misaligned Person Search
  223. Partial Feature Aggregation Network for Real-Time Object Counting
  224. Partially Overlapped Inference for Long-Form Speech Recognition
  225. Particle Gibbs Sampling for Regime-Switching State-Space Models
  226. Patch Decoder-Side Depth Estimation In Mpeg Immersive Video
  227. Patnet : A Phoneme-Level Autoregressive Transformer Network for Speech Synthesis
  228. Pause-Encoded Language Models for Recognition of Alzheimer's Disease and Emotion
  229. Perceptual Loss Based Speech Denoising with an Ensemble of Audio Pattern Recognition and Self-Supervised Models
  230. Perceptual Quality Assessment for Recognizing True and Pseudo 4k Content
  231. Performance Analysis of Spatial and Frequency Domain Index-Modulated Reconfigurable Intelligent Metasurfaces
  232. Periodic Signal Denoising: An Analysis-Synthesis Framework Based on Ramanujan Filter Banks and Dictionaries
  233. Periodnet: A Non-Autoregressive Waveform Generation Model with a Structure Separating Periodic and Aperiodic Components
  234. Personalization Strategies for End-to-End Speech Recognition Systems
  235. Personalized HRTF Modeling Using DNN-Augmented BEM
  236. Phase Recovery with Bregman Divergences for Audio Source Separation
  237. Phase Transitions for One-Vs-One and One-Vs-All Linear Separability in Multiclass Gaussian Mixtures
  238. Phone Distribution Estimation for Low Resource Languages
  239. Phoneme Based Neural Transducer for Large Vocabulary Speech Recognition
  240. Phoneme-Based Distribution Regularization for Speech Enhancement
  241. Physical-Layer Security via Distributed Beamforming in the Presence of Adversaries with Unknown Locations
  242. Pipeline Safety Early Warning Method for Distributed Signal using Bilinear CNN and LightGBM
  243. Pitch-Timbre Disentanglement Of Musical Instrument Sounds Based On Vae-Based Metric Learning
  244. Planar Array Geometry Optimization for Region Sound Acquisition
  245. Playing a Part: Speaker Verification at the movies
  246. Plug-And-Play Learned Gaussian-mixture Approximate Message Passing
  247. Point of Care Image Analysis for COVID-19
  248. Pointer Networks for Arbitrary-Shaped Text Spotting
  249. Policy Augmentation: An Exploration Strategy For Faster Convergence of Deep Reinforcement Learning Algorithms
  250. Polynomial Matrix Eigenvalue Decomposition of Spherical Harmonics for Speech Enhancement
  251. Portable Photoglottography for Monitoring Vocal Fold Vibrations in Speech Production
  252. Positnn: Training Deep Neural Networks with Mixed Low-Precision Posit
  253. Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Text Data
  254. Prediction of Egfr Mutation Status in Lung Adenocarcinoma Using Multi-Source Feature Representations
  255. Prediction of Object Geometry from Acoustic Scattering Using Convolutional Neural Networks
  256. Predictive Coding for Lossless Dataset Compression
  257. Preventing Early Endpointing for Online Automatic Speech Recognition
  258. Privacy-Accuracy Trade-Off of Inference as Service
  259. Privacy-Preserving Cloud-Based DNN Inference
  260. Privacy-Preserving Optimal Insulin Dosing Decision
  261. Privacy-Preserving near Neighbor Search via Sparse Coding with Ambiguation
  262. Private Wireless Federated Learning with Anonymous Over-the-Air Computation
  263. Probabilistic Graph Neural Networks for Traffic Signal Control
  264. Probability of Resolution of G-MUSIC: An Asymptotic Approach
  265. Probing Acoustic Representations for Phonetic Properties
  266. Processing Pipelines for Efficient, Physically-Accurate Simulation of Microphone Array Signals in Dynamic Sound Scenes
  267. Progressive Co-Teaching for Ambiguous Speech Emotion Recognition
  268. Progressive Multi-Stage Feature Mix for Person Re-Identification
  269. Progressive Spatio-Temporal Graph Convolutional Network for Skeleton-Based Human Action Recognition
  270. Progressive Voice Trigger Detection: Accuracy vs Latency
  271. Prosodic Clustering for Phoneme-Level Prosody Control in End-to-End Speech Synthesis
  272. Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech
  273. Prosody and Voice Factorization for Few-Shot Speaker Adaptation in the Challenge M2voc 2021
  274. Prototype-Based Personalized Pruning
  275. Prototypical Networks for Domain Adaptation in Acoustic Scene Classification
  276. Provably Fast Asynchronous And Distributed Algorithms For Pagerank Centrality Computation
  277. Pruning of Convolutional Neural Networks using ising Energy Model
  278. Pushing The Limit of Type I Codebook For Fdd Massive Mimo Beamforming: A Channel Covariance Reconstruction Approach
  279. Pushing the Limit of Phase Offset for Contactless Sensing Using Commodity Wifi
  280. Pyramid U-Net for Retinal Vessel Segmentation
  281. QUERYD: A Video Dataset with High-Quality Text and Audio Narrations
  282. QoE-Driven and Tile-Based Adaptive Streaming for Point Clouds
  283. Query-By-Example Keyword Spotting System Using Multi-Head Attention and Soft-triple Loss
  284. Quickest Change Detection With Time Inconsistent Anticipatory Agents In Cyber-Physical Systems
  285. Quickest Joint Detection and Classification of Faults in Statistically Periodic Processes
  286. REDAT: Accent-Invariant Representation for End-To-End ASR by Domain Adversarial Training with Relabeling
  287. REPAC: Reliable Estimation of Phase-Amplitude Coupling in Brain Networks
  288. REST: Robust lEarned Shrinkage-Thresholding Network Taming Inverse Problems with Model Mismatch
  289. RGLN: Robust Residual Graph Learning Networks via Similarity-Preserving Mapping on Graphs
  290. RIS-Aided Joint Localization and Synchronization with a Single-Antenna Mmwave Receiver
  291. RNN Transducer Models for Spoken Language Understanding
  292. RNN-T Based Open-Vocabulary Keyword Spotting in Mandarin with Multi-Level Detection
  293. Radar Clutter Classification Using Expectation-Maximization Method
  294. Radio Frequency Based Heart Rate Variability Monitoring
  295. Random Projection Streams for (Weighted) Nonnegative Matrix Factorization
  296. Range Guided Depth Refinement and Uncertainty-Aware Aggregation for View Synthesis
  297. Rank-Revealing Block-Term Decomposition for Tensor Completion
  298. Rate 1 Quasi Orthogonal Universal Transmission and Combining for MIMO Systems Achieving Full Diversity
  299. Rate-Distortion Optimized Motion Estimation for on-the-Sphere Compression of 360 Videos
  300. Raw Data Processing for Practical Time-of-Flight Super-Resolution
  301. Real Image Super-Resolution Using Token Based Contextual Attention
  302. Real Number Signal Processing can Detect Denial-of-Service Attacks
  303. Real Versus Fake 4k - Authentic Resolution Assessment
  304. Real-Time Denoising and Dereverberation wtih Tiny Recurrent U-Net
  305. Real-Time Interaural Time Delay Estimation via Onset Detection
  306. Real-Time Radio Modulation Classification With An LSTM Auto-Encoder
  307. Real-Time Speech Enhancement for Mobile Communication Based on Dual-Channel Complex Spectral Mapping
  308. Real-Time Speech Frequency Bandwidth Extension
  309. Real-Time Synchronization in Neural Networks for Multivariate Time Series Anomaly Detection
  310. Recent Advances in Arabic Syntactic Diacritics Restoration
  311. Recent Developments on Espnet Toolkit Boosted By Conformer
  312. Recognition of Dynamic Hand Gesture Based on Mm-Wave Fmcw Radar Micro-Doppler Signatures
  313. Recurrent Phase Reconstruction Using Estimated Phase Derivatives from Deep Neural Networks
  314. Recursive Input and State Estimation: a General Framework for Learning from Time Series With Missing Data
  315. Reduced-Complexity Channel Estimation by Hierarchical Interpolation Exploiting Sparsity for Massive MIMO Systems with Uniform Rectangular Array
  316. Reduced-Complexity Modular Polynomial Multiplication for R-LWE Cryptosystems
  317. Reducing Modal Error Propagation through Correcting Mismatched Microphone Gains Using Rapid
  318. Reducing Spelling Inconsistencies in Code-Switching ASR Using Contextualized CTC Loss
  319. Refinement of Direction of Arrival Estimators by Majorization-Minimization Optimization on the Array Manifold
  320. Refining Automatic Speech Recognition System for Older Adults
  321. Reflectance-Oriented Probabilistic Equalization for Image Enhancement
  322. Regression or classification? New methods to evaluate no-reference picture and video quality models
  323. Regularized Recovery by Multi-Order Partial Hypergraph Total Variation
  324. Reinforcement Stacked Learning with Semantic-Associated Attention for Visual Question Answering
  325. Relaxed Wasserstein with Applications to GANs
  326. Reliability Assessment of Singing Voice F0-Estimates Using Multiple Algorithms
  327. Relying on a Rate Constraint to Reduce Motion Estimation Complexity
  328. Replacing Human Audio with Synthetic Audio for on-Device Unspoken Punctuation Prediction
  329. Replay and Synthetic Speech Detection with Res2Net Architecture
  330. Replay-Attack Detection Using Features With Adaptive Spectro-Temporal Resolution
  331. Representation Learning for Speech Recognition Using Feedback Based Relevance Weighting
  332. Representation Learning with Spectro-Temporal-Channel Attention for Speech Emotion Recognition
  333. Representative Local Feature Mining for Few-Shot Learning
  334. Resolution Limits of 20 Questions Search Strategies for Moving Targets
  335. Respipe: Resilient Model-Distributed DNN Training at Edge Networks
  336. Rethinking The Separation Layers In Speech Separation Networks
  337. Reverb Conversion Of Mixed Vocal Tracks Using An End-To-End Convolutional Deep Neural Network
  338. Reversible Data Hiding in Jpeg Images for Privacy Protection
  339. Reweighted Dynamic Group Convolution
  340. Riemannian Geometric Optimization Methods for Joint Design of Transmit Sequence and Receive Filter of MIMO Radar
  341. Riemannian Geometry on Connectivity for Clinical BCI
  342. Riemannian Geometry-Based Decoding of the Directional Focus of Auditory Attention Using EEG
  343. Robust Binary Loss for Multi-Category Classification with Label Noise
  344. Robust Deep Reinforcement Learning for Underwater Navigation with Unknown Disturbances
  345. Robust Device-Free Proximity Detection Using Wifi
  346. Robust Domain-Free Domain Generalization with Class-Aware Alignment
  347. Robust Graph Autoencoder for Hyperspectral Anomaly Detection
  348. Robust Graph-Filter Identification with Graph Denoising Regularization
  349. Robust Latent Representations Via Cross-Modal Translation and Alignment
  350. Robust Maml: Prioritization Task Buffer with Adaptive Learning Process for Model-Agnostic Meta-Learning
  351. Robust PCA Through Maximum Correntropy Power Iterations
  352. Robust Recursive Least M-Estimate Adaptive Filter for the Identification of Low-Rank Acoustic Systems
  353. Robust STFT Domain Multi-Channel Acoustic Echo Cancellation with Adaptive Decorrelation of the Reference Signals
  354. Robust Spatial-Temporal Correlation Model for Background Initialization in Severe Scene
  355. Robust Steerable Differential Beamformers with Null Constraints for Concentric Circular Microphone Arrays
  356. Robust Voice Activity Detection Using a Masked Auditory Encoder Based Convolutional Neural Network
  357. Robust estimation of high-order phase dynamics using Variational Bayes inference
  358. Robustness and Diversity Seeking Data-Free Knowledge Distillation
  359. Role Aware Multi-Party Dialogue Question Answering
  360. Room Adaptive Conditioning Method for Sound Event Classification in Reverberant Environments
  361. Room Impulse Response Interpolation from a Sparse Set of Measurements Using a Modal Architecture
  362. Rotation Invariance Analysis of Local Convolutional Features in Image Retrieval
  363. Rotation-Robust Beamforming Based on Sound Field Interpolation with Regularly Circular Microphone Array
  364. Routinggan: Routing Age Progression and Regression with Disentangled Learning
  365. Rule-Embedded Network for Audio-Visual Voice Activity Detection in Live Musical Video Streams
  366. SA-Net: Shuffle Attention for Deep Convolutional Neural Networks
  367. SANet++: Enhanced Scale Aggregation with Densely Connected Feature Fusion for Crowd Counting
  368. SEP-28k: A Dataset for Stuttering Event Detection from Podcasts with People Who Stutter
  369. SEQ-CPC : Sequential Contrastive Predictive Coding for Automatic Speech Recognition
  370. SERN: Stance Extraction and Reasoning Network for Fake News Detection
  371. SESQA: Semi-Supervised Learning for Speech Quality Assessment
  372. SIML: Sieved Maximum Likelihood for Array Signal Processing
  373. SLAP: a Split Latency Adaptive VLIW Pipeline Architecture Which Enables on-The-Fly Variable SIMD Vector-Length
  374. SM+: Refined Scale Match for Tiny Person Detection
  375. SNR-Adaptive Deep Joint Source-Channel Coding for Wireless Image Transmission
  376. SQWA: Stochastic Quantized Weight Averaging For Improving The Generalization Capability Of Low-Precision Deep Neural Networks
  377. SRF-Net: Selective Receptive Field Network for Anchor-Free Temporal Action Detection
  378. SSFENet: Spatial and Semantic Feature Enhancement Network for Object Detection
  379. SSLIDE: Sound Source Localization for Indoors Based on Deep Learning
  380. STEP-GAN: A One-Class Anomaly Detection Model with Applications to Power System Security
  381. Safe Screening for Sparse Regression with the Kullback-Leibler Divergence
  382. Saga: Sparse Adversarial Attack on EEG-Based Brain Computer Interface
  383. Saliency-Driven Versatile Video Coding for Neural Object Detection
  384. Sample Efficient Subspace-Based Representations for Nonlinear Meta-Learning
  385. Sandglasset: A Light Multi-Granularity Self-Attentive Network for Time-Domain Speech Separation
  386. SapAugment: Learning A Sample Adaptive Policy for Data Augmentation
  387. Sar Image Autofocusing Using Wirtinger Calculus and Cauchy Regularization
  388. Scalable Discriminative Discrete Hashing For Large-Scale Cross-Modal Retrieval
  389. Scalable Multilevel Quantization for Distributed Detection
  390. Scalable Privacy-Preserving Distributed Extremely Randomized Trees for Structured Data With Multiple Colluding Parties
  391. Scalable Reinforcement Learning For Routing In Ad-Hoc Networks Based On Physical-Layer Attributes
  392. Scalable and Distributed MMSE Algorithms for Uplink Receive Combining in Cell-Free Massive MIMO Systems
  393. Scaled Fast Nested Key Equation Solver for Generalized Integrated Interleaved BCH Decoders
  394. Scene Completeness-Aware Lidar Depth Completion for Driving Scenario
  395. Score-Based Change Detection For Gradient-Based Learning Machines
  396. Searching for Anomalies with Multiple Plays under Delay and Switching Costs
  397. Secret Key Generation Over Wireless Channels using short Blocklength Multilevel Source Polar Coding
  398. Secure UAV Communications Under Uncertain Eavesdroppers Locations
  399. SeeHear: Signer Diarisation and a New Dataset
  400. Seen and Unseen Emotional Style Transfer for Voice Conversion with A New Emotional Speech Dataset
  401. Segmental Dtw: A Parallelizable Alternative to Dynamic Time Warping
  402. Segregation in Social Networks: MARKOV Bridge Models and Estimation
  403. Seizure Detection Using Power Spectral Density via Hyperdimensional Computing
  404. Selection Based on Statistical Characteristics for Object Detection
  405. Self-Attention Generative Adversarial Network for Speech Enhancement
  406. Self-Attentive VAD: Context-Aware Detection of Voice from Noise
  407. Self-Augmented Multi-Modal Feature Embedding
  408. Self-Convolution: A Highly-Efficient Operator for Non-Local Image Restoration
  409. Self-Inference Of Others' Policies For Homogeneous Agents In Cooperative Multi-Agent Reinforcement Learning
  410. Self-Supervised Depth Estimation Via Implicit Cues from Videos
  411. Self-Supervised Learning Based Domain Adaptation for Robust Speaker Verification
  412. Self-Supervised Learning for Few-Shot Image Classification
  413. Self-Supervised Learning for Sleep Stage Classification with Predictive and Discriminative Contrastive Coding
  414. Self-Supervised Text-Independent Speaker Verification Using Prototypical Momentum Contrastive Learning
  415. Self-Supervised VQ-VAE for One-Shot Music Style Transfer
  416. Self-Training and Pre-Training are Complementary for Speech Recognition
  417. Self-Training for Sound Event Detection in Audio Mixtures
  418. Selfgait: A Spatiotemporal Representation Learning Method for Self-Supervised Gait Recognition
  419. Semantic Image Synthesis from Inaccurate and Coarse Masks
  420. Semantic-Aware Context Aggregation for Image Inpainting
  421. Semantic-Aware Unpaired Image-to-Image Translation for Urban Scene Images
  422. Semi-Supervised Batch Active Learning Via Bilevel Optimization
  423. Semi-Supervised Feature Embedding for Data Sanitization in Real-World Events
  424. Semi-Supervised Learning for Singing Synthesis Timbre
  425. Semi-Supervised Multimodal Image Translation for Missing Modality Imputation
  426. Semi-Supervised Singing Voice Separation With Noisy Self-Training
  427. Semi-Supervised Skin Lesion Segmentation with Learning Model Confidence
  428. Semi-Supervised Speech Recognition Via Graph-Based Temporal Classification
  429. Semi-Supervised Spoken Language Understanding via Self-Supervised Speech and Language Model Pretraining
  430. Semi-Supervised Time Series Classification by Temporal Relation Prediction
  431. Senone-Aware Adversarial Multi-Task Training for Unsupervised Child to Adult Speech Adaptation
  432. Sensor Networks TDOA Self-Calibration: 2D Complexity Analysis and Solutions
  433. Sentence Boundary Augmentation for Neural Machine Translation Robustness
  434. Sentiment Injected Iteratively Co-Interactive Network for Spoken Language Understanding
  435. SepNet: A Deep Separation Matrix Prediction Network for Multichannel Audio Source Separation
  436. Sequence-Level Self-Teaching Regularization
  437. Sequence-To-Sequence Singing Voice Synthesis With Perceptual Entropy Loss
  438. Sequential Adversarial Anomaly Detection with Deep Fourier Kernel
  439. Shapelet Based Visual Assessment of Cluster Tendency in Analyzing Complex Upper Limb Motion
  440. Short-Time Spectral Aggregation for Speaker Embedding
  441. Show and Speak: Directly Synthesize Spoken Description of Images
  442. Siamese Capsule Network for End-to-End Speaker Recognition in the Wild
  443. Sig2Sig: Signal Translation Networks to Take the Remains of the Past
  444. Sign Language Segmentation with Temporal Convolutional Networks
  445. Signature Feature Marking Enhanced IRM Framework for Drone Image Analysis in Precision Agriculture
  446. Similarity Analysis of Self-Supervised Speech Representations
  447. Simpleflat: A Simple Whole-Network Pre-Training Approach for RNN Transducer-Based End-to-End Speech Recognition
  448. Singer Identification Using Deep Timbre Feature Learning with KNN-NET
  449. Singing Language Identification Using a Deep Phonotactic Approach
  450. Singing Melody Extraction from Polyphonic Music based on Spectral Correlation Modeling
  451. Single Channel Voice Separation for Unknown Number of Speakers Under Reverberant and Noisy Settings
  452. Single-Point Array Response Control with Minimum Pattern Deviation
  453. Skip Attention GAN for Remote Sensing Image Synthesis
  454. Sliding-Capon Based Convolutional Beamspace for Linear Arrays
  455. Slow-Fast Auditory Streams for Audio Recognition
  456. Small Footprint Text-Independent Speaker Verification For Embedded Systems
  457. Social Learning Under Inferential Attacks
  458. Solving a Class of Non-Convex Min-Max Games Using Adaptive Momentum Methods
  459. Sound Event Detection Based on Curriculum Learning Considering Learning Difficulty of Events
  460. Sound Event Detection and Separation: A Benchmark on Desed Synthetic Soundscapes
  461. Sound Event Detection by Consistency Training and Pseudo-Labeling With Feature-Pyramid Convolutional Recurrent Neural Networks
  462. Sound Event Detection in Urban Audio with Single and Multi-Rate Pcen
  463. Sound Recovery From Radio Signals
  464. Source-Aware Neural Speech Coding for Noisy Speech Compression
  465. Sparse Array Transceiver Design for Enhanced Adaptive Beamforming in MIMO Radar
  466. Sparse Bayesian Learning for Acoustic Source Localization
  467. Sparse Factorization-Based Detection of Off-the-Grid Moving Targets Using FMCW Radars
  468. Sparse Flow Adversarial Model For Robust Image Compression
  469. Sparse Graph Based Sketching for Fast Numerical Linear Algebra
  470. Sparse High-Order Portfolios Via Proximal Dca And Sca
  471. Sparse Parameter Estimation for PMCW MIMO Radar Using Few-Bit ADCs
  472. Sparse Recovery Beamforming and Upscaling in the Ray Space
  473. Sparse Representation of Complex-Valued fMRI Data Based on Hard Thresholding of Spatial Source Phase
  474. Sparse Time-Frequency Representation Via Atomic Norm Minimization
  475. Sparse-Coded Dynamic Mode Decomposition on Graph for Prediction of River Water Level Distribution
  476. Sparsification via Compressed Sensing for Automatic Speech Recognition
  477. Sparsity And Nonnegativity Constrained Krylov Approach For Direction Of Arrival Estimation
  478. Sparsity Driven Latent Space Sampling for Generative Prior Based Compressive Sensing
  479. Sparsity in Max-Plus Algebra and Applications in Multivariate Convex Regression
  480. Spatial Equalization Before Reception: Reconfigurable Intelligent Surfaces for Multi-Path Mitigation
  481. Spatiotemporal Attention for Multivariate Time Series Prediction and Interpretation
  482. Speaker Activity Driven Neural Speech Extraction
  483. Speaker Embeddings for Diarization of Broadcast Data In The Allies Challenge
  484. Speaker and Direction Inferred Dual-Channel Speech Separation
  485. Speaker-Independent Brain Enhanced Speech Denoising
  486. Speaking Rate and Tonal Realization in Mandarin Chinese: What Can We Learn From Large Speech Corpora?
  487. Specialized Embedding Approximation for Edge Intelligence: A Case Study in Urban Sound Classification
  488. Spectral Domain Convolutional Neural Network
  489. Spectral Folding And Two-Channel Filter-Banks On Arbitrary Graphs
  490. Speech Acoustic Modelling from Raw Phase Spectrum
  491. Speech Bert Embedding for Improving Prosody in Neural TTS
  492. Speech Dereverberation Using Variational Autoencoders
  493. Speech Emotion Recognition Based on Listener Adaptive Models
  494. Speech Emotion Recognition Using Quaternion Convolutional Neural Networks
  495. Speech Emotion Recognition Using Semantic Information
  496. Speech Emotion Recognition with Multiscale Area Attention and Data Augmentation
  497. Speech Enhancement Aided End-To-End Multi-Task Learning for Voice Activity Detection
  498. Speech Enhancement Autoencoder with Hierarchical Latent Structure
  499. Speech Enhancement with Mixture of Deep Experts with Clean Clustering Pre-Training
  500. Speech Prediction in Silent Videos Using Variational Autoencoders
  501. Speech Recognition by Simply Fine-Tuning Bert
  502. Speech-Based Depression Prediction Using Encoder-Weight-Only Transfer Learning and a Large Corpus
  503. Speech-Language Pre-Training for End-to-End Spoken Language Understanding
  504. Speeding Up of Kernel-Based Learning for High-Order Tensors
  505. Spherical Harmonic Representation for Dynamic Sound-Field Measurements
  506. Spoken Language Identification in Unseen Target Domain Using Within-Sample Similarity Loss
  507. Squeezing Value of Cross-Domain Labels: A Decoupled Scoring Approach for Speaker Verification
  508. St-Bert: Cross-Modal Language Model Pre-Training for End-to-End Spoken Language Understanding
  509. Stability Analysis of the RC-PLMS Adaptive Beamformer Using a Simple Transfer Function Approximation
  510. Stability of Algebraic Neural Networks to Small Perturbations
  511. Stable Checkpoint Selection and Evaluation in Sequence to Sequence Speech Synthesis
  512. Stable and Effective One-Step Method for Person Search
  513. Statistical Correction of Transcribed Melody Notes Based on Probabilistic Integration of a Music Language Model and a Transcription Error Model
  514. Statistical Distance Metric Learning for Image Set Retrieval
  515. Statistical Properties of a Modified Welch Method That Uses Sample Percentiles
  516. Stereo Rectification Based on Epipolar Constrained Neural Network
  517. Stochastic Deep Unfolding for Imaging Inverse Problems
  518. Stochastic Successive Weighted Sum-Rate Maximization for Multiuser MIMO Systems with Finite-Alphabet Inputs
  519. Stock Movement Prediction and Portfolio Management via Multimodal Learning with Transformer
  520. Streaming End-to-End Speech Recognition with Jointly Trained Neural Feature Enhancement
  521. Streaming Multi-Speaker ASR with RNN-T
  522. Streaming Simultaneous Speech Translation with Augmented Memory Transformer
  523. Strong Data Augmentation Sanitizes Poisoning and Backdoor Attacks Without an Accuracy Tradeoff
  524. Structure-Aware Audio-to-Score Alignment Using Progressively Dilated Convolutional Neural Networks
  525. Structure-Enhanced Attentive Learning For Spine Segmentation From Ultrasound Volume Projection Images
  526. Structured Support Exploration for Multilayer Sparse Matrix Factorization
  527. StyleMelGAN: An Efficient High-Fidelity Adversarial Vocoder with Temporal Adaptive Normalization
  528. Sub-Band Grouping Spectral Feature-Attention Block for Hyperspectral Image Classification
  529. Sub-NYQUIST Multichannel Blind Deconvolution
  530. Subject-Invariant Eeg Representation Learning For Emotion Recognition
  531. Subjective and Objective Evaluation of Deepfake Videos
  532. Subspace Oddity - Optimization on Product of Stiefel Manifolds for EEG Data
  533. Subspectral Normalization for Neural Audio Data Processing
  534. Super-Resolution Of Periodic Signals From Short Sequences Of Samples
  535. Super-Resolution and Infection Edge Detection Co-Guided Learning for Covid-19 Ct Segmentation
  536. Supervised Chorus Detection for Popular Music Using Convolutional Neural Network and Multi-Task Learning
  537. Supervised Direct-Path Relative Transfer Function Learning for Binaural Sound Source Localization
  538. Suremap: Predicting Uncertainty in Cnn-Based Image Reconstructions Using Stein's Unbiased Risk Estimate
  539. Surrogate Source Model Learning for Determined Source Separation
  540. Switched Hawkes Processes
  541. Switching Variational Auto-Encoders for Noise-Agnostic Audio-Visual Speech Enhancement
  542. Symmetric Sub-graph Spatio-Temporal Graph Convolution and its application in Complex Activity Recognition
  543. SynAug: Synthesis-Based Data Augmentation for Text-Dependent Speaker Verification
  544. Synchronous Multi-Bit Audio Watermarking Based on Phase Shifting
  545. Synergic Feature Attention for Image Restoration
  546. Syntactic Representation Learning For Neural Network Based TTS with Syntactic Parse Tree Traversal
  547. Synthesis of New Words for Improved Dysarthric Speech Recognition on an Expanded Vocabulary
  548. Synthetic Aperture Acoustic Imaging with Deep Generative Model Based Source Distribution Prior
  549. Synthetic Data For Dnn-Based Doa Estimation of Indoor Speech
  550. TCLA Array: A New Sparse Array Design with Less Mutual Coupling
  551. TSTNN: Two-Stage Transformer Based Neural Network for Speech Enhancement in the Time Domain
  552. TTS-by-TTS: TTS-Driven Data Augmentation for Fast and High-Quality Speech Synthesis
  553. Tabular Transformers for Modeling Multivariate Time Series
  554. Taking A Closer Look at Synthesis: Fine-Grained Attribute Analysis for Person Re-Identification
  555. Taming Voting Algorithms on Gpus for an Efficient Connected Component Analysis Algorithm
  556. Target Detection from Distributed Passive Sensors: Semi-Labeled Data Quantization
  557. Target Detection in Frequency Hopping MIMO Dual-Function Radar-Communication Systems
  558. Task Aware Multi-Task Learning for Speech to Text Tasks
  559. Task-Aware Neural Architecture Search
  560. Task-Related Self-Supervised Learning For Remote Sensing Image Change Detection
  561. Teacher-Assisted Mini-Batch Sampling for Blind Distillation Using Metric Learning
  562. Teacher-Student Learning With Multi-Granularity Constraint Towards Compact Facial Feature Representation
  563. Teacher-Student Learning for Low-Latency Online Speech Enhancement Using Wave-U-Net
  564. Temporal Exemplar Channels In High-Multipath Environments
  565. Temporal Link Prediction Via Reinforcement Learning
  566. Temporal Rain Decomposition with Spatial Structure Guidance for Video Deraining
  567. Tensor Decomposition Via Core Tensor Networks
  568. Tensor Reordering for CNN Compression
  569. Text-to-Audio Grounding: Building Correspondence Between Captions and Sound Events
  570. The Accented English Speech Recognition Challenge 2020: Open Datasets, Tracks, Baselines, Results and Methods
  571. The Benefit of Temporally-Strong Labels in Audio Event Classification
  572. The Far-Field Equatorial Array for Binaural Rendering
  573. The Huya Multi-Speaker and Multi-Style Speech Synthesis System for M2voc Challenge 2020
  574. The Idlab Voxsrc-20 Submission: Large Margin Fine-Tuning and Quality-Aware Score Calibration in DNN Based Speaker Verification
  575. The Multi-Speaker Multi-Style Voice Cloning Challenge 2021
  576. The Role of Task and Acoustic Similarity in Audio Transfer Learning: Insights from the Speech Emotion Recognition Case
  577. The Thinkit System for Icassp2021 M2voc Challenge
  578. The in-the-Wild Speech Medical Corpus
  579. The ins and outs of speaker recognition: lessons from VoxSRC 2020
  580. The use of Voice Source Features for Sung Speech Recognition
  581. Time-Domain Concentration and Approximation of Computable Bandlimited Signals
  582. Time-Domain Loss Modulation Based on Overlap Ratio for Monaural Conversational Speaker Separation
  583. Time-Domain Speaker Verification Using Temporal Convolutional Networks
  584. Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning Mechanism
  585. Time-Varying Graph Signal Inpainting Via Unrolling Networks
  586. Tiny Transducer: A Highly-Efficient Speech Recognition Model on Edge Devices
  587. Top-Down Attention in End-to-End Spoken Language Understanding
  588. Topic Sequence Embedding for User Identity Linkage from Heterogeneous Behavior Data
  589. Topic-Aware Dialogue Generation with Two-Hop Based Graph Attention
  590. Topological Volterra Filters
  591. Toward Skills Dialog Orchestration with Online Learning
  592. Towards Adversarial Robustness Via Compact Feature Representations
  593. Towards An ASR Approach Using Acoustic and Language Models for Speech Enhancement
  594. Towards Data Selection on TTS Data for Children's Speech Recognition
  595. Towards Efficient Age Estimation by Embedding Potential Gender Features
  596. Towards Efficient Models for Real-Time Deep Noise Suppression
  597. Towards Efficiently Diversifying Dialogue Generation Via Embedding Augmentation
  598. Towards Explaining Expressive Qualities in Piano Recordings: Transfer of Explanatory Features Via Acoustic Domain Adaptation
  599. Towards Immediate Backchannel Generation Using Attention-Based Early Prediction Model
  600. Towards Listening to 10 People Simultaneously: An Efficient Permutation Invariant Training of Audio Source Separation Using Sinkhorn's Algorithm
  601. Towards Low-Resource Stargan Voice Conversion Using Weight Adaptive Instance Normalization
  602. Towards Natural and Controllable Cross-Lingual Voice Conversion Based on Neural TTS Model and Phonetic Posteriorgram
  603. Towards Parkinson's Disease Prognosis Using Self-Supervised Learning and Anomaly Detection
  604. Towards Practical Lipreading with Distilled and Efficient Models
  605. Towards Practical Near-Maximum-Likelihood Decoding of Error-Correcting Codes: An Overview
  606. Towards Robust Speaker Verification with Target Speaker Enhancement
  607. Towards Robust Training of Multi-Sensor Data Fusion Network Against Adversarial Examples in Semantic Segmentation
  608. Towards The Development of Subject-Independent Inverse Metabolic Models
  609. Towards an Intrinsic Definition of Robustness for a Classifier
  610. Traffic Speed Forecasting Via Spatio-Temporal Attentive Graph Isomorphism Network
  611. Train Your Classifier First: Cascade Neural Networks Training from Upper Layers to Lower Layers
  612. Training Logical Neural Networks by Primal-Dual Methods for Neuro-Symbolic Reasoning
  613. Training Neural Networks with Domain Pattern-Aware Auxiliary Task for Sleep Staging
  614. Training Noisy Single-Channel Speech Separation with Noisy Oracle Sources: A Large Gap and a Small Step
  615. Training Real-Time Panoramic Object Detectors with Virtual Dataset
  616. Training Speech Recognition Models with Federated Learning: A Quality/Cost Framework
  617. Training a Bank of Wiener Models with a Novel Quadratic Mutual Information Cost Function
  618. TransMask: A Compact and Fast Speech Separation Model Based on Transformer
  619. Transcription Is All You Need: Learning To Separate Musical Mixtures With Score As Supervision
  620. Transfer Learning for Input Estimation of Vehicle Systems
  621. Transformer Based Unsupervised Pre-Training for Acoustic Representation Learning
  622. Transformer Language Models with LSTM-Based Cross-Utterance Information Representation
  623. Transformer in Action: A Comparative Study of Transformer-Based Acoustic Models for Large Scale Speech Recognition Applications
  624. Transformer-Based End-to-End Speech Recognition with Local Dense Synthesizer Attention
  625. Transformer-Transducers for Code-Switched Speech Recognition
  626. Transitive Transfer Sparse Coding for Distant Domain
  627. Transmittance Regularizer for Binary coded Aperture Design in a Computational Imaging end-to-end Approach
  628. Treatment Effect Estimation Using Invariant Risk Minimization
  629. Triple Sequence Generative Adversarial Nets for Unsupervised Image Captioning
  630. Tucker Decomposition for Extracting Shared and Individual Spatial Maps from Multi-Subject Resting-State fMRI Data
  631. Two-Stage Adaptive Pooling with RT-QPCR for Covid-19 Screening
  632. Two-Stage Framework for Seasonal Time Series Forecasting
  633. Two-Stage Graph-Constrained Group Testing: Theory and Application
  634. Two-Stage Textual Knowledge Distillation for End-to-End Spoken Language Understanding
  635. Typingwristband: A Human Slight Motion Sensing System Based on Vibration Detection
  636. U-Convolution Based Residual Echo Suppression with Multiple Encoders
  637. UTDN: An Unsupervised Two-Stream Dirichlet-Net for Hyperspectral Unmixing
  638. Ultra-Lightweight Speech Separation Via Group Communication
  639. Ultra-Low Bitrate Video Conferencing Using Deep Image Animation
  640. Ultrasound Elasticity Imaging Using Physics-Based Models and Learning-Based Plug-and-Play Priors
  641. Uncertainty-Based Biological Age Estimation of Brain MRI Scans
  642. Unfolding Neural Networks for Compressive Multichannel Blind Deconvolution
  643. Unidirectional Memory-Self-Attention Transducer for Online Speech Recognition
  644. Unified Clustering and Outlier Detection on Specialized Hardware
  645. Unified Gradient Reweighting for Model Biasing with Applications to Source Separation
  646. Unit Selection Synthesis Based Data Augmentation for Fixed Phrase Speaker Verification
  647. Universal Neural Vocoding with Parallel Wavenet
  648. Unrolling of Deep Graph Total Variation for Image Denoising
  649. Unsupervised Audio-Visual Subspace Alignment for High-Stakes Deception Detection
  650. Unsupervised Clustering of Time Series Signals Using Neuromorphic Energy-Efficient Temporal Neural Networks
  651. Unsupervised Common Particular Object Discovery and Localization by Analyzing a Match Graph
  652. Unsupervised Contrastive Learning of Sound Event Representations
  653. Unsupervised Discriminative Learning of Sounds for Audio Event Classification
  654. Unsupervised Domain Adaptation for Speech Recognition via Uncertainty Driven Self-Training
  655. Unsupervised Heart Abnormality Detection Based on Phonocardiogram Analysis with Beta Variational Auto-Encoders
  656. Unsupervised Learning for Asynchronous Resource Allocation In Ad-Hoc Wireless Networks
  657. Unsupervised Learning for Multi-Style Speech Synthesis with Limited Data
  658. Unsupervised Motion Representation Enhanced Network for Action Recognition
  659. Unsupervised Multimodal Image Registration with Adaptative Gradient Guidance
  660. Unsupervised Musical Timbre Transfer for Notification Sounds
  661. Unsupervised Neural Adaptation Model Based on Optimal Transport for Spoken Language Identification
  662. Unsupervised Reconstruction of Sea Surface Currents from AIS Maritime Traffic Data Using Learnable Variational Models
  663. Unsupervised Stacked Capsule Autoencoder for Hyperspectral Image Classification
  664. Unsupervised and Semi-Supervised Few-Shot Acoustic Event Classification
  665. Unveiling Anomalous Nodes Via Random Sampling and Consensus on Graphs
  666. Upsampling Artifacts in Neural Audio Synthesis
  667. UserReg: A Simple but Strong Model for Rating Prediction
  668. Using Deep Image Priors to Generate Counterfactual Explanations
  669. Using Synthetic Audio to Improve the Recognition of Out-of-Vocabulary Words in End-to-End Asr Systems
  670. VGAI: End-to-End Learning of Vision-Based Decentralized Controllers for Robot Swarms
  671. VK-Net: Category-Level Point Cloud Registration with Unsupervised Rotation Invariant Keypoints
  672. Validating the Inspired Sinewave Technique to Measure Lung Heterogeneity Compared to Atelectasis & Over-Distended Volume in Computed Tomography Images
  673. Variance-Constrained Learning for Stochastic Graph Neural Networks
  674. Variation-Stable Fusion for PPG-Based Biometric System
  675. Variational Autoencoder for Speech Enhancement with a Noise-Aware Encoder
  676. Variational Autoencoders for Hyperspectral Unmixing with Endmember Variability
  677. Variational Dialogue Generation with Normalizing Flows
  678. Variational Parameter Learning in Sequential State-Space Model Via Particle Filtering
  679. Vehicle 3d Localization in Road Scenes VIA a Monocular Moving Camera
  680. Video Quality Prediction Using Voxel-Wise fMRI Models of the Visual Cortex
  681. Violence Detection in Videos Based on Fusing Visual and Audio Information
  682. Visual Privacy Protection via Mapping Distortion
  683. Visualizing Association in Exemplar-Based Classification
  684. Voting-Based Ensemble Model for Network Anomaly Detection
  685. Vowel Non-Vowel Based Spectral Warping and Time Scale Modification for Improvement in Children's ASR
  686. Vset: A Multimodal Transformer for Visual Speech Enhancement
  687. Wake Word Detection with Streaming Transformers
  688. Warp-Q: Quality Prediction for Generative Neural Speech Codecs
  689. Wase: Learning When to Attend for Speaker Extraction in Cocktail Party Environments
  690. Wasserstein Barycenter Transport for Acoustic Adaptation
  691. Wave-Tacotron: Spectrogram-Free End-to-End Text-to-Speech Synthesis
  692. Waveform Design for the Joint MIMO Radar and Communications with Low Integrated Sidelobe Levels and Accurate Information Embedding
  693. Weakly Supervised Patch Label Inference Network with Image Pyramid for Pavement Diseases Recognition in the Wild
  694. Wearing A Mask: Compressed Representations of Variable-Length Sequences Using Recurrent Neural Tangent Kernels
  695. Webly Supervised Deep Attentive Quantization
  696. Weight Identification Through Global Optimization in a New Hysteretic Neural Network Model
  697. Weighted Magnitude-Phase Loss for Speech Dereverberation
  698. Weighted Recursive Least Square Filter and Neural Network Based Residual ECHO Suppression for the AEC-Challenge
  699. What And Where To Focus In Person Search
  700. What's all the Fuss about Free Universal Sound Separation Data?
  701. When Face Recognition Meets Occlusion: A New Benchmark
  702. Wide and Deep Graph Neural Networks with Distributed Online Learning
  703. Wiener Filter on Meet/Join Lattices
  704. Wifi-Based Device-Free Gesture Recognition Through-the-Wall
  705. Window Beamformer for Sparse Concentric Circular Array
  706. Word-Level ASL Recognition and Trigger Sign Detection with RF Sensors
  707. Yapa: Accelerated Proximal Algorithm for Convex Composite Problems
  708. Zero-Gradient Constraints for Destriping of Remote-Sensing Data
  709. Zero-Shot Audio Classification with Factored Linear and Nonlinear Acoustic-Semantic Projections
  710. Zero-Shot Voice Conversion with Adjusted Speaker Embeddings and Simple Acoustic Features
  711. m-Activity: Accurate and Real-Time Human Activity Recognition Via Millimeter Wave Radar
  712. t-k-means: A ROBUST AND STABLE k-means VARIANT

Looking for submission deadlines instead? See the conference deadline calendar.