← All conferences

ICASSP 2021 Accepted Papers

The full list of 1,712 papers accepted at ICASSP 2021 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. "You Should Probably Read This": Hedge Detection in Text
  2. (W)Earable Microphone Array and Ultrasonic Echo Localization for Coarse Indoor Environment Mapping
  3. 2D-FRFT Based Frequency Shift-Invariant Digital Image Encryption
  4. 3D Multizone Soundfield Reproduction in a Reverberant Environment Using Intensity Matching Method
  5. A Bayesian Inference Approach for Location-Based Micro Motions using Radio Frequency Sensing
  6. A Bayesian Interpretation of the Light Gated Recurrent Unit
  7. A Better and Faster end-to-end Model for Streaming ASR
  8. A Bias-Reducing Loss Function for CT Image Denoising
  9. A Capsule Network Based Approach for Detection of Audio Spoofing Attacks
  10. A Causal Deep Learning Framework for Classifying Phonemes in Cochlear Implants
  11. A Chapter-Wise Understanding System for Text-To-Speech in Chinese Novels
  12. A Classifier for Improving Cause and Effect in SSVEP-based BCIs for Individuals with Complex Communication Disorders
  13. A Closed-Loop Gain-Control Feedback Model for The Medial Efferent System of The Descending Auditory Pathway
  14. A Closer Look at Audio-Visual Multi-Person Speech Recognition and Active Speaker Selection
  15. A Co-Interactive Transformer for Joint Slot Filling and Intent Detection
  16. A Color Doppler Processing Engine with an Adaptive Clutter Filter for Portable Ultrasound Imaging Devices
  17. A Compact Joint Distillation Network for Visual Food Recognition
  18. A Comparative Study of Acoustic and Linguistic Features Classification for Alzheimer's Disease Detection
  19. A Comparison Study on Infant-Parent Voice Diarization
  20. A Comparison of Convolutional Neural Networks for Glottal Closure Instant Detection from Raw Speech
  21. A Comparison of Discrete Latent Variable Models for Speech Representation Learning
  22. A Comparison of Methods for OOV-Word Recognition on a New Public Dataset
  23. A Consensus Equilibrium Solution For Deep Image Prior Powered By Red
  24. A Convex Penalty for Block-Sparse Signals with Unknown Structures
  25. A Correntropy Based Algorithm for Robust Localization in Wireless Networks
  26. A Curated Dataset of Urban Scenes for Audio-Visual Scene Analysis
  27. A DNN Autoencoder for Automotive Radar Interference Mitigation
  28. A Decentralized Variance-Reduced Method for Stochastic Optimization Over Directed Graphs
  29. A Deep Reinforcement Learning Approach To Audio-Based Navigation In A Multi-Speaker Environment
  30. A Deep Spatio-Temporal Model for EEG-Based Imagined Speech Recognition
  31. A Diffusion FXLMS Algorithm for Multi-Channel Active Noise Control and Variable Spatial Smoothing
  32. A Dynamical Systems Perspective on Online Bayesian Nonparametric Estimators with Adaptive Hyperparameters
  33. A Fast Randomized Adaptive CP Decomposition For Streaming Tensors
  34. A Fast and Efficient Network for Single Image Deraining
  35. A Features Decoupling Method for Multiple Manipulations Identification in Image Operation Chains
  36. A Flow-Based Neural Network for Time Domain Speech Enhancement
  37. A Framework for Pruning Deep Neural Networks Using Energy-Based Models
  38. A Further Study of Unsupervised Pretraining for Transformer Based Speech Recognition
  39. A General Multi-Task Learning Framework to Leverage Text Data for Speech to Text Tasks
  40. A General Network Architecture for Sound Event Localization and Detection Using Transfer Learning and Recurrent Neural Network
  41. A Global Cayley Parametrization of Stiefel Manifold for Direct Utilization of Optimization Mechanisms Over Vector Spaces
  42. A Global-Local Attention Framework for Weakly Labelled Audio Tagging
  43. A Graph Learning Algorithm Based On Gaussian Markov Random Fields And Minimax Concave Penalty
  44. A Hierarchical Subspace Model for Language-Attuned Acoustic Unit Discovery
  45. A High-Frame-Rate Eye-Tracking Framework for Mobile Devices
  46. A Homogeneity-Based Multiscale Hyperspectral Image Representation for Sparse Spectral Unmixing
  47. A Hybrid Approach to Coded Compressed Sensing Where Coupling Takes Place Via the Outer Code
  48. A Hybrid CNN-BiLSTM Voice Activity Detector
  49. A Hybrid Feature Enhancement Method for Gl And Segmentation In Histopathology Images
  50. A Joint Convolutional and Spatial Quad-Directional LSTM Network for Phase Unwrapping
  51. A Joint Training Framework of Multi-Look Separator and Speaker Embedding Extractor for Overlapped Speech
  52. A Large-Dimensional Analysis of Symmetric SNE
  53. A Large-Scale Chinese Long-Text Extractive Summarization Corpus
  54. A Layered Embedding-Based Scheme to Cope with Intra-Frame Distortion Drift In IPM-Based HEVC Steganography
  55. A Low-Complexity Admm-Based Massive Mimo Detectors Via Deep Neural Networks
  56. A Low-Complexity MIMO Dual Function Radar Communication System via One-Bit Sampling
  57. A Meta-Learning Framework for Few-Shot Classification of Remote Sensing Scene
  58. A Method for Determining Periodically Time-Varying Bias and Its Applications in Acoustic Feedback Cancellation
  59. A Modulation-Domain Loss for Neural-Network-Based Real-Time Speech Enhancement
  60. A Multi-Channel Temporal Attention Convolutional Neural Network Model for Environmental Sound Classification
  61. A Multi-Layer Multi-Channel Attentive Network for Gender and Age Recognition
  62. A Multi-Stage Progressive Learning Strategy for Covid-19 Diagnosis Using Chest Computed Tomography with Imbalanced Data
  63. A Multi-View Approach to Audio-Visual Speaker Verification
  64. A Multiple Access Channel Game Using Latency Metric
  65. A Neural Acoustic Echo Canceller Optimized Using An Automatic Speech Recognizer and Large Scale Synthetic Data
  66. A Neural Text-to-Speech Model Utilizing Broadcast Data Mixed with Background Music
  67. A New Automotive Radar 4D Point Clouds Detector by Using Deep Learning
  68. A New DCASE 2017 Rare Sound Event Detection Benchmark Under Equal Training Data: CRNN With Multi-Width Kernels
  69. A New Framework Based on Transfer Learning for Cross-Database Pneumonia Detection
  70. A New High Quality Trajectory Tiling Based Hybrid TTS In Real Time
  71. A New Tubular Structure Tracking Algorithm Based On Curvature-Penalized Perceptual Grouping
  72. A Noise-Robust Signal Processing Strategy for Cochlear Implants Using Neural Networks
  73. A Novel Attention-Based Gated Recurrent Unit and its Efficacy in Speech Emotion Recognition
  74. A Novel Bayesian Approach for the Two-Dimensional Harmonic Retrieval Problem
  75. A Novel Convolutional Neural Network Model to Remove Muscle Artifacts from EEG
  76. A Novel NMF-HMM Speech Enhancement Algorithm Based on Poisson Mixture Model
  77. A Novel Viewport-Adaptive Motion Compensation Technique for Fisheye Video
  78. A Novel end-to-end Speech Emotion Recognition Network with Stacked Transformer Layers
  79. A Parallel Algorithm for Phase Retrieval with Dictionary Learning
  80. A Parallelizable Lattice Rescoring Strategy with Neural Language Models
  81. A Parametric Unconstrained Binaural Beamformer Based Noise Reduction and Spatial Cue Preservation for Hearing-Assistive Devices
  82. A Partially Collapsed Gibbs Sampler for Unsupervised Nonnegative Sparse Signal Restoration
  83. A Partially-Relaxed Robust DOA Estimator Under Non-Gaussian Low-Rank Interference and Noise
  84. A Patient-Invariant Model for Freezing of Gait Detection Aided by Wavelet Decomposition
  85. A Periodic Frame Learning Approach for Accurate Landmark Localization in M-Mode Echocardiography
  86. A Plug and Play Fast Intersection Over Union Loss for Boundary Box Regression
  87. A Plug-and-Play Deep Image Prior
  88. A Probabilistic Model for Segmentation of Ambiguous 3D Lung Nodule
  89. A Progressive Learning Approach to Adaptive Noise and Speech Estimation for Speech Enhancement and Noisy Speech Recognition
  90. A Quantitative Analysis Of The Robustness Of Neural Networks For Tabular Data
  91. A Quantitative Metric for Privacy Leakage in Federated Learning
  92. A Quaternion-Valued Variational Autoencoder
  93. A Rank-Constrained Clustering Algorithm with Adaptive Embedding
  94. A Ranked Similarity Loss Function with pair Weighting for Deep Metric Learning
  95. A ReLU Dense Layer to Improve the Performance of Neural Networks
  96. A Real-Time Speaker Diarization System Based on Spatial Spectrum
  97. A Robust Copula Model for Radar-Based Landmine Detection
  98. A Robust and Efficient Multi-Scale Seasonal-Trend Decomposition
  99. A Robust to Noise Adversarial Recurrent Model for Non-Intrusive Load Monitoring
  100. A Sample-Efficient Scheme for Channel Resource Allocation in Networked Estimation
  101. A Scale Invariant Measure of Flatness for Deep Network Minima
  102. A Secure Searchable Image Retrieval Scheme with Correct Retrieval Identity
  103. A Sequential Contrastive Learning Framework for Robust Dysarthric Speech Recognition
  104. A Short Tutorial on The Weisfeiler-Lehman Test And Its Variants
  105. A Simplified Wiener Beamformer Based on Covariance Matrix Modelling
  106. A Sparse Coding Approach to Automatic Diet Monitoring with Continuous Glucose Monitors
  107. A Stage Match for Query-by-Example Spoken Term Detection Based On Structure Information of Query
  108. A Structure-Guided and Sparse-Representation-Based 3d Seismic Inversion Method
  109. A Time-Domain Convolutional Recurrent Network for Packet Loss Concealment
  110. A Triplet Appearance Parsing Network for Person Re-Identification
  111. A Two-Stage Approach to Device-Robust Acoustic Scene Classification
  112. A Two-Stage Deep Modeling Approach to Articulatory Inversion
  113. A Tyler-Type Estimator of Location and Scatter Leveraging Riemannian Optimization
  114. A Unified Approach to Translate Classical Bandit Algorithms to Structured Bandits
  115. A Universal Bert-Based Front-End Model for Mandarin Text-To-Speech Synthesis
  116. A Wireless Reference Active Noise Control Headphone Using Coherence Based Selection Technique
  117. ADAPT-Then-Combine Full Waveform Inversion for Distributed Subsurface Imaging In Seismic Networks
  118. ADL-MVDR: All Deep Learning MVDR Beamformer for Target Speech Separation
  119. ADMM-Based ML Decoding: from Theory to Practice
  120. AEC in A Netshell: on Target and Topology Choices for FCRN Acoustic Echo Cancellation
  121. AISpeech-SJTU ASR System for the Accented English Speech Recognition Challenge
  122. AISpeech-SJTU Accent Identification System for the Accented English Speech Recognition Challenge
  123. ASR N-Best Fusion Nets
  124. ASV-SUBTOOLS: Open Source Toolkit for Automatic Speaker Verification
  125. ATVIO: Attention Guided Visual-Inertial Odometry
  126. Absolute 3d Pose Estimation and Length Measurement of Severely Deformed Fish from Monocular Videos in Longline Fishing
  127. Accdoa: Activity-Coupled Cartesian Direction of Arrival Representation for Sound Event Localization And Detection
  128. Accelerating Auxiliary Function-Based Independent Vector Analysis
  129. Accelerating Frank-Wolfe with Weighted Average Gradients
  130. Acoustic Analysis and Dataset of Transitions Between Coupled Rooms
  131. Acoustic Echo Cancellation with the Dual-Signal Transformation LSTM Network
  132. Acoustic Reflectors Localization from Stereo Recordings Using Neural Networks
  133. Acoustic and Linguistic Analyses to Assess Early-Onset and Genetic Alzheimer's Disease
  134. Acoustic-to-Articulatory Inversion for Dysarthric Speech by Using Cross-Corpus Acoustic-Articulatory Data
  135. Acoustics Based Intent Recognition Using Discovered Phonetic Units for Low Resource Languages
  136. Action State Update Approach to Dialogue Management
  137. Active Estimation From Multimodal Data
  138. Active Privacy-Utility Trade-Off Against A Hypothesis Testing Adversary
  139. Acute Lymphoblastic Leukemia Detection Based on Adaptive Unsharpening and Deep Learning
  140. Ada-Sise: Adaptive Semantic Input Sampling for Efficient Explanation of Convolutional Neural Networks
  141. Adaptable Ensemble Distillation
  142. Adaptable Multi-Domain Language Model for Transformer ASR
  143. Adaptive Bi-Directional Attention: Exploring Multi-Granularity Representations for Machine Reading Comprehension
  144. Adaptive Contention Window Design Using Deep Q-Learning
  145. Adaptive Dual Tree Structure For Screen Content Coding
  146. Adaptive Feature Weight Learning For Robust Clustering Problem with Sparse Constraint
  147. Adaptive GOP Size Decision for Multi-Pass Video Coding Based on Hidden Markov Model
  148. Adaptive Importance Sampling Via Auto-Regressive Generative Models and Gaussian Processes
  149. Adaptive Multi-Domain Learning for Outdoor 3d Human Pose and Shape Estimation
  150. Adaptive Quantization of Model Updates for Communication-Efficient Federated Learning
  151. Adaptive RF Fingerprint Decomposition in Micro UAV Detection based on Machine Learning
  152. Adaptive Re-Balancing Network with Gate Mechanism for Long-Tailed Visual Question Answering
  153. Adaptive Real-Time Filter for Partially-Observed Boolean Dynamical Systems
  154. Adaptive Subsampling of Multidomain Signals with Product Graphs
  155. Adaspeech 2: Adaptive Text to Speech with Untranscribed Data
  156. Admm-Based Fast Algorithm for Robust Multi-Group Multicast Beamforming
  157. Advances in Morphological Neural Networks: Training, Pruning and Enforcing Shape Constraints
  158. Advancing RNN Transducer Technology for Speech Recognition
  159. Adversarial Attacks on Audio Source Separation
  160. Adversarial Attacks on Coarse-to-Fine Classifiers
  161. Adversarial Attacks on Object Detectors with Limited Perturbations
  162. Adversarial Defense for Automatic Speaker Verification by Cascaded Self-Supervised Learning Models
  163. Adversarial Defense for Deep Speaker Recognition Using Hybrid Adversarial Training
  164. Adversarial Examples Detection Beyond Image Space
  165. Adversarial Generative Distance-Based Classifier for Robust Out-of-Domain Detection
  166. Adversarial Learning via Probabilistic Proximity Analysis
  167. Adversarially Robust Classification Based on GLRT
  168. Affine Projection Subspace Tracking
  169. Again-VC: A One-Shot Voice Conversion Using Activation Guidance and Adaptive Instance Normalization
  170. Age-VOX-Celeb: Multi-Modal Corpus for Facial and Speech Estimation
  171. Agent-Environment Network for Temporal Action Proposal Generation
  172. Aggregation Architecture and all-to-one Network for Real-Time Semantic Segmentation
  173. Align or attend? Toward More Efficient and Accurate Spoken Word Discovery Using Speech-to-Image Retrieval
  174. Aligning Sets of Temporal Signals with Riemannian Geometry and Koopman Operator
  175. Aligning the training and evaluation of unsupervised text style Transfer
  176. All For One And One For All: Improving Music Separation By Bridging Networks
  177. Allocating DNN Layers Computation Between Front-End Devices and The Cloud Server for Video Big Data Processing
  178. Alternating Projections Gridless Covariance-Based Estimation For DOA
  179. Amplitude Matching: Majorization-Minimization Algorithm for Sound Field Control Only with Amplitude Constraint
  180. An ADMM Based Network for Hyperspectral Unmixing Tasks
  181. An Accuracy Network Anomaly Detection Method Based on Ensemble Model
  182. An Actor-Critic Reinforcement Learning Approach to Minimum age of Information Scheduling in Energy Harvesting Networks
  183. An Adaptive Discriminant and Sparsity Feature Descriptor for Finger Vein Recognition
  184. An Adaptive Multi-Scale and Multi-Level Features Fusion Network with Perceptual Loss for Change Detection
  185. An Adaptive Non-Linear Process for Under-Determined Virtual Microphone Beamforming
  186. An Adaptive Part-Based Model For Person Re-Identification
  187. An Adaptive Pyramid Single-View Depth Lookup Table Coding Method
  188. An Adaptive Regularization Approach to Portfolio Optimization
  189. An Asymptotically Pointwise Optimal Procedure For Sequential Joint Detection And Estimation
  190. An Asynchronous WFST-Based Decoder for Automatic Speech Recognition
  191. An Attention Based Wavelet Convolutional Model for Visual Saliency Detection
  192. An Attention Model for Hypernasality Prediction in Children with Cleft Palate
  193. An Attention-Seq2Seq Model Based on CRNN Encoding for Automatic Labanotation Generation from Motion Capture Data
  194. An Effective Deep Embedding Learning Method Based on Dense-Residual Networks for Speaker Verification
  195. An Efficient Active Set Algorithm for Covariance Based Joint Data and Activity Detection for Massive Random Access with Massive MIMO
  196. An Efficient Algorithm For Device Detection And Channel Estimation In Asynchronous IOT Systems
  197. An Efficient Alternating Direction Method for Graph Learning from Smooth Signals
  198. An Efficient Linear Programming Rounding-and-Refinement Algorithm for Large-Scale Network Slicing Problem
  199. An Efficient Paper Anti-Counterfeiting Method Based on Microstructure Orientation Estimation
  200. An Empirical Study of End-To-End Simultaneous Speech Translation Decoding Strategies
  201. An Empirical Study of Visual Features for DNN Based Audio-Visual Speech Enhancement in Multi-Talker Environments
  202. An Empirical Study on Task-Oriented Dialogue Translation
  203. An End-To-End Actor-Critic-Based Neural Coreference Resolution System
  204. An End-To-End Non-Intrusive Model for Subjective and Objective Real-World Speech Assessment Using a Multi-Task Framework
  205. An End-to-End Speech Accent Recognition Method Based on Hybrid CTC/Attention Transformer ASR
  206. An Extension of Sparse Audio Declipper to Multiple Measurement Vectors
  207. An F-Test for Polynomial Frequency Modulation
  208. An Hrnet-Blstm Model With Two-Stage Training For Singing Melody Extraction
  209. An Improved Data Driven Dynamic SIRD Model for Predictive Monitoring of COVID-19
  210. An Improved Deep Relation Network for Action Recognition in Still Images
  211. An Improved Event-Independent Network for Polyphonic Sound Event Localization and Detection
  212. An Improved Mean Teacher Based Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection
  213. An Investigation of End-to-End Models for Robust Speech Recognition
  214. An Investigation of Using Hybrid Modeling Units for Improving End-to-End Speech Recognition System
  215. An Iterative Framework for Self-Supervised Deep Speaker Representation Learning
  216. An Optimal Stochastic Compositional Optimization Method with Applications to Meta Learning
  217. An Order-Optimal Adaptive Test Plan for Noisy Group Testing Under Unknown Noise Models
  218. Analog Beamforming With Antenna Selection For Large-Scale Antenna Arrays
  219. Analysing Bias in Spoken Language Assessment Using Concept Activation Vectors
  220. Analysis of X-Vectors for Low-Resource Speech Recognition
  221. Analysis of the but Diarization System for Voxconverse Challenge
  222. Angle-of-Arrival (AoA) Factorization in Multipath Channels
  223. Antenna Selection for Massive MIMO Systems Based on POMDP Framework
  224. Any-to-One Sequence-to-Sequence Voice Conversion Using Self-Supervised Discrete Speech Representations
  225. Application-Layer DDOS Attacks with Multiple Emulation Dictionaries
  226. Approximate Weighted C R Coded Matrix Multiplication
  227. Arrays of First-Order Steerable Differential Microphones
  228. Arrhythmia Classification with Heartbeat-Aware Transformer
  229. Artificially Synthesising Data for Audio Classification and Segmentation to Improve Speech and Music Detection in Radio Broadcast
  230. Assessment of Bipolar Disorder Using Heterogeneous Data of Smartphone-Based Digital Phenotyping
  231. Assisted Learning: Cooperative AI with Autonomy
  232. Asymptotic Distribution of Generalized Likelihood Ratio Test Under Model Misspecification With Application to Cooperative Radar-Communications
  233. Asynchronous Acoustic Echo Cancellation Over Wireless Channels
  234. Attack on Practical Speaker Verification System Using Universal Adversarial Perturbations
  235. Attacking and Defending Behind A Psychoacoustics-Based Captcha
  236. Attention Enhanced Spatial Temporal Neural Network For HRRP Recognition
  237. Attention Is All You Need In Speech Separation
  238. Attention on Attention Sparse Dense Convolutional Network for Financial Signal Processing
  239. Attention-Based Multi-Encoder Automatic Pronunciation Assessment
  240. Attention-Embedded Decomposed Network with Unpaired CT Images Prior for Metal Artifact Reduction
  241. Attention-Guided Second-Order Pooling Convolutional Networks
  242. AttentionLite: Towards Efficient Self-Attention Models for Vision
  243. Attentive Semantic Exploring for Manipulated Face Detection
  244. Attribute Decomposition for Flow-Based Domain Mapping
  245. Audio Dequantization Using (Co)Sparse (Non)Convex Methods
  246. Audio-Visual Event Recognition Through the Lens of Adversary
  247. Audio-Visual Speech Enhancement Method Conditioned in the Lip Motion and Speaker-Discriminative Embeddings
  248. Audio-Visual Speech Inpainting with Deep Learning
  249. Audio-Visual Speech Separation Using Cross-Modal Correspondence Loss
  250. Audiovisual Highlight Detection in Videos
  251. Auditory Filterbanks Benefit Universal Sound Source Separation
  252. Augmented Gaussian Linear Mixture Model for Spectral Variability in Hyperspectral Unmixing
  253. Augmenting Transferred Representations for Stock Classification
  254. AutoKWS: Keyword Spotting with Differentiable Architecture Search
  255. Autoencoder for Vibrotactile Signal Compression
  256. Automated Multi-Organ Segmentation in Pet Images Using Cascaded Training of a 3d U-Net and Convolutional Autoencoder
  257. Automatic And Perceptual Discrimination Between Dysarthria, Apraxia of Speech, and Neurotypical Speech
  258. Automatic Dysarthric Speech Detection Exploiting Pairwise Distance-Based Convolutional Neural Networks
  259. Automatic Elicitation Compliance for Short-Duration Speech Based Depression Detection
  260. Automatic Fine-Grained Localization of Utility Pole Landmarks on Distributed Acoustic Sensing Traces Based on Bilinear Resnets
  261. Automatic Multitrack Mixing With A Differentiable Mixing Console Of Neural Audio Effects
  262. Automatic Order Selection in Autoregressive Modeling with Application in EEG Sleep-Stage Classification
  263. Automatic Registration and Clustering of Time Series
  264. Autoregressive Fast Multichannel Nonnegative Matrix Factorization For Joint Blind Source Separation And Dereverberation
  265. B-Small: A Bayesian Neural Network Approach to Sparse Model-Agnostic Meta-Learning
  266. BLSTM-Based Confidence Estimation for End-to-End Speech Recognition
  267. BW-EDA-EEND: streaming END-TO-END Neural Speaker Diarization for a Variable Number of Speakers
  268. Backdoor Attack Against Speaker Verification
  269. Baitradar: A Multi-Model Clickbait Detection Algorithm Using Deep Learning
  270. Bandwidth Extension is All You Need
  271. Banraw: Band-Limited Radar Waveform Design Via Phase Retrieval
  272. Bayes-Optimal Methods for Finding the Source of a Cascade
  273. Bayesian Estimation of a Tail-Index with Marginalized Threshold
  274. Bayesian Massive MIMO Channel Estimation with Parameter Estimation Using Low-Resolution ADCs
  275. Bayesian Multiple Change-Point Detection of Propagating Events
  276. Bayesian Transformer Language Models for Speech Recognition
  277. Beam Focusing for Multi-User MIMO Communications with Dynamic Metasurface Antennas
  278. Beamforming for Bidirectional Mimo Full Duplex Under the Joint Sum Power and Per Antenna Power Constraints
  279. Benign Overfitting in Binary Classification of Gaussian Mixtures
  280. Bi-APC: Bidirectional Autoregressive Predictive Coding for Unsupervised Pre-Training and its Application to Children's ASR
  281. Bi-Level Style and Prosody Decoupling Modeling for Personalized End-to-End Speech Synthesis
  282. Bidirectional Focused Semantic Alignment Attention Network for Cross-Modal Retrieval
  283. Bifocal Neural ASR: Exploiting Keyword Spotting for Inference Optimization
  284. Binary Control and Digital-to-Analog Conversion Using Composite NUV Priors and Iterative Gaussian Message Passing
  285. Bishift-Net for Image Inpainting
  286. Bit Constrained Communication Receivers In Joint Radar Communications Systems
  287. Blend-Res2net: Blended Representation Space by Transformation of Residual Mapping with Restrained Learning for Time Series Classification
  288. Blind Amplitude Estimation of Early Room Reflections Using Alternating Least Squares
  289. Blind Carbon Copy on Dirty Paper: Seamless Spectrum Underlay via Canonical Correlation Analysis
  290. Blind Deinterleaving of Signals in Time Series with Self-Attention Based Soft Min-Cost Flow Learning
  291. Blind Extraction of Moving Audio Source in a Challenging Environment Supported by Speaker Identification Via X-Vectors
  292. Blind Extraction of Moving Sources via Independent Component and Vector Analysis: Examples
  293. Blind Image Quality Evaluator with Scale Robustness
  294. Blind and Neural Network-Guided Convolutional Beamformer for Joint Denoising, Dereverberation, and Source Separation
  295. Block Kalman Filter: An Asymptotic Block Particle Filter in the Linear Gaussian Case
  296. Bluetooth Low Energy and CNN-Based Angle of Arrival Localization in Presence of Rayleigh Fading
  297. Boosting Low-Resource Intent Detection with in-Scope Prototypical Networks
  298. Branchy-GNN: A Device-Edge Co-Inference Framework for Efficient Point Cloud Processing
  299. Bridging Unpaired Facial Photos and Sketches by Line-Drawings
  300. Bytecover: Cover Song Identification Via Multi-Loss Training
  301. Byzantine-Resilient Decentralized TD Learning with Linear Function Approximation
  302. CASS-NAT: CTC Alignment-Based Single Step Non-Autoregressive Transformer for Speech Recognition
  303. CDPAM: Contrastive Learning for Perceptual Audio Similarity
  304. CMIM: Cross-Modal Information Maximization For Medical Imaging
  305. CNN-Based Spoken Term Detection and Localization without Dynamic Programming
  306. CNR-IEMN: A Deep Learning Based Approach to Recognise Covid-19 from CT-Scan
  307. COOPNet: Multi-Modal Cooperative Gender Prediction in Social Media User Profiling
  308. CSPN: Multi-Scale Cascade Spatial Pyramid Network for Object Detection
  309. Cam: Context-Aware Masking for Robust Speaker Verification
  310. Camera Calibration with Pose Guidance
  311. Camp: A Two-Stage Approach to Modelling Prosody in Context
  312. Canet: Context-Aware Loss for Descriptor Learning
  313. Canonical Polyadic Tensor Decomposition With Low-Rank Factor Matrices
  314. Capturing Banding in Images: Database Construction and Objective Assessment
  315. Capturing Multi-Resolution Context by Dilated Self-Attention
  316. Capturing Temporal Dependencies Through Future Prediction for CNN-Based Audio Classifiers
  317. Cascade Attention Fusion for Fine-Grained Image Captioning Based on Multi-Layer LSTM
  318. Cascaded All-Pass Filters with Randomized Center Frequencies and Phase Polarity for Acoustic and Speech Measurement and Data Augmentation
  319. Cascaded Encoders for Unifying Streaming and Non-Streaming ASR
  320. Cascaded Models with Cyclic Feedback for Direct Speech Translation
  321. Cascaded Time + Time-Frequency Unet For Speech Enhancement: Jointly Addressing Clipping, Codec Distortions, And Gaps
  322. Catiloc: Camera Image Transformer for Indoor Localization
  323. Centrality Based Number of Cluster Estimation in Graph Clustering
  324. Cgan-Net: Class-Guided Asymmetric Non-Local Network for Real-Time Semantic Segmentation
  325. Channel Attention Residual U-Net for Retinal Vessel Segmentation
  326. Channel-Wise Mix-Fusion Deep Neural Networks for Zero-Shot Learning
  327. Characterization of Mems Microphone Sensitivity and Phase Distributions with Applications in Array Processing
  328. Checking PRNU Usability on Modern Devices
  329. Cif-Based Collaborative Decoding for End-to-End Contextual Speech Recognition
  330. Class Aware Robust Training
  331. Class-Conditional Defense GAN Against End-To-End Speech Attacks
  332. Class-Imbalanced Classifiers Using Ensembles of Gaussian Processes And Gaussian Process Latent Variable Models
  333. Classification of Expert-Novice Level Using Eye Tracking And Motion Data via Conditional Multimodal Variational Autoencoder
  334. Classifying Speech Intelligibility Levels of Children in Two Continuous Speech Styles
  335. Close-Talking Recording with Planarly Distributed Microphones
  336. Clustering A Collection of Networks With Mixtures of L1-Sparse Graphical Models
  337. Co-Attentional Transformers for Story-Based Video Understanding
  338. Co-Capsule Networks Based Knowledge Transfer for Cross-Domain Recommendation
  339. Coarse-To-Careful: Seeking Semantic-Related Knowledge for Open-Domain Commonsense Question Answering
  340. Code-Switch Speech Rescoring with Monolingual Data
  341. Codebook Design for Dual-Polarized Ultra-Massive Mimo Communications at Millimeter Wave and Terahertz Bands
  342. Cognitive Memory Constrained Human Decision Making based on Multi-source Information
  343. Cold Start Revisited: A Deep Hybrid Recommender with Cold-Warm Item Harmonization
  344. Collaborative Inference via Ensembles on the Edge
  345. Collaborative Intelligence: Challenges and Opportunities
  346. Collaborative Learning to Generate Audio-Video Jointly
  347. Combined Differential Beamforming With Uniform Linear Microphone Arrays
  348. Combining Adaptive Filtering And Complex-Valued Deep Postfiltering For Acoustic Echo Cancellation
  349. Combining Dynamic Image and Prediction Ensemble for Cross-Domain Face Anti-Spoofing
  350. Communication Over Block Fading Channels - An Algorithmic Perspective On Optimal Transmission Schemes
  351. Communication-Cost Aware Microphone Selection for Neural Speech Enhancement with Ad-Hoc Microphone Arrays
  352. Compact Graph Architecture for Speech Emotion Recognition
  353. Comparative Study of Different Epoch Extraction Methods for Speech Associated with Voice Disorders
  354. Comparison of Deep Co-Training and Mean-Teacher Approaches for Semi-Supervised Audio Tagging
  355. Complex Ratio Masking For Singing Voice Separation
  356. Complex-Valued Vs. Real-Valued Neural Networks for Classification Perspectives: An Example on Non-Circular Data
  357. Compositional Embedding Models for Speaker Identification and Diarization with Simultaneous Speech From 2+ Speakers
  358. Compressed Representation of Cepstral Coefficients via Recurrent Neural Networks for Informed Speech Enhancement
  359. Compressing Deep Neural Networks for Efficient Speech Enhancement
  360. Compressing Local Descriptor Models for Mobile Applications
  361. Compressive Signal Recovery Under Sensing Matrix Errors Combined With Unknown Measurement Gains
  362. Compressive Wideband Spectrum Sensing and Carrier Frequency Estimation with Unknown Mimo Channels
  363. Computationally Efficient DNN-Based Approximation of an Auditory Model for Applications in Speech Processing
  364. Confidence Estimation for Attention-Based Sequence-to-Sequence Models for Speech Recognition
  365. Constant Approximation Algorithm for Minimizing Concave Impurity
  366. Constrained Tensor Decomposition for 2d DOA Estimation In Transmit Beamspace Mimo Radar with Subarrays
  367. Construction of Unit-Norm Tight Frame Based Preconditioner for Sparse Coding
  368. Construction of a Large-Scale Japanese ASR Corpus on TV Recordings
  369. Contact Tracing Enhances the Efficiency of Covid-19 Group Testing
  370. Content-Aware Speaker Embeddings for Speaker Diarisation
  371. Context-Aware Prosody Correction for Text-Based Speech Editing
  372. Context-Aware Speech Stress Detection in Hospital Workers Using Bi-LSTM Classifiers
  373. Continuous Cnn For Nonuniform Time Series
  374. Continuous Face Aging Generative Adversarial Networks
  375. Continuous Speech Separation with Conformer
  376. Continuous-Time Self-Attention in Neural Differential Equation
  377. Contrastive Embeddind Learning Method for Respiratory Sound Classification
  378. Contrastive Learning of General-Purpose Audio Representations
  379. Contrastive Predictive Coding Supported Factorized Variational Autoencoder For Unsupervised Learning Of Disentangled Speech Representations
  380. Contrastive Self-Supervised Learning for Text-Independent Speaker Verification
  381. Contrastive Self-Supervised Learning for Wireless Power Control
  382. Contrastive Semi-Supervised Learning for ASR
  383. Contrastive Separative Coding for Self-Supervised Representation Learning
  384. Contrastive Unsupervised Learning for Speech Emotion Recognition
  385. Control Architecture of the Double-Cross-Correlation Processor for Sampling-Rate-Offset Estimation in Acoustic Sensor Networks
  386. Controlled Testing and Isolation for Suppressing Covid-19
  387. Convergence Analysis of the Graph-Topology-Inference Kernel LMS Algorithm
  388. Conversational Query Rewriting with Self-Supervised Learning
  389. Convex Neural Autoregressive Models: Towards Tractable, Expressive, and Theoretically-Backed Models for Sequential Forecasting and Generation
  390. Convolutional Dropout and Wordpiece Augmentation for End-to-End Speech Recognition
  391. Convolutional Neural Network-Aided Bit-Flipping for Belief Propagation Decoding of Polar Codes
  392. Convolutive Transfer Function Invariant SDR Training Criteria for Multi-Channel Reverberant Speech Separation
  393. Cooperative Parameter Tracking on the Unit Sphere Using Distributed Adapt-Then-Combine Particle Filters and Parallel Transport
  394. Cooperative Scenarios for Multi-Agent Reinforcement Learning in Wireless Edge Caching
  395. CopyPaste: An Augmentation Method for Speech Emotion Recognition
  396. Correlation-Based Robust Linear Regression with Iterative Outlier Removal
  397. Corrupted Contextual Bandits: Online Learning with Corrupted Context
  398. Cost Affinity Learning Network for Stereo Matching
  399. Coughwatch: Real-World Cough Detection using Smartwatches
  400. Count And Separate: Incorporating Speaker Counting For Continuous Speaker Separation
  401. Count Sketch with Zero Checking: Efficient Recovery of Heavy Components
  402. Covid-19 Diagnostic Using 3d Deep Transfer Learning for Classification of Volumetric Computerised Tomography Chest Scans
  403. Crank: An Open-Source Software for Nonparallel Voice Conversion Based on Vector-Quantized Variational Autoencoder
  404. Cross Scene Video Foreground Segmentation Via Co-Occurrence Probability Oriented Supervised and Unsupervised Model Interaction
  405. Cross-Corpus Speech Emotion Recognition Using Joint Distribution Adaptive Regression
  406. Cross-Domain Semi-Supervised Deep Metric Learning for Image Sentiment Analysis
  407. Cross-Domain Sentiment Classification with Contrastive Learning and Mutual Information Maximization
  408. Cross-Modal Knowledge Distillation For Fine-Grained One-Shot Classification
  409. Cross-Modal Representation Reconstruction for Zero-Shot Classification
  410. Cross-Modal Spectrum Transformation Network for Acoustic Scene Classification
  411. Cross-Silo Federated Training in the Cloud with Diversity Scaling and Semi-Supervised Learning
  412. Cross-Teager Energy Cepstral Coefficients for Replay Spoof Detection on Voice Assistants
  413. Crowd Counting Via Multi-Level Regression With Latent Gaussian Maps
  414. Crowdsourcing Approach for Subjective Evaluation of Echo Impairment
  415. Crypto-Oriented Neural Architecture Design
  416. Ct-Caps: Feature Extraction-Based Automated Framework for Covid-19 Disease Identification From Chest Ct Scans Using Capsule Networks
  417. Cue-Preserving MMSE Filter with Bayesian SNR Marginalization for Binaural Speech Enhancement
  418. Cycle Generative Adversarial Network Approaches to Produce Novel Portable Chest X-Rays Images for Covid-19 Diagnosis
  419. D-VDAMP: Denoising-Based Approximate Message Passing for Compressive MRI
  420. DAG-GAN: Causal Structure Learning with Generative Adversarial Nets
  421. DBnet: Doa-Driven Beamforming Network for end-to-end Reverberant Sound Source Separation
  422. DCASENET: An Integrated Pretrained Deep Neural Network for Detecting and Classifying Acoustic Scenes and Events
  423. DEAAN: Disentangled Embedding and Adversarial Adaptation Network for Robust Speaker Representation Learning
  424. DEEPTALK: Vocal Style Encoding for Speaker Recognition and Speech Synthesis
  425. DFDM: A Deep Feature Decoupling Module for Lung Nodule Segmentation
  426. DHASP: Differentiable Hearing Aid Speech Processing
  427. DHCN: Deep Hierarchical Context Networks For Image Annotation
  428. DNANet: Dense Nested Attention Network for Single Image Dehazing
  429. DO as I Mean, Not as I Say: Sequence Loss Training for Spoken Language Understanding
  430. DP-SIGNSGD: When Efficiency Meets Privacy and Robustness
  431. DP-VTON: Toward Detail-Preserving Image-Based Virtual Try-on Network
  432. DURAS: Deep Unfolded Radar Sensing Using Doppler Focusing
  433. Data Augmentation with Signal Companding for Detection of Logical Access Attacks
  434. Data Discovery Using Lossless Compression-Based Sparse Representation
  435. Data Fusion for Audiovisual Speaker Localization: Extending Dynamic Stream Weights to the Spatial Domain
  436. Data-Driven Adaptive Network Resource Slicing for Multi-Tenant Networks
  437. Data-Efficient Framework for Real-World Multiple Sound Source 2d Localization
  438. Decentralized Deep Learning Using Momentum-Accelerated Consensus
  439. Decentralized Motion Inference and Registration of Neuropixel Data
  440. Decentralized Optimization Over Noisy, Rate-Constrained Networks: How We Agree By Talking About How We Disagree
  441. Decentralized Optimization on Time-Varying Directed Graphs Under Communication Constraints
  442. Decentralizing Feature Extraction with Quantum Convolutional Neural Network for Automatic Speech Recognition
  443. Decision Tree Based Inter Partition Termination For Av1 Encoding
  444. Decoding Music Attention from "EEG Headphones": A User-Friendly Auditory Brain-Computer Interface
  445. Decoding Neural Representations of Rhythmic Sounds From Magnetoencephalography
  446. Decomposing Textures using Exponential Analysis
  447. Decouple the High-Frequency and Low-Frequency Information of Images for Semantic Segmentation
  448. Decoupling Pronunciation and Language for End-to-End Code-Switching Automatic Speech Recognition
  449. Deep Active Learning Approach to Adaptive Beamforming for mmWave Initial Alignment
  450. Deep Adversarial Quantization Network for Cross-Modal Retrieval
  451. Deep Auto-Encoding and Biohashing for Secure Finger Vein Recognition
  452. Deep Color Constancy Using Temporal Gradient Under Ac Light Sources
  453. Deep Convolutional Gaussian Processes for Mmwave Outdoor Localization
  454. Deep Convolutional and Recurrent Networks for Polyphonic Instrument Classification from Monophonic Raw Audio Waveforms
  455. Deep Deterministic Information Bottleneck with Matrix-Based Entropy Functional
  456. Deep Ensemble Siamese Network For Incremental Signal Classification
  457. Deep Generative Demixing: Error Bounds for Demixing Subgaussian Mixtures of Lipschitz Signals
  458. Deep Generative Model Learning For Blind Spectrum Cartography with NMF-Based Radio Map Disaggregation
  459. Deep Hashing for Motion Capture Data Retrieval
  460. Deep Learning Architectural Designs for Super-Resolution Of Noisy Images
  461. Deep Learning Based Hybrid Precoding in Dual-Band Communication Systems
  462. Deep Learning for Linear Inverse Problems Using the Plug-and-Play Priors Framework
  463. Deep Learning-Based Cross-Layer Resource Allocation for Wired Communication Systems
  464. Deep Lung Auscultation Using Acoustic Biomarkers for Abnormal Respiratory Sound Event Detection
  465. Deep Multi-Frame MVDR Filtering for Single-Microphone Speech Enhancement
  466. Deep Multiway Canonical Correlation Analysis For Multi-Subject Eeg Normalization
  467. Deep Neural Network Based Cough Detection Using Bed-Mounted Accelerometer Measurements
  468. Deep Neural Network Embeddings for the Estimation of the Degree of Sleepiness
  469. Deep Neural Networks with Flexible Complexity While Training Based on Neural Ordinary Differential Equations
  470. Deep Residual Echo Suppression With A Tunable Tradeoff Between Signal Distortion And Echo Suppression
  471. Deep S3PR: Simultaneous Source Separation and Phase Retrieval Using Deep Generative Models
  472. Deep Semi-Supervised Metric Learning Via Identification of Manifold Memberships
  473. Deep Transform and Metric Learning Networks
  474. Deep Unfolding Network for Block-Sparse Signal Recovery
  475. Deep Weighted MMSE Downlink Beamforming
  476. DeepF0: End-To-End Fundamental Frequency Estimation for Music and Speech Signals
  477. Deepemocluster: a Semi-Supervised Framework for Latent Cluster Representation of Speech Emotions
  478. Deepnodule: Multi-Task Learning of Segmentation Bootstrap for Pulmonary Nodule Detection
  479. Deficient Basis Estimation of Noise Spatial Covariance Matrix for Rank-Constrained Spatial Covariance Matrix Estimation Method in Blind Speech Extraction
  480. Demystifying Model Averaging for Communication-Efficient Federated Matrix Factorization
  481. Denoispeech: Denoising Text to Speech with Frame-Level Noise Modeling
  482. Dense Attention Module for Accurate Pulmonary Nodule Detection
  483. Dense Feature Pyramid Grids Network for Single Image Deraining
  484. Densely Connected Multi-Stage Model with Channel Wise Subband Feature for Real-Time Speech Enhancement
  485. Dependence-Guided Multi-View Clustering
  486. Depression Detection by Analysing Eye Movements on Emotional Images
  487. Design of Graph Signal Sampling Matrices for Arbitrary Signal Subspaces
  488. Designing Random FM Radar Waveforms with Compact Spectrum
  489. Detecting Acoustic Reflectors Using A Robot's Ego-Noise
  490. Detecting Adversarial Attacks on Audiovisual Speech Recognition
  491. Detecting Alzheimer's Disease from Speech Using Neural Networks with Bottleneck Features and Data Augmentation
  492. Detecting Covid-19 and Community Acquired Pneumonia Using Chest CT Scan Images With Deep Learning
  493. Detecting Signal Corruptions in Voice Recordings For Speech Therapy
  494. Detection Of Malicious DNS and Web Servers using Graph-Based Approaches
  495. Detection of Audio-Video Synchronization Errors Via Event Detection
  496. Detection of Covid-19 Through the Analysis of Vocal Fold Oscillations
  497. Detection of Post-Traumatic Stress Disorder Using Learned Time-Frequency Representations from Pupillometry
  498. Developing Real-Time Streaming Transformer Transducer for Speech Recognition on Large-Scale Dataset
  499. Development of the Cuhk Elderly Speech Recognition System for Neurocognitive Disorder Detection Using the Dementiabank Corpus
  500. Diagnosing Covid-19 from CT Images Based on an Ensemble Learning Framework
  501. Dian: Duration Informed Auto-Regressive Network for Voice Cloning
  502. Didispeech: A Large Scale Mandarin Speech Corpus
  503. Differentiable Signal Processing With Black-Box Audio Effects
  504. Differential Chaos Shift Keying-Based Wireless Power Transfer
  505. Differential Convolution Feature Guided Deep Multi-Scale Multiple Instance Learning for Aerial Scene Classification
  506. Dimension Selected Subspace Clustering
  507. Direction Of Arrival Estimation For Non-Coherent Sub-Arrays Via Joint Sparse And Low-Rank Signal Recovery
  508. Direction Preserving Wind Noise Reduction Of B-Format Signals
  509. Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source Localization
  510. Directional Sparse Filtering Using Weighted Lehmer Mean for Blind Separation of Unbalanced Speech Mixtures
  511. Discrete Cosine Transform Based Causal Convolutional Neural Network for Drift Compensation in Chemical Sensors
  512. Discriminability of Single-Layer Graph Neural Networks
  513. Disentangled Speaker and Language Representations Using Mutual Information Minimization and Domain Adaptation for Cross-Lingual TTS
  514. Disentanglement for Audio-Visual Emotion Recognition Using Multitask Setup
  515. Disentangling Subject-Dependent/-Independent Representations for 2D Motion Retargeting
  516. Distributed Scheduling Using Graph Neural Networks
  517. Distributed Speech Separation in Spatially Unconstrained Microphone Arrays
  518. Distribution-Aware Hierarchical Weighting Method for Deep Metric Learning
  519. Divide and Conquer: One-bit MIMO-OFDM Detection by Inexact Expectation Maximization
  520. Dnsmos: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors
  521. DoA estimation of a hidden RF source exploiting simple backscatter radio tags
  522. Domain Adaptation for Learning Generator From Paired Few-Shot Data
  523. Domain-Adversarial Autoencoder with Attention Based Feature Level Fusion for Speech Emotion Recognition
  524. Domain-Aware Neural Language Models for Speech Recognition
  525. Domestic Activities Clustering From Audio Recordings Using Convolutional Capsule Autoencoder Network
  526. Don't Look Back: An Online Beat Tracking Method Using RNN and Enhanced Particle Filtering
  527. Don't Shoot Butterfly with Rifles: Multi-Channel Continuous Speech Separation with Early Exit Transformer
  528. Double Multi-Head Attention for Speaker Verification
  529. Double-DCCCAE: Estimation of Body Gestures From Speech Waveform
  530. Double-Linear Thompson Sampling for Context-Attentive Bandits
  531. Drawgan: Text to Image Synthesis with Drawing Generative Adversarial Networks
  532. Drawing Order Recovery from Trajectory Components
  533. Dual Metric Discriminator for Open Set Video Domain Adaptation
  534. Dual-Path Modeling for Long Recording Speech Separation in Meetings
  535. Dual-Stream Network Based On Global Guidance for Salient Object Detection
  536. Dualformer: A Unified Bidirectional Sequence-to-Sequence Learning
  537. Dynamic Curriculum Learning via Data Parameters for Noise Robust Keyword Spotting
  538. Dynamic Graph Learning Based on Graph Laplacian
  539. Dynamic Graph Modeling Of Simultaneous EEG And Eye-Tracking Data For Reading Task Identification
  540. Dynamic Point Cloud Compression Using A Cuboid Oriented Discrete Cosine Based Motion Model
  541. Dynamic Resource Optimization for Adaptive Federated Learning at the Wireless Network Edge
  542. Dynamic Sparsity Neural Networks for Automatic Speech Recognition
  543. Dynamic Texture Recognition via Nuclear Distances on Kernelized Scattering Histogram Spaces
  544. EADNet: Efficient Asymmetric Dilated Network For Semantic Segmentation
  545. ECCL: Explicit Correlation-Based Convolution Boundary Locator for Moment Localization
  546. ECG Heart-Beat Classification Using Multimodal Image Fusion
  547. EEG-Based Emotion Classification Using Graph Signal Processing
  548. EKFNet: Learning System Noise Statistics from Measurement Data
  549. Eat: Enhanced ASR-TTS for Self-Supervised Speech Recognition
  550. Echo State Speech Recognition
  551. Edge-Aware Multi-Scale Progressive Colorization
  552. Effect of Language Proficiency on Subjective Evaluation of Noise Suppression Algorithms
  553. Effect of Noise and Model Complexity on Detection of Amyotrophic Lateral Sclerosis and Parkinson's Disease Using Pitch and MFCC
  554. Effect of Video Pixel-Binning on Source Attribution of Mixed Media
  555. Effective Rank-Based Estimation of the Coherent-to-Diffuse Power Ratio
  556. Efficient Adversarial Audio Synthesis VIA Progressive Upsampling
  557. Efficient Client Contribution Evaluation for Horizontal Federated Learning
  558. Efficient End-to-End Audio Embeddings Generation for Audio Classification on Target Applications
  559. Efficient Face Manipulation Via Deep Feature Disentanglement And Reintegration Net
  560. Efficient Knowledge Distillation for RNN-Transducer Models
  561. Efficient Long Periodic Binary Sequence Designs for Automotive Radar
  562. Efficient Migration to the Next Generation of Networks Based on Digital Annealing
  563. Efficient Multi-Objective GANs for Image Restoration
  564. Efficient Network Protection Games Against Multiple Types Of Strategic Attackers
  565. Efficient Power Allocation Using Graph Neural Networks and Deep Algorithm Unfolding
  566. Efficient Real-Time Video Stabilization with a Novel Least Squares Formulation
  567. Efficient Speech Emotion Recognition Using Multi-Scale CNN and Attention
  568. Efficient Training Data Generation for Phase-Based DOA Estimation
  569. Efficient Use of End-to-End Data in Spoken Language Processing
  570. Ego-Based Entropy Measures for Structural Representations on Graphs
  571. Ego-GNNs: Exploiting Ego Structures in Graph Neural Networks
  572. Elbert: Fast Albert with Confidence-Window Based Early Exit
  573. Elliptical Shape Recovery from Blurred Pixels Using Deep Learning
  574. Embedding Semantic Hierarchy in Discrete Optimal Transport for Risk Minimization
  575. Emformer: Efficient Memory Transformer Based Acoustic Model for Low Latency Streaming Speech Recognition
  576. Emotion Controllable Speech Synthesis Using Emotion-Unlabeled Dataset with the Assistance of Cross-Domain Speech Emotion Recognition
  577. Emotion Recognition by Fusing Time Synchronous and Time Asynchronous Representations
  578. Empirically Accelerating Scaled Gradient Projection Using Deep Neural Network for Inverse Problems in Image Processing
  579. Enabling Efficient and Expressive Spatial Keyword Queries On Encrypted Data
  580. Encoder-Decoder Based Pitch Tracking and Joint Model Training for Mandarin Tone Classification
  581. End To End Learning For Convolutive Multi-Channel Wiener Filtering
  582. End-2-End Modeling of Speech and Gait from Patients with Parkinson's Disease: Comparison Between High Quality Vs. Smartphone Data
  583. End-To-End Audio-Visual Speech Recognition with Conformers
  584. End-To-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings
  585. End-To-End Multi-Accent Speech Recognition with Unsupervised Accent Modelling
  586. End-To-End Speaker Diarization as Post-Processing
  587. End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced Frontend
  588. End-to-End Learning of Variational Models and Solvers for the Resolution of Interpolation Problems
  589. End-to-End Lyrics Recognition with Voice to Singing Style Transfer
  590. End-to-End Multi-Channel Transformer for Speech Recognition
  591. End-to-End Multilingual Automatic Speech Recognition for Less-Resourced Languages: The Case of Four Ethiopian Languages
  592. End-to-End Spoken Language Understanding Using Transformer Networks and Self-Supervised Pre-Trained Features
  593. End-to-End Text-to-Speech Using Latent Duration Based on VQ-VAE
  594. End-to-End anti-spoofing with RawNet2
  595. End2End Acoustic to Semantic Transduction
  596. Energy Efficiency Optimization Technique for SWIPT-Enabled Multi-Group Multicasting Systems with Heterogeneous Users
  597. Energy Minimization for Federated Learning with IRS-Assisted Over-the-Air Computation
  598. Enhanced Automotive Target Detection through Radar and Communications Sensor Fusion
  599. Enhanced Blind Calibration of Uniform Linear Arrays with One-Bit Quantization by Kullback-Leibler Divergence Covariance Fitting
  600. Enhanced Standard Esprit For Overcoming Imperfections In DOA Estimation
  601. Enhancing Audio Augmentation Methods with Consistency Learning
  602. Enhancing Data-Free Adversarial Distillation with Activation Regularization and Virtual Interpolation
  603. Enhancing Deep Paraphrase Identification via Leveraging Word Alignment Information
  604. Enhancing Image Steganography Via Stego Generation And Selection
  605. Enhancing Model Robustness by Incorporating Adversarial Knowledge into Semantic Representation
  606. Enhancing Multi-Channel Eeg Classification with Gramian Temporal Generative Adversarial Networks
  607. Enhancing into the Codec: Noise Robust Speech Coding with Vector-Quantized Autoencoders
  608. Ensemble Combination between Different Time Segmentations
  609. Ensemble Distillation Approaches for Grammatical Error Correction
  610. Ensure: Ensemble Stein's Unbiased Risk Estimator for Unsupervised Learning
  611. Environment-Independent Wi-Fi Human Activity Recognition with Adversarial Network
  612. Error Estimates in Second-Order Continuous-Time Sigma-Delta Modulators
  613. Error-Driven Fixed-Budget ASR Personalization for Accented Speakers
  614. Error-Driven Pruning of Language Models for Virtual Assistants
  615. Estimating Fiedler Value on Large Networks Based on Random Walk Observations
  616. Estimating Severity of Depression From Acoustic Features and Embeddings of Natural Speech
  617. Estimation of Groundwater Storage Variations in Indus River Basin Using Grace Data
  618. Estimation of Microphone Clusters in Acoustic Sensor Networks Using Unsupervised Federated Learning
  619. Estimation of Visual Features of Viewed Image From Individual and Shared Brain Information Based on FMRI Data Using Probabilistic Generative Model
  620. Evaluation and Comparison of Three Source Direction-of-Arrival Estimators Using Relative Harmonic Coefficients
  621. Event-Driven Modulo Sampling
  622. Evolutionary Quantization of Neural Networks with Mixed-Precision
  623. Evolving Quantized Neural Networks for Image Classification Using A Multi-Objective Genetic Algorithm
  624. Exact Linear Convergence Rate Analysis for Low-Rank Symmetric Matrix Completion via Gradient Descent
  625. Expediting discovery in Neural Architecture Search by Combining Learning with Planning
  626. Exploiting Non-Negative Matrix Factorization for Binaural Sound Localization in the Presence of Directional Interference
  627. Exploiting the Dual-Tree Complex Wavelet Transform for Ship Wake Detection in SAR Imagery
  628. Exploring Automatic COVID-19 Diagnosis via Voice and Symptoms from Crowdsourced Data
  629. Exploring Visual-Audio Composition Alignment Network for Quality Fashion Retrieval in Video
  630. Exploring the application of synthetic audio in training keyword spotters
  631. Exploring the use of Common Label Set to Improve Speech Recognition of Low Resource Indian Languages
  632. Exposing GAN-Generated Faces Using Inconsistent Corneal Specular Highlights
  633. Extended Object Tracking With Automotive Radar Using B-Spline Chained Ellipses Model
  634. Extending Music Based On Emotion And Tonality Via Generative Adversarial Network
  635. Extending Parrotron: An End-to-End, Speech Conversion and Speech Recognition Model for Atypical Speech
  636. Extending the Reverse JPEG Compatibility Attack to Double Compressed Images
  637. F-Net: Fusion Neural Network for Vehicle Trajectory Prediction in Autonomous Driving
  638. FC2RN: A Fully Convolutional Corner Refinement Network for Accurate Multi-Oriented Scene Text Detection
  639. FMA-ETA: Estimating Travel Time Entirely Based on FFN with Attention
  640. FPGA Hardware Design for Plenoptic 3D Image Processing Algorithm Targeting a Mobile Application
  641. FWB-Net: Front White Balance Network for Color Shift Correction in Single Image Dehazing Via Atmospheric Light Estimation
  642. Factorized CRF with Batch Normalization Based on the Entire Training Data
  643. Failure Prediction by Confidence Estimation of Uncertainty-Aware Dirichlet Networks
  644. Fast DCTTS: Efficient Deep Convolutional Text-to-Speech
  645. Fast Decentralized Linear Functions Via Successive Graph Shift Operators
  646. Fast Graph Kernel with Optical Random Features
  647. Fast Hierarchy Preserving Graph Embedding via Subspace Constraints
  648. Fast Inverse Mapping of Face GANs
  649. Fast Local Representation Learning with Adaptive Anchor Graph
  650. Fast Manifold Landmarking Using Extreme Eigen-Pairs
  651. Fast Threshold Optimization for Multi-Label Audio Tagging Using Surrogate Gradient Learning
  652. Fast and Provable Robust PCA VIA Normalized Coherence Pursuit
  653. Fast and Robust ADMM for Blind Super-Resolution
  654. Fast and Robust Stratified Self-Calibration Using Time-Difference-Of-Arrival Measurements
  655. Fast: Feature Aggregation for Detecting Salient Object in Real-Time
  656. FastEmit: Low-Latency Streaming ASR with Sequence-Level Emission Regularization
  657. Fastpitch: Parallel Text-to-Speech with Pitch Prediction
  658. Fcl-Taco2: Towards Fast, Controllable and Lightweight Text-to-Speech Synthesis
  659. Fden: Mining Effective Information of Features in Detecting Network Anomalies
  660. Feature Integration via Semi-Supervised Ordinally Multi-Modal Gaussian Process Latent Variable Model
  661. Feature Redundancy Mining: Deep Light-Weight Image Super-Resolution Model
  662. Feature Reuse for a Randomization Based Neural Network
  663. Federated Acoustic Modeling for Automatic Speech Recognition
  664. Federated Algorithm with Bayesian Approach: Omni-Fedge
  665. Federated Dropout Learning for Hybrid Beamforming with Spatial Path Index Modulation in Multi-User Mmwave-Mimo Systems
  666. Federated Learning from Big Data Over Networks
  667. Federated Learning with Local Differential Privacy: Trade-Offs Between Privacy, Utility, and Communication
  668. Federated Marginal Personalization for ASR Rescoring
  669. Few-Shot Continual Learning for Audio Classification
  670. Few-Shot Image Classification with Multi-Facet Prototypes
  671. Few-Shot Learning for Ct Scan Based Covid-19 Diagnosis
  672. Few-Shot Learning for Decoding Surface Electromyography for Hand Gesture Recognition
  673. Fiber-Sampled Stochastic Mirror Descent for Tensor Decomposition with β-Divergence
  674. Fine-Grained Mri Reconstruction Using Attentive Selection Generative Adversarial Networks
  675. Fine-Grained Pose Temporal Memory Module for Video Pose Estimation and Tracking
  676. Fine-Tuning of Pre-Trained End-to-End Speech Recognition with Generative Adversarial Networks
  677. First-Order Fast Algorithm for Structurally Optimal Multi-Group Multicast Beamforming in Large-Scale Systems
  678. Flow-Based Self-Supervised Density Estimation for Anomalous Sound Detection
  679. Focus on the Present: A Regularization Method for the ASR Source-Target Attention Layer
  680. Focusing-Based Wideband Adaptive Beamforming Using Covariance Matrix Reconstruction
  681. Fontnet: On-Device Font Understanding and Prediction Pipeline
  682. FoolHD: Fooling Speaker Identification by Highly Imperceptible Adversarial Disturbances
  683. Forensicability of Deep Neural Network Inference Pipelines
  684. Four-Dimensional High-Resolution Automotive Radar Imaging Exploiting Joint Sparse-Frequency and Sparse-Array Design
  685. Fourier Transformation Autoencoders for Anomaly Detection
  686. Foveal Avascular Zone Segmentation of Octa Images Using Deep Learning Approach with Unsupervised Vessel Segmentation
  687. Fragmentvc: Any-To-Any Voice Conversion by End-To-End Extracting and Fusing Fine-Grained Voice Fragments with Attention
  688. Frame Rate Up-Conversion Using Key Point Agnostic Frequency-Selective Mesh-to-Grid Resampling
  689. Frame-Rate-Aware Aggregation for Efficient Video Super-Resolution
  690. Frequency-Temporal Attention Network for Singing Melody Extraction
  691. Full-Duplex Multifunction Transceiver with Joint Constant Envelope Transmission and Wideband Reception
  692. Fullsubnet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
  693. Fully-Neural Approach to Vehicle Weighing and Strain Prediction on Bridges Using Wireless Accelerometers
  694. Fundamental Frequency Feature Normalization and Data Augmentation for Child Speech Recognition
  695. Fundamental Trade-Offs in Noisy Super-Resolution with Synthetic Apertures
  696. Fusing Information Streams in End-to-End Audio-Visual Speech Recognition
  697. Fusing Multitask Models by Recursive Least Squares
  698. Fusion-Based Digital Image Correlation Framework for Strain Measurement
  699. G-Arrays: Geometric Arrays for Efficient Point Cloud Processing
  700. GAN-Based Out-of-Domain Detection Using Both In-Domain and Out-of-Domain Samples
  701. GDTW: A Novel Differentiable DTW Loss for Time Series Tasks
  702. GTA-Net: Gradual Temporal Aggregation Network for Fast Video Deraining
  703. Gate Trimming: One-Shot Channel Pruning for Efficient Convolutional Neural Networks
  704. Gating Feature Dense Network for Single Anisotropic Mr Image Super-Resolution
  705. Gaussian Kernelized Self-Attention for Long Sequence Data and its Application to CTC-Based Speech Recognition
  706. Gaussian Process Temporal-Difference Learning with Scalability and Worst-Case Performance Guarantees
  707. General Total Variation Regularized Sparse Bayesian Learning for Robust Block-Sparse Signal Recovery
  708. Generalized Knowledge Distillation from an Ensemble of Specialized Teachers Leveraging Unsupervised Neural Clustering
  709. Generalized Polytopic Matrix Factorization
  710. Generalized Thinned Coprime Array for DOA Estimation
  711. Generating Empathetic Responses by Injecting Anticipated Emotion
  712. Generating Human Readable Transcript for Automatic Speech Recognition with Pre-Trained Language Model
  713. Generating Natural Questions from Images for Multimodal Assistants
  714. Generative Information Fusion
  715. Generative Speech Coding with Predictive Variance Regularization
  716. Geom-Spider-EM: Faster Variance Reduced Stochastic Expectation Maximization for Nonconvex Finite-Sum Optimization
  717. Geometric Scattering Attention Networks
  718. Geometry Consistency Of Augmented Reality Based On Semantics
  719. Global-Localized Agent Graph Convolution for Multi-Agent Reinforcement Learning
  720. Globally Optimal Beamforming for Rate Splitting Multiple Access
  721. Gps-Denied Navigation Using Sar Images And Neural Networks
  722. Gradual Federated Learning Using Simulated Annealing
  723. Gramian-Based Adaptive Combination Policies for Diffusion Learning Over Networks
  724. Granger Causality Based Directional Phase-Amplitude Coupling Measure
  725. Graph Attention Networks for Speaker Verification
  726. Graph Attention and Interaction Network With Multi-Task Learning for Fact Verification
  727. Graph Embedding using Multi-Layer Adjacent Point Merging Model
  728. Graph Enhanced Query Rewriting for Spoken Language Understanding System
  729. Graph Frequency Analysis of COVID-19 Incidence to Identify County-Level Contagion Patterns in the United States
  730. Graph Learning Under Spectral Sparsity Constraints
  731. Graph Neural Network for Large-Scale Network Localization
  732. Graph Neural Networks for Decentralized Controllers
  733. Graph Signal Compression via Task-Based Quantization
  734. Graph Signal Denoising Using Nested-Structured Deep Algorithm Unrolling
  735. Graph Signal Denoising Via Unrolling Networks
  736. Graph-Adaptive Incremental Learning Using an Ensemble of Gaussian Process Experts
  737. Graph-Based Pyramid Global Context Reasoning With a Saliency- Aware Projection for Covid-19 Lung Infections Segmentation
  738. Graph-Homomorphic Perturbations for Private Decentralized Learning
  739. Graphcomm: A Graph Neural Network Based Method for Multi-Agent Reinforcement Learning
  740. Graphnet: Graph Clustering with Deep Neural Networks
  741. Graphon and Graph Neural Network Stability
  742. Graphspeech: Syntax-Aware Graph Attention Network for Neural Speech Synthesis
  743. Grid Optimization for Matrix-Based Source Localization Under Inhomogeneous Sensor Topology
  744. Guaranteed Reconstruction from Integrate-and-Fire Neurons with Alpha Synaptic Activation
  745. Guided Variational Autoencoder for Speech Enhancement with a Supervised Classifier
  746. H-GPR: A Hybrid Strategy for Large-Scale Gaussian Process Regression
  747. HCAG: A Hierarchical Context-Aware Graph Attention Model for Depression Detection
  748. HCGM-Net: A Deep Unfolding Network for Financial Index Tracking
  749. HFGCNET: High-Frequency Graph Reasoning for Finer Semantic Image Segmentation
  750. HIGCNN: Hierarchical Interleaved Group Convolutional Neural Networks for Point Clouds Analysis
  751. HOCA: Higher-Order Channel Attention for Single Image Super-Resolution
  752. HSAN: A Hierarchical Self-Attention Network for Multi-Turn Dialogue Generation
  753. HVS-Based Perceptual Color Compression of Image Data
  754. Handling Class Imbalance in Low-Resource Dialogue Systems by Combining Few-Shot Classification and Interpolation
  755. Handwritten Digits Reconstruction from Unlabelled Embeddings
  756. Hardware Implementation of Iterative Projection-Aggregation Decoding of Reed-Muller Codes
  757. Head-Synchronous Decoding for Transformer-Based Streaming ASR
  758. HebbNet: A Simplified Hebbian Learning Framework to do Biologically Plausible Learning
  759. Heterogeneous two-Stream Network with Hierarchical Feature Prefusion for Multispectral Pan-Sharpening
  760. Hidden Markov Model Diarisation with Speaker Location Information
  761. Hide Chopin in the Music: Efficient Information Steganography Via Random Shuffling
  762. Hierarchical Attention Fusion for Geo-Localization
  763. Hierarchical Attention-Based Temporal Convolutional Networks for Eeg-Based Emotion Recognition
  764. Hierarchical Bit-Wise Differential Coding (HBDC) of Point Cloud Attributes
  765. Hierarchical Coded Elastic Computing
  766. Hierarchical Context Guided Aggregation Network for Stereo Matching
  767. Hierarchical Network Based on the Fusion of Static and Dynamic Features for Speech Emotion Recognition
  768. Hierarchical Pose Classification for Infant Action Analysis and Mental Development Assessment
  769. Hierarchical Recurrent Neural Network for Handwritten Strokes Classification
  770. Hierarchical Refined Attention for Scene Text Recognition
  771. Hierarchical Similarity Learning for Language-Based Product Image Retrieval
  772. Hierarchical Speaker-Aware Sequence-to-Sequence Model for Dialogue Summarization
  773. Hierarchical Transformer-Based Large-Context End-To-End ASR with Large-Context Knowledge Distillation
  774. High Accuracy Tracking of Targets Using Massive MIMO
  775. High Fidelity Speech Regeneration with Application to Speech Enhancement
  776. High-Frequency Adversarial Defense for Speech and Audio
  777. High-Intelligibility Speech Synthesis for Dysarthric Speakers with LPCNet-Based TTS and CycleVAE-Based VC
  778. High-Throughput VLSI Architecture for Soft-Decision Decoding with ORBGRAND
  779. Highly Efficient Protection of Biometric Face Samples with Selective JPEG2000 Encryption
  780. History Utterance Embedding Transformer LM for Speech Recognition
  781. How Convolutional Neural Networks Deal with Aliasing
  782. How Phonotactics Affect Multilingual and Zero-Shot ASR Performance
  783. How Similar or Different is Rakugo Speech Synthesizer to Professional Performers?
  784. How to Make Text-to-Speech System Pronounce "Voldemort": an Experimental Approach of Foreign Word Phonemization in Vietnamese
  785. How to Use Time Information Effectively? Combining with Time Shift Module for Lipreading
  786. Hubert: How Much Can a Bad Teacher Benefit ASR Pre-Training?
  787. Human-Aware Coarse-to-Fine Online Action Detection
  788. Human-Centered Favorite Music Classification Using EEG-Based Individual Music Preference Via Deep Time-Series CCA
  789. Human-Expert-Level Brain Tumor Detection Using Deep Learning with Data Distillation And Augmentation
  790. Humanacgan: Conditional Generative Adversarial Network with Human-Based Auxiliary Classifier and its Evaluation in Phoneme Perception
  791. Hybrid Analog-Digital MIMO Radar Receivers With Bit-Limited ADCs
  792. Hybrid Beamforming for Wideband OFDM Dual Function Radar Communications
  793. Hybrid Model for Network Anomaly Detection with Gradient Boosting Decision Trees and Tabtransformer
  794. Hyperspectral Image Super-Resolution Via Adjacent Spectral Fusion Strategy
  795. Hypothesis Stitcher for End-to-End Speaker-Attributed ASR on Long-Form Multi-Talker Recordings
  796. ICA with Orthogonality Constraint: Identifiability And A New Efficient Algorithm
  797. ICASSP 2021 Acoustic Echo Cancellation Challenge: Datasets, Testing Framework, and Results
  798. ICASSP 2021 Acoustic Echo Cancellation Challenge: Integrated Adaptive Echo Cancellation with Time Alignment and Deep Learning-Based Residual Echo Plus Noise Suppression
  799. ICASSP 2021 Deep Noise Suppression Challenge
  800. ICASSP 2021 Deep Noise Suppression Challenge: Decoupling Magnitude and Phase Optimization with a Two-Stage Deep Network
  801. ICI-Aware Parameter Estimation for Mimo-Ofdm Radar via Apes Spatial Filtering
  802. Identification of Deep Breath While Moving Forward Based on Multiple Body Regions and Graph Signal Analysis
  803. Identification of Uterine Contractions by An Ensemble of Gaussian Processes
  804. Identifying First-Order Lowpass Graph Signals Using Perron Frobenius Theorem
  805. Identifying Spammers to Boost Crowdsourced Classification
  806. Image Coding For Machines: an End-To-End Learned Approach
  807. Image Coding with Neural Network-Based Colorization
  808. Image Denoising Based on Correlation Adaptive Sparse Modeling
  809. Image Generation Based on Texture Guided VAE-AGAN for Regions of Interest Detection in Remote Sensing Images
  810. Image Steganography Based on Iterative Adversarial Perturbations Onto a Synchronized-Directions Sub-Image
  811. Image Super-Resolution Using Multi-Resolution Attention Network
  812. Image-Assisted Transformer in Zero-Resource Multi-Modal Translation
  813. Impact of Sound Duration and Inactive Frames on Sound Event Detection Performance
  814. Impact of Speaking Rate on the Source Filter Interaction in Speech: A Study
  815. Implicit HRTF Modeling Using Temporal Convolutional Networks
  816. Improved Atomic Norm Based Channel Estimation for Time-Varying Narrowband Leaked Channels
  817. Improved Data Selection for Domain Adaptation in ASR
  818. Improved Intra Mode Coding Beyond Av1
  819. Improved Mask-CTC for Non-Autoregressive End-to-End ASR
  820. Improved Neural Language Model Fusion for Streaming Recurrent Neural Network Transducer
  821. Improved Probabilistic Context-Free Grammars for Passwords Using Word Extraction
  822. Improved Robustness to Disfluencies in Rnn-Transducer Based Speech Recognition
  823. Improved Step-Size Schedules for Noisy Gradient Methods
  824. Improved Supervised Training of Physics-Guided Deep Learning Image Reconstruction with Multi-Masking
  825. Improvements to Prosodic Alignment for Automatic Dubbing
  826. Improving Audio Anomalies Recognition Using Temporal Convolutional Attention Networks
  827. Improving Automatic Drum Transcription Using Large-Scale Audio-to-Midi Aligned Data
  828. Improving Cross-Domain Slot Filling with Common Syntactic Structure
  829. Improving Deep Learning Sound Events Classifiers Using Gram Matrix Feature-Wise Correlations
  830. Improving Dialogue Response Generation Via Knowledge Graph Filter
  831. Improving Entity Recall in Automatic Speech Recognition with Neural Embeddings
  832. Improving Event Detection by Exploiting Label Hierarchy
  833. Improving Identification of System-Directed Speech Utterances by Deep Learning of ASR-Based Word Embeddings and Confidence Metrics
  834. Improving Intraoperative Liver Registration in Image-Guided Surgery with Learning-Based Reconstruction
  835. Improving Memory Banks for Unsupervised Learning with Large Mini-Batch, Consistency and Hard Negative Mining
  836. Improving Multimodal Speech Enhancement by Incorporating Self-Supervised and Curriculum Learning
  837. Improving NER in Social Media via Entity Type-Compatible Unknown Word Substitution
  838. Improving Naturalness and Controllability of Sequence-to-Sequence Speech Synthesis by Learning Local Prosody Representations
  839. Improving Neural Text Normalization with Partial Parameter Generator and Pointer-Generator Network
  840. Improving Pronunciation Assessment Via Ordinal Regression with Anchored Reference Samples
  841. Improving Prosody Modelling with Cross-Utterance Bert Embeddings for End-to-End Speech Synthesis
  842. Improving RNN Transducer Modeling for Small-Footprint Keyword Spotting
  843. Improving RNN Transducer with Target Speaker Extraction and Neural Uncertainty Estimation
  844. Improving Reconstruction Loss Based Speaker Embedding in Unsupervised and Semi-Supervised Scenarios
  845. Improving Sound Event Detection Metrics: Insights from DCASE 2020
  846. Improving Speaker Verification in Reverberant Environments
  847. Improving Stability of Adversarial Li-ion Cell Usage Data Generation using Generative Latent Space Modelling
  848. Improving Streaming Automatic Speech Recognition with Non-Streaming Model Distillation on Unsupervised Data
  849. Improving The Robustness Of Right Whale Detection In Noisy Conditions Using Denoising Autoencoders And Augmented Training
  850. Improving Ultrasound Tongue Contour Extraction Using U-Net and Shape Consistency-Based Regularizer
  851. Improving the Classification of Rare Chords With Unlabeled Data
  852. Improving the Energy-Efficiency of a Kalman Filter Using Unreliable Memories
  853. Imrnet: An Iterative Motion Compensation and Residual Reconstruction Network for Video Compressed Sensing
  854. In Situ Calibration of Cross-Sensitive Sensors in Mobile Sensor Arrays Using Fast Informed Non-Negative Matrix Factorization
  855. In-Bed Pressure-Based Pose Estimation Using Image Space Representation Learning
  856. Incomplete Multi-View Subspace Clustering with Low-Rank Tensor
  857. Incorporate Maximum Mean Discrepancy in Recurrent Latent Space for Sequential Generative Model
  858. Incorporating Syntactic and Phonetic Information into Multimodal Word Embeddings Using Graph Convolutional Networks
  859. Incorporating Uncertainty In Data Labeling Into Detection of Brain Interictal Epileptiform Discharges From EEG Using Weighted optimization
  860. Independent Sign Language Recognition with 3d Body, Hands, and Face Reconstruction
  861. Independent Vector Analysis Using Semi-Parametric Density Estimation via Multivariate Entropy Maximization
  862. Inertial Proximal Deep Learning Alternating Minimization for Efficient Neutral Network Training
  863. Inferring High-Resolutional Urban Flow With Internet Of Mobile Things
  864. Information Decoding and SDR Implementation of DFRC Systems without Training Signals
  865. Information and Regularization in Restricted Boltzmann Machines
  866. Injecting Word Information with Multi-Level Word Adapter for Chinese Spoken Language Understanding
  867. Instance Segmentation with the Number of Clusters Incorporated in Embedding Learning
  868. Instrument Classification of Solo Sheet Music Images
  869. Integer Carrier Frequency Offset Estimation in OFDM with Zadoff-Chu Sequences
  870. Integrated Classification and Localization of Targets Using Bayesian Framework In Automotive Radars
  871. Integrated Grad-Cam: Sensitivity-Aware Visual Explanation of Deep Convolutional Networks Via Integrated Gradient-Based Scoring
  872. Integrating Deep Learning with First-Order Logic Programmed Constraints for Zero-Day Phishing Attack Detection
  873. Integrating End-to-End Neural and Clustering-Based Diarization: Getting the Best of Both Worlds
  874. Integrating Subgraph-Aware Relation and Direction Reasoning for Question Answering
  875. Interference Analysis in Reconfigurable Intelligent Surface-Assisted Multiple-Input Multiple-Output Systems
  876. Intermediate Loss Regularization for CTC-Based Speech Recognition
  877. Internal Language Model Training for Domain-Adaptive End-To-End Speech Recognition
  878. Interpolation of Irregularly Sampled Frequency Response Functions Using Convolutional Neural Networks
  879. Interpreting Glottal Flow Dynamics for Detecting Covid-19 From Voice
  880. Introducing Deep Reinforcement Learning to Nlu Ranking Tasks
  881. Investigating Local and Global Information for Automated Audio Captioning with Transfer Learning
  882. Investigating on Incorporating Pretrained and Learnable Speaker Representations for Multi-Speaker Multi-Style Text-to-Speech
  883. Investigating the Efficacy of Music Version Retrieval Systems for Setlist Identification
  884. Investigation of Fast and Efficient Methods for Multi-Speaker Modeling and Speaker Adaptation
  885. Iterative Geometry Calibration from Distance Estimates for Wireless Acoustic Sensor Networks
  886. Iterative Reweighted Algorithms for Joint User Identification and Channel Estimation in Spatially Correlated Massive MTC
  887. Jamming Strategy Generation for Hidden Communication Modes Via Graph Convolution Networks
  888. Joint ASR and Language Identification Using RNN-T: An Efficient Approach to Dynamic Language Switching
  889. Joint Alignment Learning-Attention Based Model for Grapheme-to-Phoneme Conversion
  890. Joint Channel, Data, and Phase-Noise Estimation in MIMO-OFDM Systems Using a Tensor Modeling Approach
  891. Joint Communications with FH-MIMO Radar Systems: An Extended Signaling Strategy
  892. Joint Coupled Transform Learning Framework for Multimodal Image Super-Resolution
  893. Joint Dereverberation and Separation With Iterative Source Steering
  894. Joint Intent Detection and Slot Filling Based on Continual Learning Model
  895. Joint Learning of Image Aesthetic Quality Assessment and Semantic Recognition Based on Feature Enhancement
  896. Joint Localization and Predictive Beamforming in Vehicular Networks: Power Allocation Beyond Water-Filling
  897. Joint Masked CPC And CTC Training For ASR
  898. Joint Maximum Likelihood Estimation of Power Spectral Densities and Relative Acoustic Transfer Functions for Acoustic Beamforming
  899. Joint Multi-Pitch Detection and Score Transcription for Polyphonic Piano Music
  900. Joint Optimization for Full-Duplex Cellular Communications Via Intelligent Reflecting Surface
  901. Joint Optimization of Spectrally Co-Existing Multi-Carrier Radar and Communication Systems in Cluttered Environments
  902. Joint Reinforcement Learning and Game Theory Bitrate Control Method for 360-Degree Dynamic Adaptive Streaming
  903. Jointly Trained Transformers Models for Spoken Language Translation
  904. KAN: Knowledge-Augmented Networks for Few-Shot Learning
  905. Kalman Filter Based MIMO CSI Phase Recovery for COTS Wifi Devices
  906. Kalman Optimizer for Consistent Gradient Descent
  907. Kalmannet: Data-Driven Kalman Filtering
  908. Karaoke Key Recommendation Via Personalized Competence-Based Rating Prediction
  909. Kernearl-Based Lifelong Policy Gradient Reinforcement Learning
  910. Kernel Orthogonal Nonnegative Matrix Factorization: Application to Multispectral Document Image Decomposition
  911. Kernel Regression on Graphs in Random Fourier Features Space
  912. Kernel-Interpolation-Based Filtered-X Least Mean Square for Spatial Active Noise Control In Time Domain
  913. Kld Minimization-Based Constrained Measurement Filtering For Two-Step TDOA Indoor Tracking
  914. Knowledge Distillation for Improved Accuracy in Spoken Question Answering
  915. Knowledge Reasoning for Semantic Segmentation
  916. Knowledge Transfer for Efficient on-Device False Trigger Mitigation
  917. Knowledge-Based Chat Detection with False Mention Discrimination
  918. L-Red: Efficient Post-Training Detection of Imperceptible Backdoor Attacks Without Access to the Training Set
  919. LIFI: Towards Linguistically Informed Frame Interpolation
  920. LSSED: A Large-Scale Dataset and Benchmark for Speech Emotion Recognition
  921. LVCNet: Efficient Condition-Dependent Modeling Network for Waveform Generation
  922. Label-Aware Text Representation for Multi-Label Text Classification
  923. Language Model is all You Need: Natural Language Understanding as Question Answering
  924. Language-Sensitive Music Emotion Recognition Models: are We Really There Yet?
  925. Laplacian Regularized Tensor Low-Rank Minimization for Hyperspectral Snapshot Compressive Imaging
  926. Large Margin Training Improves Language Models for ASR
  927. Lasaft: Latent Source Attentive Frequency Transformation For Conditioned Source Separation
  928. Latent Space Motion Analysis for Collaborative Intelligence
  929. Lattice-Free Mmi Adaptation of Self-Supervised Pretrained Acoustic Models
  930. Layer-Wise Interpretation of Deep Neural Networks using Identity Initialization
  931. Leaky Integrator Dynamical Systems and Reachable Sets
  932. Learned Decimation for Neural Belief Propagation Decoders : Invited Paper
  933. Learned Transferable Architectures Can Surpass Hand-Designed Architectures for Large Scale Speech Recognition
  934. Learning Audio Embeddings with User Listening Data for Content-Based Music Recommendation
  935. Learning Audio-Visual Correlations From Variational Cross-Modal Generation
  936. Learning Binary Semantic Embedding for Breast Histology Image Classification and Retrieval
  937. Learning Bollobás-Riordan Graphs Under Partial Observability
  938. Learning Contextual Tag Embeddings for Cross-Modal Alignment of Audio and Tags
  939. Learning Discriminative Features for Semi-Supervised Anomaly Detection
  940. Learning Disentangled Feature Representations for Speech Enhancement Via Adversarial Training
  941. Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE Paradigm
  942. Learning Double-Compression Video Fingerprints Left From Social-Media Platforms
  943. Learning From Heterogeneous Eeg Signals with Differentiable Channel Reordering
  944. Learning Integrodifferential Models for Image Denoising
  945. Learning Mixed Membership from Adjacency Graph Via Systematic Edge Query: Identifiability and Algorithm
  946. Learning Model-Blind Temporal Denoisers without Ground Truths
  947. Learning On Heterogeneous Graphs Using High-Order Relations
  948. Learning Optimal Lattice Codes for MIMO Communications
  949. Learning Pose-Adaptive Lip Sync with Cascaded Temporal Convolutional Network
  950. Learning Representation of Multi-Scale Object for Fine-Grained Image Retrieval
  951. Learning Separable Time-Frequency Filterbanks for Audio Classification
  952. Learning Sparse Graph Laplacian with K Eigenvector Prior via Iterative Glasso and Projection
  953. Learning Sparsifying Transforms for Image Reconstruction in Electrical Impedance Tomography
  954. Learning Word-Level Confidence for Subword End-To-End ASR
  955. Learning a Sparse Generative Non-Parametric Supervised Autoencoder
  956. Learning a Tree of Neural Nets
  957. Learning the Relevant Substructures for Tasks on Graph Data
  958. Learning to Continuously Optimize Wireless Resource in Episodically Dynamic Environment
  959. Learning to Estimate Kernel Scale and Orientation of Defocus Blur with Asymmetric Coded Aperture
  960. Learning to Select Context in a Hierarchical and Global Perspective for Open-Domain Dialogue Generation
  961. Learning to Select for Mimo Radar Based on Hybrid Analog-Digital Beamforming
  962. Learning-Based Lossless Compression of 3D Point Cloud Geometry
  963. Length No Longer Matters: A Real Length Adaptive Arrhythmia Classification Model with Multi-Scale Convolution
  964. Less is More: Improved RNN-T Decoding Using Limited Label Context and Path Merging
  965. Leveraging A Multiple-Strain Model with Mutations in Analyzing the Spread of Covid-19
  966. Leveraging Acoustic and Linguistic Embeddings from Pretrained Speech and Language Models for Intent Classification
  967. Leveraging the Structure of Musical Preference in Content-Aware Music Recommendation
  968. Light Field Style Transfer with Local Angular Consistency
  969. Light-TTS: Lightweight Multi-Speaker Multi-Lingual Text-to-Speech
  970. Lightspeech: Lightweight and Fast Text to Speech with Neural Architecture Search
  971. Lightweight Dual-Task Networks For Crowd Counting In Aerial Images
  972. Lightweight Human Pose Estimation under Resource-Limited Scenes
  973. Lightweight Non-Local Network for Image Super-Resolution
  974. Lightweight and Accurate Single Image Super-Resolution with Channel Segregation Network
  975. Lightweight and Interpretable Neural Modeling of an Audio Distortion Effect Using Hyperconditioned Differentiable Biquads
  976. Linear Computation Coding
  977. Linear Multichannel Blind Source Separation based on Time-Frequency Mask Obtained by Harmonic/Percussive Sound Separation
  978. Litesing: Towards Fast, Lightweight and Expressive Singing Voice Synthesis
  979. Locally Optimal Detection of Stochastic Targeted Universal Adversarial Perturbations
  980. Long-Short Temporal Modeling for Efficient Action Recognition
  981. Looking Through Walls: Inferring Scenes from Video-Surveillance Encrypted Traffic
  982. Loopnet: Musical Loop Synthesis Conditioned on Intuitive Musical Parameters
  983. Low Complexity SLM for OFDMA System with Implicit Side Information
  984. Low Complexity Secure P-Tensor Product Compressed Sensing Reconstruction Outsourcing and Identity Authentication in Cloud
  985. Low Latency Online Blind Source Separation Based on Joint Optimization with Blind Dereverberation
  986. Low Mutual Coupling Sparse Array Design Using ULA Fitting
  987. Low Resource Audio-To-Lyrics Alignment from Polyphonic Music Recordings
  988. Low-Complexity Parameter Learning for OTFS Modulation Based Automotive Radar
  989. Low-Complexity, Real-Time Joint Neural Echo Control and Speech Enhancement Based On Percepnet
  990. Low-Dimensional Denoising Embedding Transformer for ECG Classification
  991. Low-Latency Polar Decoder Using Overlapped SCL Processing
  992. Low-Rank and Sparse Decomposition for Joint DOA Estimation and Contaminated Sensors Detection with Sparsely Contaminated Arrays
  993. Low-Rank on Graphs Plus Temporally Smooth Sparse Decomposition for Anomaly Detection in Spatiotemporal Data
  994. Low-Resource Expressive Text-To-Speech Using Data Augmentation
  995. Ltaf-Net: Learning Task-Aware Adaptive Features and Refining Mask for Few-Shot Semantic Segmentation
  996. MAEC: Multi-Instance Learning with an Adversarial Auto-Encoder-Based Classifier for Speech Emotion Recognition
  997. MAPGN: Masked Pointer-Generator Network for Sequence-to-Sequence Pre-Training
  998. MBNET: MOS Prediction for Synthesized Speech with Mean-Bias Network
  999. MCR-NET: A Multi-Step Co-Interactive Relation Network for Unanswerable Questions on Machine Reading Comprehension
  1000. MPDNet: A 3D Missing Part Detection Network Based on Point Cloud Segmentation

Looking for submission deadlines instead? See the conference deadline calendar.