← All conferences

ICASSP 2020 Accepted Papers

The full list of 1,849 papers accepted at ICASSP 2020 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. Language-Agnostic Multilingual Modeling
  2. Laplace State Space Filter with Exact Inference and Moment Matching
  3. Large Dimensional Asymptotics of Multi-Task Learning
  4. Large-Context Pointer-Generator Networks for Spoken-to-Written Style Conversion
  5. Large-Scale Fading Precoding for Maximizing the Product of SINRs
  6. Large-Scale Time Series Clustering with k-ARs
  7. Large-Scale Unsupervised Pre-Training for End-to-End Spoken Language Understanding
  8. Large-Scale Weakly-Supervised Content Embeddings for Music Recommendation and Tagging
  9. Latency-Minimized Design of secure transmissions in UAV-Aided Communications
  10. Latent Fused Lasso
  11. Lattice-Based Improvements for Voice Triggering Using Graph Neural Networks
  12. Layer-Normalized LSTM for Hybrid-Hmm and End-To-End ASR
  13. Learn-By-Calibrating: Using Calibration As A Training Objective
  14. Learned Lossless Image Compression with A Hyperprior and Discretized Gaussian Mixture Likelihoods
  15. Learning A Common Granger Causality Network Using A Non-Convex Regularization
  16. Learning Asr-Robust Contextualized Embeddings for Spoken Language Understanding
  17. Learning Based Reconfigurable Sub-nyquist Sampling Framework for Ultra-wideband Angular Sensing
  18. Learning Blind Denoising Network for Noisy Image Deblurring
  19. Learning Data Representation and Emotion Assessment from Physiological Data
  20. Learning Differentiable Sparse and Low Rank Networks for Audio-Visual Object Localization
  21. Learning Diverse Sub-Policies via a Task-Agnostic Regularization on Action Distributions
  22. Learning Domain Invariant Representations for Child-Adult Classification from Speech
  23. Learning Eating Environments Through Scene Clustering
  24. Learning Endmember Dynamics in Multitemporal Hyperspectral Data Using A State-Space Model Formulation
  25. Learning Geometric Features with Dual-stream CNN for 3D Action Recognition
  26. Learning Graph Influence from Social Interactions
  27. Learning Local Structure of Representative Points for Point Cloud Classification and Semantic Segmentation
  28. Learning Multi-Scale Attentive Features for Series Photo Selection
  29. Learning Network Representation Through Reinforcement Learning
  30. Learning Noise Invariant Features Through Transfer Learning For Robust End-to-End Speech Recognition
  31. Learning Partial Differential Equations From Data Using Neural Networks
  32. Learning Perception and Planning With Deep Active Inference
  33. Learning Plug-And-Play Proximal Quasi-Newton Denoisers
  34. Learning Product Graphs from Multidomain Signals
  35. Learning Recurrent Neural Network Language Models With Context-Sensitive Label Smoothing for Automatic Speech Recognition
  36. Learning Sampling and Model-Based Signal Recovery for Compressed Sensing MRI
  37. Learning Semi-Supervised Anonymized Representations by Mutual Information
  38. Learning Signed Graphs from Data
  39. Learning Spatio-Temporal Convolutional Network for Real-Time Object Tracking
  40. Learning Spatio-Temporal Representations With Temporal Squeeze Pooling
  41. Learning Spectral-Spatial Prior Via 3DDNCNN for Hyperspectral Image Deconvolution
  42. Learning Task-Based Analog-to-Digital Conversion for MIMO Receivers
  43. Learning With Out-of-Distribution Data for Audio Classification
  44. Learning a Generic Adaptive Wavelet Shrinkage Function for Denoising
  45. Learning a Representation for Cover Song Identification Using Convolutional Neural Network
  46. Learning a Subword Inventory Jointly with End-to-End Automatic Speech Recognition
  47. Learning connectivity and higher-order interactions in radial distribution grids
  48. Learning from Dances: Pose-Invariant Re-Identification for Multi-Person Tracking
  49. Learning the Helix Topology of Musical Pitch
  50. Learning the Spatio-Temporal Dynamics of Physical Processes from Partial Observations
  51. Learning to Characterize Adversarial Subspaces
  52. Learning to Detect Keyword Parts and Whole by Smoothed Max Pooling
  53. Learning to Estimate Driver Drowsiness from Car Acceleration Sensors Using Weakly Labeled Data
  54. Learning to Fool the Speaker Recognition
  55. Learning to Generate Diverse Questions from Keywords
  56. Learning to Rank Music Tracks Using Triplet Loss
  57. Learning to Separate Sounds from Weakly Labeled Scenes
  58. Learning-Aided Content Placement in Caching-Enabled fog Computing Systems Using Thompson Sampling
  59. Learning-Based Content Caching and User Clustering: A Deep Deterministic Policy Gradient Approach
  60. Least-Squares DOA Estimation with an Informed Phase Unwrapping and Full Bandwidth Robustness
  61. Levenberg-Marquardt and Line-Search Extended Kalman Smoothers
  62. Leveraging Cuboids for Better Motion Modeling in High Efficiency Video Coding
  63. Leveraging Gans to Improve Continuous Path Keyboard Input Models
  64. Leveraging Ordinal Regression With Soft Labels For 3d Head Pose Estimation From Point Sets
  65. Leveraging Unpaired Text Data for Training End-To-End Speech-to-Intent Systems
  66. Libri-Adapt: a New Speech Dataset for Unsupervised Domain Adaptation
  67. Libri-Light: A Benchmark for ASR with Limited or No Supervision
  68. Lie Group State Estimation via Optimal Transport
  69. Lifter Training and Sub-Band Modeling for Computationally Efficient and High-Quality Voice Conversion Using Spectral Differentials
  70. Light-Field Reconstruction and Depth Estimation from Focal Stack Images Using Convolutional Neural Networks
  71. Lightdet: A Lightweight and Accurate Object Detection Network
  72. Lightweight Hardware Implementation of VVC Transform Block for ASIC Decoder
  73. Lightweight V-Net for Liver Segmentation
  74. Lightweight and Efficient End-To-End Speech Recognition Using Low-Rank Transformer
  75. Limitations of Weak Labels for Embedding and Tagging
  76. Line Spectral Estimation with Palindromic Kernels
  77. Linear Model-Based Intra Prediction in VVC Test Model
  78. Linear Speedup in Saddle-Point Escape for Decentralized Non-Convex Optimization
  79. Linear Thompson Sampling Under Unknown Linear Constraints
  80. Lipreading Using Temporal Convolutional Networks
  81. Load Management with Predictions of Solar Energy Production for Cloud Data Centers
  82. Local Key Estimation In Classical Music Recordings: A Cross-Version Study on Schubert's Winterreise
  83. Local-Global Feature for Video-Based One-Shot Person Re-Identification
  84. Location-Relative Attention Mechanisms for Robust Long-Form Speech Synthesis
  85. Look Globally, Age Locally: Face Aging With an Attention Mechanism
  86. Lookahead Converges to Stationary Points of Smooth Non-convex Functions
  87. Looking Enhances Listening: Recovering Missing Speech Using Images
  88. Low Complexity NLMS for Multiple Loudspeaker Acoustic ECHO Canceller Using Relative Loudspeaker Transfer Functions
  89. Low Complexity Single Image Super-Resolution with Channel Splitting and Fusion Network
  90. Low Mutual and Average Coherence Dictionary Learning Using Convex Approximation
  91. Low Rank Activations for Tensor-Based Convolutional Sparse Coding
  92. Low-Complexity 5g Slam with CKF-PHD Filter
  93. Low-Complexity Accurate Mmwave Positioning for Single-Antenna Users Based on Angle-of-Departure and Adaptive Beamforming
  94. Low-Complexity Compressed Alignment-Aided Compressive Analysis for Real-Time Electrocardiography Telemonitoring
  95. Low-Complexity Fixed-Point Convolutional Neural Networks For Automatic Target Recognition
  96. Low-Complexity LSTM-Assisted Bit-Flipping Algorithm For Successive Cancellation List Polar Decoder
  97. Low-Complexity Levenberg-Marquardt Algorithm for Tensor Canonical Polyadic Decomposition
  98. Low-Complexity and Reliable Transforms for Physical Unclonable Functions
  99. Low-Frequency Compensated Synthetic Impulse Responses For Improved Far-Field Speech Recognition
  100. Low-Latency Lightweight Streaming Speech Recognition with 8-Bit Quantized Simple Gated Convolutional Neural Networks
  101. Low-Latency Single Channel Speech Enhancement Using U-Net Convolutional Neural Networks
  102. Low-Rank Approximation of Matrices Via A Rank-Revealing Factorization with Randomization
  103. Low-Rank Gradient Approximation for Memory-Efficient on-Device Training of Deep Neural Network
  104. Low-Rank MMWAVE MIMO Channel Estimation in One-Bit Receivers
  105. Low-Rank Tensor Ring Model for Completing Missing Visual Data
  106. Low-Rank Toeplitz Matrix Estimation Via Random Ultra-Sparse Rulers
  107. Low-Tubal-Rank Tensor Recovery From One-Bit Measurements
  108. Low-bit Quantization of Recurrent Neural Network Language Models Using Alternating Direction Methods of Multipliers
  109. Lqaid: Localized Quality Aware Image Denoising Using Deep Convolutional Neural Networks
  110. Lupulus: A Flexible Hardware Accelerator for Neural Networks
  111. M-Estimators of Scatter with Eigenvalue Shrinkage
  112. MDR-SURV: A Multi-Scale Deep Learning-Based Radiomics for Survival Prediction in Pulmonary Malignancies
  113. ML and EM Estimation of Sampling Intervals of Sensor Devices
  114. MMSE-Based Channel Estimation for Hybrid Beamforming Massive MIMO with Correlated Channels
  115. MSPNET: Multi-Supervised Parallel Network for Crowd Counting
  116. Mahalanobis Distance Based Adversarial Network for Anomaly Detection
  117. Manet: Multi-Scale Aggregated Network For Light Field Depth Estimation
  118. Mango: A Python Library for Parallel Hyperparameter Tuning
  119. Manifold Gradient Descent Solves Multi-Channel Sparse Blind Deconvolution Provably and Efficiently
  120. Many-To-Many Voice Conversion Using Conditional Cycle-Consistent Adversarial Networks
  121. Mask-Dependent Phase Estimation for Monaural Speaker Separation
  122. Masking and Inpainting: A Two-Stage Speech Enhancement Approach for Low SNR and Non-Stationary Noise
  123. Matching Pursuit Based Dynamic Phase-Amplitude Coupling Measure
  124. Maximally Energy-Concentrated Differential Window for Phase-Aware Signal Processing Using Instantaneous Frequency
  125. Maximum Likelihood Estimation of the Interference-Plus-Noise Cross Power Spectral Density Matrix for Own Voice Retrieval
  126. Maximum Likelihood Multi-Speaker Direction of Arrival Estimation Utilizing a Weighted Histogram
  127. Maxpolynomial Division with Application To Neural Network Simplification
  128. Media Classification with Bayesian Optimization and Vapnik-Chervonenkis (VC) Bounds
  129. Mellotron: Multispeaker Expressive Voice Synthesis by Conditioning on Rhythm, Pitch and Global Style Tokens
  130. Mental Fatigue Prediction from Multi-Channel ECOG Signal
  131. Message Transmission ThroughUnderspread Time-Varying Linear Channels
  132. Meta Learning for End-To-End Low-Resource Speech Recognition
  133. Meta Metric Learning for Highly Imbalanced Aerial Scene Classification
  134. Meta-Learning Extractors for Music Source Separation
  135. Meta-Learning for Robust Child-Adult Classification from Speech
  136. Meta-Learning to Communicate: Fast End-to-End Training for Fading Channels
  137. Metric Learning with Background Noise Class for Few-Shot Detection of Rare Sound Events
  138. Metric Representations of Networks: A Uniqueness Result
  139. Minimal Adversarial Perturbations in Mobile Health Applications: The Epileptic Brain Activity Case Study
  140. Minimum Latency Training Strategies for Streaming Sequence-to-Sequence ASR
  141. Mining Effective Negative Training Samples for Keyword Spotting
  142. Mirrored Arrays for Direction-of-Arrival Estimation
  143. Misspecified Cramer-Rao Bound For Delay Estimation with a Mismatched Waveform: A Case Study
  144. Mixture Factorized Auto-Encoder for Unsupervised Hierarchical Deep Factorization of Speech Signal
  145. Mixup Multi-Attention Multi-Tasking Model for Early-Stage Leukemia Identification
  146. Mixup-breakdown: A Consistency Training Method for Improving Generalization of Speech Separation Models
  147. MoGA: Searching Beyond Mobilenetv3
  148. Mobility-Aware Beam Steering in Metasurface-Based Programmable Wireless Environments
  149. Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders
  150. Model Order Selection in DoA Scenarios via Cross-entropy Based Machine Learning Techniques
  151. Modeling Behavior as Mutual Dependency between Physiological Signals and Indoor Location in Large-Scale Wearable Sensor Study
  152. Modeling Behavioral Consistency in Large-Scale Wearable Recordings of Human Bio-Behavioral Signals
  153. Modeling Piece-Wise Stationary Time Series
  154. Modeling Plate and Spring Reverberation Using A DSP-Informed Deep Neural Network
  155. Modeling Uncertainty in Predicting Emotional Attributes from Spontaneous Speech
  156. Modeling the Environment in Deep Reinforcement Learning: The Case of Energy Harvesting Base Stations
  157. Modelling Sea Clutter In Sar Images Using Laplace-Rician Distribution
  158. Monaural Speech Enhancement Using Intra-Spectral Recurrent Layers in the Magnitude and Phase Responses
  159. Motion Dynamics Improve Speaker-Independent Lipreading
  160. Motion Feedback Design for Video Frame Interpolation
  161. Mspec-Net : Multi-Domain Speech Conversion Network
  162. Mt-Gcn For Multi-Label Audio Tagging With Noisy Labels
  163. Multi Image Depth from Defocus Network with Boundary Cue for Dual Aperture Camera
  164. Multi-Agent Deep Reinforcement Learning For Distributed Handover Management In Dense MmWave Networks
  165. Multi-Branch Learning for Weakly-Labeled Sound Event Detection
  166. Multi-Channel Speech Source Separation and Dereverberation With Sequential Integration of Determined and Underdetermined Models
  167. Multi-Conditioning and Data Augmentation Using Generative Noise Model for Speech Emotion Recognition in Noisy Conditions
  168. Multi-Depth Computational Periscopy with an Ordinary Camera
  169. Multi-Head Attention for Speech Emotion Recognition with Auxiliary Learning of Gender Recognition
  170. Multi-Label Consistent Convolutional Transform Learning: Application to Non-Intrusive Load Monitoring
  171. Multi-Label Sound Event Retrieval Using A Deep Learning-Based Siamese Structure With A Pairwise Presence Matrix
  172. Multi-Layer Content Interaction Through Quaternion Product for Visual Question Answering
  173. Multi-Level Deep Neural Network Adaptation for Speaker Verification Using MMD and Consistency Regularization
  174. Multi-Microphone Complex Spectral Mapping for Speech Dereverberation
  175. Multi-Modal Self-Supervised Pre-Training for Joint Optic Disc and Cup Segmentation in Eye Fundus Images
  176. Multi-MotifGAN (MMGAN): Motif-Targeted Graph Generation And Prediction
  177. Multi-Patch Aggregation Models for Resampling Detection
  178. Multi-Polarization Information Fusion for Object Contour Display in Passive Millimeter-Wave and Terahertz Security Imaging
  179. Multi-Resolution Multi-Head Attention in Deep Speaker Embedding
  180. Multi-Resolution Overlapping Stripes Network for Person Re-Identification
  181. Multi-Scale Deep Feature Fusion for Vehicle Re-Identification
  182. Multi-Scale Octave Convolutions for Robust Speech Recognition
  183. Multi-Scale Residual Network for Image Classification
  184. Multi-Speaker and Multi-Domain Emotional Voice Conversion Using Factorized Hierarchical Variational Autoencoder
  185. Multi-Stage Residual Hiding for Image-Into-Audio Steganography
  186. Multi-Step Online Unsupervised Domain Adaptation
  187. Multi-Task Center-Of-Pressure Metrics Estimation from Skeleton Using Graph Convolutional Network
  188. Multi-Task Learning Via SA-FPN and EJ-Head
  189. Multi-Task Learning for Speaker Verification and Voice Trigger Detection
  190. Multi-Task Learning for Voice Trigger Detection
  191. Multi-Task Learning in Autonomous Driving Scenarios Via Adaptive Feature Refinement Networks
  192. Multi-Task Self-Supervised Learning for Robust Speech Recognition
  193. Multi-Time-Scale Convolution for Emotion Recognition from Speech Audio Signals
  194. Multi-View Bayesian Generative Model for Multi-Subject FMRI Data on Brain Decoding of Viewed Image Categories
  195. Multi-View Clustering Via Mixed Embedding Approximation
  196. Multi-View Shape Estimation of Transparent Containers
  197. Multi-View Wasserstein Discriminant Analysis with Entropic Regularized Wasserstein Distance
  198. Multi-Way Multi-View Deep Autoencoder for Image Feature Learning with Multi-Level Graph Regularization
  199. Multi-constraint Spectral Co-design for Colocated MIMO Radar and MIMO Communications
  200. Multichannel Active Noise Control with Spatial Derivative Constraints to Enlarge the Quiet Zone
  201. Multichannel Signal Classification Using Vector Autoregression
  202. Multichannel Signal Processing for Road Surface Identification
  203. Multigraph Spectral Clustering for Joint Content Delivery and Scheduling in Beam-Free Satellite Communications
  204. Multilinear Generalized Singular Value Decomposition (Ml-gsvd) with Application to Coordinated Beamforming in Multi-user Mimo Systems
  205. Multilingual Acoustic Word Embedding Models for Processing Zero-resource Languages
  206. Multilingual Grapheme-To-Phoneme Conversion with Byte Representation
  207. Multimodal Active Speaker Detection and Virtual Cinematography for Video Conferencing
  208. Multimodal Learning for Classroom Activity Detection
  209. Multimodal Speaker Diarization of Real-World Meetings Using D-Vectors With Spatial Features
  210. Multimodal Transformer Fusion for Continuous Emotion Recognition
  211. Multimodal Violence Detection in Videos
  212. Multiple Points Input For Convolutional Neural Networks in Replay Attack Detection
  213. Multispectral Fusion of RGB and NIR Images Using Weighted Least Squares and Alternating Guidance
  214. Multistate Encoding with End-To-End Speech RNN Transducer Network
  215. Multitaper Spectral Granger Causality with Application to Ssvep
  216. Multitask Learning and Multistage Fusion for Dimensional Audiovisual Emotion Recognition
  217. Multitask Learning with Capsule Networks for Speech-to-Intent Applications
  218. Multiuser Massive Mimo Downlink Precoding Using Second-Order Spatial Sigma-Delta Modulation
  219. Multivariate Tropical Regression and Piecewise-Linear Surface Fitting
  220. Mutual-Information-Based Sensor Placement for Spatial Sound Field Recording
  221. Nasil: Neural Architecture Search with Imitation Learning
  222. Near Capacity RCQD Constellations for PAPR Reduction of OFDM Systems
  223. Near-Optimal Interference Exploitation 1-Bit Massive MIMO Precoding Via Partial Branch-and-Bound
  224. Nearest Kronecker Product Decomposition Based Normalized Least Mean Square Algorithm
  225. Neural Attentive Multiview Machines
  226. Neural Coding Strategies for Event-Based Vision Data
  227. Neural Lattice Search for Speech Recognition
  228. Neural Network Training with Approximate Logarithmic Computations
  229. Neural Network Wiretap Code Design for Multi-Mode Fiber Optical Channels
  230. Neural Oracle Search on N-BEST Hypotheses
  231. Neural Percussive Synthesis Parameterised by High-Level Timbral Features
  232. Neural Time Warping for Multiple Sequence Alignment
  233. Neutral to Lombard Speech Conversion with Deep Learning
  234. New Metrics for Evaluating the Accuracy of Fundamental Frequency Estimation Approaches in Musical Signals
  235. No-Regret Non-Convex Online Meta-Learning
  236. Node-Asynchronous Spectral Clustering On Directed Graphs
  237. Noise-Robust Key-Phrase Detectors for Automated Classroom Feedback
  238. Non-Experts or Experts? Statistical Analyses of MOS using DSIS method
  239. Non-Gaussian BLE-Based Indoor Localization Via Gaussian Sum Filtering Coupled with Wasserstein Distance
  240. Non-Griffin-Lim Type Signal Recovery from Magnitude Spectrogram
  241. Non-Local Nested Residual Attention Network for Stereo Image Super-Resolution
  242. Non-Uniform Video Time-Lapse Method Based on Motion Scenario and Stabilization Constraint
  243. Non-parametric Community Change-points Detection in Streaming Graph Signals
  244. Noncoherent Maximum-Likelihood Detection for Ambient Backscattering Communications Over Ambient OFDM Signals
  245. Nonlinear Spatial Filtering for Multichannel Speech Enhancement in Inhomogeneous Noise Fields
  246. Normalized Least-Mean-Square Algorithms with Minimax Concave Penalty
  247. OH, JEEZ! or UH-HUH? A Listener-Aware Backchannel Predictor on ASR Transcriptions
  248. OOV Recovery with Efficient 2nd Pass Decoding and Open-vocabulary Word-level RNNLM Rescoring for Hybrid ASR
  249. Object Detection and 3d Estimation Via an FMCW Radar Using a Fully Convolutional Network
  250. Object Detection with Color and Depth Images with Multi-Reduced Region Proposal Network and Multi-Pooling
  251. Object Surface Estimation from Radar Images
  252. Objective Bayesian Detection Under Spatially Correlated Gaussian Observations for Multi-Antenna Cognitive Radio Network
  253. On Binary Sequence Set Design with Applications to Automotive Radar
  254. On Cramér-Rao Lower Bounds with Random Equality Constraints
  255. On Design of Optimal Smart Meter Privacy Control Strategy Against Adversarial Map Detection
  256. On Distributed Stochastic Gradient Algorithms for Global Optimization
  257. On Distributed Stochastic Gradient Descent for Nonconvex Functions in the Presence of Byzantines
  258. On Divergence Approximations for Unsupervised Training of Deep Denoisers Based on Stein's Unbiased Risk Estimator
  259. On End-to-end Multi-channel Time Domain Speech Separation in Reverberant Environments
  260. On Exponentially Consistency of Linkage-Based Hierarchical Clustering Algorithm Using Kolmogrov-Smirnov Distance
  261. On Harmonic Approximations of Inharmonic Signals
  262. On Measuring Doppler Shifts between Tags in a Backscattering Tag-to-Tag Network with Applications in Tracking
  263. On Modeling ASR Word Confidence
  264. On Network Science and Mutual Information for Explaining Deep Neural Networks
  265. On Polar Coding For Finite Blocklength Secret Key Generation Over Wireless Channels
  266. On Regularization Parameter for L0-Sparse Covariance Fitting Based DOA Estimation
  267. On Robust Variance Filtering and Change of Variance Detection
  268. On The Choice of Graph Neural Network Architectures
  269. On The Degrees Of Freedom in Total Variation Minimization
  270. On The Frequency Domain Detection of High Dimensional Time Series
  271. On The Impact of Language Familiarity in Talker Change Detection
  272. On The Stability of Polynomial Spectral Graph Filters
  273. On Throughput of Millimeter Wave MIMO Systems with Low Resolution ADCs
  274. On the Byzantine Robustness of Clustered Federated Learning
  275. On the Effect of Reflectance on Phasor Field Non-Line-of-Sight Imaging
  276. On the Importance of Vocal Tract Constriction for Speaker Characterization: The Whispered Speech Study
  277. On the Limit Distribution of the Canonical Correlation Coefficients Between the Past and the Future of a High-Dimensional White Noise
  278. On the Opportunistic use of Commercial Ku and Ka Band Satcom Networks for Rain Rate Estimation: Potentials and Critical Issues
  279. On the Use of Rényi Entropy for Optimal Window Size Computation in the Short-Time Fourier Transform
  280. On-The-Fly Feature Selection and Classification with Application to Civic Engagement Platforms
  281. One-Bit Compressed Sensing Using Generative Models
  282. One-Bit DoA Estimation via Sparse Linear Arrays
  283. One-Bit Normalized Scatter Matrix Estimation For Complex Elliptically Symmetric Distributions
  284. One-Bit Sampling in Fractional Fourier Domain
  285. One-Shot Parametric Audio Production Style Transfer with Application to Frequency Equalization
  286. One-Shot Voice Conversion Using Star-Gan
  287. One-Shot Voice Conversion by Vector Quantization
  288. Online Channel Estimation for Hybrid Beamforming Architectures
  289. Online Community Detection by Spectral Cusum
  290. Online Graph Topology Inference with Kernels For Brain Connectivity Estimation
  291. Online Positron Emission Tomography By Online Portfolio Selection
  292. Online Tensor Completion and Free Submodule Tracking With The T-SVD
  293. Open Set Video Camera Model Verification
  294. Opendenoising: An Extensible Benchmark for Building Comparative Studies of Image Denoisers
  295. Opportunistic use of GNSS Signals to Characterize the Environment by Means of Machine Learning Based Processing
  296. Optimal Design of Energy-Efficient Cell-Free Massive Mimo: Joint Power Allocation and Load Balancing
  297. Optimal Joint Channel Estimation and Data Detection by L1-norm PCA for Streetscape IoT
  298. Optimal Laplacian Regularization for Sparse Spectral Community Detection
  299. Optimal Power Flow Using Graph Neural Networks
  300. Optimal Sampling Rate and Bandwidth of Bandlimited Signals - an Algorithmic Perspective
  301. Optimal Transport Based Change Point Detection and Time Series Segment Clustering
  302. Optimal Transport Structure of CycleGAN for Unsupervised Learning for Inverse Problems
  303. Optimal Window Design for Joint Spatial-Spectral Domain Filtering of Signals on the Sphere
  304. Optimal Window Design for W-OFDM
  305. Optimized Sensor Selection for Joint Radar-communication Systems
  306. Optimized Single Carrier Transceiver for Future Sub-TeraHertz Applications
  307. Optimizing Backscattering Coefficient Design for Minimizing BER at Monostatic MIMO reader
  308. Optimizing Bayesian Hmm Based X-Vector Clustering for the Second Dihard Speech Diarization Challenge
  309. Optimum Kernel Particle Filter for Asymmetric Laplace Noise
  310. Ordinal Learning for Emotion Recognition in Customer Service Calls
  311. Orthogonal Training for Text-Independent Speaker Verification
  312. Overcoming High Nanopore Basecaller Error Rates for DNA Storage via Basecaller-Decoder Integration and Convolutional Codes
  313. Overdetermined Independent Vector Analysis
  314. Overlap Local-SGD: An Algorithmic Approach to Hide Communication Delays in Distributed SGD
  315. Overlap-Aware Diarization: Resegmentation Using Neural End-to-End Overlapped Speech Detection
  316. Overlapped State Hidden Semi-Markov Model for Grouped Multiple Sequences
  317. PAGAN: A Phase-Adapted Generative Adversarial Networks for Speech Enhancement
  318. PEVD-Based Speech Enhancement in Reverberant Environments
  319. Paco and Paco-Dct: Patch Consensus and Its Application To Inpainting
  320. Pan: Phoneme-Aware Network for Monaural Speech Enhancement
  321. Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram
  322. Parallelizing Adam Optimizer with Blockwise Model-Update Filtering
  323. Parameter Estimation of In-City Frontal Rainfall Propagation
  324. Parsing Map Guided Multi-Scale Attention Network For Face Hallucination
  325. Partial AUC Optimization Based Deep Speaker Embeddings with Class-Center Learning for Text-Independent Speaker Verification
  326. Particle Filter with Rejection Control and Unbiased Estimator of the Marginal Likelihood
  327. Particle Filtering on the Complex Stiefel Manifold with Application to Subspace Tracking
  328. Particle Group Metropolis Methods for Tracking the Leaf Area Index
  329. Passive Intelligent Surface Assisted MIMO Powered Sustainable IoT
  330. Patch-Level Selection and Breadth-First Prediction Strategy for Reversible Data Hiding
  331. Pathloss Prediction using Deep Learning with Applications to Cellular Optimization and Efficient D2D Link Scheduling
  332. Peer To Peer Offloading With Delayed Feedback: An Adversary Bandit Approach
  333. Perception-Distortion Trade-Off with Restricted Boltzmann Machines
  334. Perceptual loss function for neural modeling of audio systems
  335. Performance Analysis for Path Attenuation Estimation of Microwave Signals Due to Rainfall and Beyond
  336. Performance Bounds for Displaced Sensor Automotive Radar Imaging
  337. Performance Comparison of Lossless Compression Strategies for Dynamic Vision Sensor Data
  338. Performance Study of a Convolutional Time-Domain Audio Separation Network for Real-Time Speech Denoising
  339. Person Identification Using Deep Convolutional Neural Networks on Short-Term Signals from Wearable Sensors
  340. Phase Reconstruction Based On Recurrent Phase Unwrapping With Deep Neural Networks
  341. Phoneme Boundary Detection Using Learnable Segmental Features
  342. Phonetic Feedback for Speech Enhancement with and Without Parallel Speech Data
  343. Phylogenetic Minimum Spanning Tree Reconstruction Using Autoencoders
  344. Pitch Estimation Via Self-Supervision
  345. Pitchnet: Unsupervised Singing Voice Conversion with Pitch Adversarial Network
  346. Pixel-Level Self-Paced Learning For Super-Resolution
  347. Pixel-Wise Linear/Nonlinear Nonnegative Matrix Factorization for Unmixing of Hyperspectral Data
  348. Playing Technique Recognition by Joint Time-Frequency Scattering
  349. Polarization Parameters Estimation with Scalar Sensor Arrays
  350. Polarizing Front Ends for Robust Cnns
  351. Polyphonic Sound Event Detection Using Transposed Convolutional Recurrent Neural Network
  352. Portfolio Cuts: A Graph-Theoretic Framework to Diversification
  353. Position Constraint Loss For Fashion Landmark Estimation
  354. Positive Semidefinite Matrix Factorization: A Link to Phase Retrieval And A Block Gradient Algorithm
  355. Power Spectrum Optimization for Capacity of the Extended Spectrum Hybrid Fiber Coax Network
  356. Pre-Training for Query Rewriting in a Spoken Language Understanding System
  357. Preconditioned Ghost Imaging Via Sparsity Constraint
  358. Preconditioning ADMM for Fast Decentralized Optimization
  359. Predicting Performance Outcome with a Conversational Graph Convolutional Network for Small Group Interactions
  360. Predicting Word Error Rate for Reverberant Speech
  361. Prediction of Individual Progression Rate in Parkinson's Disease Using Clinical Measures and Biomechanical Measures of Gait and Postural Stability
  362. Prediction of Voicing and the F0 Contour from Electromagnetic Articulography Data for Articulation-to-Speech Synthesis
  363. Prediction oof Vessel Trajectories From AIS Data Via Sequence-To-Sequence Recurrent Neural Networks
  364. Preference-Aware Mask for Session-Based Recommendation with Bidirectional Transformer
  365. Preservation of Anomalous Subgroups On Variational Autoencoder Transformed Data
  366. Primal-Dual Stochastic Subgradient Method For Log-Determinant Optimization
  367. Primary Path Estimator Based on Individual Secondary Path for ANC Headphones
  368. Principal Angle Detector for Subspace Signal with Structured Unknown Interference
  369. Principle-Inspired Multi-Scale Aggregation Network for Extremely Low-Light Image Enhancement
  370. Privacy Aware Acoustic Scene Synthesis Using Deep Spectral Feature Inversion
  371. Privacy-Aware Quickest Change Detection
  372. Privacy-Preserving Image Sharing Via Sparsifying Layers on Convolutional Groups
  373. Privacy-Preserving Pattern Recognition Using Encrypted Sparse Representations in L0 Norm Minimization
  374. Privacy-Preserving Phishing Web Page Classification Via Fully Homomorphic Encryption
  375. Private FL-GAN: Differential Privacy Synthetic Data Generation Based on Federated Learning
  376. Probabilistic Filter and Smoother for Variational Inference of Bayesian Linear Dynamical Systems
  377. Processing Convolutional Neural Networks on Cache
  378. Programmable Dataflow Accelerators: A 5G OFDM Modulation/Demodulation Case Study
  379. Progressive Multi-Target Network Based Speech Enhancement with Snr-Preselection for Robust Speaker Diarization
  380. Projected Weight Regularization to Improve Neural Network Generalization
  381. Projection Free Dynamic Online Learning
  382. Propeller Noise Detection with Deep Learning
  383. Prototypical Networks for Small Footprint Text-Independent Speaker Verification
  384. Proximal Distance Algorithm for Nonconvex QCQP with Beamforming Applications
  385. Proximal Multitask Learning Over Distributed Networks with Jointly Sparse Structure
  386. Pseudo Labeling and Negative Feedback Learning for Large-Scale Multi-Label Domain Classification
  387. Pseudo Likelihood Correction Technique for Low Resource Accented ASR
  388. Pyannote.Audio: Neural Building Blocks for Speaker Diarization
  389. Q-GADMM: Quantized Group ADMM for Communication Efficient Decentralized Machine Learning
  390. Q-Learning Based Predictive Relay Selection for Optimal Relay Beamforming
  391. QOS-Aware Flow Control for Power-Efficient Data Center Networks with Deep Reinforcement Learning
  392. Quality-of-Service Prediction for Physical-layer Security via Secrecy Maps
  393. Quantized Tensor Robust Principal Component Analysis
  394. Quartznet: Deep Automatic Speech Recognition with 1D Time-Channel Separable Convolutions
  395. Quickest Change Detection In Anonymous Heterogeneous Sensor Networks
  396. Quickest Detection of Growing Dynamic Anomalies in Networks
  397. REV-AE: A Learned Frame Set for Image Reconstruction
  398. ROIMIX: Proposal-Fusion Among Multiple Images for Underwater Object Detection
  399. Rate Assignment in 360-Degree Video Tiled Streaming Using Random Forest Regression
  400. Rate-Invariant Autoencoding of Time-Series
  401. Raw Waveform Based End-to-end Deep Convolutional Network for Spatial Localization of Multiple Acoustic Sources
  402. Ray Separation and Source Depth Estimation Based on Sound Pressure Field Transformation
  403. Re-Translation Strategies for Long Form, Simultaneous, Spoken Language Translation
  404. Real-Time Binaural Speech Separation with Preserved Spatial Cues
  405. Real-Time Hand Gesture Recognition Using Temporal Muscle Activation Maps of Multi-Channel Semg Signals
  406. Real-Time Implementation Aspects of Large Intelligent Surfaces
  407. Real-Time Speech Enhancement Using Equilibriated RNN
  408. Real-Time Task Offloading for Large-Scale Mobile Edge Computing
  409. Real-Time, Universal, and Robust Adversarial Attacks Against Speaker Recognition Systems
  410. Realizability of Planar Point Embeddings from Angle Measurements
  411. Receiver Design and AGC optimization with Self Interference Induced Saturation
  412. Receptive Field Pyramid Network for Object Detection
  413. Reconstruction of Fri Signals Using Deep Neural Network Approaches
  414. Recurrent Neural Audiovisual Word Embeddings for Synchronized Speech and Real-Time Mri
  415. Recursive Prediction of Graph Signals With Incoming Nodes
  416. Reduced-Complexity Singular Value Decomposition For Tucker Decomposition: Algorithm And Hardware
  417. Redundant Convolutional Network With Attention Mechanism For Monaural Speech Enhancement
  418. Reflectance-Guided, Contrast-Accumulated Histogram Equalization
  419. Regression Before Classification for Temporal Action Detection
  420. Regularized Beamformer for the Spherical Microphone Array to Cope with the White Noise Amplification
  421. Regularized Fast Multichannel Nonnegative Matrix Factorization with ILRMA-Based Prior Distribution of Joint-Diagonalization Process
  422. Regularized Partial Phase Synchrony Index Applied to Dynamical Functional Connectivity Estimation
  423. Reinforced Depth-Aware Deep Learning for Single Image Dehazing
  424. Relative Cost Based Model Selection for Sparse High-Dimensional Linear Regression Models
  425. Reliable and Secure Transmission for Future Networks
  426. Residual Attention Network for Wavelet Domain Super-Resolution
  427. Residual Recurrent Neural Network for Speech Enhancement
  428. Resilient Distributed Recovery of Large Fields
  429. Resilient to Byzantine Attacks Finite-Sum Optimization Over Networks
  430. Resource Management in the Multibeam NOMA-based Satellite Downlink
  431. Resting-State EEG-Based Biometrics with Signals Features Extracted by Multivariate Empirical Mode Decomposition
  432. Rethinking Retinal Landmark Localization as Pose Estimation: Naïve Single Stacked Network for Optic Disk and Fovea Detection
  433. Retinal Vessel Segmentation via a Semantics and Multi-Scale Aggregation Network
  434. Retrieving Vocal-Tract Resonance and anti-Resonance From High-Pitched Vowels Using a Rahmonic Subtraction Technique
  435. Revealing Backdoors, Post-Training, in DNN Classifiers via Novel Inference on Optimized Perturbations Inducing Group Misclassification
  436. Revealing Hidden Drawings in Leonardo's 'the Virgin of the Rocks' from Macro X-Ray Fluorescence Scanning Data through Element Line Localisation
  437. Reversal No Longer Matters: Attention-Based Arrhythmia Detection with Lead-Reversal ECG Data
  438. Revisit of Estimate Sequence for Accelerated Gradient Methods
  439. Revisiting Fast Spectral Clustering with Anchor Graph
  440. Rgb-D Based Multi-Modal Deep Learning for Face Identification
  441. Riemannian Framework for Robust Covariance Matrix Estimation in Spiked Models
  442. Riemannian Geometry and Cramér-rao Bound for Blind Separation of Gaussian Sources
  443. Risk Convergence of Centered Kernel Ridge Regression with Large Dimensional Data
  444. Rnn-Transducer with Stateless Prediction Network
  445. Robust CFAR Radar Detection Using a K-nearest Neighbors Rule
  446. Robust Covariance Matrix Estimation and Portfolio Allocation: The Case of Non-Homogeneous Assets
  447. Robust Frequency-Domain Recursive Least M-Estimate Adaptive Filter For Acoustic System Identification
  448. Robust Full-Fov Depth Estimation in Tele-Wide Camera System
  449. Robust Fundamental Frequency Estimation in Coloured Noise
  450. Robust Global Optimized Affine Registration Method for Microscopic Images of Biological Tissue
  451. Robust Hybrid Beamforming for Satellite-Terrestrial Integrated Networks
  452. Robust Likelihood Ratio Test Using α-Divergence
  453. Robust Low Rate Speech Coding Based on Cloned Networks and Wavenet
  454. Robust Marine Buoy Placement for Ship Detection Using Dropout K-Means
  455. Robust Matrix Completion via ℓP-Greedy Pursuits
  456. Robust Multi-Channel Speech Recognition Using Frequency Aligned Network
  457. Robust Music Estimation Under Array Response Uncertainty
  458. Robust Online Matrix Completion with Gaussian Mixture Model
  459. Robust Online Mirror Saddle-Point Method for Constrained Resource Allocation
  460. Robust Parameter Estimation of Contaminated Damped Exponentials
  461. Robust Phase Retrieval with Outliers
  462. Robust Pricing Mechanism for Resource Sustainability Under Privacy Constraint in Competitive Online Learning Multi-Agent Systems
  463. Robust Rank Constrained Sparse Learning: A Graph-Based Method for Clustering
  464. Robust Speaker Recognition Using Unsupervised Adversarial Invariance
  465. Robust Symbol-Level Precoding Via Autoencoder-Based Deep Learning
  466. Robust Tdoa Indoor Tracking Using Constrained Measurement Filtering and Grid-Based Filtering
  467. Robust Transmission Over Channels with Channel Uncertainty: an Algorithmic Perspective
  468. Robust Unsupervised Audio-Visual Speech Enhancement Using a Mixture of Variational Autoencoders
  469. Robust Visual Tracking with Context-Based Active Occlusion Recognition
  470. Robust and Computationally-Efficient Anomaly Detection Using Powers-Of-Two Networks
  471. Robust and steerable kronecker product differential beamforming With rectangular microphone arrays
  472. Robustness Assessment of Automatic Reinke's Edema Diagnosis Systems
  473. Robustness of Sparse Bayesian Learning in Correlated Environments
  474. S-DOD-CNN: Doubly Injecting Spatially-Preserved Object Information for Event Recognition
  475. SDTCN: Similarity Driven Transmission Computing Network for Image Dehazing
  476. SECL-UMons Database for Sound Event Classification and Localization
  477. SED-MDD: Towards Sentence Dependent End-To-End Mispronunciation Detection and Diagnosis
  478. SLOGD: Speaker Location Guided Deflation Approach to Speech Separation
  479. SNDCNN: Self-Normalizing Deep CNNs with Scaled Exponential Linear Units for Speech Recognition
  480. SPIDERnet: Attention Network For One-Shot Anomaly Detection In Sounds
  481. SSGD: Sparsity-Promoting Stochastic Gradient Descent Algorithm for Unbiased Dnn Pruning
  482. SSTNet: Detecting Manipulated Faces Through Spatial, Steganalysis and Temporal Features
  483. Saliency-Based Image Contrast Enhancement with Reversible Data Hiding
  484. Salient Object Detection Based On Image Bit-Map
  485. Sampling Classes of Non-Bandlimited Signals Using Integrate-and-Fire Devices: Average Case Analysis
  486. Sampling Strategies for GAN Synthetic Data
  487. Sampling of Surfaces and Learning Functions in High Dimensions
  488. Scalable Detection and Tracking of Extended Objects
  489. Scalable Kernel Learning Via the Discriminant Information
  490. Scalable Learning-Based Sampling Optimization for Compressive Dynamic MRI
  491. Scalable Multilingual Frontend for TTS
  492. Scalpnet: Detection of Spatiotemporal Abnormal Intervals in Epileptic EEG Using Convolutional Neural Networks
  493. Scene Text Recognition with Temporal Convolutional Encoder
  494. Scene-Dependent Acoustic Event Detection with Scene Conditioning and Fake-Scene-Conditioned Loss
  495. SeCoST: : Sequential Co-Supervision for Large Scale Weakly Labeled Audio Event Detection
  496. Secure Face Recognition in Edge and Cloud Networks: From the Ensemble Learning Perspective
  497. Secure Identification for Gaussian Channels
  498. Secure Symbol-Level Miso Precoding
  499. Selection-Channel-Aware Reverse JPEG Compatibility for Highly Reliable Steganalysis of JPEG Images
  500. Selective Attention Encoders by Syntactic Graph Convolutional Networks for Document Summarization
  501. Selective Convolutional Network: An Efficient Object Detector with Ignoring Background
  502. Self-Adaptive Feature Fool
  503. Self-Attention and Retrieval Enhanced Neural Networks for Essay Generation
  504. Self-Attentive Sentimental Sentence Embedding for Sentiment Analysis
  505. Self-Driven Graph Volterra Models for Higher-Order Link Prediction
  506. Self-Paced Probabilistic Principal Component Analysis For Data With Outliers
  507. Self-Supervised Adversarial Training
  508. Self-Supervised Deep Learning for Fisheye Image Rectification
  509. Self-Supervised Denoising Autoencoder with Linear Regression Decoder for Speech Enhancement
  510. Self-Supervised Learning for Audio-Visual Speaker Diarization
  511. Self-Supervised Learning for ECG-Based Emotion Recognition
  512. Self-Training for End-to-End Speech Recognition
  513. Semantic Augmentation Hashing for Zero-Shot Image Retrieval
  514. Semanticgan: Generative Adversarial Networks For Semantic Image To Photo-Realistic Image Translation
  515. Semi-Implicit Stochastic Recurrent Neural Networks
  516. Semi-Regular Geometric Kernel Encoding & Reconstruction for Video Compression
  517. Semi-Supervised Learning Based on Hierarchical Generative Models for End-to-End Speech Synthesis
  518. Semi-Supervised Learning for Text Classification by Layer Partitioning
  519. Semi-Supervised Learning of Processes Over Multi-Relational Graphs
  520. Semi-Supervised Optimal Transport Methods for Detecting Anomalies
  521. Semi-Supervised Sentence Classification Based on User Polarity in the Social Scenarios
  522. Semi-Supervised Speaker Adaptation for End-to-End Speech Synthesis with Pretrained Models
  523. Sensor Selection for Model-Free Source Localization: where Less is More
  524. Separable Optimization for Joint Blind Deconvolution and Demixing
  525. Sequence-Level Consistency Training for Semi-Supervised End-to-End Automatic Speech Recognition
  526. Sequence-To-Subsequence Learning With Conditional Gan For Power Disaggregation
  527. Sequence-to-Sequence Automatic Speech Recognition with Word Embedding Regularization and Fused Decoding
  528. Sequence-to-Sequence Labanotation Generation Based on Motion Capture Data
  529. Sequence-to-Sequence Singing Synthesis Using the Feed-Forward Transformer
  530. Sequential Deep Unrolling With Flow Priors For Robust Video Deraining
  531. Sequential IoT Data Augmentation Using Generative Adversarial Networks
  532. Sequential Joint Detection and Estimation with an Application to Joint Symbol Decoding and Noise Power Estimation
  533. Sequential Methods for Detecting a Change in the Distribution of an Episodic Process
  534. Sequential Semi-Orthogonal Multi-Level NMF with Negative Residual Reduction for Network Embedding
  535. Sequential Vessel Trajectory Identification Using Truncated Viterbi Algorithm
  536. Shadow Removal of Text Document Images by Estimating Local and Global Background Colors
  537. Shape From Bandwidth: Central Projection Case
  538. Short and Squeezed: Accelerating the Computation of Antisparse Representations with Safe Squeezing
  539. Sight to Sound: An End-to-End Approach for Visual Piano Transcription
  540. Signal Clustering With Class-Independent Segmentation
  541. Signal Sensing and Reconstruction Paradigms for a Novel Multi-Source Static Computed Tomography System
  542. Signal-Aware Broadband DOA Estimation Using Attention Mechanisms
  543. Similarity Learning For Cover Song Identification Using Cross-Similarity Matrices of Multi-Level Deep Sequences
  544. Simple Caching Schemes for Non-homogeneous MISO Cache-Aided Communication via Convexity
  545. Simplified Dynamic SC-Flip Polar Decoding
  546. Simultaneous Separation and Transcription of Mixtures with Multiple Polyphonic and Percussive Instruments
  547. Singing Voice Conversion with Disentangled Representations of Singer and Vocal Technique Using Variational Autoencoders
  548. Single Frequency Filter Bank Based Long-Term Average Spectra for Hypernasality Detection and Assessment in Cleft Lip and Palate Speech
  549. Single-Channel Speech Separation Integrating Pitch Information Based on a Multi Task Learning Framework
  550. Single-Shot Real-Time Multiple-Path Time-of-Flight Depth Imaging for Multi-Aperture and Macro-Pixel Sensors
  551. Sketchppnet: A Joint Pixel and Point Convolutional Neural Network For Low Resolution Sketch Image Recognition
  552. SkinAugment: Auto-Encoding Speaker Conversions for Automatic Speech Translation
  553. Slicenet: Slice-Wise 3D Shapes Reconstruction from Single Image
  554. Slow-Time MIMO-FMCW Automotive Radar Detection with Imperfect Waveform Separation
  555. Small Energy Masking for Improved Neural Network Training for End-To-End Speech Recognition
  556. Small-Footprint Keyword Spotting on Raw Audio Data with Sinc-Convolutions
  557. Smoothing Graph Signals via Random Spanning Forests
  558. Snorer Diarisation Based On Deep Neural Network Embeddings
  559. Social Data Assisted Multi-Modal Video Analysis For Saliency Detection
  560. Social Learning with Partial Information Sharing
  561. Soft-Output Finite Alphabet Equalization for mmWave Massive MIMO
  562. Solving Missing-Annotation Object Detection with Background Recalibration Loss
  563. Solving Non-Convex Non-Differentiable Min-Max Games Using Proximal Gradient Method
  564. Sound Event Detection Via Dilated Convolutional Recurrent Neural Networks
  565. Sound Event Detection by Multitask Learning of Sound Events and Scenes with Soft Scene Labels
  566. Sound Event Detection in Synthetic Domestic Environments
  567. Sound Event Localization Based on Sound Intensity Vector Refined by Dnn-Based Denoising and Source Separation
  568. Sound Texture Synthesis Using RI Spectrograms
  569. Source Coding of Audio Signals with a Generative Model
  570. Source Domain Data Selection for Improved Transfer Learning Targeting Dysarthric Speech Recognition
  571. Source Enumeration via Toeplitz Matrix Completion
  572. Source Separation with Weakly Labelled Data: an Approach to Computational Auditory Scene Analysis
  573. Space Filling Curves for MRI Sampling
  574. Sparse Beamspace Equalization for Massive MU-MIMO MMWave Systems
  575. Sparse Branch and Bound for Exact Optimization of L0-Norm Penalized Least Squares
  576. Sparse CSP Algorithm via Joint Spatio-Temporal Filtering
  577. Sparse Convolutional Beamforming for Wireless Ultrasound
  578. Sparse Directed Graph Learning for Head Movement Prediction in 360 Video Streaming
  579. Sparse Low-redundancy Linear Array with Uniform Sum Co-array
  580. Sparse Modeling on Distributed Encryption Data
  581. Sparse Recovery with Non-Linear Fourier Features
  582. Spatial Active Noise Control Based on Kernel Interpolation with Directional Weighting
  583. Spatial Attention for Far-Field Speech Recognition with Deep Beamforming Neural Networks
  584. Spatial Attentional Bilinear 3D Convolutional Network for Video-Based Autism Spectrum Disorder Detection
  585. Spatial Gating Strategies for Graph Recurrent Neural Networks
  586. Spatial and Temporal Smoothing for Covariance Estimation in Super-Resolution Angle Estimation in Automotive Radars
  587. Spatial-Temporal Feature Aggregation Network For Video Object Detection
  588. Spatially Adaptive Intra Mode Pre-Selection for ERP 360 Video Coding
  589. Spatially Guided Independent Vector Analysis
  590. Spatio-Temporal and Geometry Constrained Network for Automobile Visual Odometry
  591. Speaker Adaptation of a Multilingual Acoustic Model for Cross-Language Synthesis
  592. Speaker Augmentation for Low Resource Speech Recognition
  593. Speaker Diarization Using Latent Space Clustering in Generative Adversarial Network
  594. Speaker Diarization with Region Proposal Network
  595. Speaker Diarization with Session-Level Speaker Embedding Refinement Using Graph Neural Networks
  596. Speaker Embeddings Incorporating Acoustic Conditions for Diarization
  597. Speaker Independence of Neural Vocoders and Their Effect on Parametric Resynthesis Speech Enhancement
  598. Speaker-Aware Target Speaker Enhancement by Jointly Learning with Speaker Embedding Extraction
  599. Speaker-Aware Training of Attention-Based End-to-End Speech Recognition Using Neural Speaker Embeddings
  600. Speaker-Invariant Affective Representation Learning via Adversarial Training
  601. Speakerfilter: Deep Learning-Based Target Speaker Extraction Using Anchor Speech
  602. Specaugment on Large Scale Datasets
  603. Spectrogram Analysis Via Self-Attention for Realizing Cross-Model Visual-Audio Generation
  604. Spectrograms Fusion with Minimum Difference Masks Estimation for Monaural Speech Dereverberation
  605. Spectrum Allocation in Wireless Networks for Crowd Labelling
  606. Speech Breathing Estimation Using Deep Learning Methods
  607. Speech Emotion Recognition with Dual-Sequence LSTM Architecture
  608. Speech Emotion Recognition with Local-Global Aware Deep Representation Learning
  609. Speech Enhancement Using Self-Adaptation and Multi-Head Self-Attention
  610. Speech Intelligibility Enhancement by Equalization for in-Car Applications
  611. Speech Recognition Model Compression
  612. Speech Sentiment Analysis via Pre-Trained Features from End-to-End ASR Models
  613. Speech Synthesis Using EEG
  614. Speech-Based Parameter Estimation of an Asymmetric Vocal Fold Oscillation Model and its Application in Discriminating Vocal Fold Pathologies
  615. Speech-Driven Facial Animation Using Polynomial Fusion of Features
  616. Speech-To-Singing Conversion in an Encoder-Decoder Framework
  617. Spherical Large Intelligent Surfaces
  618. Spherical Video Coding with Geometry and Region Adaptive Transform Domain Temporal Prediction
  619. Spiking Neural Networks Trained With Backpropagation for Low Power Neuromorphic Implementation of Voice Activity Detection
  620. Spoken Document Retrieval Leveraging Bert-Based Modeling and Query Reformulation
  621. Spoken Language Acquisition Based on Reinforcement Learning and Word Unit Segmentation
  622. Srzoo: An Integrated Repository For Super-Resolution Using Deep Learning
  623. Stability of Graph Neural Networks to Relative Perturbations
  624. Stabilizing Multi-Agent Deep Reinforcement Learning by Implicitly Estimating Other Agents' Behaviors
  625. Stable Training of Dnn for Speech Enhancement Based on Perceptually-Motivated Black-Box Cost Function
  626. Stacked Pooling for Boosting Scale Invariance of Crowd Counting
  627. Staged Training Strategy and Multi-Activation for Audio Tagging with Noisy and Sparse Multi-Label Data
  628. Stargan for Emotional Speech Conversion: Validated by Data Augmentation of End-To-End Emotion Recognition
  629. State-Based Transcription of Components of Carnatic Music
  630. State-Space Gaussian Process for Drift Estimation in Stochastic Differential Equations
  631. Static Visual Spatial Priors for DoA Estimation
  632. Statistical Signal Processing Approach for Rain Estimation Based on Measurements from Network Management Systems
  633. Statistics Pooling Time Delay Neural Network Based on X-Vector for Speaker Verification
  634. Steepening Squared Error Function Facilitates Online Adaptation of Gaussian Scales
  635. Steganography and its Detection in JPEG Images Obtained with the "TRUNC" Quantizer
  636. Stochastic Admm For Byzantine-Robust Distributed Learning
  637. Stochastic Geometry Planning of Electric Vehicles Charging Stations
  638. Stochastic Graph Neural Networks
  639. Stochastic Ml Estimation for Hyperspectral Unmixing Under Endmember Variability and Nonlinear Models
  640. Stochastic Multi-Scale Aggregation Network for Crowd Counting
  641. Stock Movement Prediction That Integrates Heterogeneous Data Sources Using Dilated Causal Convolution Networks with Attention
  642. Storing Digital Data Into DNA: A Comparative Study Of Quaternary Code Construction
  643. Strategic Attention Learning for Modality Translation
  644. Streaming Automatic Speech Recognition with the Transformer Model
  645. Structural Sparsification for Far-Field Speaker Recognition with Intel® Gna
  646. Structured Citation Trend Prediction Using Graph Neural Networks
  647. Structured Sparse Attention for end-to-end Automatic Speech Recognition
  648. Study of Closed Phase Resonance Bandwidths for Oral and Nasal Tracts Using Zero Time Windowing
  649. Study of Formant Modification for Children ASR
  650. Sub-Dip: Optimization On A Subspace With Deep Image Prior Regularization And Application To Superresolution
  651. Subject Transfer Framework Based on Source Selection and Semi-Supervised Style Transfer Mapping for Semg Pattern Recognition
  652. Subjective Quality Estimation Using PESQ For Hands-Free Terminals
  653. Submodular Rank Aggregation on Score-Based Permutations for Distributed Automatic Speech Recognition
  654. Subspace-Based Speech Correlation Vector Estimation for Single-Microphone Multi-Frame MVDR Filtering
  655. Super-Resolution of 3D Color Point Clouds Via Fast Graph Total Variation
  656. Super-Resolution with Noisy Measurements: Reconciling Upper and Lower Bounds
  657. Superpixel Segmentation Via Convolutional Neural Networks with Regularized Information Maximization
  658. Supervised Canonical Correlation Analysis of Data on Symmetric Positive Definite Manifolds by Riemannian Dimensionality Reduction
  659. Supervised Deep Hashing for Efficient Audio Event Retrieval
  660. Supervised Encoding for Discrete Representation Learning
  661. Supervised Graph Representation Learning for Modeling the Relationship between Structural and Functional Brain Connectivity
  662. Supervised Online Diarization with Sample Mean Loss for Multi-Domain Data
  663. Synchronous Transformers for end-to-end Speech Recognition
  664. Synthesizing Engaging Music Using Dynamic Models of Statistical Surprisal
  665. Synthetic Crowd and Pedestrian Generator for Deep Learning Problems
  666. Synthetic Data Generation Through Statistical Explosion: Improving Classification Accuracy of Coronary Artery Disease Using PPG
  667. Synthetic Speech References for Automatic Pathological Speech Intelligibility Assessment
  668. T-GSA: Transformer with Gaussian-Weighted Self-Attention for Speech Enhancement
  669. TDMF: Task-Driven Multilevel Framework for End-to-End Speaker Verification
  670. TOSO: Student's-T Distribution Aided One-Stage Orientation Target Detection in Remote Sensing Images
  671. Tackling Real Noisy Reverberant Meetings with All-Neural Source Separation, Counting, and Diarization System
  672. Talker-Independent Speaker Separation in Reverberant Conditions
  673. Target Parameter Estimation via One-Bit PMCW Radar
  674. Task-Aware Mean Teacher Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection
  675. Teacher-Student Training For Robust Tacotron-Based TTS
  676. Teaching Signals and Systems - A First Course in Signal Processing
  677. Temporal Coding in Spiking Neural Networks with Alpha Synaptic Function
  678. Tensor Decomposition-based Beamspace Esprit Algorithm for Multidimensional Harmonic Retrieval
  679. Tensor-To-Vector Regression for Multi-Channel Speech Enhancement Based on Tensor-Train Network
  680. Tensorflow Audio Models in Essentia
  681. Texception: A Character/Word-Level Deep Learning Model for Phishing URL Detection
  682. Text Adaptation for Speaker Verification with Speaker-Text Factorized Embeddings
  683. Text-Independent Speaker Verification with Adversarial Learning on Short Utterances
  684. Text-To-Image Synthesis Method Evaluation Based On Visual Patterns
  685. The Compressed Nested Array for Underdetermined DOA Estimation by Fourth-order Difference Coarrays
  686. The Discrete Stockwell Transforms for Infinite-Length Signals and Their Real-Time Implementations
  687. The Effect of Data Augmentation on Classification of Atrial Fibrillation in Short Single-Lead ECG Signals Using Deep Neural Networks
  688. The Effect of Power Allocation on Visible Light Communication Using Commercial Phosphor-Converted Led Lamp for Indirect Illumination
  689. The Empirical Duality Gap of Constrained Statistical Learning
  690. The Fifthnet Chroma Extractor
  691. The Fractional Quaternion Fourier Number Transform
  692. The Graphon Fourier Transform
  693. The Matched Reassigned Cross-Spectrogram for Phase Estimation
  694. The Open Brands Dataset: Unified Brand Detection and Recognition at Scale
  695. The Picasso Algorithm for Bayesian Localization Via Paired Comparisons in a Union of Subspaces Model
  696. The Processing of Mandarin Chinese Tonal Alternations in Contexts: An Eye-Tracking Study
  697. The Role of Annotation Fusion Methods in the Study of Human-Reported Emotion Experience During Music Listening
  698. The Rwth Asr System for Ted-Lium Release 2: Improving Hybrid Hmm With Specaugment
  699. The Sound of My Voice: Speaker Representation Loss for Target Voice Separation
  700. Theoretical Analysis of Multi-Carrier Agile Phased Array Radar
  701. Theoretical Performance Bound of Uplink Channel Estimation Accuracy in Massive MIMO
  702. This Dataset Does Not Exist: Training Models from Generated Images
  703. Threshold-Adjusted ORB Strategies with Genetic Algorithm and Protective Closing Strategy on Taiwan Futures Market
  704. Time Difference of Arrival Estimation from Frequency-Sliding Generalized Cross-Correlations Using Convolutional Neural Networks
  705. Time Domain Velocity Vector for Retracing the Multipath Propagation
  706. Time Reversal Based Robust Gesture Recognition Using Wifi
  707. Time-Domain Audio Source Separation Based on Wave-U-Net Combined with Discrete Wavelet Transform
  708. Time-Domain Neural Network Approach for Speech Bandwidth Extension
  709. Time-Frequency Analysis of Unimodal Sensory Processing In Autism Spectrum Disorder
  710. Time-Frequency Feature Decomposition Based on Sound Duration for Acoustic Scene Classification
  711. Time-Frequency Loss for CNN Based Speech Super-Resolution
  712. Time-Predictable Software-Defined Architecture with Sdf-Based Compiler Flow for 5g Baseband Processing
  713. Time-Scale Synthesis for Locally Stationary Signals
  714. Toward Better Speaker Embeddings: Automated Collection of Speech Samples From Unknown Distinct Speakers
  715. Towards Blind Quality Assessment of Concert Audio Recordings Using Deep Neural Networks
  716. Towards Data-Efficient Modeling for Wake Word Spotting
  717. Towards Decoding Selective Attention from Single-Trial EEG Data in Cochlear Implant users based on Deep Neural Networks
  718. Towards Fast and Accurate Streaming End-To-End ASR
  719. Towards High-Performance Object Detection: Task-Specific Design Considering Classification and Localization Separation
  720. Towards Linking the Lakh and IMSLP Datasets
  721. Towards Multilingual Sign Language Recognition
  722. Towards Pose-Invariant Lip-Reading
  723. Towards Real-Time Single-Channel Singing-Voice Separation with Pruned Multi-Scaled Densenets
  724. Towards Real-Time, Multi-View Video Stereopsis
  725. Towards Unsupervised Speech Recognition and Synthesis with Quantized Speech Representation Learning
  726. Towards a New Understanding of the Training of Neural Networks with Mislabeled Training Data
  727. Towards an Efficient and General Framework of Robust Training for Graph Neural Networks
  728. Towards an Intelligent Microscope: Adaptively Learned Illumination for Optimal Sample Classification
  729. Trace Norm Generative Adversarial Networks for Sensor Generation and Feature Extraction
  730. Tracing Network Evolution Using The Parafac2 Model
  731. Track-Before-Detect for Sub-Nyquist Radar
  732. Tracking to Improve Detection Quality in Lidar For Autonomous Driving
  733. Training ASR Models By Generation of Contextual Information
  734. Training Code-Switching Language Model with Monolingual Data
  735. Training Deep Spiking Neural Networks for Energy-Efficient Neuromorphic Computing
  736. Training Keyword Spotters with Limited and Synthesized Speech Data
  737. Training LSTM for Unsupervised Anomaly Detection Without A Priori Knowledge
  738. Training Spoken Language Understanding Systems with Non-Parallel Speech and Text
  739. Transfer Learning from Youtube Soundtracks to Tag Arctic Ecoacoustic Recordings
  740. Transferable Policies for Large Scale Wireless Networks with Graph Neural Networks
  741. Transferring Neural Speech Waveform Synthesizers to Musical Instrument Sounds Generation
  742. Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss
  743. Transformer VAE: A Hierarchical Model for Structure-Aware and Interpretable Music Representation Learning
  744. Transformer-Based Acoustic Modeling for Hybrid Speech Recognition
  745. Transformer-Based Online CTC/Attention End-To-End Speech Recognition Architecture
  746. Transformer-Based Text-to-Speech with Weighted Forced Attention
  747. Transforming Seismocardiograms Into Electrocardiograms by Applying Convolutional Autoencoders
  748. Translation of a Higher Order Ambisonics Sound Scene Based on Parametric Decomposition
  749. Transmit Beamforming Design with Received-Interference Power Constraints: The Zero-Forcing Relaxation
  750. Transmit Beampattern Shaping via Waveform Design in Cognitive Mimo Radar
  751. Trapezoidal Segment Sequencing: A Novel Approach for Fusion of Human-Produced Continuous Annotations
  752. Tree of Shapes Cut for Material Segmentation Guided by a Design
  753. Triggerless Random Interleaved Sampling
  754. Trilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention Mechanisms
  755. Triplet Loss Feature Aggregation for Scalable Hash
  756. Truth-to-Estimate Ratio Mask: A Post-Processing Method for Speech Enhancement Direct at Low Signal-to-Noise Ratios
  757. Ts-Fen: Probing Feature Selection Strategy for Face Anti-Spoofing
  758. Two-Element Biomimetic Antenna Array Design and Performance
  759. Two-Step Acoustic Model Adaptation for Dysarthric Speech Recognition
  760. Two-Step Sound Source Separation: Training On Learned Latent Targets
  761. Two-dimensional DOA Estimation for Coprime Planar Array: A Coarray Tensor-based Solution
  762. UNet 3+: A Full-Scale Connected UNet for Medical Image Segmentation
  763. Uncertainties in Short Commercial Microwave Links Fading Due to Rain
  764. Uncertainty Quantification for Remaining Useful Lifetime Prediction with Multi-Channel Sensory Data
  765. Underwater Tracking Based on the Sum-Product Algorithm Enhanced by a Neural Network Detections Classifier
  766. Unified Signal Compression Using Generative Adversarial Networks
  767. Universal Phone Recognition with a Multilingual Allophone System
  768. Unseen Face Presentation Attack Detection with Hypersphere Loss
  769. Unsupervised Auto-Encoding Multiple-Object Tracker for Constraint-Consistent Combinatorial Problem
  770. Unsupervised Change Detection for Multimodal Remote Sensing Images via Coupled Dictionary Learning and Sparse Coding
  771. Unsupervised Content-Preserved Adaptation Network for Classification of Pulmonary Textures from Different CT Scanners
  772. Unsupervised Domain Adaptation for Semantic Segmentation with Symmetric Adaptation Consistency
  773. Unsupervised Feature Enhancement for Speaker Verification
  774. Unsupervised Image-to-Image Translation Via Fair Representation of Gender Bias
  775. Unsupervised Key Hand Shape Discovery of Sign Language Videos with Correspondence Sparse Autoencoders
  776. Unsupervised Multiple Source Localization Using Relative Harmonic Coefficients
  777. Unsupervised Neural Mask Estimator for Generalized Eigen-Value Beamforming Based Asr
  778. Unsupervised Person Re-Identification Using Multi-Branch Feature Compensation Network and Link-Based Cluster Dissimilarity Metric
  779. Unsupervised Pre-Training of Bidirectional Speech Encoders via Masked Reconstruction
  780. Unsupervised Pretraining Transfers Well Across Languages
  781. Unsupervised Speaker Adaptation Using Attention-Based Speaker Memory for End-to-End ASR
  782. Unsupervised Style and Content Separation by Minimizing Mutual Information for Speech Synthesis
  783. Unsupervised Training for Deep Speech Source Separation with Kullback-Leibler Divergence Based Probabilistic Loss Function
  784. Unsupervised Variational Bayesian Kalman Filtering For Large-Dimensional Gaussian Systems
  785. Upgrade Methods for Stratified Sensor Network Self-Calibration
  786. Upgrading CRFS to JRFS and its Benefits to Sequence Modeling and Labeling
  787. Upscaling Vector Approximate Message Passing
  788. Urtis: a Small 3d Imaging Sonar Sensor for Robotic Applications
  789. Using Automatic Speech Recognition and Speech Synthesis to Improve the Intelligibility of Cochlear Implant users in Reverberant Listening Environments
  790. Using Intelligent Reflecting Surfaces for Rank Improvement in MIMO Communications
  791. Using Panoramic Videos for Multi-Person Localization and Tracking In A 3D Panoramic Coordinate
  792. Using Personalized Speech Synthesis and Neural Language Generator for Rapid Speaker Adaptation
  793. Using Separate Losses for Speech and Noise in Mask-Based Speech Enhancement
  794. Using Speech Synthesis to Train End-To-End Spoken Language Understanding Models
  795. Using Vaes and Normalizing Flows for One-Shot Text-To-Speech Synthesis of Expressive Speech
  796. Using X-Vectors to Automatically Detect Parkinson's Disease from Speech
  797. Utterance-Level Sequential Modeling for Deep Gaussian Process Based Speech Synthesis Using Simple Recurrent Unit
  798. VAMP with Vector-Valued Diagonalization
  799. Vapar Synth - A Variational Parametric Model for Audio Synthesis
  800. Variable Bitrate Image Compression with Quality Scaling Factors
  801. Variable Metric Proximal Gradient Method with Diagonal Barzilai-Borwein Stepsize
  802. Variable Projection for Multiple Frequency Estimation
  803. Variational Student: Learning Compact and Sparser Networks In Knowledge Distillation Framework
  804. Versatile Video Coding and Super-Resolution for Efficient Delivery of 8k Video with 4k Backward-Compatibility
  805. Vggsound: A Large-Scale Audio-Visual Dataset
  806. ViMo: Vital Sign Monitoring Using Commodity Millimeter Wave Radio
  807. Video Deblurring Via 3d CNN and Fourier Accumulation Learning
  808. Video Frame Interpolation Via Exceptional Motion-Aware Synthesis
  809. Video Frame Interpolation Via Residue Refinement
  810. Video Question Generation via Semantic Rich Cross-Modal Self-Attention Networks Learning
  811. View-Angle Invariant Object Monitoring Without Image Registration
  812. Visually Guided Self Supervised Learning of Speech Representations
  813. Vocal Tract Articulatory Contour Detection in Real-Time Magnetic Resonance Images Using Spatio-Temporal Context
  814. Voice Conversion with Transformer Network
  815. Voice based classification of patients with Amyotrophic Lateral Sclerosis, Parkinson's Disease and Healthy Controls with CNN-LSTM using transfer learning
  816. Voiceai Systems to NIST Sre19 Evaluation: Robust Speaker Recognition on Conversational Telephone Speech
  817. Volume Reconstruction for Light Field Microscopy
  818. WHAMR!: Noisy and Reverberant Single-Channel Speech Separation
  819. WaveFFJORD: FFJORD-Based Vocoder for Statistical Parametric Speech Synthesis
  820. Wawenets: A No-Reference Convolutional Waveform-Based Approach to Estimating Narrowband and Wideband Speech Quality
  821. Weakly Labelled Audio Tagging Via Convolutional Networks with Spatial and Channel-Wise Attention
  822. Weakly Supervised Crowd-Wise Attention For Robust Crowd Counting
  823. Weakly Supervised Segmentation Guided Hand Pose Estimation During Interaction with Unknown Objects
  824. Weakly Supervised Semantic Segmentation For Remote Sensing Hyperspectral Imaging
  825. Weakly-Supervised Sound Event Detection with Self-Attention
  826. Weight Sharing and Deep Learning for Spectral Data
  827. Weighted Gradient Coding with Leverage Score Sampling
  828. Weighted Krylov-Levenberg-Marquardt Method for Canonical Polyadic Tensor Decomposition
  829. Weighted Null Vector Initialization and its Application to Phase Retrieval
  830. Weighted Speech Distortion Losses for Neural-Network-Based Real-Time Speech Enhancement
  831. What Does a Network Layer Hear? Analyzing Hidden Representations of End-to-End ASR Through Speech Synthesis
  832. What Makes the Sound?: A Dual-Modality Interacting Network for Audio-Visual Event Localization
  833. What did your adversary believeƒ Optimal Filtering and Smoothing in Counter-Adversarial Autonomous Systems
  834. What is best for spoken language understanding: small but task-dependant embeddings or huge but out-of-domain embeddings?
  835. Whosecough: In-the-Wild Cougher Verification Using Multitask Learning
  836. Wideband Channel Tracking for Millimeter Wave Massive Mimo Systems with Hybrid Beamforming Reception
  837. Wideband Direction of Arrival Estimation with Sparse Linear Arrays
  838. Wind: Wasserstein Inception Distance For Evaluating Generative Adversarial Network Performance
  839. Wirtinger Flow Algorithms for Phase Retrieval from Binary Measurements
  840. Witchcraft: Efficient PGD Attacks with Random Step Size
  841. Within-Sample Variability-Invariant Loss for Robust Speaker Recognition Under Noisy Environments
  842. X-Vectors Meet Emotions: A Study On Dependencies Between Emotion and Speaker Recognition
  843. XceptionTime: Independent Time-Window Xceptiontime Architecture for Hand Gesture Classification
  844. Xpsnr: A Low-Complexity Extension of The Perceptually Weighted Peak Signal-To-Noise Ratio For High-Resolution Video Quality Assessment
  845. Y-Net: Multi-Scale Feature Aggregation Network With Wavelet Structure Similarity Loss Function For Single Image Dehazing
  846. Zero-Crossing Precoding with Maximum Distance to the Decision Threshold for Channels with 1-Bit Quantization and Oversampling
  847. Zero-Shot Multi-Speaker Text-To-Speech with State-Of-The-Art Neural Speaker Embeddings
  848. dMazeRunner: Optimizing Convolutions on Dataflow Accelerators
  849. β-NMF and Sparsity Promoting Regularizations for Complex Mixture Unmixing. Application to 2D HSQC NMR

Looking for submission deadlines instead? See the conference deadline calendar.