← All conferences

ICASSP 2019 Accepted Papers

The full list of 1,730 papers accepted at ICASSP 2019 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. Learning from Multiview Correlations in Open-domain Videos
  2. Learning from the Best: A Teacher-student Multilingual Framework for Low-resource Languages
  3. Learning the Spiral Sharing Network with Minimum Salient Region Regression for Saliency Detection
  4. Learning to Dequantize Speech Signals by Primal-dual Networks: an Approach for Acoustic Sensor Networks
  5. Learning to Detect Dysarthria from Raw Speech
  6. Learning to Fuse Latent Representations for Multimodal Data
  7. Learning to Match Transient Sound Events Using Attentional Similarity for Few-shot Sound Recognition
  8. Learning to Rank: A Progressive Neural Network Learning Approach
  9. Learning-Based Pricing for Privacy-Preserving Job Offloading in Mobile Edge Computing
  10. Lessons from Building Acoustic Models with a Million Hours of Speech
  11. Leveraging Image-to-image Translation Generative Adversarial Networks for Face Aging
  12. Leveraging Weakly Supervised Data to Improve End-to-end Speech-to-text Translation
  13. Leveraging mmWave Imaging and Communications for Simultaneous Localization and Mapping
  14. Light Field Denoising Using 4D Anisotropic Diffusion
  15. Light Field Image Compression Using Depth-based CNN in Intra Prediction
  16. Linear Prediction-based Part-defined Auto-encoder Used for Speech Enhancement
  17. Linearized Kernel Representation Learning from Video Tensors by Exploiting Manifold Geometry for Gesture Recognition
  18. Local Convergence of the Heavy Ball Method in Iterative Hard Thresholding for Low-rank Matrix Completion
  19. Local Phase U-net for Fundus Image Segmentation
  20. Localization of an Unknown Number of Speakers in Adverse Acoustic Conditions Using Reliability Information and Diarization
  21. Localized Random Sampling for Robust Compressive Beam Alignment
  22. Long Term Background Reference Based Satellite Video Coding
  23. Long Text Analysis Using Sliced Recurrent Neural Networks with Breaking Point Information Enrichment
  24. Look, Listen, and Learn More: Design Choices for Deep Audio Embeddings
  25. Lora Digital Receiver Analysis and Implementation
  26. Loss and Double-edge-triggered Detector for Robust Small-footprint Keyword Spotting
  27. Low Bit-rate Speech Coding with VQ-VAE and a WaveNet Decoder
  28. Low Frequency Crosstalk Cancellation and Its Relationship to Amplitude Panning
  29. Low Power Pilot Aided Sub-sample Based Channel Estimation for Mmwave Cellular Systems
  30. Low Power Ultrasonic Gesture Recognition for Mobile Handsets
  31. Low-Complexity Compressive Analysis in Sub-Eigenspace for ECG Telemonitoring System
  32. Low-complexity Detection and Performance Analysis for Decode-and-forward Relay Networks
  33. Low-complexity Recurrent Neural Network-based Polar Decoder with Weight Quantization Mechanism
  34. Low-cost Measurement of Industrial Shock Signals via Deep Learning Calibration
  35. Low-latency Deep Clustering for Speech Separation
  36. Low-latency Speaker-independent Continuous Speech Separation
  37. Low-pass Filtering as Bayesian Inference
  38. Low-power Continuous Heart and Respiration Rates Monitoring on Wearable Devices
  39. Low-power Programmable Processor for Fast Fourier Transform Based on Transport Triggered Architecture
  40. Low-rank Embedding of Kernels in Convolutional Neural Networks under Random Shuffling
  41. Low-rank Estimation Based Evolutionary Clustering for Community Detection in Temporal Networks
  42. Low-rank Matrix Approximation Based on Intermingled Randomized Decomposition
  43. Low-rankness of Complex-valued Spectrogram and Its Application to Phase-aware Audio Processing
  44. Low-resolution Visual Recognition via Deep Feature Distillation
  45. Lung Nodule Detection with a 3D ConvNet via IoU Self-normalization and Maxout Unit
  46. M-vectors: Sub-band Based Energy Modulation Features for Multi-stream Automatic Speech Recognition
  47. MIMO Radar Transmit Beampattern Synthesis via Waveform Design for Target Localization
  48. MSE Based Precoding Schemes for Partially Correlated Transmissions in Interference Channels
  49. Machine Learning for Condition Monitoring and Innovation
  50. Magnetic Resonance Fingerprinting Using a Residual Convolutional Neural Network
  51. Majorization-minimization Algorithms for Convolutive NMF with the Beta-divergence
  52. Making Decisions with Shuffled Bits
  53. Mask-based MVDR Beamformer for Noisy Multisource Environments: Introduction of Time-varying Spatial Covariance Model
  54. Massive MIMO Multicast Beamforming via Accelerated Random Coordinate Descent
  55. Massive Mimo Channel Estimation with 1-Bit Spatial Sigma-delta ADCS
  56. Material Identification Using RF Sensors and Convolutional Neural Networks
  57. Material Segmentation in Hyperspectral Images with a Spatio-spectral Texture Descriptor
  58. Matrix Completion with Variational Graph Autoencoders: Application in Hyperlocal Air Quality Inference
  59. Maximally Separated Averages Prediction for High Fidelity Reversible Data Hiding
  60. Maximally Smooth Dirichlet Interpolation from Complete and Incomplete Sample Points on the Unit Circle
  61. Maximum-entropy Scattering Models for Financial Time Series
  62. Measuring the Spherical-harmonic Representation of a Sound Field Using a Cylindrical Array
  63. Measuring the Task Induced Oscillatory Brain Activity Using Tensor Decomposition
  64. Median Activation Functions for Graph Neural Networks
  65. Memorization Capacity of Deep Neural Networks under Parameter Quantization
  66. Methodical Design and Trimming of Deep Learning Networks: Enhancing External BP Learning with Internal Omnipresent-supervision Training Paradigm
  67. Mid-depth Based Block Structure Determination for AV1
  68. Mid-level Chord Transition Features for Musical Style Analysis
  69. Minimax Magnitude Response Approximation of Pole-radius Constrained IIR Digital Filters
  70. Minimum-volume Rank-deficient Nonnegative Matrix Factorizations
  71. Mirage: 2D Source Localization Using Microphone Pair Augmentation with Echoes
  72. Missing Data in Traffic Estimation: A Variational Autoencoder Imputation Method
  73. Misspecified CRB on Parameter Estimation for a Coupled Mixture of Polynomial Phase and Sinusoidal FM Signals
  74. Mitigating the Impact of Speech Recognition Errors on Spoken Question Answering by Adversarial Domain Adaptation
  75. Modality Attention for End-to-end Audio-visual Speech Recognition
  76. Model Change Detection with Application to Machine Learning
  77. Model Selection for Nonnegative Matrix Factorization by Support Union Recovery
  78. Modeling Characteristics of Real Loudspeakers Using Various Acoustic Models: Modal-domain Approaches
  79. Modeling Melodic Feature Dependency with Modularized Variational Auto-encoder
  80. Modeling Nonlinear Audio Effects with End-to-end Deep Neural Networks
  81. Modeling and Estimation of Interactions of Yule-Simon Processes
  82. Modelling Sample Informativeness for Deep Affective Computing
  83. Models of Visually Grounded Speech Signal Pay Attention to Nouns: A Bilingual Experiment on English and Japanese
  84. Monitoring of Trees' Health Condition Using a UAV Equipped with Low-cost Digital Camera
  85. Motion Artefact Removal in Functional Near-infrared Spectroscopy Signals Based on Robust Estimation
  86. Motion-adapted Three-dimensional Frequency Selective Extrapolation
  87. Multi Label Restricted Boltzmann Machine for Non-intrusive Load Monitoring
  88. Multi-attention Network for Thoracic Disease Classification and Localization
  89. Multi-band PIT and Model Integration for Improved Multi-channel Speech Separation
  90. Multi-channel Itakura Saito Distance Minimization with Deep Neural Network
  91. Multi-channel Time Encoding for Improved Reconstruction of Bandlimited Signals
  92. Multi-channel Wind Noise Reduction Using the Corcos Model
  93. Multi-classification of Breast Cancer Histology Images by Using Gravitation Loss
  94. Multi-feature Fusion Based on Supervised Multi-view Multi-label Canonical Correlation Projection
  95. Multi-frame Super-resolution for Time-of-flight Imaging
  96. Multi-geometry Spatial Acoustic Modeling for Distant Speech Recognition
  97. Multi-level Supervised Network for Person Re-identification
  98. Multi-modal Blind Source Separation with Microphones and Blinkies
  99. Multi-modal Image Stitching with Nonlinear Optimization
  100. Multi-objective Optimization Training of PLDA for Speaker Verification
  101. Multi-scale Dense Network for Single-image Super-resolution
  102. Multi-scale Spatial-temporal Network for Person Re-identification
  103. Multi-scale Vehicle Re-identification Using Self-adapting Label Smoothing Regularization
  104. Multi-speaker Emotional Acoustic Modeling for CNN-based Speech Synthesis
  105. Multi-speaker Sequence-to-sequence Speech Synthesis for Data Augmentation in Acoustic-to-word Speech Recognition
  106. Multi-spectral Image Denoising with Shared Dictionaries and Low-rank Representation
  107. Multi-step Self-attention Network for Cross-modal Retrieval Based on a Limited Text Space
  108. Multi-target Motion Parameter Estimation Exploiting Collaborative UAV Network
  109. Multi-task Adaptive Matching Pursuit for Sparse Signal Recovery Exploiting Signal Structures
  110. Multi-teacher Knowledge Distillation for Compressed Video Action Recognition on Deep Neural Networks
  111. Multi-user Communication in Difficult Interference
  112. Multi-view Networks for Multi-channel Audio Classification
  113. Multicarrier Radar-communications Waveform Design for RF Convergence and Coexistence
  114. Multicast Beamforming Using Semidefinite Relaxation and Bounded Perturbation Resilience
  115. Multichannel Quaternion Least Mean Square Algorithm
  116. Multichannel Sparse Blind Deconvolution on the Sphere
  117. Multimodal Grounding for Sequence-to-sequence Speech Recognition
  118. Multimodal One-shot Learning of Speech and Images
  119. Multimodal Retinal Image Registration and Fusion Based on Sparse Regularization via a Generalized Minimax-concave Penalty
  120. Multimodal Speaker Adaptation of Acoustic Model and Language Model for Asr Using Speaker Face Embedding
  121. Multipath-enabled Private Audio with Noise
  122. Multiple Agents Representation Using Motion Fields
  123. Multiple Linear Regression for High Efficiency Video Intra Coding
  124. Multiple Sound Source Localization with Rigid Spherical Microphone Arrays via Residual Energy Test
  125. Multiple Subspace Alignment Improves Domain Adaptation
  126. Multiple Temporal Scales Based Speaker Embeddings Learning for Text-dependent Speaker Recognition
  127. Multiple-graph Recurrent Graph Convolutional Neural Network Architectures for Predicting Disease Outcomes
  128. Multiresolution Time-of-arrival Estimation from Multiband Radio Channel Measurements
  129. Multiscale Directional Fusion for Depth Map Super Resolution with Denoising
  130. Multiscale Structure Tensor Total Variation for Image Recovery
  131. Multisource Remote Sensing Data Classification Using Deep Hierarchical Random Walk Networks
  132. Multisource Surveillance Video Coding by Exploiting 3D and 2D Knolwedge
  133. Multitask Learning for Frame-level Instrument Recognition
  134. Multiview Canonical Correlation Analysis over Graphs
  135. Muse-ing on the Impact of Utterance Ordering on Crowdsourced Emotion Annotations
  136. Music Boundary Detection Based on a Hybrid Deep Model of Novelty, Homogeneity, Repetition and Duration
  137. Mvdr Robust Adaptive Beamforming Design with Direction of Arrival and Generalized Similarity Constraints
  138. NN-based Ordinal Regression for Assessing Fluency of ESL Speech
  139. Native Language and Stimuli Signal Prediction from EEG
  140. Near-infrared Image Guided Neural Networks for Color Image Denoising
  141. Near-optimal Coded Apertures for Imaging via Nazarov's Theorem
  142. Negative Correlation, Non-linear Filtering, and Discovering of Repetitiveness for Cache Timing Channel Detection
  143. Network Adaptation Strategies for Learning New Classes without Forgetting the Original Ones
  144. Neural Approaches to Automated Speech Scoring of Monologue and Dialogue Responses
  145. Neural CRF Transducers for Sequence Labeling
  146. Neural Codes to Factor Language in Multilingual Speech Recognition
  147. Neural Music Synthesis for Flexible Timbre Control
  148. Neural Networks Sequential Training Using Variational Gaussian Particle Filter
  149. Neural Source-filter-based Waveform Model for Statistical Parametric Speech Synthesis
  150. Neural Variational Identification and Filtering for Stochastic Non-linear Dynamical Systems with Application to Non-intrusive Load Monitoring
  151. Neuromorphic Vision Sensing for CNN-based Action Recognition
  152. Node-asynchronous Implementation of Rational Filters on Graphs
  153. Noise-tolerant Audio-visual Online Person Verification Using an Attention-based Neural Network Fusion
  154. Noisy 1-Bit Compressed Sensing with Heterogeneous Side-information
  155. Non-coherent Sensor Fusion via Entropy Regularized Optimal Mass Transport
  156. Non-harmonic Analysis Based Instantaneous Heart Rate Estimation from Photoplethysmography
  157. Non-intrusive Speech Quality Assessment Using Neural Networks
  158. Non-intrusive Speech Quality Assessment for Super-wideband Speech Communication Networks
  159. Non-local Self-attention Structure for Function Approximation in Deep Reinforcement Learning
  160. Non-negative Matrix Factorization Using Bregman Monotone Operator Splitting
  161. Nonlinear Acceleration of Constrained Optimization Algorithms
  162. Nonlinear Multi-scale Super-resolution Using Deep Learning
  163. Nonlinear Prediction of Multidimensional Signals via Deep Regression with Applications to Image Coding
  164. Nonlinear State Estimation Using Particle Filters on the Stiefel Manifold
  165. Nonnegative Low-rank Sparse Component Analysis
  166. Nose, Eyes and Ears: Head Pose Estimation by Locating Facial Keypoints
  167. Novel Detection Methods for Zero-padded Single Carrier Spatial Modulation in Doubly Selective Channels
  168. Novel Lower Bound on the Performance of a Partial Zero Forcing Receiver in a Mimo Cellular Network
  169. Novel Metric Learning for Non-parallel Voice Conversion
  170. OMP and Continuous Dictionaries: Is k-step Recovery Possible?
  171. Object Counting in Video Surveillance Using Multi-scale Density Map Regression
  172. Object Detection in Curved Space for 360-Degree Camera
  173. Object and Text-guided Semantics for CNN-based Activity Recognition
  174. Objective Assessment of Spatial Audio Quality Using Directional Loudness Maps
  175. Objective Assessment of Vocal Tremor
  176. Objective Comparison of Speech Enhancement Algorithms with Hearing Loss Simulation
  177. Objective Measures of Plosive Nasalization in Hypernasal Speech
  178. Obtaining Narrow Transition Region in STFT Domain Processing Using Subband Filters
  179. Occupancy Pattern Recognition with Infrared Array Sensors: A Bayesian Approach to Multi-body Tracking
  180. On Achievable Rates for Massive Mimo System with Imperfect Channel Covariance Information
  181. On Evaluating CNN Representations for Low Resource Medical Image Classification
  182. On Massive MIMO Cellular Systems Resilience to Radar Interference
  183. On Modified Squared Givens Rotations for Sphere Decoder Preprocessing
  184. On Nonparametric Identification of Wiener Systems with Deterministic Inputs
  185. On Optimal Beam Steering Directions in Millimeter Wave Systems
  186. On Radar Privacy in Shared Spectrum Scenarios
  187. On Reducing the Effect of Speaker Overlap for Chime-5
  188. On Role and Location of Normalization before Model-based Data Augmentation in Residual Blocks for Classification Tasks
  189. On Self-assessment of Proficiency of Autonomous Systems
  190. On Training Targets and Objective Functions for Deep-learning-based Audio-visual Speech Enhancement
  191. On Using 2D Sequence-to-sequence Models for Speech Recognition
  192. On the Accuracy Limit of Time-delay Estimation with a Band-limited Signal
  193. On the Adversarial Robustness of Subspace Learning
  194. On the Computability of the Secret Key Capacity under Rate Constraints
  195. On the Design of Flexible Kronecker Product Beamformers with Linear Microphone Arrays
  196. On the Equivalence of Semidifinite Relaxations for MIMO Detection with General Constellations
  197. On the Fourier Representation of Computable Continuous Signals
  198. On the Move: Localization with Kinetic Euclidean Distance Matrices
  199. On the Performance of DIBR Methods When Using Depth Maps from State-of-the-art Stereo Matching Algorithms
  200. On the Sensitivity of Spectral Initialization for Noisy Phase Retrieval
  201. On the Transferability of Adversarial Examples against CNN-based Image Forensics
  202. On the Usefulness of Statistical Normalisation of Bottleneck Features for Speech Recognition
  203. One-bit Unlimited Sampling
  204. One-dimensional Edge-preserving Spline Smoothing for Estimation of Piecewise Smooth Functions
  205. Online Deep Attractor Network for Real-time Single-channel Speech Separation
  206. Online Estimation and Smoothing of a Target Trajectory in Mixed Stationary/moving Conditions
  207. Online Learning for Computation Peer Offloading with Semi-bandit Feedback
  208. Online Learning with Self-tuned Gaussian Kernels: Good Kernel-initialization by Multiscale Screening
  209. Online Radio Map Update Based on a Marginalized Particle Gaussian Process
  210. Online Singing Voice Separation Using a Recurrent One-dimensional U-NET Trained with Deep Feature Losses
  211. Online Single Person Tracking for Unmanned Aerial Vehicles: Benchmark and New Baseline
  212. Online Variational Bayesian Subspace Filtering
  213. Optimal Feature Selection for Blind Super-resolution Image Quality Evaluation
  214. Optimal ROC Curves from Score Variable Threshold Tests
  215. Optimal Sensor Placement for Signal Extraction
  216. Optimal Trilateration Is an Eigenvalue Problem
  217. Optimization of Speaker Extraction Neural Network with Magnitude and Temporal Spectrum Approximation Loss
  218. Optimization of a Moving Colored Coded Aperture in Compressive Spectral Imaging
  219. Optimized Color-guided Filter for Depth Image Denoising
  220. Optimized Quantization in Distributed Graph Signal Processing
  221. Optimizing QoE of Multiple Users over DASH: A Meta-learning Approach
  222. Optimum Sampling for Packet Assisted Round Trip Time Measurement
  223. Outphasing Elements for Hybrid Analogue Digital Beamforming and Single-RF MIMO
  224. Overlap-add Windows with Maximum Energy Concentration for Speech and Audio Processing
  225. PGR-Net: A Parallel Network Based on Group and Regression for Age Estimation
  226. PPSAN: Perceptual-aware 3D Point Cloud Segmentation via Adversarial Learning
  227. Pairwise Approximate K-SVD
  228. Parallel Coordinate Descent Algorithms for Sparse Phase Retrieval
  229. Parameter Uncertainty for End-to-end Speech Recognition
  230. Parametric Cepstral Mean Normalization for Robust Speech Recognition
  231. Parametric Hear through Equalization for Augmented Reality Audio
  232. Particle Filtering: the First 25 Years and beyond
  233. Passive Detection and Discrimination of Body Movements in the sub-THz Band: A Case Study
  234. Pathological Speech Intelligibility Assessment Based on the Short-time Objective Intelligibility Measure
  235. Peak Detection and Baseline Correction Using a Convolutional Neural Network
  236. Perceptual Audio Coding with Adaptive Non-uniform Time/frequency Tilings Using Subband Merging and Time Domain Aliasing Reduction
  237. Perceptual Quality Preserving Image Super-resolution via Channel Attention
  238. Perceptual Soundfield Reconstruction in Three Dimensions via Sound Field Extrapolation
  239. Perceptually Enhanced Single Frequency Filtering for Dysarthric Speech Detection and Intelligibility Assessment
  240. Perceptually-motivated Environment-specific Speech Enhancement
  241. Perfect Match: Improved Cross-modal Embeddings for Audio-visual Synchronisation
  242. Performance Advantages of Deep Neural Networks for Angle of Arrival Estimation
  243. Performance Analysis of Convex Data Detection in MIMO
  244. Performance Analysis of Discrete-valued Vector Reconstruction Based on Box-constrained Sum of L1 Regularizers
  245. Performance Analysis of One-bit Group-sparse Signal Reconstruction
  246. Performance Bound for Blind Extraction of Non-gaussian Complex-valued Vector Component from Gaussian Background
  247. Performance Enhancement of the Measure-transformed Music Algorithm via Mse Based Optimization
  248. Performance of Jensen Shannon Divergence in Incipient Fault Detection and Estimation
  249. Perturbed Projected Gradient Descent Converges to Approximate Second-order Points for Bound Constrained Nonconvex Problems
  250. PhaST: Model-free Phaseless Subspace Tracking
  251. Phase-aware Harmonic/percussive Source Separation via Convex Optimization
  252. Phase-only Robust Minimum Dispersion Beamforming
  253. Phoebe: Pronunciation-aware Contextualization for End-to-end Speech Recognition
  254. Phoneme Dependent Speaker Embedding and Model Factorization for Multi-speaker Speech Synthesis and Adaptation
  255. Phoneme Level Language Models for Sequence Based Low Resource ASR
  256. Phoneme Specific Modelling and Scoring Techniques for Anti Spoofing System
  257. Phonemic-level Duration Control Using Attention Alignment for Natural Speech Synthesis
  258. Phonespoof: A New Dataset for Spoofing Attack Detection in Telephone Channel
  259. Phonetic Analysis of Dysarthric Speech Tempo and Applications to Robust Personalised Dysarthric Speech Recognition
  260. Phylogenetic Analysis of Software Using Cache Miss Statistics
  261. Piano Sustain-pedal Detection Using Convolutional Neural Networks
  262. Pigment Unmixing of Hyperspectral Images of Paintings Using Deep Neural Networks
  263. Pixel Level Data Augmentation for Semantic Image Segmentation Using Generative Adversarial Networks
  264. Pixel-level Texture Segmentation Based AV1 Video Compression
  265. Pliable Data Shuffling for On-device Distributed Learning
  266. Point Cloud Segmentation Using Hierarchical Tree for Architectural Models
  267. Polynomial Networks Representation of Nonlinear Mixtures with Application in Underdetermined Blind Source Separation
  268. Polyphonic Music Transcription with Semantic Segmentation
  269. Polyphonic Sound Event Detection Using Convolutional Bidirectional Lstm and Synthetic Data-based Transfer Learning
  270. Post-stitching Depth Adjustment for Stereoscopic Panorama
  271. Postfiltering Using an Adversarial Denoising Autoencoder with Noise-aware Training
  272. Potential Games for Distributed Parameter Estimation in Networks with Ambiguous Measurements
  273. Power Minimization in Multi-tier Networks with Flexible Duplexing
  274. Power Network Parameter Correction via Sparse Unsupervised Regression
  275. Power System State Forecasting via Deep Recurrent Neural Networks
  276. Power-efficient Beam Pattern Synthesis via Sequential Outer Approximation Procedure
  277. Practical Concentric Open Sphere Cardioid Microphone Array Design for Higher Order Sound Field Capture
  278. Pre-training of Speaker Embeddings for Low-latency Speaker Change Detection in Broadcast News
  279. Precoding Design for the MIMO-RoC Downlink
  280. Predicting Tongue Motion in Unlabeled Ultrasound Videos Using Convolutional Lstm Neural Networks
  281. Predicting Video-frames Using Encoder-convlstm Combination
  282. Predicting the Precision of Elevation Localization Based on Head Related Transfer Functions
  283. Predicting the Secret Parameters of a Chaotic Random Number Generator from Time Series
  284. Prediction of Multi-target Dynamics Using Discrete Descriptors: an Interactive Approach
  285. Prediction-correction for Nonsmooth Time-varying Optimization via Forward-backward Envelopes
  286. Prediction-error-ordering for High-fidelity Reversible Data Hiding
  287. Prewarping Siamese Network: Learning Local Representations for Online Signature Verification
  288. Price-aware Renewable Energy Management with Transmission Losses
  289. Privacy-aware Feature Extraction for Gender Discrimination versus Speaker Identification
  290. Privacy-cost Trade-off in a Smart Meter System with a Renewable Energy Source and a Rechargeable Battery
  291. Privacy-preserving Online Human Behaviour Anomaly Detection Based on Body Movements and Objects Positions
  292. Privacy-preserving Paralinguistic Tasks
  293. Progressive Filtering for Feature Matching
  294. Promising Accurate Prefix Boosting for Sequence-to-sequence ASR
  295. Proper Guidance Image Generation Based on Saliency Factor for Better Transmission Refinement in Image Dehazing
  296. Properties and Limits of the Minimum-norm Differential Beamformers with Circular Microphone Arrays
  297. Provable Memory-efficient Online Robust Matrix Completion
  298. Provably Accelerated Randomized Gossip Algorithms
  299. Proximal Deep Recurrent Neural Network for Monaural Singing Voice Separation
  300. Prune Your Neurons Blindly: Neural Network Compression through Structured Class-blind Pruning
  301. Pruning SIFT & SURF for Efficient Clustering of Near-duplicate Images
  302. PyHTK: Python Library and ASR Pipelines for HTK
  303. Quadratic Envelope Regularization for Structured Low Rank Approximation
  304. Quality Control of Voice Recordings in Remote Parkinson's Disease Monitoring Using the Infinite Hidden Markov Model
  305. Quantized Event-triggered Sampled-data Average Consensus with Guaranteed Rate of Convergence
  306. Quantized Gaussian Embedding Steganography
  307. Quasi Black Hole Effect of Gradient Descent in Large Dimension: Consequence on Neural Network Learning
  308. Quasi-fully Convolutional Neural Network with Variational Inference for Speech Synthesis
  309. Quaternion Convolutional Neural Networks for Detection and Localization of 3D Sound Events
  310. Quaternion Convolutional Neural Networks for Heterogeneous Image Processing
  311. Quaternion-Valued Adaptive Filtering via Nesterov's Extrapolation
  312. Question Answering for Spoken Lecture Processing
  313. Quickest Detection of Deviations from Periodic Statistical Behavior
  314. Quickest Detection of Time-varying False Data Injection Attacks in Dynamic Smart Grids
  315. RF-based Analytics Generated by Tag-to-tag Networks
  316. RHFCN: : Fully CNN-based Steganalysis of MP3 with Rich High-pass Filtering
  317. RTF-steered Binaural MVDR Beamforming Incorporating an External Microphone for Dynamic Acoustic Scenarios
  318. Radar Stationary and Moving Indoor Target Localization with Low-rank and Sparse Regularizations
  319. Radial Loss for Learning Fine-grained Video Similarity Metric
  320. Rain Streak Removal via Multi-scale Mixture Exponential Power Model
  321. Random Forest Oriented Fast QTBT Frame Partitioning
  322. Random Infinite Tree and Dependent Poisson Diffusion Process for Nonparametric Bayesian Modeling in Multiple Object Tracking
  323. Random Sampling for Distributed Coded Matrix Multiplication
  324. Randomized Tensor Ring Decomposition and Its Application to Large-scale Data Reconstruction
  325. Randomly Weighted CNNs for (Music) Audio Classification
  326. Real-time Object Detection via Pruning and a Concatenated Multi-feature Assisted Region Proposal Network
  327. Real-time Passive Acoustic 3D Tracking of Deep Diving Cetacean by Small Non-uniform Mobile Surface Antenna
  328. Real-time Prediction for Fine-grained Air Quality Monitoring System with Asynchronous Sensing
  329. Real-time Speech Enhancement Using an Efficient Convolutional Recurrent Network for Dual-microphone Mobile Phones in Close-talk Scenarios
  330. Real-time Tracker with Fast Recovery from Target Loss
  331. Receiver Design for Doppler Positioning with Leo Satellites
  332. Recognition of Online Handwriting with Variability on Smart Devices
  333. Reconfigurable Multitask Audio Dynamics Processing Scheme
  334. Reconstruction-cognizant Graph Sampling Using Gershgorin Disc Alignment
  335. Recurrent 3D Convolutional Network for Rodent Behavior Recognition
  336. Recurrent Deep Divergence-based Clustering for Simultaneous Feature Learning and Clustering of Variable Length Time Series
  337. Recurrent Neural Network Language Model Training Using Natural Gradient
  338. Recurrent Neural Networks with Stochastic Layers for Acoustic Novelty Detection
  339. Reduced Complexity Image Clustering Based on Camera Fingerprints
  340. Reduced-complexity Deep Neural Network-aided Channel Code Decoder: A Case Study for BCH Decoder
  341. Reducing the Search Space for Hyperparameter Optimization Using Group Sparsity
  342. Referential Vowel Duration Ratio as a Feature for Automatic Assessment of L2 Word Prosody
  343. Reflection Symmetry Detection by Embedding Symmetry in a Graph
  344. Reflection Tomographic Imaging of Highly Scattering Objects Using Incremental Frequency Inversion
  345. Regular Sampling of Tensor Signals: Theory and Application to FMRI
  346. Regularized Fourier Ptychography Using an Online Plug-and-play Algorithm
  347. Reinforcement Learning Based Speech Enhancement for Robust Speech Recognition
  348. Reinforcement Learning with Safe Exploration for Network Security
  349. Reinforcing Self-expressive Representation with Constraint Propagation for Face Clustering in Movies
  350. Relationships between Deep Learning and Linear Adaptive Systems
  351. Reliability of the Most Common Objective Metrics for Light Field Quality Assessment
  352. Remote State Preparation for Multiple Parties
  353. Replay Attack Detection Using Magnitude and Phase Information with Attention-based Adaptive Filters
  354. Representation Learning Using Convolution Neural Network for Acoustic-to-articulatory Inversion
  355. Representation Mixing for TTS Synthesis
  356. Residual Integration Neural Network
  357. Resource Optimization in Quantum Access Networks
  358. Rethinking Super-resolution: the Bandwidth Selection Problem
  359. Rethinking Teaching Practices for Signal Processing Education
  360. Retrieving Speech Samples with Similar Emotional Content Using a Triplet Loss Function
  361. Revisiting Hidden Markov Models for Speech Emotion Recognition
  362. Revisiting and Improving Semi-supervised Learning: A Large Dimensional Approach
  363. Robust Approximate Message Passing for Nonzero-mean Sensing Matrices
  364. Robust Audio-visual Speech Recognition Using Bimodal Dfsmn with Multi-condition Training and Dropout Regularization
  365. Robust Bayesian Beamforming for Sources at Different Distances with Applications in Urban Monitoring
  366. Robust Beamspace Design for Direct Localization
  367. Robust Capon Beamforming via ADMM
  368. Robust Common Spatial Patterns Estimation Using Dynamic Time Warping to Improve BCI Systems
  369. Robust Detection for Cluster Analysis
  370. Robust Dictionary Learning Using α-Divergence
  371. Robust Freeway Accident Detection: A Two-Stage Approach
  372. Robust Full-sphere Binaural Sound Source Localization Using Interaural and Spectral Cues
  373. Robust Graph Signal Sampling
  374. Robust Gridless Sound Field Decomposition Based on Structured Reciprocity Gap Functional in Spherical Harmonic Domain
  375. Robust Least Mean Squares Estimation of Graph Signals
  376. Robust Linear Discriminant Analysis Using Tyler's Estimator: Asymptotic Performance Characterization
  377. Robust Low-tubal-rank Tensor Completion
  378. Robust M-estimation Based Matrix Completion
  379. Robust Molecular Dynamics Simulations Using Coded FFT Algorithm
  380. Robust Recognition of Reverberant and Noisy Speech Using Coherence-based Processing
  381. Robust Room Equalization Using Sparse Sound-field Reconstruction
  382. Robust Secure Precoding and Antenna Selection: A Probabilistic Optimization Approach for Interference Exploitation
  383. Robust Self-calibration of Constant Offset Time-difference-of-arrival
  384. Robust Sparse Multichannel Active Noise Control
  385. Robust Speech Activity Detection in Movie Audio: Data Resources and Experimental Evaluation
  386. Robust Subspace Clustering by Learning an Optimal Structured Bipartite Graph via Low-rank Representation
  387. Robust Unsupervised Flexible Auto-weighted Local-coordinate Concept Factorization for Image Clustering
  388. Robust View Synthesis in Wide-baseline Complex Geometric Environments
  389. Robust Visual Tracking via Adaptive Occlusion Detection
  390. Robust and Fine-grained Prosody Control of End-to-end Speech Synthesis
  391. Rodent Sleep Assessment with a Trainable Video-based Approach
  392. Role Specific Lattice Rescoring for Speaker Role Recognition from Speech Recognition Outputs
  393. SAMIR: Sparsity Amplified Iteratively-reweighted Beamforming for High-rsolution Ultrasound Imaging
  394. SDR - Half-baked or Well Done?
  395. SLiQA-I: Towards Cold-start Development of End-to-end Spoken Language Interface for Question Answering
  396. SNIPER: Few-shot Learning for Anomaly Detection to Minimize False-negative Rate with Ensured True-positive Rate
  397. SPFEMD: Super-pixel Based Finger Earth Mover's Distance for Hand Gesture Recognition
  398. STFT Spectral Loss for Training a Neural Speech Waveform Model
  399. SURE-TISTA: A Signal Recovery Network for Compressed Sensing
  400. SVD-PHAT: A Fast Sound Source Localization Method
  401. SVM-based Seal Imprint Verification Using Edge Difference
  402. Safety in the Face of Unknown Unknowns: Algorithm Fusion in Data-driven Engineering Systems
  403. Saliency Aware: Weakly Supervised Object Localization
  404. Saliency Map on Cnns for Protein Secondary Structure Prediction
  405. Saliency Prediction for Omnidirectional Images Considering Optimization on Sphere Domain
  406. Salient Object Detection on Hyperspectral Images Using Features Learned from Unsupervised Segmentation Task
  407. Sample Complexity of Joint Structure Learning
  408. Sample Space-time Covariance Matrix Estimation
  409. Sampling Schemes for Accurate Reconstruction and Computation of Performance Parameters of Antenna Radiation Pattern
  410. Scalable Gaussian Process Using Inexact Admm for Big Data
  411. Scalable MCMC in Degree Corrected Stochastic Block Model
  412. Scalable Mutual Information Estimation Using Dependence Graphs
  413. Scaling up MIMO Radar for Target Detection
  414. Scanet: Spatial-channel Attention Network for 3D Object Detection
  415. Scattering Multi-connectivity Estimation for Indoor mmWave Small Cells under Limited Training Steps
  416. Scene Privacy Protection
  417. Scene-dependent Anomalous Acoustic-event Detection Based on Conditional Wavenet and I-vector
  418. Score-based Learning for Relevance Prediction in Image Similarity Search
  419. Score-specific Non-maximum Suppression and Coexistence Prior for Multi-scale Face Detection
  420. Second Order Sequential Best Rotation Algorithm with Householder Reduction for Polynomial Matrix Eigenvalue Decomposition
  421. Secure Analytics and Resilient Inference for the Internet of Things
  422. Secure MIMO Interference Channel with Confidential Messages and Delayed CSIT
  423. Securing Smartphone Handwritten Pin Codes with Recurrent Neural Networks
  424. Seeing through Sounds: Predicting Visual Semantic Segmentation Results from Multichannel Audio Signals
  425. Segment-level Training of ANNs Based on Acoustic Confidence Measures for Hybrid HMM/ANN Speech Recognition
  426. Segmentation, Classification, and Visualization of Orca Calls Using Deep Learning
  427. Seizure Detection Using Least Eeg Channels by Deep Convolutional Neural Network
  428. Selecting Optimal Proposal Number for Image-based Object Detection
  429. Selective Jpeg2000 Encryption of Iris Data: Protecting Sample Data vs. Normalised Texture
  430. Selective Virtual Sensing Technique for Multi-channel Feedforward Active Noise Control Systems
  431. Self-attention Aligner: A Latency-control End-to-end Model for ASR Using Self-attention Network and Chunk-hopping
  432. Self-attention Based Model for Punctuation Prediction Using Word and Speech Embeddings
  433. Self-attention Based Prosodic Boundary Prediction for Chinese Speech Synthesis
  434. Self-attention Networks for Connectionist Temporal Classification in Speech Recognition
  435. Self-supervised Audio-visual Co-segmentation
  436. Sell-corpus: an Open Source Multiple Accented Chinese-english Speech Corpus for L2 English Learning Assessment
  437. Semantic Query-by-example Speech Search Using Visual Grounding
  438. Semantic Super-resolution for Extremely Low-resolution Vehicle License Plate
  439. Semi-supervised Acoustic Event Detection Based on Tri-training
  440. Semi-supervised Depth Estimation from a Single Image Based on Confidence Learning
  441. Semi-supervised End-to-end Speech Recognition Using Text-to-speech and Autoencoders
  442. Semi-supervised Learning with Generative Adversarial Networks for Arabic Dialect Identification
  443. Semi-supervised Monaural Singing Voice Separation with a Masking Network Trained on Synthetic Mixtures
  444. Semi-supervised Multichannel Speech Enhancement with Variational Autoencoders and Non-negative Matrix Factorization
  445. Semi-supervised Multiclass Clustering Based on Signed Total Variation
  446. Semi-supervised Nuisance-attribute Networks for Domain Adaptation
  447. Semi-supervised Training for End-to-end Models via Weak Distillation
  448. Semi-supervised Training for Improving Data Efficiency in End-to-end Speech Synthesis
  449. Semi-supervised Transfer Learning for Convolutional Neural Networks for Glaucoma Detection
  450. Semi-supervised Triplet Loss Based Learning of Ambient Audio Embeddings
  451. Semi-supervised and Population Based Training for Voice Commands Recognition
  452. Sensor-Assisted Global Motion Estimation for Efficient UAV Video Coding
  453. Sentiment Aware Fake News Detection on Online Social Networks
  454. Separable Simplex-structured Matrix Factorization: Robustness of Combinatorial Approaches
  455. Seq2Seq Attentional Siamese Neural Networks for Text-dependent Speaker Verification
  456. Sequence Noise Injected Training for End-to-end Speech Recognition
  457. Sequence-level Knowledge Distillation for Model Compression of Attention-based Sequence-to-sequence Speech Recognition
  458. Sequence-to-sequence Modelling of F0 for Speech Emotion Conversion
  459. Sequential Matching Model for End-to-end Multi-turn Response Selection
  460. Sequential Structured Dictionary Learning for Block Sparse Representations
  461. Sergan: Speech Enhancement Using Relativistic Generative Adversarial Networks with Gradient Penalty
  462. Shadow Removal Detection and Localization for Forensics Analysis
  463. Sharpening Sparse Regularizers
  464. Sharpening of Angular Spectra Based on a Directional Re-assignment Approach for Ambisonic Sound-field Visualisation
  465. Shift-invariant Subspace Tracking with Missing Data
  466. Ship Wake Detection in X-band SAR Images Using Sparse GMC Regularization
  467. Short-segment Heart Sound Classification Using an Ensemble of Deep Convolutional Neural Networks
  468. Shot Type Feasibility in Autonomous UAV Cinematography
  469. Sign Language Detection "in the Wild" with Recurrent Neural Networks
  470. SignProx: One-bit Proximal Algorithm for Nonconvex Stochastic Optimization
  471. Signals and Systems: Casting It as an Action-adventure Rather than a Horror Genre
  472. Similarity Learning for Authorship Verification in Social Media
  473. Similarity Metric Based on Siamese Neural Networks for Voice Casting
  474. Similarity Search-based Blind Source Separation
  475. Simple Cooperative Transmission Schemes for Underlay Spectrum Sharing Using Symbol-level Precoding and Load-controlled Arrays
  476. Simulation or Real-time?
  477. Simultaneous Blind Deconvolution and Phase Retrieval with Tensor Iterative Hard Thresholding
  478. Simultaneous DFT and IDFT through Widely Linear CLMS
  479. Simultaneous Optimization of Forgetting Factor and Time-frequency Mask for Block Online Multi-channel Speech Enhancement
  480. Singing Voice Separation: A Study on Training Data
  481. Singing Voice Synthesis Based on Generative Adversarial Networks
  482. Single Image Interpolation Exploiting Semi-local Similarity
  483. Single-channel Speech Extraction Using Speaker Inventory and Attention Network
  484. Skin Lesion Classification Using Hybrid Deep Neural Networks
  485. Sleep Gesture Detection in Classroom Monitor System
  486. Small Array Reproduction Method for Ambisonic Encodings Using Headtracking
  487. Smart DSP for a Smarter Power Grid: Teaching Power System Analysis through Signal Processing
  488. Smooth Signal Recovery on Product Graphs
  489. Solving Complex Quadratic Equations with Full-rank Random Gaussian Matrices
  490. Solving Continuous-domain Problems Exactly with Multiresolution B-splines
  491. Solving Memory Access Conflicts in LTE-4G Standard
  492. Solving Quadratic Equations via Amplitude-based Nonconvex Optimization
  493. Sound Event Detection Using Graph Laplacian Regularization Based on Event Co-occurrence
  494. Sound Event Detection with Sequentially Labelled Data Based on Connectionist Temporal Classification and Unsupervised Clustering
  495. Sound Event Envelope Estimation in Polyphonic Mixtures
  496. Sound Source Localization in a Reverberant Room Using Harmonic Based Music
  497. Sound-based Transportation Mode Recognition with Smartphones
  498. Soundfield Reconstruction in Reverberant Environments Using Higher-order Microphones and Impulse Response Measurements
  499. Space Alternating Variational Estimation and Kronecker Structured Dictionary Learning
  500. Space Warping Based Dimensionality Reduction of Higher Order Ambisonics Signals
  501. Sparse Bayesian Learning for Robust PCA
  502. Sparse Blind Demixing for Low-latency Signal Recovery in Massive Iot Connectivity
  503. Sparse Fractal Array Design with Increased Degrees of Freedom
  504. Sparse Gaussian Process Audio Source Separation Using Spectrum Priors in the Time-domain
  505. Sparse Learning of Parsimonious Reproducing Kernel Hilbert Space Models
  506. Sparse Recovery and Non-stationary Blind Demodulation
  507. Sparse Recovery over Nonlinear Dictionaries
  508. Sparse Signal Recovery Using MPDR Estimation
  509. Sparse Subspace Clustering for Evolving Data Streams
  510. Sparsity-based Blind Deconvolution of Neural Activation Signal in FMRI
  511. Spatial Audio Coding without Recourse to Background Signal Compression
  512. Spatial Constraint on Multi-channel Deep Clustering
  513. Spatial and Channel Attention Based Convolutional Neural Networks for Modeling Noisy Speech
  514. Spatial-fourier Retrieval of Head-related Impulse Responses from Fast Continuous-azimuth Recordings in the Time-domain
  515. Spatially Adaptive Losses for Video Super-resolution with GANs
  516. Spatio-spectral Modulation Using a Binary Photomask for Compressive Chromotomography
  517. Speaker Agnostic Foreground Speech Detection from Audio Recordings in Workplace Settings from Wearable Recorders
  518. Speaker Change Detection Using Fundamental Frequency with Application to Multi-talker Segmentation
  519. Speaker Characterization Using TDNN-LSTM Based Speaker Embedding
  520. Speaker Diarisation Using 2D Self-attentive Combination of Embeddings
  521. Speaker Recognition for Multi-speaker Conversations Using X-vectors
  522. Speaker Verification Using End-to-end Adversarial Language Adaptation
  523. Speaker-dependent Wavenet-based Delay-free Adpcm Speech Coding
  524. Speaker-independent Classification of Phonetic Segments from Raw Ultrasound in Child Speech
  525. Spectral Efficiency of Noncooperative Uplink Massive MIMO Systems with Joint Decoding
  526. Spectral Graph Wavelet Transform as Feature Extractor for Machine Learning in Neuroimaging
  527. Spectral Method for Multiplexed Phase Retrieval and Application in Optical Imaging in Complex Media
  528. Spectral Partitioning of Time-varying Networks with Unobserved Edges
  529. Spectrum-adapted Polynomial Approximation for Matrix Functions
  530. Speech Artifact Removal from Eeg Recordings of Spoken Word Production with Tensor Decomposition
  531. Speech Augmentation Using Wavenet in Speech Recognition
  532. Speech Denoising by Parametric Resynthesis
  533. Speech Emotion Recognition Using Capsule Networks
  534. Speech Emotion Recognition Using Deep Neural Network Considering Verbal and Nonverbal Speech Sounds
  535. Speech Emotion Recognition Using Multi-hop Attention Mechanism
  536. Speech Enhancement with Variational Autoencoders and Alpha-stable Distributions
  537. Speech Landmark Bigrams for Depression Detection from Naturalistic Smartphone Speech
  538. Speech Markers for Clinical Assessment of Cocaine Users
  539. Speech Recognition with No Speech or with Noisy Speech
  540. Speech Super Resolution Generative Adversarial Network
  541. Speech Waveform Reconstruction Using Convolutional Neural Networks with Noise and Periodic Inputs
  542. Speech as a Biomarker for Obstructive Sleep Apnea Detection
  543. Spherical Clustering of Users Navigating 360° Content
  544. Spike Detection in Axonal-synaptic Channels with Multiple Synapses
  545. Spoofing Attack Detection by Anomaly Detection
  546. Squared-Loss Mutual Information via High-Dimension Coherence Matrix Estimation
  547. Static and Dynamic State Predictions for Acoustic Model Combination
  548. Statistical Persistent Homology of Brain Signals
  549. Statistical Rank Selection for Incomplete Low-rank Matrices
  550. Stereo Source Separation in the Frequency Domain: Solving the Permutation Problem by a Sliding K-means Method
  551. Stochastic Adaptive Neural Architecture Search for Keyword Spotting
  552. Stochastic Data-driven Hardware Resilience to Efficiently Train Inference Models for Stochastic Hardware Implementations
  553. Stochastic Gradient Descent for Spectral Embedding with Implicit Orthogonality Constraint
  554. Stochastic Markov Recurrent Neural Network for Source Separation
  555. Stochastic Ml Simplex-structured Matrix Factorization under the Dirichlet Mixture Model
  556. Stream Attention-based Multi-array End-to-end Speech Recognition
  557. Streaming End-to-end Speech Recognition for Mobile Devices
  558. Strip the Stripes: Artifact Detection and Removal for Scanning Electron Microscopy Imaging
  559. Structural Recurrent Neural Network for Traffic Speed Prediction
  560. SubSpectralNet - Using Sub-spectrogram Based Convolutional Neural Networks for Acoustic Scene Classification
  561. Subband Optimization and Filtering Technique for Practical Personal Audio Systems
  562. Subband Temporal Envelope Features and Data Augmentation for End-to-end Recognition of Distant Conversational Speech
  563. Subword Regularization and Beam Search Decoding for End-to-end Automatic Speech Recognition
  564. Sum Throughput Maximization for Multi-tag MISO Backscattering
  565. Super-gaussianity of Speech Spectral Coefficients as a Potential Biomarker for Dysarthric Speech Detection
  566. Super-resolution DOA Estimation for Arbitrary Array Geometries Using a Single Noisy Snapshot
  567. Super-resolution Results for a 1D Inverse Scattering Problem
  568. Super-resolution Using Flow Estimation in Contrast Enhanced Ultrasound Imaging
  569. Supervised Kernel Change Point Detection with Partial Annotations
  570. Supervised Speech Enhancement with Real Spectrum Approximation
  571. Support Tensor Machine for Financial Forecasting
  572. Surgical Activities Recognition Using Multi-scale Recurrent Networks
  573. System and VLSI Implementation of Phase-based View Synthesis
  574. TCN: Transferable Coupled Network for Cross-Resolution Face Recognition*
  575. TCNN: Temporal Convolutional Neural Network for Real-time Speech Enhancement in the Time Domain
  576. TS-MC: Two Stage Matrix Completion Algorithm for Wireless Sensor Networks
  577. TV-DCT: Method to Impute Gene Expression Data Using DCT Based Sparsity and Total Variation Denoising
  578. Target Localization and Mutual Information Improvement for Cooperative MIMO Radar and MIMO Communication Systems
  579. Target and Non-target Speaker Discrimination by Humans and Machines
  580. Task-Based Quantization for Massive MIMO Channel Estimation
  581. Teach an All-rounder with Experts in Different Domains
  582. Teacher-student Deep Clustering for Low-delay Single Channel Speech Separation
  583. Teacher-student Training for Acoustic Event Detection Using Audioset
  584. Teaching Practical DSP with Off-the-shelf Hardware and Free Software
  585. Teaching Signal Processing Concepts to Digital Natives
  586. Team Policy Learning for Multi-agent Reinforcement Learning
  587. Temporal Salience Based Human Action Recognition
  588. Tensor Matched Kronecker-structured Subspace Detection for Missing Information
  589. Tensor Robust PCA on Graphs
  590. Tensor Super-resolution for Seismic Data
  591. Tensor-Train Discriminant Analysis
  592. Tensor-based Estimation of mmWave MIMO Channels with Carrier Frequency Offset
  593. Tensor-ring Nuclear Norm Minimization and Application for Visual : Data Completion
  594. The CORAL+ Algorithm for Unsupervised Domain Adaptation of PLDA
  595. The Design of Personal Audio Systems for Speech Transmission Using Analytical and Measured Responses
  596. The Direction Cosine Matrix Algorithm in Fixed-point: Implementation and Analysis
  597. The Discrete Cosine Transform on Triangles
  598. The Effect of Spatio-temporal Inconsistency on the Subjective Quality Evaluation of Omnidirectional Videos
  599. The Generalization Effect for Multilingual Speech Emotion Recognition across Heterogeneous Languages
  600. The Geometry of Equality-constrained Global Consensus Problems
  601. The Good, the Bad, Algorithmic Noise Tolerance (Ant), the Ugly
  602. The Impact of Stalling on the Perceptual Quality of HTTP-based Omnidirectional Video Streaming
  603. The Leap Speaker Recognition System for NIST SRE 2018 Challenge
  604. The Limitation and Practical Acceleration of Stochastic Gradient Algorithms in Inverse Problems
  605. The Matched Reassignment Applied to Echolocation Data
  606. The Phasebook: Building Complex Masks via Discrete Representations for Source Separation
  607. The Pytorch-kaldi Speech Recognition Toolkit
  608. The Speechtransformer for Large-scale Mandarin Chinese Speech Recognition
  609. The Universal Manifold Embedding for Estimating Rigid Transformations of Point Clouds
  610. Tied Normal Variance-Mean Mixtures for Linear Score Calibration
  611. Time Difference of Arrival Estimation of Speech Signals Using Deep Neural Networks with Integrated Time-frequency Masking
  612. Time Domain Spherical Harmonic Analysis for Adaptive Noise Cancellation over a Spatial Region
  613. Time Series Prediction for Kernel-based Adaptive Filters Using Variable Bandwidth, Adaptive Learning-rate, and Dimensionality Reduction
  614. Time Signal Classification Using Random Convolutional Features
  615. Time-based Sampling and Reconstruction of Non-bandlimited Signals
  616. Time-frequency-bin-wise Switching of Minimum Variance Distortionless Response Beamformer for Underdetermined Situations
  617. Time-frequency-masking-based Determined BSS with Application to Sparse IVA
  618. Time-varying Graph Learning Based on Sparseness of Temporal Variation
  619. Timescalenet : A Multiresolution Approach for Raw Audio Recognition
  620. To Reverse the Gradient or Not: an Empirical Comparison of Adversarial and Multi-task Learning in Speech Recognition
  621. Toa Source Node Self-positioning with Unknown Clock Skew in Wireless Sensor Networks
  622. Toeplitz Matrix Completion for Direction Finding Using a Modified Nested Linear Array
  623. Token-wise Training for Attention Based End-to-end Speech Recognition
  624. Topic Detection in Conversational Telephone Speech Using CNN with Multi-stream Inputs
  625. Total-variation-regularized Tensor Ring Completion for Remote Sensing Image Reconstruction
  626. Toward Robust Interpretable Human Movement Pattern Analysis in a Workplace Setting
  627. Toward Subjective Violence Detection in Videos
  628. Toward the Quantum Internet: A Directional-dependent Noise Model for Quantum Signal Processing
  629. Towards Audio to Scene Image Synthesis Using Generative Adversarial Network
  630. Towards Automatic Methods to Detect Errors in Transcriptions of Speech Recordings
  631. Towards Better Confidence Estimation for Neural Models
  632. Towards Code-switching ASR for End-to-end CTC Models
  633. Towards Cross-modality Topic Modelling via Deep Topical Correlation Analysis
  634. Towards Disease-specific Speech Markers for Differential Diagnosis in Parkinsonism
  635. Towards End-to-end Speech-to-text Translation with Two-pass Decoding
  636. Towards Generating Ambisonics Using Audio-visual Cue for Virtual Reality
  637. Towards Learned Color Representations for Image Splicing Detection
  638. Towards Perceptually Optimized Sound Zones: A Proof-of-concept Study
  639. Towards Unsupervised Single-channel Blind Source Separation Using Adversarial Pair Unmix-and-remix
  640. Towards Unsupervised Speech-to-text Translation
  641. Towards Visually Grounded Sub-word Speech Unit Discovery
  642. Tracking Dynamic Systems in α-Stable Environments
  643. Tracking Multiple Image Sharing on Social Networks
  644. Tracking a Cluster of Space Debris in Low Orbit by Filtering on Lie Groups
  645. Trainable Adaptive Window Switching for Speech Enhancement
  646. Trainable Time Warping: Aligning Time-series in the Continuous-time Domain
  647. Training Dynamic Exponential Family Models with Causal and Lateral Dependencies for Generalized Neuromorphic Computing
  648. Training Multi-task Adversarial Network for Extracting Noise-robust Speaker Embedding
  649. Training Neural Audio Classifiers with Few Data
  650. Transdrums: A Drum Pattern Transfer System Preserving Global Pattern Structure
  651. Transfer Learning Using Raw Waveform Sincnet for Robust Speaker Diarization
  652. Transfer Learning of Language-independent End-to-end ASR with Language Model Fusion
  653. Transfer and Collaborative Learning Method for Personalized Noninvasive Blood Glucose Measurement Modeling
  654. Transferability of Neural Network Approaches for Low-rate Energy Disaggregation
  655. Transferable Positive/negative Speech Emotion Recognition via Class-wise Adversarial Domain Adaptation
  656. Transferring Piano Performance Control across Environments
  657. Transform Coefficient Coding for Screen Content in Versatile Video Coding (VVC)
  658. Transform Domain Based Medical Image Super-resolution via Deep Multi-scale Network
  659. Transmission Line Cochlear Model Based AM-FM Features for Replay Attack Detection
  660. Transmit Beampattern Design for MIMO Radar with One-bit DACs
  661. Triggered Attention for End-to-end Speech Recognition
  662. Trigonometric Interpolation Beamforming for a Circular Microphone Array
  663. Tropical Modeling of Weighted Transducer Algorithms on Graphs
  664. Truly Unsupervised Acoustic Word Embeddings Using Weak Top-down Constraints in Encoder-decoder Models
  665. Tuning Frequency Dependency in Music Classification
  666. Tuplemax Loss for Language Identification
  667. Turning a Vulnerability into an Asset: Accelerating Facial Identification with Morphing
  668. Two-B-real Net: Two-branch Network for Real-time Salient Object Detection
  669. Two-stream Multi-focus Image Fusion Based on the Latent Decision Map
  670. Type and Leak Your Ethnicity on Smartphones
  671. UTD-CRSS Systems for 2018 NIST Speaker Recognition Evaluation
  672. Understanding Deep Neural Networks through Input Uncertainties
  673. Unified Framework for Minimax MIMO Transmit Beampattern Matching under Waveform Constraints
  674. Unifying Isolated and Overlapping Audio Event Detection with Multi-label Multi-task Convolutional Recurrent Neural Networks
  675. Unifying Probabilistic Models for Time-frequency Analysis
  676. Universal Acoustic Modeling Using Neural Mixture Models
  677. Universal Adversarial Attacks on Text Classifiers
  678. Unmixing Dynamic Pet Images: Combining Spatial Heterogeneity and Non-gaussian Noise
  679. Unrolled Projected Gradient Descent for Multi-spectral Image Fusion
  680. Unsupervised Deep Clustering for Source Separation: Direct Learning from Mixtures Using Spatial Information
  681. Unsupervised Feature Ranking and Selection Based on Autoencoders
  682. Unsupervised Feature Selection Based on Reconstruction Error Minimization
  683. Unsupervised Learning of Deep Features for Music Segmentation
  684. Unsupervised Melody Style Conversion
  685. Unsupervised Person Re-identification Using Reliable and Soft Labels
  686. Unsupervised Polyglot Text-to-speech
  687. Unsupervised Training of a Deep Clustering Model for Multichannel Blind Source Separation
  688. Unsupervised User Clustering in Non-orthogonal Multiple Access
  689. Updates in Bayesian Filtering by Continuous Projections on a Manifold of Densities
  690. Uplink Multi-user MIMO Detection via Parallel Access
  691. User Constrained Thumbnail Generation Using Adaptive Convolutions
  692. Using 3D Residual Network for Spatio-temporal Analysis of Remote Sensing Data
  693. Using Deep-Q Network to Select Candidates from N-best Speech Recognition Hypotheses for Enhancing Dialogue State Tracking
  694. Using Extreme Gradient Boosting to Detect Glottal Closure Instants in Speech Signal
  695. Using RFID Technology to Introduce Properties of LMS
  696. Using Recurrences in Time and Frequency within U-net Architecture for Speech Enhancement
  697. Utterance-level Aggregation for Speaker Recognition in the Wild
  698. Utterance-level End-to-end Language Identification Using Attention-based CNN-BLSTM
  699. Variance Preserving Initialization for Training Deep Neuromorphic Photonic Networks with Sinusoidal Activations
  700. Variational and Hierarchical Recurrent Autoencoder
  701. Vehicle Pose Estimation Using Mask Matching
  702. Video Quality Assessment for Encrypted HTTP Adaptive Streaming: Attention-based Hybrid RNN-HMM Model
  703. Video-based, Occlusion-robust Multi-view Stereo Using Inner-boundary Depths of Textureless Areas
  704. View-invariant Action Recognition from RGB Data via 3D Pose Estimation
  705. Visual Relationship Recognition via Language and Position Guided Attention
  706. Vocal Melody Extraction via DNN-based Pitch Estimation and Salience-based Pitch Refinement
  707. Voice Conversion with Cyclic Recurrent Neural Network and Fine-tuned Wavenet Vocoder
  708. Voice Trigger Detection from Lvcsr Hypothesis Lattices Using Bidirectional Lattice Recurrent Neural Networks
  709. Wav2Letter++: A Fast Open-source Speech Recognition System
  710. Wav2Pix: Speech-conditioned Face Generation Using Generative Adversarial Networks
  711. Waveform Generation for Text-to-speech Synthesis Using Pitch-synchronous Multi-scale Generative Adversarial Networks
  712. Waveform Modeling by Adaptive Weighted Hermite Functions
  713. Waveglow: A Flow-based Generative Network for Speech Synthesis
  714. Wavelength-resolved Neutron Tomography for Crystalline Materials
  715. Wavenilm: A Causal Neural Network for Power Disaggregation from the Complex Power Signal
  716. Weakly Standard Interference Mappings: Existence of Fixed Points and Applications to Power Control in Wireless Networks
  717. Weakly Supervised Instance Segmentation Using Hybrid Networks
  718. West: Word Encoded Sequence Transducers
  719. When CTC Training Meets Acoustic Landmarks
  720. When Can a System of Subnetworks Be Registered Uniquely?
  721. When Not to Classify: Detection of Reverse Engineering Attacks on DNN Image Classifiers
  722. Who Do I Sound like? Showcasing Speaker Recognition Technology by Youtube Voice Search
  723. Why Do Neural Dialog Systems Generate Short and Meaningless Replies? a Comparison between Dialog and Translation
  724. Widely Linear Kernels for Complex-valued Kernel Activation Functions
  725. Windowed Attention Mechanisms for Speech Recognition
  726. Word Characters and Phone Pronunciation Embedding for ASR Confidence Classifier
  727. Word and Class Common Space Embedding for Code-switch Language Modelling
  728. Workload-aware Automatic Parallelization for Multi-GPU DNN Training
  729. Zero Resource Speaking Rate Estimation from Change Point Detection of Syllable-like Units
  730. Zero-mean Convolutional Network with Data Augmentation for Sound Level Invariant Singing Voice Separation

Looking for submission deadlines instead? See the conference deadline calendar.