← All conferences

ICASSP 2016 Accepted Papers

The full list of 1,322 papers accepted at ICASSP 2016 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. RD-SVM: A resilient distributed support vector machine
  2. RSS-based sensor localization in underwater acoustic sensor networks
  3. Radar imaging of stationary indoor targets using joint low-rank and sparsity constraints
  4. Radial filters for near field source separation in spherical harmonic domain
  5. Radioastronomical image reconstruction with regularized least squares
  6. Random access for massive MIMO systems with intra-cell pilot contamination
  7. Random matrix based method for joint DOD and DOA estimation for large scale MIMO radar in non-Gaussian noise
  8. Random projections through multiple optical scattering: Approximating Kernels at the speed of light
  9. Randomized requantization with local differential privacy
  10. Range azimuth indication using a random frequency diverse array
  11. Rank-one tensor injection: A novel method for canonical polyadic tensor decomposition
  12. Ranking the parameters of deep neural networks using the fisher information
  13. Rate analysis for detection of sparse mixtures
  14. Rate analysis of spatial multiplexing in MIMO heterogeneous networks with wireless backhaul
  15. Rate optimization for massive MIMO relay networks: A minorization-maximization approach
  16. Readability enhancement of low light images based on dual-tree complex wavelet transform
  17. Real-time data selection and ordering for cognitive bias mitigation
  18. Real-time integration of statistical model-based speech enhancement with unsupervised noise PSD estimation using microphone array
  19. Real-time joint energy storage management and load scheduling with renewable integration
  20. Real-time multi-candidates fusion based head tracking on Kinect depth sequence
  21. Realistic human action recognition: When deep learning meets VLAD
  22. Recognition of occluded facial expressions based on CENTRIST features
  23. Reconstructing non-point sources of diffusion fields using sensor measurements
  24. Recovering K-sparse N-length vectors in O(K log N) time: Compressed sensing using sparse-graph codes
  25. Recurrent neural network training with dark knowledge transfer
  26. Recurrent neural networks for polyphonic sound event detection in real life recordings
  27. Recurrent support vector machines for speech recognition
  28. Recursive versions of the Levenberg-Marquardt reassigned spectrogram and of the synchrosqueezed STFT
  29. Reduced complexity FFT-based DOA and DOD estimation for moving target in bistatic MIMO radar
  30. Reduced-order modeling of hidden dynamics
  31. Reference-based compressed sensing: A sample complexity approach
  32. Region matching and similarity enhancing for image retrieval
  33. Regression, the periodogram, and the Lomb-Scargle periodogram
  34. Relative location for light field saliency detection
  35. Relative-gradient Bussgang-type blind equalization algorithms
  36. Reliably detecting humans with RGB-D camera with physical blob detector followed by learning-based filtering
  37. Removal of EEG artifacts for BCI applications using fully Bayesian tensor completion
  38. Representations of piecewise smooth signals on graphs
  39. Resilient decentralized consensus-based state estimation for smart grid in presence of false data
  40. Resource allocation for asynchronous cognitive radio networks with FBMC/OFDM under statistical CSI
  41. Retiming and dual-supply voltage based energy optimization for DSP applications
  42. Retinal vessel enhancement using multi-dictionary and sparse coding
  43. Retrieving audio recordings using musical themes
  44. Reversible data hiding in encrypted image based on block histogram shifting
  45. Revertible deep convolutional networks with iterated directional filter bank
  46. Risk assessment for RGBD scans in real time
  47. Risk-sensitive decision making via constrained expected returns
  48. Robust CDMA receiver design under disguised jamming
  49. Robust MVDR beamforming using time-frequency masks for online/offline ASR in noise
  50. Robust TTS duration modelling using DNNS
  51. Robust adaptive beamforming based on DOA support using decomposed coprime subarrays
  52. Robust artificial-noise aided transmit design for multi-user MISO systems with integrated services
  53. Robust audiovisual speech recognition using noise-adaptive linear discriminant analysis
  54. Robust blind source separation in a reverberant room based on beamforming with a large-aperture microphone array
  55. Robust blind spikes deconvolution
  56. Robust dictionary learning: Application to signal disaggregation
  57. Robust geographical load balancing for sustainable data centers
  58. Robust image hashing based on low-rank and sparse decomposition
  59. Robust lane marking detection using boundary-based inverse perspective mapping
  60. Robust multiple speech source localization using time delay histogram
  61. Robust pilot decontamination: A joint angle and power domain approach
  62. Robust pitch tracking in noisy speech using speaker-dependent deep neural networks
  63. Robust receiver design based on FEC code diversity in pilot-contaminated multi-user massive MIMO systems
  64. Robust saliency propagation based on random walks
  65. Robust sparse recovery for compressive sensing in impulsive noise using ℓp-norm model fitting
  66. Robust sparsity-promoting acoustic multi-channel equalization for speech dereverberation
  67. Robust speaker DOA estimation with single AVS in bispectrum domain
  68. Robust speech recognition from ratio masks
  69. Robust speech recognition using multivariate copula models
  70. Robust submodular data partitioning for distributed speech recognition
  71. Robust transmit precoding for underlay MIMO cognitive radio with interference leakage rate limit
  72. Robust visual tracking via inverse nonnegative matrix factorization
  73. Robust volume minimization-based matrix factorization via alternating optimization
  74. Robust waveform design of wideband cognitive radar for extended target detection
  75. Room geometry estimation from acoustic echoes using graph-based echo labeling
  76. Rotating coded aperture for depth from defocus
  77. SAR image target recognition using kernel sparse representation based on reconstruction coefficient energy maximization rule
  78. SAT-LHUC: Speaker adaptive training for learning hidden unit contributions
  79. SBL-based joint target imaging and Doppler frequency estimation in monostatic MIMO radar systems
  80. SIMD-based datapath with efficient operation structure for motion estimation
  81. SINR performance of matched illumination signals with dynamic target models
  82. SNR-invariant PLDA with multiple speaker subspaces
  83. STC anti-spoofing systems for the ASVspoof 2015 challenge
  84. SVR based double-scale regression for dynamic emotion prediction in music
  85. Safe screening tests for LASSO based on firmly non-expansiveness
  86. SalSi: A new seismic attribute for salt dome detection
  87. Saliency & structure preserving multi-operator image retargeting
  88. Saliency analysis based on depth contrast increased
  89. Saliency detection based on integration of central bias, reweighting and multi-scale for superpixels
  90. Saliency detection using tensor sparse reconstruction residual analysis
  91. Saliency preprocessing for person re-identification images
  92. Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering
  93. Scaling and occlusion robust athlete tracking in sports videos
  94. Scanned document enhancement based on fast text detection
  95. Scene text recognition with high performance CNN classifier and efficient word inference
  96. Secrecy degrees of freedom of a MIMO Gaussian wiretap channel with a cooperative jammer
  97. Secure M-PSK communication via directional modulation
  98. Secure performance analysis of buffer-aided cognitive relay networks under delay unconstraint case
  99. Segment-oriented evaluation of speaker diarisation performance
  100. Selection and combination of hypotheses for dialectal speech recognition
  101. Self-stabilized deep neural network
  102. Semantic word embedding neural network language models for automatic speech recognition
  103. Semi-autonomous data enrichment based on cross-task labelling of missing targets for holistic speech analysis
  104. Semi-non-negative matrix factorization using alternating direction method of multipliers for voice conversion
  105. Semi-supervised learning in the presence of novel class instances
  106. Sequence design to minimize the peak sidelobe level
  107. Sequence summarizing neural network for speaker adaptation
  108. Sequence training of multi-task acoustic models using meta-state labels
  109. Sequential Monte Carlo sampling for correlated latent long-memory time-series
  110. Shadow detection using double-threshold pulse coupled neural networks
  111. Shape initialization without ground truth for face alignment
  112. Shape: Linear-time camera pose estimation with quadratic error-decay
  113. Shifted and convolutive source-filter non-negative matrix factorization for monaural audio source separation
  114. Ship wake detection for SAR images with complex backgrounds based on morphological dictionary learning
  115. Siamese neural network based gait recognition for human identification
  116. Signal detection in para complex normal noise
  117. Signal detection of ambient backscatter system with differential modulation
  118. Signal processing concepts help teach optical engineering
  119. Signal processing on graphs: Performance of graph structure estimation
  120. Signal reconstruction in the presence of side information: The impact of projection kernel design
  121. Signal sparsity estimation from compressive noisy projections via γ-sparsified random matrices
  122. Signal-adaptive switching of overlap ratio in audio transform coding
  123. Signer-independent fingerspelling recognition with deep neural network adaptation
  124. Significance of Pseudo-syllables in building better acoustic models for Indian English TTS
  125. Simple multi frame analysis methods for estimation of amplitude spectral envelope estimation in singing voice
  126. Simplified learning with binary orthogonal constraints
  127. Simplified multi-bit SC list decoding for polar codes
  128. Simplifying long short-term memory acoustic models for fast training and decoding
  129. Single image brightening via exposure fusion
  130. Single underwater image restoration by blue-green channels dehazing and red channel correction
  131. Single-microphone speech enhancement using MVDR filtering and Wiener post-filtering
  132. Sketching for large-scale learning of mixture models
  133. Smartphone-based real-time classification of noise signals using subband features and random forest classifier
  134. Smooth talking: Articulatory join costs for unit selection
  135. Social force model aided robust particle PHD filter for multiple human tracking
  136. Soft linear discriminant analysis (SLDA) for pattern recognition with ambiguous reference labels: Application to social signal processing
  137. Song recommendation with non-negative matrix factorization and graph total variation
  138. Sound field decomposition in reverberant environment using sparse and low-rank signal models
  139. Sound source localization based on deep neural networks with directional activate function exploiting phase information
  140. Source cell phone matching from speech recordings by sparse representation and KISS metric
  141. Source localization on solids utilizing logistic modeling of energy transition in vibration signals
  142. Source modeling for HMM based speech synthesis using integrated LP residual
  143. Source-specific system identification
  144. Space-shift sampling of graph signals
  145. Sparse Bayesian dictionary learning with a Gaussian hierarchical model
  146. Sparse PCA via hard thresholding for blind source separation
  147. Sparse attacking strategies in multi-sensor dynamic systems maximizing state estimation errors
  148. Sparse canonical correlation analysis based on rank-1 matrix approximation and its application for FMRI signals
  149. Sparse coding with fast image alignment via large displacement optical flow
  150. Sparse complex FxLMS for active noise cancellation over spatial regions
  151. Sparse deconvolution for moving-source localization
  152. Sparse phase retrieval with near minimal measurements: A structured sampling based approach
  153. Sparse reconstruction of quantized speech signals
  154. Sparse reconstruction-based angle-range-polarization-dependent beamforming with polarization sensitive frequency diverse array
  155. Sparse recovery of multiple measurement vectors in impulsive noise: A smooth block successive minimization algorithm
  156. Sparse signal recovery methods for variant detection in next-generation sequencing data
  157. Sparse sound field decomposition with multichannel extension of complex NMF
  158. Sparsity based multi-target tracking using mobile sensors
  159. Sparsity-based direction-of-arrival estimation for strictly non-circular sources
  160. Sparsity-based localization of spatially coherent distributed sources
  161. Sparsity-based reconstruction method for signals with finite rate of innovation
  162. Sparsity-promoting sensor selection with energy harvesting constraints
  163. Spatial correlation model based observation vector clustering and MVDR beamforming for meeting recognition
  164. Spatial feature learning for robust binaural sound source localization using a composite feature vector
  165. Spatio-temporal mid-level feature bank for action recognition in low quality video
  166. Speaker adaptation OF RNN-BLSTM for speech recognition based on speaker code
  167. Speaker adaptive model based on Boltzmann machine for non-parallel training in voice conversion
  168. Speaker adaptive training in deep neural networks using speaker dependent bottleneck features
  169. Speaker age estimation on conversational telephone speech using senone posterior based i-vectors
  170. Speaker and language factorization in DNN-based TTS synthesis
  171. Speaker cluster-based speaker adaptive training for deep neural network acoustic modeling
  172. Speaker diarization with unsupervised training framework
  173. Speaker recognition using matched filters
  174. Speaker-aware training of LSTM-RNNS for acoustic modelling
  175. Speech analysis of sung-speech and lyric recognition in monophonic singing
  176. Speech dereverberation using linear prediction with estimation of early speech spectral variance
  177. Speech emotion recognition using transfer non-negative matrix factorization
  178. Speech enhancement based on neural networks applied to cochlear implant coding strategies
  179. Speech enhancement using an MMSE spectral amplitude estimator based on a modulation domain Kalman filter with a Gamma prior
  180. Speech recognition robust against speech overlapping in monaural recordings of telephone conversations
  181. Spherical microphone array acoustic rake receivers
  182. Spoofing detection from a feature representation perspective
  183. Spread spectrum compressed sensing MRI using chirp radio frequency pulses
  184. Stability analysis of the least-mean-magnitude-phase algorithm
  185. Stabilization of adaptive eigenvector extraction by continuation in nested orthogonal complement structure
  186. Stable and symmetric filter convolutional neural network
  187. Stable dysphonia measures selection for Parkinson speech rehabilitation via diversity regularized ensemble
  188. Stacked correlation filters for biometric verification
  189. Statistical F0 prediction for electrolaryngeal speech enhancement considering generative process of F0 contours within product of experts framework
  190. Statistical analysis of neuronal population codes for encoding acute pain
  191. Statistical near-far detection techniques for GNSS snapshot receivers
  192. Steganalysis of AAC using calibrated Markov model of adjacent codebook
  193. Stochastic energy management in distribution grids
  194. Stochastic load scheduling for risk-limiting economic dispatch in smart microgrids
  195. Stochastic online control for smart-grid powered MIMO downlink transmissions
  196. Stochastic proximal gradient consensus over time-varying networks
  197. Stochastic thermodynamic integration: Efficient Bayesian model selection via stochastic gradient MCMC
  198. Structural maximum a posteriori speaker adaptation of speaking rate-dependent hierarchical prosodic model for Mandarin TTS
  199. Structural segmentation with the Variable Markov Oracle and boundary adjustment
  200. Structural spatio-temporal transform for robust visual tracking
  201. Structurally-constrained gradient descent for matrix factorization in haplotype assembly problems
  202. Structure-guided image completion via regularity statistics
  203. Student's T nonnegative matrix factorization and positive semidefinite tensor factorization for single-channel audio source separation
  204. Study of attenuation due to wet antenna in microwave radio communication
  205. Style retrieval from natural images
  206. Style-centric image summarization from photographic views of a city
  207. Subspace clustering with a learned dimensionality reduction projection
  208. Subspace fitting via sparse representation of signal covariance for DOA estimation
  209. Subspace superdirective beamformers based on joint diagonalization
  210. Subspace-based adaptive widely linear blind channel estimation for constrained minimum variance CDMA receiver
  211. Sum secrecy rate maximization for full-duplex two-way relay networks
  212. Super nested arrays: Sparse arrays with less mutual coupling than nested arrays
  213. Super-resolution DOA estimation via continuous group sparsity in the covariance domain
  214. Super-resolution spectral analysis for ultrasound scatter characterization
  215. Super-resolved time-of-flight sensing via FRI sampling theory
  216. Superimposed pilots: An alternative pilot structure to mitigate pilot contamination in massive MIMO
  217. Supervised and unsupervised active learning for automatic speech recognition of low-resource languages
  218. Supervised speech dereverberation in noisy environments using exemplar-based sparse representations
  219. Supervised subspace learning based on deep randomized networks
  220. Supervised-learning based face hallucination for enhancing face recognition
  221. Symmetric matrix perturbation for differentially-private principal component analysis
  222. Synaptic depression in deep neural networks for speech processing
  223. Synthesis of Volterra filters for the parametric array loudspeaker
  224. System architectures for communication-aware multi-robot navigation
  225. System combination with log-linear models
  226. System fusion and speaker linking for longitudinal diarization of TV shows
  227. System-compatible robustness improvement for new generation dect decoders by G.722 soft-decision decoding
  228. TC: Throughput centric successive cancellation decoder hardware implementation for polar codes
  229. Tag recommendation via robust probabilistic discriminative matrix factorization
  230. Taking meredith out of Grey's anatomy: Automating hospital ICU emergency signaling
  231. Target detection for depth imaging using sparse single-photon data
  232. Task-driven deep transfer learning for image classification
  233. Template based techniques for automatic segmentation of TTS unit database
  234. Tensor beamforming for multilinear translation invariant arrays
  235. Tensor completion via adaptive sampling of tensor fibers: Application to efficient indoor RF fingerprinting
  236. Tensor completion via functional smooth component deflation
  237. Tensor-based subspace learning for tracking salt-dome boundaries constrained by seismic attributes
  238. Terrain-scattered jammer suppression in MIMO radar using space-(fast) time adaptive processing
  239. Testing for impropriety of multivariate complex random processes
  240. Testing the consistency assumption: Pronunciation variant forced alignment in read and spontaneous speech synthesis
  241. The Rao test for testing handedness of complex-valued covariance matrix
  242. The divergence behavior of adaptive signal processing algorithms with finite search horizon
  243. The effect of vocal fry on pitch perception
  244. The first-order high-pass filter influences the automatic measurements of the electrocardiogram
  245. The graph FRI framework-spline wavelet theory and sampling on circulant graphs
  246. The intrinsic value of HFO features as a biomarker of epileptic activity
  247. The matching-minimization algorithm, the INCA algorithm and a mathematical framework for voice conversion with unaligned corpora
  248. The method for defocusing selfie taken by mobile frontal camera using burst shot
  249. The multiple-point variogram of images for robust texture classification
  250. The recursive hessian sketch for adaptive filtering
  251. The relationship of voice onset time and Voice Offset Time to physical age
  252. The spherical harmonics root-music
  253. The steady-state of the (Normalized) LMS is schur convex
  254. The use of unit norm tight measurement matrices for one-bit compressed sensing
  255. Theoretical guarantees for poisson disk sampling using pair correlation function
  256. Time domain acoustic contrast control implementation of sound zones for low-frequency input signals
  257. Time-resolved image demixing
  258. Time-varying frequency modes of resting fMRI brain networks reveal significant gender differences
  259. Title assignment for automatic topic segments in TV broadcast news
  260. Tomographic reconstruction of atmospheric density with Mumford-Shah functionals
  261. Towards PLDA-RBM based speaker recognition in mobile environment: Designing stacked/deep PLDA-RBM systems
  262. Towards a behaviorally-validated computational audiovisual saliency model
  263. Towards a characterization of the uncertainty curve for graphs
  264. Towards an automatic monitoring of the neurological state of Parkinson's patients from speech
  265. Towards implicit complexity control using variable-depth deep neural networks for automatic speech recognition
  266. Towards information-based feedback control for binaural active localization
  267. Towards multi-rigid body localization
  268. Towards optimal vlad for human action recognition from still images
  269. Towards robust close-talking microphone arrays for noise reduction in mobile phones
  270. Track selection in multifunction radars for multi-target tracking: An anti-coordination game
  271. Tradeoff between quality and quantity of emotional annotations to characterize expressive behaviors
  272. Trading accuracy for numerical stability: Orthogonalization, biorthogonalization and regularization
  273. Traffic-aware association in heterogeneous networks
  274. Training deep neural-networks based on unreliable labels
  275. Trajectory training considering global variance for speech synthesis based on neural networks
  276. Transform domain temporal prediction with extended blocks
  277. Transient model of EEG using Gini Index-based matching pursuit
  278. Tree-structured probabilistic model of monophonic written music based on the generative theory of tonal music
  279. Triple-based analysis of music alignments without the need of ground-truth annotations
  280. True time delay beamspace wideband source localization
  281. Turbo compressed sensing using message passing de-quantization
  282. Twin-HMM-based non-intrusive speech intelligibility prediction
  283. Two-dimensional correlated topic models
  284. Two-dimensional positive spline smoothing and its application to probability density estimation
  285. Two-stage noise aware training using asymmetric deep denoising autoencoder
  286. Type-2 fuzzy GMM for text-independent speaker verification under unseen noise conditions
  287. Typically developed adults and adults with autism spectrum disorder classification using centre of pressure measurements
  288. UTD-CRSS system for the NIST 2015 language recognition i-vector machine learning challenge
  289. Uniform expected likelihood solution for interference rejection combining regularization
  290. Universal encoding of multispectral images
  291. Universal outlying sequence detection for continuous observations
  292. Unmixing multitemporal hyperspectral images with variability: An online algorithm
  293. Unsupervised diffusion-based LMS for node-specific parameter estimation over wireless sensor networks
  294. Unsupervised neighbor dependent nonlinear unmixing
  295. Unsupervised spatiotemporal video clustering a versatile mean-shift formulation robust to total object occlusions
  296. Unsupervised speaker adaptation for DNN-based TTS synthesis
  297. Unsupervised time-series clustering of distorted and asynchronous temporal patterns
  298. Unsupervised user intent modeling by feature-enriched matrix factorization
  299. Using conditional restricted Boltzmann machines for spectral envelope modeling in speech bandwidth extension
  300. Using continuous lexical embeddings to improve symbolic-prosody prediction in a text-to-speech front-end
  301. Using hydrodynamical simulations of stellar atmospheres for periodogram standardization: Application to exoplanet detection
  302. VMF-SNE: Embedding for spherical data
  303. Variable span filters for speech enhancement
  304. Variational Bayesian image fusion based on combined sparse representations
  305. Variational inference for infinite mixtures of sparse Gaussian processes through KL-correction
  306. Vectorial total variation based on arranged structure tensor for multichannel image restoration
  307. Very deep multilingual convolutional neural networks for LVCSR
  308. Vibration parameter estimation using FMCW radar
  309. View synthesis based on temporal prediction via warped motion vector fields
  310. Visual tracking via multi-task non-negative matrix factorization
  311. Visual tracking via robust multi-task multi-feature joint sparse representation
  312. Visualizations relevant to the user by multi-view latent variable factorization
  313. Voice Morphing that improves TTS quality using an optimal dynamic frequency warping-and-weighting transform
  314. Wavelet features for classification of vote snore sounds
  315. Wavelet-based decomposition of F0 as a secondary task for DNN-based speech synthesis with multi-task learning
  316. What to do about noisy consensus?
  317. WiFi action recognition via vision-based methods
  318. Wide matching - An approach to improving noise robustness for speech enhancement
  319. Wideband multilinear array processing through tensor decomposition
  320. Wind speed estimation of low-altitude wind-shear based on multiple Doppler channels joint adaptive processing
  321. Work-efficient parallel non-maximum suppression for embedded GPU architectures
  322. Zero-shot learning of intent embeddings for expansion by convolutional deep structured semantic models

Looking for submission deadlines instead? See the conference deadline calendar.