← All conferences

ICASSP 2016 Accepted Papers

The full list of 1,322 papers accepted at ICASSP 2016 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. 1-Bit compressed sensing of positive semi-definite matrices via rank-1 measurement matrices
  2. 2.5D higher order ambisonics for a sound field described by angular spectrum coefficients
  3. 3D DOA estimation of multiple sound sources based on spatially constrained beamforming driven by intensity vectors
  4. 3D acoustic source localization in the spherical harmonic domain based on optimized grid search
  5. 3D mesh steganalysis using local shape features
  6. 3D panorama reconstruction based on sitemap joining
  7. 3D pseudolinear Kalman filter with own-ship path optimization for AOA target tracking
  8. 3D video frame interpolation via adaptive hybrid motion estimation and compensation
  9. A Bayesian framework for the multifractal analysis of images using data augmentation and a whittle approximation
  10. A GMM-based stair quality model for human perceived JPEG images
  11. A Gaussian mixture regression approach toward modeling the affective dynamics between acoustically-derived vocal arousal score (VC-AS) and internal brain fMRI bold signal response
  12. A KL divergence and DNN approach to cross-lingual TTS
  13. A MIL-based interactive approach for hotspot segmentation from bone scintigraphy
  14. A Metropolis-within-Gibbs sampler to infer task-based functional brain connectivity
  15. A behavior-based evaluation of product quality
  16. A benchmark for robustness analysis of visual tracking algorithms
  17. A comparative study of multi-channel processing methods for noisy automatic speech recognition in urban environments
  18. A comparative study of recurrent neural network models for lexical domain classification
  19. A comparative study of robustness of deep learning approaches for VAD
  20. A comparison between deep neural nets and kernel acoustic models for speech recognition
  21. A comparison of ASR and human errors for transcription of non-native spontaneous speech
  22. A cost-effective minutiae disk code for fingerprint recognition and its implementation
  23. A data set providing synthetic and real-world fisheye video sequences
  24. A deep auto-encoder based low-dimensional feature extraction from FFT spectral envelopes for statistical parametric speech synthesis
  25. A deep bidirectional long short-term memory based multi-scale approach for music dynamic emotion prediction
  26. A deep scattering spectrum - Deep Siamese network pipeline for unsupervised acoustic modeling
  27. A deterministic plus noise model of excitation signal using principal component analysis for parametric speech synthesis
  28. A diffusion kernel LMS algorithm for nonlinear adaptive networks
  29. A distributed algorithm for robust LCMV beamforming
  30. A divide-and-conquer dictionary learning algorithm and its performance analysis
  31. A dynamic Bayesian network approach for device-free radio vision: Modeling, learning and inference for body motion recognition
  32. A fast 3D face reconstruction method from a single image using adjustable model
  33. A fast direct source localization approach for acoustic sensor array
  34. A fast dual iterative algorithm for convexly constrained spline smoothing
  35. A fast method of fog and haze removal
  36. A frame rate up-conversion method with quadruple motion vector post-processing
  37. A framework for globally optimal energy-efficient resource allocation in wireless networks
  38. A full training framework of cross-stream dependence modelling for HMM-based singing voice synthesis
  39. A general baseband volterra model for dual-band predistortion
  40. A general framework for reconstruction and classification from compressive measurements with side information
  41. A generalized Bayesian model for tracking long metrical cycles in acoustic music signals
  42. A generalized LDPC framework for robust and sublinear compressive sensing
  43. A generative-discriminative hybrid approach to multi-channel noise reduction for robust automatic speech recognition
  44. A generator of memory-based, runtime-reconfigurable 2N3M5K FFT engines
  45. A glottal chink model for the synthesis of voiced fricatives
  46. A hierarchical algorithm for causality discovery among atrial fibrillation electrograms
  47. A hierarchical framework for language identification
  48. A highly parallel coding unit size selection for HEVC
  49. A joint approach to vector road map registration and vehicle tracking for wide area motion imagery
  50. A joint design approach for spectrum sharing between radar and communication systems
  51. A joint learning approach for cross domain age estimation
  52. A largest matching area approach to image denoising
  53. A lattice algorithm for optimal phase unwrapping in noise
  54. A least square approach for distributed sensor fusion in bandwidth-constrained sensor networks
  55. A linear operator for the computation of soundfield maps
  56. A linear sensor array with self-bending sensitivity
  57. A low complexity weighted least squares narrowband DOA estimator for arbitrary array geometries
  58. A low-cost solution to 3D pinna modeling for HRTF prediction
  59. A machine learning approach for computationally and energy efficient speech enhancement in binaural hearing aids
  60. A map-based NMF approach to hyperspectral image unmixing using a linear-quadratic mixture model
  61. A maximum likelihood-based unscented Kalman filter for multipath mitigation in a multi-correlator based GNSS receiver
  62. A method for predicting the intelligibility of noisy and non-linearly enhanced binaural speech
  63. A method to reconstruct coverage loss maps based on matrix completion and adaptive sampling
  64. A multi-scale approach to extract meaningful annotations from document images
  65. A multimodal analysis of synchrony during dyadic interaction using a metric based on sequential pattern mining
  66. A multimodal mixture-of-experts model for dynamic emotion prediction in movies
  67. A new approach for heart rate monitoring using photoplethysmography signals contaminated by motion artifacts
  68. A new array geometry for DOA estimation with enhanced degrees of freedom
  69. A new haze image database with detailed air quality information and a novel no-reference image quality assessment method for haze images
  70. A new low-rank solution result for a semidefinite program problem subclass with applications to transmit beamforming optimization
  71. A new time-frequency approach for underdetermined convolutive blind speech separation
  72. A new uncertainty decoding scheme for DNN-HMM hybrid systems with multichannel speech enhancement
  73. A normalized spatial spectrum for DOA estimation with uniform linear arrays in the presence of unknown mutual coupling
  74. A novel DNN-HMM-based approach for extracting single loads from aggregate power signals
  75. A novel array processing method for precise depth detection of ultrasound point scatter
  76. A novel color space based on RGB color barycenter
  77. A novel feedforward noise shaping for word-length reduction
  78. A novel generalized assignment framework for the classification of hyperspectral image
  79. A novel image classifier based on Gaussian mixture language model
  80. A novel sub-Nyquist Fourier transform estimator based on alias-free hybrid stratified sampling
  81. A novel time-frequency feature extraction algorithm based on dictionary learning
  82. A novel video-based smoke detection method based on color invariants
  83. A parameter-free Cauchy-Schwartz information measure for independent component analysis
  84. A partial least squares based ranker for fast and accurate age estimation
  85. A partitioned approach to signal separation with microphone ad hoc arrays
  86. A penalty-BSUM approach for rate optimization in full-duplex MIMO relay networks with relay processing delay
  87. A phonetically aware system for speech activity detection
  88. A polynomial optimization approach for robust beamforming design in a device-to-device two-hop one-way relay network
  89. A practical clock synchronization algorithm for UWB positioning systems
  90. A precision-improved processing architecture of physical computing for energy-efficient SIFT feature extraction
  91. A rate-splitting approach to robust multiuser MISO transmission
  92. A real-time example-based single-image super-resolution algorithm via cross-scale high-frequency components self-learning
  93. A recursive predictive risk estimate for proximal algorithms
  94. A risk-unbiased approach to a new Cramér-Rao bound
  95. A robust Gaussian approximate filter for nonlinear systems with heavy tailed measurement noises
  96. A robust speech rate estimation based on the activation profile from the selected acoustic unit dictionary
  97. A score-informed shift-invariant extension of complex matrix factorization for improving the separation of overlapped partials in music recordings
  98. A segment-sliding reconstruction scheme for pulsed radar echoes with sub-Nyquist sampling
  99. A semi-global matching method for large-scale light field images
  100. A semidefinite relaxation approach to the geolocation of two unknown co-channel emitters by a cluster of formation-flying satellites using both TDOA and FDOA measurements
  101. A short-graph fourier transform via personalized pagerank vectors
  102. A single-channel noise cancelation filter in the short-time-fourier-transform domain
  103. A source/filter model with adaptive constraints for NMF-based speech separation
  104. A sparse regression based approach for cuff-less blood pressure measurement
  105. A sparse-graph-coded filter bank approach to minimum-rate spectrum-blind sampling
  106. A speaker adaptation technique for Gaussian process regression based speech synthesis using feature space transform
  107. A speech enhancement system using binaural hearing aids and an external microphone
  108. A study of different weighting schemes for spoken language understanding based on convolutional neural networks
  109. A study of rank-constrained multilingual DNNS for low-resource ASR
  110. A subjective listening test of six different artificial bandwidth extension approaches in English, Chinese, German, and Korean
  111. A swiss army knife for finite rate of innovation sampling theory
  112. A thin-slice perception of emotion? An information theoretic-based framework to identify locally emotion-rich behavior segments for global affect recognition
  113. A topography structure used in audio steganography
  114. A transfer learning method for PLDA-based speaker verification
  115. A unified approach to the design of IIR and FIR notch filters
  116. A unified framework for atlas-based segmentation with forward deformation and label refinement
  117. A unified sparse signal decomposition and reconstruction framework for elimination of muscle artifacts from ECG signal
  118. A weakly-supervised discriminative model for audio-to-score alignment
  119. A weighted STOI intelligibility metric based on mutual information
  120. A weighted atomic norm approach to spectral super-resolution with probabilistic priors
  121. AAC encoding detection and bitrate estimation using a convolutional neural network
  122. ALADDIN: A locality aligned deep model for instance search
  123. Abnormal event detection based on sparse reconstruction in crowded scenes
  124. Abnormal sound event detection using temporal trajectories mixtures
  125. Accelerated spectral clustering using graph filtering of random signals
  126. Accelerating multi-user large vocabulary continuous speech recognition on heterogeneous CPU-GPU platforms
  127. Accelerating stochastic computation for binary classification applications
  128. Accurate asymptotic analysis for John's test in multichannel signal detection
  129. Accurate recovery of a specularity from a few samples of the reflectance function
  130. Achieving global optimality for wirelessly-powered multi-antenna TWRC with lattice codes
  131. Acoustic data-driven pronunciation lexicon generation for logographic languages
  132. Acoustic event detection based on non-negative matrix factorization with mixtures of local dictionaries and activation aggregation
  133. Acoustic scene classification with matrix factorization for unsupervised feature learning
  134. Acoustic simultaneous localization and mapping (A-SLAM) of a moving microphone array and its surrounding speakers
  135. Acoustic source separation using the short-time quaternion fourier transforms of particle velocity signals
  136. Action recognition using interest points capturing differential motion information
  137. Active eavesdropping via spoofing relay attack
  138. Active learning for magnetic resonance image quality assessment
  139. Active learning on weighted graphs using adaptive and non-adaptive approaches
  140. Active online learning of trusts in social networks
  141. Adapting ASR for under-resourced languages using mismatched transcriptions
  142. Adaptive Boolean compressive sensing by using multi-armed bandit
  143. Adaptive algorithms for hypergraph learning
  144. Adaptive consensus-based distributed detection in WSN with unreliable links
  145. Adaptive distributed compressed estimation based on recursive least squares with sensing matrix design
  146. Adaptive enhancement of luminance and details in images under ambient light
  147. Adaptive extraction of repeating non-negative temporal patterns for single-channel speech enhancement
  148. Adaptive learning for stochastic generalized Nash equilibrium problems
  149. Adaptive margin slack minimization in RKHS for classification
  150. Adaptive radar detection in the presence of Gaussian clutter with symmetric spectrum
  151. Adaptive rate control algorithm for SHVC: Application to HD/UHD
  152. Adaptive regularization for BEM channel estimation in multicarrier systems
  153. Adaptive reverberation cancelation for multizone soundfield reproduction using sparse methods
  154. Adaptive sequential optimization with applications to machine learning
  155. Adaptive sparsity tradeoff for ℓ1-constraint NLMS algorithm
  156. Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent network
  157. Advanced b-vector system based deep neural network as classifier for speaker verification
  158. Adversarial Bandit for online interactive active learning of zero-shot spoken language understanding
  159. Agreement and disagreement classification of dyadic interactions using vocal and gestural cues
  160. Algebraic solution for stationary emitter geolocation by a LEO satellite using Doppler frequency measurements
  161. Algorithm for DNA copy number variation detection with read depth and paramorphism information
  162. An acoustic keystroke transient canceler for speech communication terminals using a semi-blind adaptive filter model
  163. An adaptive fixed-point IVA algorithm applied to multi-subject complex-valued FMRI data
  164. An adaptive multi-level wavelet denoising method for 40-Hz ASSR
  165. An adaptive resolution rate control method for intra coding in HEVC
  166. An adaptive robust regression method: Application to galaxy spectrum baseline estimation
  167. An alternative approach for auditory attention tracking using single-trial EEG
  168. An alternative proof for the identifiability of independent vector analysis using second order statistics
  169. An approximate message passing approach for tensor-based seismic data interpolation with randomly missing traces
  170. An effective color space for face recognition
  171. An effective performance ranking mechanism to image dehazing methods with psychological inference benchmark
  172. An efficient anomaly detection approach in surveillance video based on oriented GMM
  173. An efficient method for polyphonic audio-to-score alignment using onset detection and constant Q transform
  174. An empirical exploration of CTC acoustic models
  175. An energy-aware auction for hybrid access in heterogeneous networks under QoS requirements
  176. An energy-efficient compressive sensing framework incorporating online dictionary learning for long-term wireless health monitoring
  177. An engineer's guide to particle filtering on matrix Lie groups
  178. An estimation method of voice timbre evaluation values using feature extraction with Gaussian mixture model based on reference singer
  179. An expectation-maximization eigenvector clustering approach to direction of arrival estimation of multiple speech sources
  180. An extensible speaker identification sidekit in Python
  181. An image smoothing operator for fast and accurate scale space approximation
  182. An improved DOA estimation algorithm for circular and non-circular signals with high resolution
  183. An improved anthropometry-based customization method of individual head-related transfer functions
  184. An improved local binary pattern operator for texture classification
  185. An information theoretic framework for order of operations forensics
  186. An introduction to hypergraph signal processing
  187. An inverse-gamma source variance prior with factorized parameterization for audio source separation
  188. An investigation into using parallel data for far-field speech recognition
  189. An iterative hard thresholding approach to ℓ0 sparse Hellinger NMF
  190. An iterative sure-let approach to sparse reconstruction
  191. An iteratively reweighted method for recovery of block-sparse signal with unknown block partition
  192. An online algorithm for throughput maximization of wireless powered communication networks
  193. An online tensor robust PCA algorithm for sequential 2D data
  194. An optimization framework for combining multiple graphs
  195. An unbiased risk estimator for Gaussian mixture noise distributions - Application to speech denoising
  196. Analog multiple descriptions: A zero-delay source-channel coding approach
  197. AnalogCast: Full linear coding and pseudo analog transmission for satellite remote-sensing images
  198. Analysis of DNN approaches to speaker identification
  199. Analysis of distributed ADMM algorithm for consensus optimization in presence of error
  200. Analysis of error resiliency of belief propagation in computer vision
  201. Analysis of natural and synthetic speech using Fujisaki model
  202. Analysis of p-norm regularized subproblem minimization for sparse photon-limited image recovery
  203. Analysis of secure communication in millimeter wave networks: Are blockages beneficial?
  204. Analytical performance assessment of esprit-type algorithms for coexisting circular and strictly non-circular signals
  205. Annealed learning based block transforms for HEVC video coding
  206. Anti-occlusion observation model and automatic recovery for multi-view ball tracking in sports analysis
  207. Applications of 3D spherical transforms to personalization of head-related transfer functions
  208. Approximate search of audio queries by using DTW with phone time boundary and data augmentation
  209. Are there approximate fast fourier transforms on graphs?
  210. Array thinning for antenna selection in millimeter wave MIMO systems
  211. Asking for a second opinion: Re-querying of noisy multi-class labels
  212. Aspect Ratio Similarity (ARS) for image retargeting quality assessment
  213. Asymptotic analysis of downlink MISO systems over Rician fading channels
  214. Asymptotic closed-loop design of error resilient predictive compression systems
  215. Asymptotic optimal quantizer design for distributed Bayesian estimation
  216. Asymptotic perfect secrecy in distributed detection against a global eavesdropper
  217. Asymptotic performance analysis for 1-bit Bayesian smoothing
  218. Asynchronous distributed alternating direction method of multipliers: Algorithm and convergence analysis
  219. Asynchronous local voltage control in power distribution networks
  220. Asynchronous systems for constraint satisfaction: Filtering and stability
  221. Atmospheric turbulence mitigation based on turbulence extraction
  222. Audio enhancing with DNN autoencoder for speaker recognition
  223. Audio watermarking based on empirical mode decomposition and beat detection
  224. Audio word similarity for clustering with zero resources based on iterative HMM classification
  225. Audio-based multimedia event detection using deep recurrent neural networks
  226. Auditory attention decoding with EEG recordings using noisy acoustic reference signals
  227. Autocalibration of lidar and optical cameras via edge alignment
  228. Automatic Chord estimation on seventhsbass Chord vocabulary using deep neural network
  229. Automatic allocation of NTF components for user-guided audio source separation
  230. Automatic composition of broadcast news summaries using rank classifiers trained with acoustic and lexical features
  231. Automatic gain control for parametric array loudspeakers
  232. Automatic human fall detection in fractional fourier domain for assisted living
  233. Automatic image region annotation through segmentation based visual semantic analysis and discriminative classification
  234. Automatic speech recognition for acoustical analysis and assessment of cantonese pathological voice and speech
  235. Auxiliary beam pair design in mmWave cellular systems with hybrid precoding and limited feedback
  236. BFGUI: An interactive tool for the synthesis and analysis of microphone array beamformers
  237. BIAS correction methods for adaptive recursive smoothing with applications in noise PSD estimation
  238. Bagging regularized common spatial pattern with hybrid motor imagery and myoelectric signal
  239. Bandlimited field reconstruction from samples obtained on a discrete grid with unknown random locations
  240. Basis compensation in non-negative matrix factorization model for speech enhancement
  241. Batch normalized recurrent neural networks
  242. Bayesian quickest detection with unknown post-change parameter
  243. Bayesian tuning for support detection and sparse signal estimation via iterative shrinkage-thresholding
  244. Benchmarking of scoring functions for bias-based fingerprinting code
  245. Benchmarking state-of-the-art visual saliency models for image quality assessment
  246. Ber analysis of the box relaxation for BPSK signal recovery
  247. Bernstein filter: A new solver for mean curvature regularized models
  248. Better acoustic normalization in subject independent acoustic-to-articulatory inversion: Benefit to recognition
  249. Beyond L2-loss functions for learning sparse models
  250. Beyond low rank + sparse: Multi-scale low rank matrix decomposition
  251. Beyond union of subspaces: Subspace pursuit on Grassmann manifold for data representation
  252. Bi-directional recurrent neural network with ranking loss for spoken language understanding
  253. Binary code learning with semantic ranking based supervision
  254. Binaural sound generation corresponding to omnidirectional video view using angular region-wise source enhancement
  255. Binaural speaker localization and separation based on a joint ITD/ILD model and head movement tracking
  256. Bipartite subgraph decomposition for critically sampled wavelet filterbanks on arbitrary graphs
  257. Bird species recognition using HMM-based unsupervised modelling of individual syllables with incorporated duration modelling
  258. Bit-depth expansion for noisy contour reduction in natural images
  259. Blind CFO estimation for multiuser OFDM uplink with large number of receive antennas
  260. Blind channel estimation in OFDM-based amplify-and-forward two-way relay networks
  261. Blind deconvolution of sparse but filtered pulses with linear state space models
  262. Blind estimation of unknown time delay in periodic non-uniform sampling: Application to desynchronized time interleaved-ADCs
  263. Blind identification of graph filters with multiple sparse inputs
  264. Blind image quality assessment for multiply distorted images via convolutional neural networks
  265. Blind mobile sensor calibration using an informed nonnegative matrix factorization with a relaxed rendezvous model
  266. Blind polychromatic X-ray CT reconstruction from poisson measurements
  267. Blind separation of underdetermined linear mixtures based on source nonstationarity and AR(1) modeling
  268. Blind speech separation based on complex spherical k-mode clustering
  269. Blind sub-Nyquist GNSS signal detection
  270. Block compressed sensing based distributed resource allocation for M2M communications
  271. Boosted classification of breast cancer by retrieval of cases having similar disease likelihood
  272. Boosted multi-scale dictionaries for image compression
  273. Boosting objectness: Semi-supervised learning for object detection and segmentation in multi-view images
  274. Bottleneck capacity of random graphs for connectomics
  275. Bottleneck linear transformation network adaptation for speaker adaptive training-based hybrid DNN-HMM speech recognizer
  276. Buffer aided distributed space time coding techniques for cooperative DS-CDMA systems
  277. CNMF-based acoustic features for noise-robust ASR
  278. CS based processing for high resolution GSM passive bistatic radar
  279. CS-based device-free localization in the presence of model errors
  280. CUED-RNNLM - An open-source toolkit for efficient training and evaluation of recurrent neural network language models
  281. Calibration of the attenuation-rain rate power-law parameters using measurements from commercial microwave networks
  282. Camera based estimation of respiration rate by analyzing shape and size variation of structured light
  283. Capacity analysis of WCC-FBMC/OQAM systems
  284. Capacity maximization for distributed broadband beamforming
  285. Carrier frequency and bandwidth estimation of cyclostationary multiband signals
  286. Channel gain prediction for multi-agent networks in the presence of location uncertainty
  287. Channel learning in indoor localization
  288. Character proposal network for robust text extraction
  289. Character-level incremental speech recognition with recurrent neural networks
  290. Choosing the diagonal loading factor for linear signal estimation using cross validation
  291. Chroma scaling for high dynamic range video compression
  292. Chute based automated fish length measurement and water drop detection
  293. Classification of bisyllabic lexical stress patterns in disordered speech using deep learning
  294. Classification of breath and snore sounds using audio data recorded with smartphones in the home environment
  295. Classification of head movement patterns to aid patients undergoing home-based cervical spine rehabilitation
  296. Classification of human cough signals using spectro-temporal Gabor filterbank features
  297. Classification of hyperspectral data with ensemble of subspace ICA and edge-preserving filtering
  298. Classification of medical images using edge-based features and sparse representation
  299. Classification of respiratory effort and disordered breathing during sleep from audio and pulse oximetry signals
  300. Classification of voices that elicit soothing effect by applying a voiced vs. unvoiced feature engineering strategy
  301. Cluster-based dictionary learning and locality-constrained sparse reconstruction for trajectory classification
  302. Clustering of interictal spikes by dynamic time warping and affinity propagation
  303. Co-segmentation of multiple images through random walk on graphs
  304. Codebook enhancement of vlad representation for visual recognition
  305. Coded excitation ultrasound: Efficient implementation via frequency domain processing
  306. Coherence regularized dictionary learning
  307. Column-wise symmetric block partitioned tensor decomposition
  308. Combining dirty-paper coding and artificial noise for secrecy
  309. Combining i-vector representation and structured neural networks for rapid adaptation
  310. Combining multiple kernel models for automatic intelligibility detection of pathological speech
  311. Combining non-negative matrix factorization and deep neural networks for speech enhancement and automatic speech recognition
  312. Combining soft decisions of several unreliable experts
  313. Comix: Joint estimation and lightspeed comparison of mixture models
  314. Common fate model for unison source separation
  315. Communication-efficient weighted ADMM for decentralized network optimization
  316. Community detection game
  317. Commuting operator of offset linear canonical transform and its applications
  318. Compact convolutional neural network transfer learning for small-scale image classification
  319. Compact kernel models for acoustic modeling via random feature selection
  320. Comparison of different development kits and its suitability in signal processing education
  321. Comparison of statistical algorithms for power system line outage detection
  322. Comparison of unsupervised sequence adaptations for deep neural networks
  323. Compensation of attacks on consensus networks
  324. Completion of structurally-incomplete matrices with reweighted low-rank and sparsity priors
  325. Complex NMF under phase constraints based on signal modeling: Application to audio source separation
  326. Complex ratio masking for joint enhancement of magnitude and phase
  327. Complexity reduction of SUMIS MIMO soft detection based on box optimization for large systems
  328. Compressed training adaptive equalization
  329. Compression and reconstruction methodology for neural signals based on patch ordering inpainting for brain monitoring
  330. Compression of dynamic 3D point clouds using subdivisional meshes and graph wavelet transforms
  331. Compressive sensing based target counting and localization exploiting joint sparsity
  332. Computational agile beam ladar imaging
  333. Computationally efficient estimation of multi-dimensional spectral lines
  334. Computed tomography reconstruction based on a hierarchical model and variational Bayesian method
  335. Conditional MMSE-based single-channel speech enhancement using inter-frame and inter-band correlations
  336. Confidence assessment for spectral estimation based on estimated covariances
  337. Connectivity for overlaid wireless networks with outage constraints
  338. Consensus inference on mobile phone sensors for activity recognition
  339. Constructive interference exploitation for downlink beamforming based on noise robustness and outage probability
  340. Content-aware local variability vector for speaker verification with short utterance
  341. Context adaptive deep neural networks for fast acoustic model adaptation in noisy conditions
  342. Context adaptive thresholding and entropy coding for very low complexity JPEG transcoding
  343. Context-dependent point process models for keyword search and detection-based ASR
  344. Continuous ultrasound based tongue movement video synthesis from speech
  345. Contour-based 3D tongue motion visualization using ultrasound image sequences
  346. Convergence analysis for Guassian belief propagation: Dynamic behaviour of marginal covariances
  347. Convergence-optimized variable node structure for stochastic LDPC decoder
  348. Convolutional neural network for robust pitch determination
  349. Convolutional neural network pre-trained with projection matrices on linear discriminant analysis
  350. Cooperative joint synchronization and localization using time delay measurements
  351. Cooperative localization based on severely quantized RSS measurements in wireless sensor network
  352. Coordinated uplink scheduling and beamforming for wireless cellular networks via sum-of-ratio programming and matching
  353. Coprime array adaptive beamforming based on compressive sensing virtual array signal
  354. Correlation-statistics-based simulator of perturbed phases triggered by the ionospheric irregularities for HF radar systems
  355. Coupled dictionary learning for multimodal data: An application to concurrent intracranial and scalp EEG
  356. Coupled rank-(Lm, Ln, •) block term decomposition by coupled block simultaneous generalized Schur decomposition
  357. Cross lingual speech emotion recognition using canonical correlation analysis on principal component subspace
  358. Cross-acoustic transfer learning for sound event classification
  359. Cross-corpus acoustic emotion recognition from singing and speaking: A multi-task learning approach
  360. Cross-lingual deep neural network based submodular unbiased data selection for low-resource keyword search
  361. Cute: A concatenative method for voice conversion using exemplar-based unit selection
  362. Cyclostationary-based detection of steady-state visually evoked potential signals recorded from EEG
  363. D-FW: Communication efficient distributed algorithms for high-dimensional sparse optimization
  364. D3M: Distributed multi-cell multigroup multicasting
  365. DCT based region log-tiedrank covariance matrices for face recognition
  366. DEMV-matchmaker: Emotional temporal course representation and deep similarity matching for automatic music video generation
  367. DNN speaker adaptation using parameterised sigmoid and ReLU hidden activation functions
  368. DNN-based enhancement of noisy and reverberant speech
  369. DOA estimation of audio sources in reverberant environments
  370. DOA estimation of closely-spaced and spectrally-overlapped sources based on time-frequency sparse representation
  371. DTM: Deformable template matching
  372. Data selection for noise robust exemplar matching
  373. Data selection from multiple ASR systems' hypotheses for unsupervised acoustic model training
  374. Data sketching for large-scale Kalman filtering
  375. Data-guided random walks for fine-structured object segmentation
  376. Data-weighted ensemble learning for privacy-preserving distributed learning
  377. Dealing with uncertain models in wireless communications
  378. Decentralized coordination of energy resources in electricity distribution networks
  379. Decoding visemes: Improving machine lip-reading
  380. Decreasing the measurement time of blood sugar tests using particle filtering
  381. Deep beamforming networks for multi-channel speech recognition
  382. Deep belief network-based post-filtering for statistical parametric speech synthesis
  383. Deep clustering: Discriminative embeddings for segmentation and separation
  384. Deep complementary bottleneck features for visual speech recognition
  385. Deep convolutional acoustic word embeddings using word-pair side information
  386. Deep discriminative manifold learning
  387. Deep kernel map networks for image annotation
  388. Deep multi-view representation learning for multi-modal features of the schizophrenia and schizo-affective disorder
  389. Deep neural network based posteriors for text-dependent speaker verification
  390. Deep neural network-guided unit selection synthesis
  391. Deep neural networks for automatic detection of screams and shouted speech in subway trains
  392. Deep unfolding for multichannel source separation
  393. Deep unfolding inference for supervised topic model
  394. Degradedness and stochastic orders of fast fading Gaussian broadcast channels with statistical channel state information at the transmitter
  395. Delay estimation between EEG and EMG via coherence with time lag
  396. Delay-Doppler estimation via structured low-rank matrix recovery
  397. Depth estimation from single images using modified stacked generalization
  398. Depth fused from intensity range and blur estimation for light-field cameras
  399. Depth guided image completion for structure and texture synthesis
  400. Depth map coding based on virtual view quality
  401. Depth map estimation using census transform for light field cameras
  402. Depth propagation in 2D-to-3D conversion based on frame clustering
  403. Depth propagation with tensor voting for 2D-to-3D video conversion
  404. Depth-aware saliency detection using discriminative saliency fusion
  405. Design space exploration for hardware-efficient stochastic computing: A case study on discrete cosine transformation
  406. Detectability prediction of hidden Markov models with cluttered observation sequences
  407. Detecting double MPEG compression with the same quantiser scale based on MBM feature
  408. Detecting occlusion from color information to improve visual tracking
  409. Detecting the instant of emotion change from speech using a martingale framework
  410. Detection of cyclostationarity in the presence of temporal or spatial structure with applications to cognitive radio
  411. Detection of drops measured by the time shift technique for spray characterization
  412. Detection of faint extended sources in hyperspectral data and application to HDF-S MUSE observations
  413. Detection of overlapping acoustic events using a temporally-constrained probabilistic model
  414. Detection of pilot contamination attack in T.D.D./S.D.M.A. systems
  415. Detection of the number of superimposed signals using modified MDL criterion: A random matrix approach
  416. Detection with phaseless measurements
  417. Deterministic maximum likelihood method for direction-of-arrival estimation of strictly noncircular signals
  418. Dictionary learning for Poisson compressed sensing
  419. Dictionary learning from phaseless measurements
  420. Diffusion LMS over multitask networks with noisy links
  421. Diffusion filtering of graph signals and its use in recommendation systems
  422. Diffusion social learning over weakly-connected graphs
  423. Diffusion stochastic optimization with non-smooth regularizers
  424. Diffusive particle filtering for distributed multisensor estimation
  425. Direction of arrival estimation based on information geometry
  426. Direction of arrival estimation in MIMO radar systems with nonlinear reflectors
  427. Direction-of-arrival estimation based on Toeplitz covariance matrix reconstruction
  428. Direction-of-arrival estimation with espar antennas using Bayesian compressive sensing
  429. Directional maximum likelihood self-estimation of the path-loss exponent
  430. Directly modeling voiced and unvoiced components in speech waveforms by neural networks
  431. Dirichlet process mixture models for time-dependent clustering
  432. Discontinuous operation for precoded G.fast
  433. Discourse connective detection in spoken conversations
  434. Discovering rāga motifs by characterizing communities in networks of melodic patterns
  435. Discriminant correlation analysis for feature level fusion with application to multimodal biometrics
  436. Discriminative deep recurrent neural networks for monaural speech separation
  437. Discriminative feature extraction from X-ray images using deep convolutional neural networks
  438. Discriminative multi-domain PLDA for speaker verification
  439. Discriminatively learned filter bank for acoustic features
  440. Discriminatively trained joint speaker and environment representations for adaptation of deep neural network acoustic models
  441. Distances between directed networks and applications
  442. Distributed LMS estimation of scaled and delayed impulse responses
  443. Distributed MIMO systems: Receiver design and ML detection
  444. Distributed beamforming in relay networks for energy harvesting multi-group multicast systems
  445. Distributed beamforming using mobile robots
  446. Distributed dyadic cyclic descent for non-negative matrix factorization
  447. Distributed estimation of latent parameters in state space models using separable likelihoods
  448. Distributed estimation via paid crowd work
  449. Distributed generalized likelihood ratio tests: Fundamental limits and tradeoffs
  450. Distributed linear blind source separation over wireless sensor networks with arbitrary connectivity patterns
  451. Distributed multi-sensor CPHD filter using pairwise gossiping
  452. Distributed nonconvex optimization over time-varying networks
  453. Distributed path optimization of multiple UAVs for AOA target localization
  454. Distributed sparse MVDR beamforming using the bi-alternating direction method of multipliers
  455. Distributional semantics for understanding spoken meal descriptions
  456. Distributionally robust chance-constrained minimum variance beamforming
  457. Divergence estimation based on deep neural networks and its use for language identification
  458. Document level semantic context for retrieving OOV proper names
  459. Domain adaptation for speech emotion recognition by sharing priors between related source and target classes
  460. Domain adaptation using maximum likelihood linear transformation for PLDA-based speaker verification
  461. Downlink SINR balancing in C-RAN under limited fronthaul capacity
  462. Dropped pronoun generation for dialogue machine translation
  463. Dual-microphone voice activity detection based on using optimally weighted maximum a posteriori probabilities
  464. Dual-stage algorithm to identify channels with poor electrode-to-neuron interface in cochlear implant users
  465. Dynamic analysis of resting state fMRI data and its applications
  466. Dynamic relative impulse response estimation using structured sparse Bayesian learning
  467. EchoSLAM: Simultaneous localization and mapping with acoustic echoes
  468. Effective utilization of multiple examples in query-by-example spoken term detection
  469. Effectiveness of fundamental frequency (F0) and strength of excitation (SOE) for spoofed speech detection
  470. Efficient algorithms for linear polyhedral bandits
  471. Efficient channel statistics estimation for millimeter-wave MIMO systems
  472. Efficient deblocking filter implementation on reconfigurable processor
  473. Efficient estimation of inter-subband speech correlations
  474. Efficient keypoint detection and description via polynomial regression of scale space
  475. Efficient near optimal joint modulation classification and detection for MU-MIMO systems
  476. Efficient neighborhood-based topic modeling for collaborative audio enhancement on massive crowdsourced recordings
  477. Efficient non-linear feature adaptation using Maxout networks
  478. Efficient object feature selection for action recognition
  479. Efficient one-vs-one kernel ridge regression for speech recognition
  480. Efficient parameter inference in general hidden Markov models using the filter derivatives
  481. Efficient sensor position selection using graph signal sampling theory
  482. Efficient stochastic detector for large-scale MIMO
  483. Efficient subspace detection for high-order MIMO systems
  484. Efficient target-response interpolation for a graphic equalizer
  485. Egocentric activity recognition with multimodal fisher vector
  486. Eigen and multimodal analysis for localizing moving sounding objects
  487. Emotion classification: How does an automated system compare to Naive human coders?
  488. Emotion recognition from peripheral physiological signals enhanced by EEG
  489. Emotion-flow guided music accompaniment generation
  490. Empirically-estimable multi-class classification bounds
  491. End-to-end attention-based large vocabulary speech recognition
  492. End-to-end text-dependent speaker verification
  493. Energy detection in ISI channels using large-scale receiver arrays
  494. Energy efficient beamforming for secure communication in cognitive radio networks
  495. Energy-efficient pilot and data power allocation in massive MIMO communication systems based on MMSE channel estimation
  496. Enhanced HEVC intra prediction with ordered dither technique
  497. Enhanced just noticeable difference model with visual regularity consideration
  498. Enhanced semi-supervised learning for multimodal emotion recognition
  499. Enhanced vote count circuit based on nor flash memory for fast similarity search
  500. Epileptiform spike detection via convolutional neural networks
  501. Equalization matching of speech recordings in real-world environments
  502. Error performance analysis of the symbol-decision SC polar decoder
  503. Error-resilient sequential cells with successive time borrowing for stochastic computing
  504. Estimating direct-to-reverberant ratio mapped from power spectral density using deep neural network
  505. Estimating ear canal geometry and eardrum reflection coefficient from ear canal input impedance
  506. Estimating high-dimensional covariance matrices with misses for Kronecker product expansion models
  507. Estimating orientation in tracking individuals of flying swarms
  508. Estimating parameters in noisy low frequency exponentially damped sinusoids and exponentials
  509. Estimation efficiency, accuracy and robustness improvement by exploiting the geometry information in SAR-GMTI system
  510. Estimation of TDOA for room reflections by iterative weighted l1 constraint
  511. Estimation of the reliability of multiple rhythm features extraction from a single descriptor
  512. Eulerian emotion magnification for subtle expression recognition
  513. Evaluating instrumental measures of speech quality using Bayesian model selection: Correlations can be misleading!
  514. Evaluation of estimated hammerstein models via normalized projection misalignment of linear and nonlinear subsystems
  515. Exemplar-based sparse representation of timbre and prosody for voice conversion
  516. Exemplar-inspired strategies for low-resource spoken keyword search in Swahili
  517. Experimental study of generalized subspace filters for the cocktail party situation
  518. Experimental validation of TOA-based methods for microphones array positions calibration
  519. Exploiting LSTM structure in deep neural networks for speech recognition
  520. Exploiting correlations among channels in distributed compressive sensing with convolutional deep stacking networks
  521. Exploiting low-dimensional structures to enhance DNN based acoustic modeling in speech recognition
  522. Exploiting sparsity for image-based object surface anomaly detection
  523. Exploiting spectro-temporal structures using NMF for DNN-based supervised speech separation
  524. Exploratory analysis of speech features related to depression in adults with Aphasia
  525. Exploring articulatory characteristics of Cantonese dysarthric speech using distinctive features
  526. Exploring deep learning architectures for automatically grading non-native spontaneous speech
  527. Exploring multidimensional lstms for large vocabulary ASR
  528. Exploring persistent local homology in topological data analysis
  529. Exploring the role of phonetic bottleneck features for speaker and language recognition
  530. Extension of SeDJoCo and its use in a combination of multicast and coordinated multi-point systems
  531. Extension of nested arrays with the fourth-order difference co-array enhancement
  532. Extension of the semi-algebraic framework for approximate CP decompositions via non-symmetric simultaneous matrix diagonalization
  533. Extensions of semidefinite programming methods for atomic decomposition
  534. Extensions of the binaural MWF with interference reduction preserving the binaural cues of the interfering source
  535. Extraction of tongue contour in real-time magnetic resonance imaging sequences
  536. F0 estimation for noisy speech by exploring temporal harmonic structures in local time frequency spectrum segment
  537. FPGA based implementation of deep neural networks using on-chip memory only
  538. Face alignment by deep convolutional network with adaptive learning rate
  539. Face hallucination via locality-constrained low-rank representation
  540. Face liveness detection and recognition using shearlet based feature descriptors
  541. Face recognition with local contourlet combined patterns
  542. Factored spatial and spectral multichannel raw waveform CLDNNs
  543. Fall detection in RGB-D videos by combining shape and motion features
  544. Fast adaptive PARAFAC decomposition algorithm with linear complexity
  545. Fast alternating projected gradient descent algorithms for recovering spectrally sparse signals
  546. Fast and easy crowdsourced perceptual audio evaluation
  547. Fast and efficient rejection of background waveforms in interictal EEG
  548. Fast and statistically efficient fundamental frequency estimation
  549. Fast anomaly detection in traffic surveillance video based on robust sparse optical flow
  550. Fast continuous HRTF acquisition with unconstrained movements of human subjects
  551. Fast depth image denoising and enhancement using a deep convolutional network
  552. Fast dynamic MRI using linear dynamical system model
  553. Fast intra mode decision and block matching for HEVC screen content compression
  554. Fast keypoint detection in video sequences
  555. Fast lossless compression of whole slide pathology images using HEVC intra-prediction
  556. Fast online orthonormal dictionary learning for efficient full waveform inversion
  557. Fast response aggregation for depth estimation using light field camera
  558. Fast sparse 2-D DFT computation using sparse-graph alias codes
  559. Fast variational Bayesian signal recovery in the presence of Poisson-Gaussian noise
  560. Fast voxel line update for time-space image reconstruction
  561. Feature adapted convolutional neural networks for downbeat tracking
  562. Feature mapping, score-, and feature-level fusion for improved normal and whispered speech speaker verification
  563. Feature-enriched word embeddings for named entity recognition in open-domain conversations
  564. Feedback of differential precoder for geometrical mean decomposition systems
  565. Fetal heart rate analysis by hierarchical dirichlet process mixture models
  566. Filterbank learning using Convolutional Restricted Boltzmann Machine for speech recognition
  567. Finding the minimum rate of innovation in the presence of noise
  568. Finding unique dense communities
  569. Fine-structured object segmentation via edge-guided graph cut with interaction simplification
  570. Fingerprint recognition with ridge features and minutiae on distortion
  571. Finite-state channel models for signal transduction in neural systems
  572. First order echo based room shape recovery using a single mobile device
  573. Fixed-complexity variants of the effective LLL algorithm with greedy convergence for MIMO detection
  574. Fixed-point performance analysis of recurrent neural networks
  575. Flat start training of CD-CTC-SMBR LSTM RNN acoustic models
  576. Flexibeam: Analytic spatial filtering by beamforming
  577. Formant shifting for speech intelligibility improvement in car noise environment
  578. Framewise speech-nonspeech classification by neural networks for voice activity detection with statistical noise suppression
  579. Frank-Wolfe works for non-Lipschitz continuous gradient objectives: Scalable poisson phase retrieval
  580. Frequency recognition of steady-state visually evoked potentials using binary subband canonical correlation analysis with reduced dimension of reference signals
  581. Frequency-based customization of multizone sound system design
  582. From HMMS to DNNS: Where do the improvements come from?
  583. From acoustic room reconstruction to slam
  584. Functional connectivity brain network analysis through network to signal transform based on the resistance distance
  585. Further results on mainlobe orientation reversal of the first-order steerable differential array due to microphone phase errors
  586. Fusion of algorithms for multiple measurement vectors
  587. Fusion of depth, skeleton, and inertial data for human action recognition
  588. Fuzzy entropy based nonnegative matrix factorization for muscle synergy extraction
  589. Gain relaxation: A useful technique for signal enhancement with an unaware local noise source targeted at speech recognition
  590. Gammatone filter based on stochastic computation
  591. Gating recurrent mixture density networks for acoustic modeling in statistical parametric speech synthesis
  592. Gauss-Seidel based non-negative matrix factorization for gene expression clustering
  593. Generalized Laplacian precision matrix estimation for graph signal processing
  594. Generalized coprime sampling of Toeplitz matrices
  595. Generalized k-level cutset sampling and reconstruction
  596. Generalized wave-domain transforms for listening room equalization with azimuthally irregularly spaced loudspeaker arrays
  597. Generating a morphable model of ears
  598. Geo-location dependent deep neural network acoustic model for speech recognition
  599. Geodesic-based pavement shadow removal revisited
  600. Geometric-guided label propagation for moving object detection
  601. Geometrical room geometry estimation from room impulse responses
  602. Geometry and radiometry invariant matched manifold detection and tracking
  603. Ghosting-free multi-exposure image fusion in gradient domain
  604. Globally optimized least-squares post-filtering for microphone array speech enhancement
  605. Gold classification of COPDGene cohort based on deep learning
  606. Grab-n-Pull: An optimization framework for fairness-achieving networks
  607. Gradient schemes for robust FFT-based motion estimation
  608. Gram Schmidt based greedy hybrid precoding for frequency selective millimeter wave MIMO systems
  609. Graph filter banks with M-channels, maximal decimation, and perfect reconstruction
  610. Graph signal recovery from incomplete and noisy information using approximate message passing
  611. Graph-based lifting transform for intra-predicted video coding
  612. Graph-based representation and coding of 3D images for interactive multiview navigation
  613. Group diffusion LMS
  614. Group nonnegative matrix factorisation with speaker and session variability compensation for speaker identification
  615. Group sparse Bayesian learning via exact and fast marginal likelihood maximization
  616. Group-blind detection with very large antenna arrays in the presence of pilot contamination
  617. Groupwise learning for ASR k-best list reranking in spoken language translation
  618. Hardware implementation of FIR/IIR digital filters using integral stochastic computation
  619. Harmonic-percussive-residual sound separation using the structure tensor on spectrograms
  620. Heart-trend: An affordable heart condition monitoring system exploiting morphological pattern
  621. Heterogeneous domain adaptation with label and structure consistency
  622. High accuracy indoor localization: A WiFi-based approach
  623. High diagnostic quality ECG compression and CS signal reconstruction in body sensor networks
  624. High dynamic range imaging via truncated nuclear norm minimization of low-rank matrix
  625. High-pitched excitation generation for glottal vocoding in statistical parametric speech synthesis using a deep neural network
  626. High-resolution sinusoidal modeling of unvoiced speech
  627. Higher-order listening room compensation with additive compensation signals
  628. Highway long short-term memory RNNS for distant speech recognition
  629. Histogram feature deblurring
  630. Honey chatting: A novel instant messaging system robust to eavesdropping over communication
  631. How neural network features and depth modify statistical properties of HMM acoustic models
  632. Hybrid beamforming with two bit RF phase shifters in single group multicasting
  633. Hybrid music recommender using content-based and social information
  634. Hyperspectral image classification using set-to-set distance
  635. IMISOUND: An unsupervised system for sound query by vocal imitation
  636. IVA for abandoned object detection: Exploiting dependence across color channels
  637. Identity association using PHD filters in multiple head tracking with depth sensors
  638. Image colorization based on ADMM with fast singular value thresholding by Chebyshev polynomial approximation
  639. Image phylogeny tree reconstruction based on region selection
  640. Image restoration using a stochastic variant of the alternating direction method of multipliers
  641. Image sentiment analysis using latent correlations among visual, textual, and sentiment views
  642. Image-assisted geometry simplification for the plenoptic sampling
  643. Imaging in radio interferometry by iterative subset scanning using a modified AMP algorithm
  644. Impact of channel access issues and packet losses on distributed outlier detection within wireless sensor networks
  645. Implementation of F0 transformation for statistical singing voice conversion based on direct waveform modification
  646. Implementation of the precoder matrix indicator selection using MMSE trace criterion for the downlink transmission in LTE
  647. Implicit kernel presentation aware object segmentation framework
  648. Importance sampling of delta-AUC: A basis for active learning for improved keyword search
  649. Improved DNN-based segmentation for multi-genre broadcast audio
  650. Improved decoding of analog modulo block codes for noise mitigation
  651. Improved forgery detection with lateral chromatic aberration
  652. Improved illumination invariant homomorphic filtering using the dual tree complex wavelet transform
  653. Improved multi-microphone noise reduction preserving binaural cues
  654. Improved set-membership partial-update affine projection algorithm
  655. Improved speaker independent lip reading using speaker adaptive training and deep neural networks
  656. Improved spoken document summarization with coverage modeling techniques
  657. Improving SHVC performance with a joint layer coding mode
  658. Improving adaptive feedback cancellation in hearing aids using an affine combination of filters
  659. Improving face detection with depth
  660. Improving non-native mispronunciation detection and enriching diagnostic feedback with DNN-based speech attribute modeling
  661. Improving resolution in supervised patch-based target detection
  662. Improving semantic video indexing: Efforts in Waseda TRECVID 2015 SIN system
  663. Improving speech privacy in personal sound zones
  664. Incorporating relative transfer function preservation into the binaural multi-channel wiener filter for hearing aids
  665. Independent versus repeated measurements: A performance quantification via state evolution
  666. Inferring depolarization of cells from 3D-electrode measurements using a bank of linear state space models
  667. Information fusion based on kernel entropy component analysis in discriminative canonical correlation space with application to audio emotion recognition
  668. Information point set registration for shape recognition
  669. Information theoretic clustering for unsupervised domain-adaptation
  670. Information theoretic multivariate change detection for multisensory information processing in Internet of Things
  671. Informed Direction of Arrival estimation using a spherical-head model for Hearing Aid applications
  672. Infrared small target detection with compressive measurements
  673. Initial investigation of speech synthesis based on complex-valued neural networks
  674. Insight into a phase modulation technique for signal decorrelation in multi-channel acoustic echo cancellation
  675. Instantaneous pitch estimation algorithm based on multirate sampling
  676. Integrated adaptation with multi-factor joint-learning for far-field speech recognition
  677. Integrated approach of feature extraction and sound source enhancement based on maximization of mutual information
  678. Integration of machine learning and human learning for training optimization in robust linear regression
  679. Integration of orthogonal feature detectors in parameter learning of artificial neural networks to improve robustness and the evaluation on hand-written digit recognition tasks
  680. Intelligible enhancement of 3D articulation animation by incorporating airflow information
  681. Intensity-only optical compressive imaging using a multiply scattering material and a double phase retrieval approach
  682. Inter-speaker variability in forensic voice comparison: A preliminary evaluation
  683. Interlaced sigma-point information filtering for distributed state estimation of multi-agent systems
  684. Interpreting the prediction process of a deep network constructed from supervised topic models
  685. Intrinsic two-dimensional local structures for micro-expression recognition
  686. Introduction to the special session on Topological Data Analysis, ICASSP 2016
  687. Intrusive howling detection methods for hearing aid evaluations
  688. Investigating gated recurrent networks for speech synthesis
  689. Investigating techniques for low resource conversational speech recognition
  690. Investigation of speaker embeddings for cross-show speaker diarization
  691. Investigation on log-linear interpolation of multi-domain neural network language model
  692. Investigations into vowel and consonant structures in articulatory and auditory spaces using Laplacian eigenmaps
  693. Investigations on speaker adaptation of LSTM RNN models for speech recognition
  694. Iterative estimation of phase using complex cepstrum representation
  695. Iterative linear regression classification for image recognition
  696. Iterative quadratic relaxation method for optimization of multiple radar waveforms
  697. Iteratively reweighted tensor SVD for robust multi-dimensional harmonic retrieval
  698. Joint ML calibration and DOA estimation with separated arrays
  699. Joint acoustic factor learning for robust deep neural network based automatic speech recognition
  700. Joint action recognition and summarization by sub-modular inference
  701. Joint device-to-device transmission activation and transceiver design for sum-rate maximization in MIMO interfering channels
  702. Joint dictionary training for bandwidth extension of speech signals
  703. Joint estimation of sound source location and boundary impedance with physics-driven cosparse regularization
  704. Joint information from nonlinear and linear features for spoofing detection: An i-vector/DNN based approach
  705. Joint instance and feature importance re-weighting for person reidentification
  706. Joint maximum likelihood estimation of late reverberant and speech power spectral density in noisy environments
  707. Joint offloading decision and resource allocation for mobile cloud with computing access point
  708. Joint sub-band based neighbor embedding for image super-resolution
  709. Joint transceiver designs for secure communications over MIMO relay
  710. Joint user association and content placement for Cache-enabled wireless access networks
  711. Joint-view Kalman-filter recovery of compressed-sensed multiview videos
  712. Jointly optimal near-end and far-end multi-microphone speech intelligibility enhancement based on mutual information
  713. Kalman filter for speech enhancement in cocktail party scenarios using a codebook-based approach
  714. Kalman filters with Bayesian quadratic game fusion in networks
  715. Keyword search using query expansion for graph-based rescoring of hypothesized detections
  716. Knowledge-aided hyperparameter-free Bayesian detection in stochastic homogeneous environments
  717. L1-L1 norms for face super-resolution with mixed Gaussian-impulse noise
  718. LCMV beamforming with subspace projection for multi-speaker speech enhancement
  719. LDADEEP+: Latent aspect discovery with deep representations
  720. Landmark of Mandarin nasal codas and its application in pronunciation error detection
  721. Language model adaptation for ASR of spoken translations using phrase-based translation models and named entity models
  722. Language recognition using deep neural networks with very limited training data
  723. Language-independent acoustic cloning of HTS voices: A preliminary study
  724. Laplacian deep kernel learning for image annotation
  725. Large region acoustic source mapping: A generalized sparse constrained deconvolution approach
  726. Large-scale l0 sparse inverse covariance estimation
  727. Latent feature representation with 3-D multi-view deep convolutional neural network for bilateral analysis in digital breast tomosynthesis
  728. Learning compact recurrent neural networks
  729. Learning compact structural representations for audio events using regressor banks
  730. Learning cross-lingual information with multilingual BLSTM for speech synthesis of low-resource languages
  731. Learning data triage: Linear decoding works for compressive MRI
  732. Learning deep neural network using max-margin minimum classification error
  733. Learning discriminative and shareable patches for scene classification
  734. Learning full-range affinity for diffusion-based saliency detection
  735. Learning in constrained stochastic dynamic potential games
  736. Learning network structures from firing patterns
  737. Learning separable fixed-point kernels for deep convolutional neural networks
  738. Learning structured dictionary based on inter-class similarity and representative margins
  739. Learning to separate vocals from polyphonic mixtures via ensemble methods and structured output prediction
  740. Learning-based fully 3D face reconstruction from a single image
  741. Least squares phase retrieval using feasible point pursuit
  742. Lightly-supervised utterance-level emotion identification using latent topic modeling of multimodal words
  743. Linear network operators using node-variant graph filters
  744. Linearly augmented deep neural network
  745. Lipreading with long short-term memory
  746. Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
  747. Local Q-linear convergence and finite-time active set identification of ADMM on a class of penalized regression problems
  748. Local fisher discriminant analysis for spoken language identification
  749. Local likelihood estimation of time-variant Hawkes models
  750. Localization of sound sources with known statistics in the presence of interferers
  751. Long short term memory recurrent neural network based encoding method for emotion recognition in video
  752. Long-CPI multi-channel SAR based ground moving target indication
  753. Long-term general rank multiuser downlink beamforming with shaping constraints using QOSTBC
  754. Low bit-rate intra coding scheme based on constrained quantization and median-type filter
  755. Low complexity tonality control in the Intelligent Gap Filling tool
  756. Low complexity transform competition for HEVC
  757. Low level descriptors based DBLSTM bottleneck feature for speech driven talking avatar
  758. Low rank approximation based hybrid precoding schemes for multi-carrier single-user massive MIMO systems
  759. Low-complexity beamforming designs of sum secrecy rate maximization for the Gaussian MISO multi-receiver wiretap channel
  760. Low-complexity recursive convolutional precoding for OFDM-based large-scale antenna systems
  761. Low-complexity robust multi-cell MISO downlink precoder design
  762. Low-rank matrices recovery via entropy function
  763. Low-rank plus diagonal adaptation for deep neural networks
  764. Low-resolution reconstruction of intensity functions on the sphere for single-particle diffraction imaging
  765. Lower bounds on the L2-norms of digital resampling filters with zero-valued input samples
  766. Lucky ranging with towed arrays in underwater environments subject to non-stationary spatial coherence loss
  767. MIMO radar waveform design for multiple extended target estimation based on greedy SINR maximization
  768. MMSE denoising of sparse and non-Gaussian AR(1) processes
  769. MMSE precoder for massive MIMO using 1-bit quantization
  770. Machine translation based data augmentation for Cantonese keyword spotting
  771. Magnetic beamforming for wireless power transfer
  772. Maintaining throughput network connectivity in ad hoc networks
  773. Manga-specific features and latent style model for manga style analysis
  774. Manifold denoising based on spectral graph wavelets
  775. Manifold-based Bayesian inference for semi-supervised source localization
  776. Masked correlation filters for partially occluded face recognition
  777. Maximally improper interference in underlay cognitive radio networks
  778. Maximum likelihood PSD estimation for speech enhancement in reverberant and noisy conditions
  779. Maximum likelihood and maximum a posteriori direction-of-arrival estimation in the presence of sirp noise
  780. Maximum likelihood rumor source detection in a star network
  781. Measure-transformed quasi likelihood ratio test
  782. Measurement partitioning and observational equivalence in state estimation
  783. Mediated experts for deep convolutional networks
  784. Medical image super-resolution with non-local embedding sparse representation and improved IBP
  785. Membrane shape and boundary conditions estimation using eigenmode decomposition
  786. Memory reduction techniques for successive cancellation decoding of polar codes
  787. Memory-restricted multiscale dynamic time warping
  788. Micro-Doppler extraction from ISAR image
  789. Microtexture inpainting through Gaussian conditional simulation
  790. Millimeter wave communications channel estimation via Bayesian group sparse recovery
  791. Minimum distance criterion for non-negative hyperspectral image deconvolution
  792. Minimum word error training of long short-term memory recurrent neural network language models for speech recognition
  793. Mining representative actions for actor identification
  794. Mitigation of sparsely sampled nonstationary jammers for multi-antenna GNSS receivers
  795. Mobile beamforming & spatially controlled relay communications
  796. Modeling audio directional statistics using a complex bingham mixture model for blind source extraction from diffuse noise
  797. Modeling deep bidirectional relationships for image classification and generation
  798. Modeling spectral envelopes using deep conditional restricted Boltzmann machines for statistical parametric speech synthesis
  799. Modelling stress in public speaking: Evolution of stress levels during conference presentations
  800. Modulation spectrum compensation for HMM-based speech synthesis using line spectral pairs
  801. Mood state prediction from speech of varying acoustic quality for individuals with bipolar disorder
  802. Multi-centrality graph spectral decompositions and their application to cyber intrusion detection
  803. Multi-channel power allocation for device-to-device communication underlaying cellular networks
  804. Multi-focus image fusion via coupled dictionary training
  805. Multi-focus pixel-based image fusion in dual domain
  806. Multi-fold Gabor filter convolution descriptor for face recognition
  807. Multi-index voting for asymmetric distance computation in a large-scale binary codes
  808. Multi-kernel based nonlinear models for connectivity identification of brain networks
  809. Multi-pair two-way AF relaying systems with massive arrays and imperfect CSI
  810. Multi-pass feature enhancement based on generative-discriminative hybrid approach for noise robust speech recognition
  811. Multi-processor approximate message passing using lossy compression
  812. Multi-stream spectral representation for statistical parametric speech synthesis
  813. Multi-view distributed source coding of binary features for visual sensor networks
  814. Multiantenna spectrum sensing for improper signals over frequency selective channels
  815. Multichannel audio declipping
  816. Multichannel blind source separation based on non-negative tensor factorization in wavenumber domain
  817. Multichannel identification of room acoustic systems with adaptive filters based on orthonormal basis functions
  818. Multicore implementation of LDPC decoders based on ADMM algorithm
  819. Multicriteria optimization for nonunitary joint block diagonalization
  820. Multilingual data selection for training stacked bottleneck features
  821. Multilingual region-dependent transforms
  822. Multimodal Kalman filtering
  823. Multimodal human action recognition in assistive human-robot interaction
  824. Multipath radar tracking with large uncertainty in the environment
  825. Multipath removal by online blind deconvolution in through-the-wall-imaging
  826. Multiple instance discriminative dictionary learning for action recognition
  827. Multiple instance learning for model ensemble and meta data transfer
  828. Multiple scattering effects on the localization of two point scatterers
  829. Multiple-kernel adaptive segmentation and tracking (MAST) for robust object tracking
  830. Multiplicative update of AR gains in codebook-driven speech enhancement
  831. Multiview learning via deep discriminative canonical correlation analysis
  832. Music emotion recognition with adaptive aggregation of Gaussian process regressors
  833. Mutual information based radar waveform design for joint radar and cellular communication systems
  834. NMF-based informed source separation
  835. NMF-based source separation utilizing prior knowledge on encoding vector
  836. Network topology adaptation and interference coordination for energy saving in heterogeneous networks
  837. Neural network based spectral mask estimation for acoustic beamforming
  838. Neural network shape: Organ shape representation with radial basis function neural networks
  839. News story clustering with fisher embedding
  840. No-reference image quality assessment for photographic images of consumer device
  841. Noise and reverberation effects on depression detection from speech
  842. Noise estimation for speech reinforcement in the presence of strong echoes
  843. Noise robust recognition method based on scatterer pattern for radar HRRP data
  844. Noise robust speech recognition using recent developments in neural networks for computer vision
  845. Noise suppression method for body-conducted soft speech enhancement based on external noise monitoring
  846. Non-asymptotic performance bounds of eigenvalue based detection of signals in non-Gaussian noise
  847. Non-cooperative cross-channel gain estimation using full-duplex amplify-and-forward relaying in cognitive radio networks
  848. Non-linear regression for bivariate self-similarity identification - application to anomaly detection in Internet traffic based on a joint scaling analysis of packet and byte counts
  849. Non-monotone quadratic potential games with single quadratic constraints
  850. Non-negative decomposition of linear relationships: Application to multi-source ocean remote sensing data
  851. Non-negative intermediate-layer DNN adaptation for a 10-KB speaker adaptation profile
  852. Non-stationary blind super-resolution
  853. Non-stationary noise power spectral density estimation based on regional statistics
  854. Non-verbal speech analysis of interviews with schizophrenic patients
  855. Nonconvex compressive sensing reconstruction for tensor using structures in modes
  856. Nonnegative matrix factorization using ADMM: Algorithm and convergence analysis
  857. Nonnegative matrix factorization-based frequency lowering technology for Mandarin-speaking hearing aid users
  858. Nonparametric detection of an anomalous disk over a two-dimensional lattice network
  859. Novel 3D-WPP algorithms for parallel HEVC encoding
  860. Novel acoustic features for automatic dialog-act tagging
  861. Novel favorite music classification using EEG-based optimal audio features selected via KDLPCCA
  862. Novel neural network based fusion for multistream ASR
  863. Novel quaternion matrix factorisations
  864. OCR-aided person annotation and label propagation for speaker modeling in TV shows
  865. Object recognition in art drawings: Transfer of a neural network
  866. Object saliency using a background prior
  867. Oligopoly dynamic pricing: A repeated game with incomplete information
  868. On Renyi's entropy estimation with one-dimensional Gaussian kernels
  869. On adaptive selection of estimation bandwidth for analysis of locally stationary multivariate processes
  870. On combining i-vectors and discriminative adaptation methods for unsupervised speaker normalization in DNN acoustic models
  871. On convexity and identifiability in 1-D Fourier phase retrieval
  872. On gridless sparse methods for multi-snapshot DOA estimation
  873. On multiple solutions of the "sequentially drilled" joint congruence transformation (SeDJoCo) problem for semi-blind source separation
  874. On parameter estimation of symmetric alpha-stable distribution
  875. On parametric lower bounds for discrete-time filtering
  876. On pilot-symbol aided channel estimation in FBMC-OQAM
  877. On privacy preference in collusion-deterrence games for secure multi-party computation
  878. On projected stochastic gradient descent algorithm with weighted averaging for least squares regression
  879. On scalable coding of hidden Markov sources
  880. On simplifying the primal-dual method of multipliers
  881. On sparse controllability of graph signals
  882. On spatio-frequential smoothing for joint angles and times of arrival estimation of multipaths
  883. On target localization with communication costs via tensor completion: A multi-modal approach
  884. On the LP-convergence of a Girsanov theorem based particle filter
  885. On the average staleness of global channel state information in wireless networks with random transmit node selection
  886. On the compression of recurrent neural networks with an application to LVCSR acoustic modeling for embedded speech recognition
  887. On the decay - and the smoothness behavior of the Fourier transform, and the construction of signals having strong divergent Shannon sampling series
  888. On the detection of non-stationary signals in the matched signal transform domain
  889. On the impact of residual CFO in UL MU-MIMO
  890. On the importance of event detection for ASR
  891. On the importance of harmonic phase modification for improved speech signal reconstruction
  892. On the influence of momentum acceleration on online learning
  893. On the influence of quantization on the identifiability of emotions from voice coding parameters
  894. On the performance of cloud radio access networks using Matérn hard-core point processes
  895. On the periodically time-varying bias in adaptive feedback cancellation systems with frequency shifting
  896. On the separability of signal and interference-plus-noise subspaces in blind pilot decontamination
  897. On time delay estimation based on multichannel spatiotemporal sparse linear prediction
  898. On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech recognition
  899. One plus two may not equal two plus one in a social sensing network with unknown parameters
  900. One-bit ADCs in wideband massive MIMO systems with OFDM transmission
  901. Online adaptation of the number of particles of SMC methods
  902. Online change detection of linear regression models
  903. Online incremental higher-order partial least squares regression for fast reconstruction of motion trajectories from tensor streams
  904. Online learning and optimization of Markov jump linear models
  905. Online least-squares one-class support vector machine for outlier detection in power grid data
  906. Online low-rank + sparse structure learning for dynamic network tracking
  907. Online low-rank tensor subspace tracking from incomplete data by CP decomposition using recursive least squares
  908. Online nonnegative matrix factorization with outliers
  909. Online speaker diarization using adapted i-vector transforms
  910. Online speaking rate estimation using recurrent neural networks
  911. Open-set microphone classification via blind channel analysis
  912. Opening big in box office? Trailer content can help
  913. Opinion dynamics in multi-agent systems with binary decision exchanges
  914. Opportunistic spectrum access with temporal-spatial reuse in cognitive radio networks
  915. Optimal UAV localisation in vision based navigation systems
  916. Optimal copula transport for clustering multivariate time series
  917. Optimal design of constant-modulus channel training sequences
  918. Optimal linear cooperation for signal classification
  919. Optimal pilot length for uplink massive MIMO systems with pilot reuse
  920. Optimal resource block allocation and muting in heterogeneous networks
  921. Optimal space signalling for intensity modulated MIMO optical wireless communications
  922. Optimal zero forcing precoder and decoder design for multi-user MIMO FBMC under strong channel selectivity
  923. Optimizing DTW-based audio-to-MIDI alignment and matching
  924. Orthogonal sparse eigenvectors: A procrustes problem
  925. Outlier-robust recovery of low-rank positive semidefinite matrices from magnitude measurements
  926. Outlying sequence detection in large datasets: Comparison of universal hypothesis testing and clustering
  927. Overlapping clustering of network data using cut metrics
  928. PCA using graph total variation
  929. PLIP based unsharp masking for medical image enhancement
  930. PROJET - Spatial audio separation using projections
  931. Pansharpening via coupled triple factorization dictionary learning
  932. Parallel metropolis chains with cooperative adaptation
  933. Parallel proximal methods for total variation minimization
  934. Parallelizing WFST speech decoders
  935. Parameter estimation of polynomial phase signal based on low-complexity LSU-EKF algorithm in entire identifiable region
  936. Parametric Frugal sensing of autoregressive power spectra
  937. Parametric analog mappings for correlated Gaussian sources over AWGN channels
  938. Partial face recognition: A sparse representation-based approach
  939. Particle filtering for slice-to-volume motion correction in EPI based functional MRI
  940. Particle filters with independent resampling
  941. Particle flow for particle filtering
  942. Partitioned successive-cancellation list decoding of polar codes
  943. Pathological speech processing: State-of-the-art, current challenges, and future directions
  944. Pattern-based 3D model compression
  945. Perceptual and instrumental evaluation of the perceived level of reverberation
  946. Perfect error compensation via algorithmic error cancellation
  947. Performance advantage of quaternion widely linear estimation: An approximate uncorrelating transform approach
  948. Performance analysis for pilot-based 1-bit channel estimation with unknown quantization threshold
  949. Performance analysis of EWF codes with intermediate feedback
  950. Performance analysis of a modified Rao test for adaptive subspace detection
  951. Performance analysis of joint-sparse recovery from multiple measurement vectors with prior information via convex optimization
  952. Performance analysis of spectral community detection in realistic graph models
  953. Performance limits of single-agent and multi-agent sub-gradient stochastic learning
  954. Persistent homology lower bounds on network distances
  955. Persistent homology of toroidal sliding window embeddings
  956. Personalized mispronunciation detection and diagnosis based on unsupervised error pattern discovery
  957. Personalized speech recognition on mobile devices
  958. Phaseless super-resolution using masks
  959. Phoneme-specific speech separation
  960. Phrase-based rĀga recognition using vector space modeling
  961. Phylogenetic analysis of near-duplicate images using processing age metrics
  962. Physical object authentication: Detection-theoretic comparison of natural and artificial randomness
  963. Physical-model based efficient data representation for many-channel microphone array
  964. Piecewise sparse signal recovery via piecewise orthogonal matching pursuit
  965. Pilot aided direction of arrival estimation for mmWave cellular systems
  966. Pilot-based channel estimation for FBMC/OQAM systems under strong frequency selectivity
  967. Pinpoint extraction of distant sound source based on DNN mapping from multiple beamforming outputs to prior SNR
  968. Policy recognition via expectation maximization
  969. Portfolio optimization with asset selection and risk parity control
  970. Posterior probabilistic modeling for inter-channel phase and time difference estimation in audio signals
  971. Practical considerations on the use of preference learning for ranking emotional speech
  972. Precise phase transition of total variation minimization
  973. Precise player segmentation in team sports videos using contrast-aware co-segmentation
  974. Predicting humor response in dialogues from TV sitcoms
  975. Predicting visual attention using gamma kernels
  976. Prediction-adaptation-correction recurrent neural networks for low-resource language speech recognition
  977. Predominant melody extraction from vocal polyphonic music signal by combined spectro-temporal method
  978. Presentation quality assessment using acoustic information and hand movements
  979. Principal components analysis-based visual saliency detection
  980. Printed document authentication using two level or code
  981. Privacy-preserving energy flow control in smart grids
  982. Privacy-preserving nonparametric decentralized detection
  983. Privacy-preserving sound to degrade automatic speaker verification performance
  984. Progress on phoneme recognition with a continuous-state HMM
  985. Proportionate affine projection algorithms for block-sparse system identification
  986. Prosparse denoise: Prony's based sparse pattern recovery in the presence of noise
  987. Proximity without consensus in online multi-agent optimization
  988. Pruning subsequence search with attention-based embedding
  989. Pushing the limit of non-rigid structure-from-motion by shape clustering
  990. Quadtree decision for depth intra coding in 3D-HEVC by good feature
  991. Quality-aware adaptive delivery of multi-view video
  992. Quantification of balance in single limb stance using kinect
  993. Quantifying cooperation in choir singing: Respiratory and cardiac synchronisation
  994. Quantization bin matching for cloud storage of JPEG images
  995. Quantized consensus ADMM for multi-agent distributed optimization
  996. Quantizer design for exploiting common information in layered coding
  997. Quasi closed phase analysis of speech signals using time varying weighted linear prediction for accurate formant tracking
  998. Question detection from acoustic features using recurrent neural network with gated recurrent unit
  999. Quickest convergence of online algorithms via data selection
  1000. Quickest search over correlated sequences with model uncertainty

Looking for submission deadlines instead? See the conference deadline calendar.