← All conferences

ICASSP 2017 Accepted Papers

The full list of 1,320 papers accepted at ICASSP 2017 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. Practical Matlab experience in lecture-based signals and systems courses
  2. Practical strategies for content-adaptive batch steganography and pooled steganalysis
  3. Practically efficient nonlinear acoustic echo cancellers using cascaded block RLS and FLMS adaptive filters
  4. Pre-echo noise reduction in frequency-domain audio codecs
  5. Pre-movement contralateral EEG low beta power is modulated with motor adaptation learning
  6. Pre-processing and classification of hyperspectral imagery via selective inpainting
  7. Precision cell boundary tracking on DIC microscopy video for patch clamping
  8. Predicting dialogue success, naturalness, and length with acoustic features
  9. Predicting error rates for unknown data in automatic speech recognition
  10. Prediction-based learning for continuous emotion recognition in speech
  11. Predistortion for power amplifier linearization in full-duplex transceivers without extra RF chain
  12. Primal-dual algorithms for non-negative matrix factorization with the Kullback-Leibler divergence
  13. Prior knowledge aided super-resolution line spectral estimation: an iterative reweighted algorithm
  14. Privacy preserving Distance computation using somewhat-trusted third parties
  15. Privacy preserving encrypted phonetic search of speech data
  16. Privacy-preserving indoor localization via light transport analysis
  17. ProSparse extension: Prony's based sparse pattern recovery with extended dictionaries
  18. Probabilistic analysis of tone reservation method for the PAPR reduction of OFDM systems
  19. Probabilistic spatial dictionary based online adaptive beamforming for meeting recognition in noisy and reverberant environments
  20. Probabilistic transcription of sung melody using a pitch dynamic model
  21. Projection-based dual averaging for stochastic sparse optimization
  22. Proportionate NLMS for adaptive feedback control in hearing aids
  23. Quality assessment of voice converted speech using articulatory features
  24. Quality estimation based multi-focus image fusion
  25. Quantifying regulation mechanisms in dating couples through a dynamical systems model of acoustic and physiological arousal
  26. Quantisation effects in PDMM: A first study for synchronous distributed averaging
  27. Quantization-aware parameter estimation for audio upmixing
  28. Quickest change detection in structured data with incomplete information
  29. Quickest change detection under transient dynamics
  30. Quickest change detection with unknown post-change distribution
  31. RGB-NIR imaging with exposure bracketing for joint denoising and deblurring of low-light color images
  32. Radio-browsing for developmental monitoring in Uganda
  33. Radioastronomical Least Squares image reconstruction with iteration regularized Krylov subspaces and beamforming-based prior conditioning
  34. Random matrices meet machine learning: A large dimensional analysis of LS-SVM
  35. Ranking emotional attributes with deep neural networks
  36. Rate-coverage analysis and optimization for joint audio-video multimedia retrieval
  37. Rate-distortion analysis of Delta-Sigma modulators
  38. Rate-distortion trade-offs in acquisition of signal parameters
  39. Real-time and parallel SHVC hybrid codec AVC to HEVC decoder
  40. Real-time distributed speech enhancement with two collaborating microphone arrays
  41. Real-time implementation of hearing aid with combined noise and acoustic feedback reduction based on smartphone
  42. Realistic human action recognition: When CNNS meet LDS
  43. Realtime binaural speech enhancement demo on raspberry Pi
  44. Realtime plane detection for projection Augmented Reality in an unknown environment
  45. Receiver-transmitter pair selection in MIMO phased array radar
  46. Reconstruction-error-based learning for continuous emotion recognition in speech
  47. Recovery of sparse signals via Branch and Bound Least-Squares
  48. Recurrent Neural Network based language modeling with controllable external Memory
  49. Recurrent convolutional neural network for speech processing
  50. Recurrent deep stacking networks for supervised speech separation
  51. Recurrent latent variable conditional heteroscedasticity
  52. Recurrent neural network language models for keyword search
  53. Recursive Bayesian estimation of the acoustic noise emitted by wind farms
  54. Recursive Least-Squares algorithms for sparse system modeling
  55. RedDots replayed: A new replay spoofing attack corpus for text-dependent speaker verification research
  56. Reduced calibration by efficient transformation of templates for high speed hybrid coded SSVEP brain-computer interfaces
  57. Reduced-complexity digital predistortion for massive MIMO
  58. Reducing total latency in online real-time inference and decoding via combined context window and model smoothing latencies
  59. Reduction of necessary data rate for neural data through exponential and sinusoidal spline decomposition using the Finite Rate of Innovation framework
  60. Reflections: An eModule for echolocation education
  61. Regional deep feature aggregation for image retrieval
  62. Registration based retargeted image quality assessment
  63. Regularization of geophysical inversion using dictionary learning
  64. Regularized tracking of shear-wave in ultrasound elastography
  65. Reinforcing signal processing theory using real-time hardware
  66. Relative error bounds for nonnegative matrix factorization under a geometric assumption
  67. Remembering what you said: Semantic personalized memory for personal digital assistants
  68. Residual memory networks: Feed-forward approach to learn long-term temporal dependencies
  69. Resolution enhancement for hyperspectral images: A super-resolution and fusion approach
  70. Respiratory airflow estimation from lung sounds based on regression
  71. Retinex-based perceptual contrast enhancement in images using luminance adaptation
  72. Returnn: The RWTH extensible training framework for universal recurrent neural networks
  73. Reverberation-based feature extraction for acoustic scene classification
  74. Revisiting the problem of audio-based hit song prediction using convolutional neural networks
  75. RoDLSR: Robust discriminative least squares regression model for multi-category classification
  76. Robust Automatic Recognition of Speech with background music
  77. Robust DOA estimation in the presence of mis-calibrated sensors
  78. Robust MIMO OFDM transmit beamformer design for large Doppler scenarios under partial CSIT
  79. Robust MMSE filtering for single-microphone speech enhancement
  80. Robust and compact video descriptor learned by deep neural network
  81. Robust audio localization with phase unwrapping
  82. Robust clustering of data collected via crowdsourcing
  83. Robust direction estimation with convolutional neural networks based steered response power
  84. Robust feature selection for block covariance Bayesian models
  85. Robust front-end processing for Speech Recognition in noisy conditions
  86. Robust linear discriminant analysis with a Laplacian assumption on projection distribution
  87. Robust multichannel TDOA estimation for speaker localization using the impulsive characteristics of speech spectrum
  88. Robust network topology inference
  89. Robust online direction of arrival estimation using low dimensional spherical harmonic features
  90. Robust online matrix completion on graphs
  91. Robust particle filter by dynamic averaging of multiple noise models
  92. Robust reconstruction of spherical signals with finite rate of innovation
  93. Robust removal of fixed pattern noise on multi-focus images
  94. Robust speaker DOA estimation based on the inter-sensor data ratio model and binary mask estimation in the bispectrum domain
  95. Robust speaker recognition based on DNN/i-vectors and speech separation
  96. Robust spherical harmonic domain interpolation of spatially sampled array manifolds
  97. Robust transform learning
  98. Robust video fingerprints using positions of salient regions
  99. Robust visual tracking via deep discriminative model
  100. Robust visual tracking with deep feature fusion
  101. Rotation invariance through structured sparsity for robust hyperspectral image classification
  102. Run-length limited codes for backscatter communication
  103. SDR approximation bounds for the robust multicast beamforming problem with interference temperature constraints
  104. SPARTA: Sparse phase retrieval via Truncated Amplitude flow
  105. Salience based lexical features for emotion recognition
  106. Sample complexity bounds for dictionary learning of tensor data
  107. Sampling and reconstruction in the 21st century
  108. Sampling without time: Recovering echoes of light via temporal phase retrieval
  109. Scalable and flexible Max-Var generalized canonical correlation analysis via alternating optimization
  110. Scalable group level probabilistic sparse factor analysis
  111. Scale selective extended local binary pattern for texture classification
  112. Scaled and square-root elastic net
  113. Schedule based self localization of asynchronous wireless nodes with experimental validation
  114. Second-order performance analysis of Standard ESPRIT
  115. Second-order tensor-based convolutive ICA: Deconvolution versus tensorization
  116. Secure genomic susceptibility testing based on lattice encryption
  117. See and listen: Score-informed association of sound tracks to players in chamber music performance videos
  118. Segment-tree based cost aggregation for stereo matching with enhanced segmentation advantage
  119. Segmentation of music signals based on explained variance ratio for applications in spectral complexity reduction
  120. Selecting optimal layer reduction factors for model reduction of deep neural networks
  121. Selective object and context tracking
  122. Semantic mapping of natural language input to database entries via convolutional neural networks
  123. Semi-supervised classification via both label and side information
  124. Semi-supervised ensemble DNN acoustic model training
  125. Sensay analyticstm: A real-time speaker-state platform
  126. Sensor scheduling for target tracking in large multistatic sonobuoy fields
  127. Sequence segmentation using joint RNN and structured prediction models
  128. Sequence-to-sequence models for punctuated transcription combining lexical and acoustic features
  129. Sequential MCMC with invertible particle flow
  130. Sequential joint signal detection and signal-to-noise ratio estimation
  131. Set-membership kernel adaptive algorithms
  132. Shape from bandwidth: The 2-D orthogonal projection case
  133. Shape parameter estimation for generalized-Gaussian-distributed frequency spectra of audio signals
  134. Shefce: A Cantonese-English bilingual speech corpus for pronunciation assessment
  135. Signal representations in modern signal processing
  136. Simultaneous coded plane wave imaging in ultrasound: Problem formulation and constraints
  137. Simultaneous low-rank component and graph estimation for high-dimensional graph signals: Application to brain imaging
  138. Simultaneous segmentation and classification of bird song using CNN
  139. Simultaneous sparsity-based binary hypothesis model for real hyperspectral target detection
  140. Simultaneous wireless information and power transfer over inductively coupled circuits
  141. Single-channel Wiener filtering of deterministic signals in stochastic noise using the panorama
  142. Single-channel enhancement of convolutive noisy speech based on a discriminative NMF algorithm
  143. Single-tap equalizer for MIMO FBMC systems under doubly selective channels
  144. Skin detection based on multi-seed propagation in a multi-layer graph for regional and color consistency
  145. Smartphone-based anywhere-anytime signals and systems laboratory
  146. Smooth graph signal recovery via efficient Laplacian solvers
  147. Smoothed optimization for sparse off-grid directions-of-arrival estimation
  148. Son of Zorn's lemma: Targeted style transfer using instance-aware semantic segmentation
  149. Sound event detection using spatial features and convolutional recurrent neural network
  150. Sound field estimation using two spherical microphone arrays
  151. Sound physical property matching between non central listening point and central listening point for NHK 22.2 system reproduction
  152. Source tracking using moving microphone arrays for robot audition
  153. Sparse Bayesian learning with uncertain sensing matrix
  154. Sparse Signal Recovery for ultrasonic detection and reconstruction of shadowed flaws
  155. Sparse eigenvectors of graphs
  156. Sparse error correction with multiple measurement vectors: Observability-aware approach
  157. Sparse inverse bilateral filters for image processing
  158. Sparse modeling for topic-oriented video summarization
  159. Sparse reconstruction-based beampattern synthesis for multi-carrier frequency diverse array antenna
  160. Sparse representation for colors of 3D point cloud via virtual adaptive sampling
  161. Sparse signal recovery using generalized approximate message passing with built-in parameter estimation
  162. Sparse spectral estimation from point process observations
  163. Sparse waveform design for all-spectrum channelization
  164. Sparsity amplified
  165. Sparsity and low-rank amplitude based blind Source Separation
  166. Sparsity based super-resolution optical imaging using correlation information
  167. Sparsity regularized Principal Component Pursuit
  168. Sparsity-assisted signal smoothing (revisited)
  169. Spatial focusing inspired 5G spectrum sharing
  170. Spatio-temporal binary video inpainting via threshold dynamics
  171. Spatio-temporal sparse sound field decomposition considering acoustic source signal characteristics
  172. Speaker diarization using deep neural network embeddings
  173. Speaker diarization: A perspective on challenges and opportunities from theory to practice
  174. Speaker localization in reverberant rooms based on direct path dominance test statistics
  175. Speaker recognition using common passphrases in RedDots
  176. Speaker segmentation using deep speaker vectors for fast speaker change scenarios
  177. Speaker segmentation using i-vector in meetings domain
  178. Spectral statistics of lattice graph structured, non-uniform percolations
  179. Spectrum attacks aimed at minimizing spectrum opportunities
  180. Speech Activity Detection in online broadcast transcription using Deep Neural Networks and Weighted Finite State Transducers
  181. Speech dereverberation and denoising using complex ratio masks
  182. Speech dereverberation using NMF with regularized room impulse response
  183. Speech emotion recognition with ensemble learning methods
  184. Speech emotion recognition with skew-robust neural networks
  185. Speech enhancement based on Deep Neural Networks with skip connections
  186. Speech polarity detection using strength of impulse-like excitation extracted from speech epochs
  187. Speech recognition in unseen and noisy channel conditions
  188. Speech temporal dynamics fusion approaches for noise-robust reverberation time estimation
  189. Speeding up softmax computations in DNN-based large vocabulary speech recognition by senone weight vector selection
  190. Stable recovery of sparse vectors from random sinusoidal feature maps
  191. Stationary graph processes: Parametric power spectral estimation
  192. Statistical normalisation of phase-based feature representation for robust speech recognition
  193. Statistics of natural fused image distortions
  194. Steady-state mean square performance of a sparsified kernel least mean square algorithm
  195. Steganography with two JPEGs of the same scene
  196. Stereo image de-fencing using smartphones
  197. Stereoscopic image quality assessment based on the binocular properties of the human visual system
  198. Stimulated training for automatic speech recognition and keyword search in limited resource conditions
  199. Stochastic Truncated Wirtinger Flow Algorithm for phase retrieval using boolean coded apertures
  200. Stochastic backpressure in energy harvesting networks
  201. Stochastic filtering of two-photon imaging using reweighted ℓ1
  202. Stochastic online control for energy-harvesting wireless networks with battery imperfections
  203. Stronger recovery guarantees for sparse signals exploiting coherence structure in dictionaries
  204. Structure of the set of signals with strong divergence of the Shannon sampling series
  205. Structure-aware classification using supervised dictionary learning
  206. Structured dictionary learning for sparse common component and innovation model
  207. Structured dropout for weak label and multi-instance learning and its application to score-informed source separation
  208. Structured estimation of time-varying narrowband wireless communication channels
  209. Student-teacher network learning with enhanced features
  210. Study of the frequency-domain multichannel noise reduction problem with the householder transformation
  211. Study-flow: Studying effective student-content interaction in signal processing education
  212. Sub-Nyquist pulse Doppler MIMO radar
  213. Subjective and objective quality assessment of Mobile Videos with In-Capture distortions
  214. Subspace projection cepstral coefficients for noise robust acoustic event recognition
  215. Summarization of human activity videos via low-rank approximation
  216. Super-resolution delay-Doppler estimation for sub-Nyquist radar via atomic norm minimization
  217. Super-resolution for differently exposed mixed-resolution multi-view images adapted by a histogram matching method
  218. Superpixel-guided CFAR detection of ships at sea in SAR imagery
  219. Supervised audio tampering detection using an autoregressive model
  220. Supervised group nonnegative matrix factorisation with similarity constraints and applications to speaker identification
  221. Supervised independent vector analysis through pilot dependent components
  222. Supervised monaural source separation based on autoencoders
  223. Supervised source enhancement composed of nonnegative auto-encoders and complementarity subtraction
  224. Surrounding adaptive tone mapping in displayed images under ambient light
  225. Synchronization for multi-perspective videos in the wild
  226. Syntax Element Partitioning for high-throughput HEVC CABAC decoding
  227. Synthesis versus analysis in patch-based image priors
  228. Taichi distance for person re-identification
  229. Target detecton and tracking via structured convex optimization
  230. Teaching image and video processing using middle-school mathematics and the Raspberry Pi
  231. Temporal localization of audio events for conflict monitoring in social media
  232. Tensor-based crowdsourced clustering via triangle queries
  233. The 2016 BBN Georgian telephone speech keyword spotting system
  234. The Power-Oja method for decentralized subspace estimation/tracking
  235. The Sheffield Search and Rescue corpus
  236. The counterintuitive mechanism of graph-based semi-supervised learning in the big data regime
  237. The geometry of random paired comparisons
  238. The group k-support norm for learning with structured sparsity
  239. The microsoft 2016 conversational speech recognition system
  240. The penalty term of Exponentially Embedded Family is estimated mutual information
  241. The second-order wavelet synchrosqueezing transform
  242. Theoretical vulnerabilities in map speaker adaptation
  243. Three dimensional ultrasound imaging of pre- and post-vocalic liquid consonants in American English: Preliminary observations
  244. Through-the-wall radar signal classification using discriminative dictionary learning
  245. Time and frequency domain long short-term memory for noise robust pitch tracking
  246. Time of arrival disambiguation using the linear Radon transform
  247. Time reversal based wireless events detection
  248. Time-domain channel estimation for wideband millimeter wave systems with hybrid architecture
  249. Time-frequency processing for sound source localization from a micro aerial vehicle
  250. Time-multiplexed / superimposed pilot selection for massive MIMO pilot decontamination
  251. Topic identification of spoken documents using unsupervised acoustic unit discovery
  252. Topology inference of directed graphs using nonlinear structural vector autoregressive models
  253. Towards a definition of local stationarity for graph signals
  254. Towards confidence measures on fundamental frequency estimations
  255. Towards decoding speech production from single-trial magnetoencephalography (MEG) signals
  256. Towards expressive instrument synthesis through smooth frame-by-frame reconstruction: From string to woodwind
  257. Towards phoneme inventory discovery for documentation of unwritten languages
  258. Towards stationary time-vertex signal processing
  259. Towards the characterization of singing styles in world music
  260. Towards wireless acoustic sensor networks for location estimation and counting of multiple speakers in real-life conditions
  261. Tracking metrical structure changes with sparse-NMF
  262. Traffic congestion analysis: A new Perspective
  263. Traffic engineering for backhaul networks with wireless link scheduling
  264. Trainable frontend for robust and far-field keyword spotting
  265. Training algorithm to deceive Anti-Spoofing Verification for DNN-based speech synthesis
  266. Training data reduction in deep neural networks with partial mutual information based feature selection and correlation matching based active learning
  267. Training variance and performance evaluation of neural networks in speech
  268. Transfer learning for EEG based BCI using LEARN++.NSE and mutual information
  269. Transfer of vignetting effect from paintings to photographs
  270. Transferring clothing parsing from fashion dataset to surveillance
  271. Transparent objects: Influence of shape and color on depth perception
  272. TristouNet: Triplet loss for speaker turn embedding
  273. Two models for fusion of medical imaging data: Comparison and connections
  274. Two-dimensional anti-jamming communication based on deep reinforcement learning
  275. Two-stage facial age prediction using group-specific features
  276. UWB radar signal processing in measurement of heartbeat features
  277. Ultra-fast robust compressive sensing based on memristor crossbars
  278. Ultrasound based gesture recognition
  279. Underdetermined source separation using time-frequency masks and an adaptive combined Gaussian-Student's t probabilistic model
  280. Unified analysis of co-array interpolation for direction-of-arrival estimation
  281. Unifying attribute splitting criteria of decision trees by Tsallis entropy
  282. Universal bounds for the sampling of graph signals
  283. Unlabeled sensing: Reconstruction algorithm and theoretical guarantees
  284. Unsupervised adaptation for deep neural networks using Alternating Direction Method of Multipliers
  285. Unsupervised adaptation of deep neural networks for sound source localization using entropy minimization
  286. Unsupervised feature extraction for hyperspectral images using combined low rank representation and locally linear embedding
  287. Unsupervised image segmentation using convolutional autoencoder with total variation regularization as preprocessing
  288. Unsupervised latent behavior manifold learning from acoustic features: Audio2behavior
  289. Unsupervised learning of asymmetric high-order autoregressive stochastic volatility model
  290. Unsupervised speaker adaptation of batch normalized acoustic models for robust ASR
  291. Unsupervised utterance-wise beamformer estimation with speech recognition-level criterion
  292. Uplink and downlink user pairing in full-duplex multi-user systems: Complexity and algorithms
  293. Use of affect based interaction classification for continuous emotion tracking
  294. User assisted separation of repeating patterns in time and frequency using magnitude projections
  295. Using optimal transport for estimating inharmonic pitch signals
  296. Using regional saliency for speech emotion recognition
  297. Variational inference for nonparametric subspace dictionary learning with hierarchical beta process
  298. Variational manifold learning for speaker recognition
  299. Vehicle tracking in Wide area motion imagery: A facility location motivated combinatorial approach
  300. Very deep convolutional networks for end-to-end speech recognition
  301. Very deep convolutional neural networks for raw waveforms
  302. Very low bitrate spatial audio coding with dimensionality reduction
  303. Vid2speech: Speech reconstruction from silent video
  304. Visual features for context-aware speech recognition
  305. Visually informed multi-pitch analysis of string ensembles
  306. Voice-transformation-based data augmentation for prosodic classification
  307. Wavelet based head movement artifact removal from electrooculography signals
  308. Wavelet-based single image super-resolution with an overall enhancement procedure
  309. Weak interference detection with signal cancellation in satellite communications
  310. Weak law of large numbers for stationary graph processes
  311. Weakly supervised spoken term discovery using cross-lingual side information
  312. Weakly-supervised audio event detection using event-specific Gaussian filters and fully convolutional networks
  313. Wearable motion sensor based phasic analysis of tennis serve for performance feedback
  314. When sparsity meets low-rankness: Transform learning with non-local low-rank constraint for image restoration
  315. Word level lyrics-audio synchronization using separated vocals
  316. X-ray Computed Tomography simultaneous image reconstruction and contour detection using a hierarchical Markovian model
  317. Xampling-enabled coexistence in spectrally crowded environments
  318. e-vectors: JFA and i-vectors revisited
  319. eAMR: Wideband speech over legacy narrowband networks
  320. i-Vector/PLDA speaker recognition using support vectors with discriminant analysis

Looking for submission deadlines instead? See the conference deadline calendar.