ICASSP 2017 Accepted Papers
The full list of 1,320 papers accepted at ICASSP 2017 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- Practical Matlab experience in lecture-based signals and systems courses
- Practical strategies for content-adaptive batch steganography and pooled steganalysis
- Practically efficient nonlinear acoustic echo cancellers using cascaded block RLS and FLMS adaptive filters
- Pre-echo noise reduction in frequency-domain audio codecs
- Pre-movement contralateral EEG low beta power is modulated with motor adaptation learning
- Pre-processing and classification of hyperspectral imagery via selective inpainting
- Precision cell boundary tracking on DIC microscopy video for patch clamping
- Predicting dialogue success, naturalness, and length with acoustic features
- Predicting error rates for unknown data in automatic speech recognition
- Prediction-based learning for continuous emotion recognition in speech
- Predistortion for power amplifier linearization in full-duplex transceivers without extra RF chain
- Primal-dual algorithms for non-negative matrix factorization with the Kullback-Leibler divergence
- Prior knowledge aided super-resolution line spectral estimation: an iterative reweighted algorithm
- Privacy preserving Distance computation using somewhat-trusted third parties
- Privacy preserving encrypted phonetic search of speech data
- Privacy-preserving indoor localization via light transport analysis
- ProSparse extension: Prony's based sparse pattern recovery with extended dictionaries
- Probabilistic analysis of tone reservation method for the PAPR reduction of OFDM systems
- Probabilistic spatial dictionary based online adaptive beamforming for meeting recognition in noisy and reverberant environments
- Probabilistic transcription of sung melody using a pitch dynamic model
- Projection-based dual averaging for stochastic sparse optimization
- Proportionate NLMS for adaptive feedback control in hearing aids
- Quality assessment of voice converted speech using articulatory features
- Quality estimation based multi-focus image fusion
- Quantifying regulation mechanisms in dating couples through a dynamical systems model of acoustic and physiological arousal
- Quantisation effects in PDMM: A first study for synchronous distributed averaging
- Quantization-aware parameter estimation for audio upmixing
- Quickest change detection in structured data with incomplete information
- Quickest change detection under transient dynamics
- Quickest change detection with unknown post-change distribution
- RGB-NIR imaging with exposure bracketing for joint denoising and deblurring of low-light color images
- Radio-browsing for developmental monitoring in Uganda
- Radioastronomical Least Squares image reconstruction with iteration regularized Krylov subspaces and beamforming-based prior conditioning
- Random matrices meet machine learning: A large dimensional analysis of LS-SVM
- Ranking emotional attributes with deep neural networks
- Rate-coverage analysis and optimization for joint audio-video multimedia retrieval
- Rate-distortion analysis of Delta-Sigma modulators
- Rate-distortion trade-offs in acquisition of signal parameters
- Real-time and parallel SHVC hybrid codec AVC to HEVC decoder
- Real-time distributed speech enhancement with two collaborating microphone arrays
- Real-time implementation of hearing aid with combined noise and acoustic feedback reduction based on smartphone
- Realistic human action recognition: When CNNS meet LDS
- Realtime binaural speech enhancement demo on raspberry Pi
- Realtime plane detection for projection Augmented Reality in an unknown environment
- Receiver-transmitter pair selection in MIMO phased array radar
- Reconstruction-error-based learning for continuous emotion recognition in speech
- Recovery of sparse signals via Branch and Bound Least-Squares
- Recurrent Neural Network based language modeling with controllable external Memory
- Recurrent convolutional neural network for speech processing
- Recurrent deep stacking networks for supervised speech separation
- Recurrent latent variable conditional heteroscedasticity
- Recurrent neural network language models for keyword search
- Recursive Bayesian estimation of the acoustic noise emitted by wind farms
- Recursive Least-Squares algorithms for sparse system modeling
- RedDots replayed: A new replay spoofing attack corpus for text-dependent speaker verification research
- Reduced calibration by efficient transformation of templates for high speed hybrid coded SSVEP brain-computer interfaces
- Reduced-complexity digital predistortion for massive MIMO
- Reducing total latency in online real-time inference and decoding via combined context window and model smoothing latencies
- Reduction of necessary data rate for neural data through exponential and sinusoidal spline decomposition using the Finite Rate of Innovation framework
- Reflections: An eModule for echolocation education
- Regional deep feature aggregation for image retrieval
- Registration based retargeted image quality assessment
- Regularization of geophysical inversion using dictionary learning
- Regularized tracking of shear-wave in ultrasound elastography
- Reinforcing signal processing theory using real-time hardware
- Relative error bounds for nonnegative matrix factorization under a geometric assumption
- Remembering what you said: Semantic personalized memory for personal digital assistants
- Residual memory networks: Feed-forward approach to learn long-term temporal dependencies
- Resolution enhancement for hyperspectral images: A super-resolution and fusion approach
- Respiratory airflow estimation from lung sounds based on regression
- Retinex-based perceptual contrast enhancement in images using luminance adaptation
- Returnn: The RWTH extensible training framework for universal recurrent neural networks
- Reverberation-based feature extraction for acoustic scene classification
- Revisiting the problem of audio-based hit song prediction using convolutional neural networks
- RoDLSR: Robust discriminative least squares regression model for multi-category classification
- Robust Automatic Recognition of Speech with background music
- Robust DOA estimation in the presence of mis-calibrated sensors
- Robust MIMO OFDM transmit beamformer design for large Doppler scenarios under partial CSIT
- Robust MMSE filtering for single-microphone speech enhancement
- Robust and compact video descriptor learned by deep neural network
- Robust audio localization with phase unwrapping
- Robust clustering of data collected via crowdsourcing
- Robust direction estimation with convolutional neural networks based steered response power
- Robust feature selection for block covariance Bayesian models
- Robust front-end processing for Speech Recognition in noisy conditions
- Robust linear discriminant analysis with a Laplacian assumption on projection distribution
- Robust multichannel TDOA estimation for speaker localization using the impulsive characteristics of speech spectrum
- Robust network topology inference
- Robust online direction of arrival estimation using low dimensional spherical harmonic features
- Robust online matrix completion on graphs
- Robust particle filter by dynamic averaging of multiple noise models
- Robust reconstruction of spherical signals with finite rate of innovation
- Robust removal of fixed pattern noise on multi-focus images
- Robust speaker DOA estimation based on the inter-sensor data ratio model and binary mask estimation in the bispectrum domain
- Robust speaker recognition based on DNN/i-vectors and speech separation
- Robust spherical harmonic domain interpolation of spatially sampled array manifolds
- Robust transform learning
- Robust video fingerprints using positions of salient regions
- Robust visual tracking via deep discriminative model
- Robust visual tracking with deep feature fusion
- Rotation invariance through structured sparsity for robust hyperspectral image classification
- Run-length limited codes for backscatter communication
- SDR approximation bounds for the robust multicast beamforming problem with interference temperature constraints
- SPARTA: Sparse phase retrieval via Truncated Amplitude flow
- Salience based lexical features for emotion recognition
- Sample complexity bounds for dictionary learning of tensor data
- Sampling and reconstruction in the 21st century
- Sampling without time: Recovering echoes of light via temporal phase retrieval
- Scalable and flexible Max-Var generalized canonical correlation analysis via alternating optimization
- Scalable group level probabilistic sparse factor analysis
- Scale selective extended local binary pattern for texture classification
- Scaled and square-root elastic net
- Schedule based self localization of asynchronous wireless nodes with experimental validation
- Second-order performance analysis of Standard ESPRIT
- Second-order tensor-based convolutive ICA: Deconvolution versus tensorization
- Secure genomic susceptibility testing based on lattice encryption
- See and listen: Score-informed association of sound tracks to players in chamber music performance videos
- Segment-tree based cost aggregation for stereo matching with enhanced segmentation advantage
- Segmentation of music signals based on explained variance ratio for applications in spectral complexity reduction
- Selecting optimal layer reduction factors for model reduction of deep neural networks
- Selective object and context tracking
- Semantic mapping of natural language input to database entries via convolutional neural networks
- Semi-supervised classification via both label and side information
- Semi-supervised ensemble DNN acoustic model training
- Sensay analyticstm: A real-time speaker-state platform
- Sensor scheduling for target tracking in large multistatic sonobuoy fields
- Sequence segmentation using joint RNN and structured prediction models
- Sequence-to-sequence models for punctuated transcription combining lexical and acoustic features
- Sequential MCMC with invertible particle flow
- Sequential joint signal detection and signal-to-noise ratio estimation
- Set-membership kernel adaptive algorithms
- Shape from bandwidth: The 2-D orthogonal projection case
- Shape parameter estimation for generalized-Gaussian-distributed frequency spectra of audio signals
- Shefce: A Cantonese-English bilingual speech corpus for pronunciation assessment
- Signal representations in modern signal processing
- Simultaneous coded plane wave imaging in ultrasound: Problem formulation and constraints
- Simultaneous low-rank component and graph estimation for high-dimensional graph signals: Application to brain imaging
- Simultaneous segmentation and classification of bird song using CNN
- Simultaneous sparsity-based binary hypothesis model for real hyperspectral target detection
- Simultaneous wireless information and power transfer over inductively coupled circuits
- Single-channel Wiener filtering of deterministic signals in stochastic noise using the panorama
- Single-channel enhancement of convolutive noisy speech based on a discriminative NMF algorithm
- Single-tap equalizer for MIMO FBMC systems under doubly selective channels
- Skin detection based on multi-seed propagation in a multi-layer graph for regional and color consistency
- Smartphone-based anywhere-anytime signals and systems laboratory
- Smooth graph signal recovery via efficient Laplacian solvers
- Smoothed optimization for sparse off-grid directions-of-arrival estimation
- Son of Zorn's lemma: Targeted style transfer using instance-aware semantic segmentation
- Sound event detection using spatial features and convolutional recurrent neural network
- Sound field estimation using two spherical microphone arrays
- Sound physical property matching between non central listening point and central listening point for NHK 22.2 system reproduction
- Source tracking using moving microphone arrays for robot audition
- Sparse Bayesian learning with uncertain sensing matrix
- Sparse Signal Recovery for ultrasonic detection and reconstruction of shadowed flaws
- Sparse eigenvectors of graphs
- Sparse error correction with multiple measurement vectors: Observability-aware approach
- Sparse inverse bilateral filters for image processing
- Sparse modeling for topic-oriented video summarization
- Sparse reconstruction-based beampattern synthesis for multi-carrier frequency diverse array antenna
- Sparse representation for colors of 3D point cloud via virtual adaptive sampling
- Sparse signal recovery using generalized approximate message passing with built-in parameter estimation
- Sparse spectral estimation from point process observations
- Sparse waveform design for all-spectrum channelization
- Sparsity amplified
- Sparsity and low-rank amplitude based blind Source Separation
- Sparsity based super-resolution optical imaging using correlation information
- Sparsity regularized Principal Component Pursuit
- Sparsity-assisted signal smoothing (revisited)
- Spatial focusing inspired 5G spectrum sharing
- Spatio-temporal binary video inpainting via threshold dynamics
- Spatio-temporal sparse sound field decomposition considering acoustic source signal characteristics
- Speaker diarization using deep neural network embeddings
- Speaker diarization: A perspective on challenges and opportunities from theory to practice
- Speaker localization in reverberant rooms based on direct path dominance test statistics
- Speaker recognition using common passphrases in RedDots
- Speaker segmentation using deep speaker vectors for fast speaker change scenarios
- Speaker segmentation using i-vector in meetings domain
- Spectral statistics of lattice graph structured, non-uniform percolations
- Spectrum attacks aimed at minimizing spectrum opportunities
- Speech Activity Detection in online broadcast transcription using Deep Neural Networks and Weighted Finite State Transducers
- Speech dereverberation and denoising using complex ratio masks
- Speech dereverberation using NMF with regularized room impulse response
- Speech emotion recognition with ensemble learning methods
- Speech emotion recognition with skew-robust neural networks
- Speech enhancement based on Deep Neural Networks with skip connections
- Speech polarity detection using strength of impulse-like excitation extracted from speech epochs
- Speech recognition in unseen and noisy channel conditions
- Speech temporal dynamics fusion approaches for noise-robust reverberation time estimation
- Speeding up softmax computations in DNN-based large vocabulary speech recognition by senone weight vector selection
- Stable recovery of sparse vectors from random sinusoidal feature maps
- Stationary graph processes: Parametric power spectral estimation
- Statistical normalisation of phase-based feature representation for robust speech recognition
- Statistics of natural fused image distortions
- Steady-state mean square performance of a sparsified kernel least mean square algorithm
- Steganography with two JPEGs of the same scene
- Stereo image de-fencing using smartphones
- Stereoscopic image quality assessment based on the binocular properties of the human visual system
- Stimulated training for automatic speech recognition and keyword search in limited resource conditions
- Stochastic Truncated Wirtinger Flow Algorithm for phase retrieval using boolean coded apertures
- Stochastic backpressure in energy harvesting networks
- Stochastic filtering of two-photon imaging using reweighted ℓ1
- Stochastic online control for energy-harvesting wireless networks with battery imperfections
- Stronger recovery guarantees for sparse signals exploiting coherence structure in dictionaries
- Structure of the set of signals with strong divergence of the Shannon sampling series
- Structure-aware classification using supervised dictionary learning
- Structured dictionary learning for sparse common component and innovation model
- Structured dropout for weak label and multi-instance learning and its application to score-informed source separation
- Structured estimation of time-varying narrowband wireless communication channels
- Student-teacher network learning with enhanced features
- Study of the frequency-domain multichannel noise reduction problem with the householder transformation
- Study-flow: Studying effective student-content interaction in signal processing education
- Sub-Nyquist pulse Doppler MIMO radar
- Subjective and objective quality assessment of Mobile Videos with In-Capture distortions
- Subspace projection cepstral coefficients for noise robust acoustic event recognition
- Summarization of human activity videos via low-rank approximation
- Super-resolution delay-Doppler estimation for sub-Nyquist radar via atomic norm minimization
- Super-resolution for differently exposed mixed-resolution multi-view images adapted by a histogram matching method
- Superpixel-guided CFAR detection of ships at sea in SAR imagery
- Supervised audio tampering detection using an autoregressive model
- Supervised group nonnegative matrix factorisation with similarity constraints and applications to speaker identification
- Supervised independent vector analysis through pilot dependent components
- Supervised monaural source separation based on autoencoders
- Supervised source enhancement composed of nonnegative auto-encoders and complementarity subtraction
- Surrounding adaptive tone mapping in displayed images under ambient light
- Synchronization for multi-perspective videos in the wild
- Syntax Element Partitioning for high-throughput HEVC CABAC decoding
- Synthesis versus analysis in patch-based image priors
- Taichi distance for person re-identification
- Target detecton and tracking via structured convex optimization
- Teaching image and video processing using middle-school mathematics and the Raspberry Pi
- Temporal localization of audio events for conflict monitoring in social media
- Tensor-based crowdsourced clustering via triangle queries
- The 2016 BBN Georgian telephone speech keyword spotting system
- The Power-Oja method for decentralized subspace estimation/tracking
- The Sheffield Search and Rescue corpus
- The counterintuitive mechanism of graph-based semi-supervised learning in the big data regime
- The geometry of random paired comparisons
- The group k-support norm for learning with structured sparsity
- The microsoft 2016 conversational speech recognition system
- The penalty term of Exponentially Embedded Family is estimated mutual information
- The second-order wavelet synchrosqueezing transform
- Theoretical vulnerabilities in map speaker adaptation
- Three dimensional ultrasound imaging of pre- and post-vocalic liquid consonants in American English: Preliminary observations
- Through-the-wall radar signal classification using discriminative dictionary learning
- Time and frequency domain long short-term memory for noise robust pitch tracking
- Time of arrival disambiguation using the linear Radon transform
- Time reversal based wireless events detection
- Time-domain channel estimation for wideband millimeter wave systems with hybrid architecture
- Time-frequency processing for sound source localization from a micro aerial vehicle
- Time-multiplexed / superimposed pilot selection for massive MIMO pilot decontamination
- Topic identification of spoken documents using unsupervised acoustic unit discovery
- Topology inference of directed graphs using nonlinear structural vector autoregressive models
- Towards a definition of local stationarity for graph signals
- Towards confidence measures on fundamental frequency estimations
- Towards decoding speech production from single-trial magnetoencephalography (MEG) signals
- Towards expressive instrument synthesis through smooth frame-by-frame reconstruction: From string to woodwind
- Towards phoneme inventory discovery for documentation of unwritten languages
- Towards stationary time-vertex signal processing
- Towards the characterization of singing styles in world music
- Towards wireless acoustic sensor networks for location estimation and counting of multiple speakers in real-life conditions
- Tracking metrical structure changes with sparse-NMF
- Traffic congestion analysis: A new Perspective
- Traffic engineering for backhaul networks with wireless link scheduling
- Trainable frontend for robust and far-field keyword spotting
- Training algorithm to deceive Anti-Spoofing Verification for DNN-based speech synthesis
- Training data reduction in deep neural networks with partial mutual information based feature selection and correlation matching based active learning
- Training variance and performance evaluation of neural networks in speech
- Transfer learning for EEG based BCI using LEARN++.NSE and mutual information
- Transfer of vignetting effect from paintings to photographs
- Transferring clothing parsing from fashion dataset to surveillance
- Transparent objects: Influence of shape and color on depth perception
- TristouNet: Triplet loss for speaker turn embedding
- Two models for fusion of medical imaging data: Comparison and connections
- Two-dimensional anti-jamming communication based on deep reinforcement learning
- Two-stage facial age prediction using group-specific features
- UWB radar signal processing in measurement of heartbeat features
- Ultra-fast robust compressive sensing based on memristor crossbars
- Ultrasound based gesture recognition
- Underdetermined source separation using time-frequency masks and an adaptive combined Gaussian-Student's t probabilistic model
- Unified analysis of co-array interpolation for direction-of-arrival estimation
- Unifying attribute splitting criteria of decision trees by Tsallis entropy
- Universal bounds for the sampling of graph signals
- Unlabeled sensing: Reconstruction algorithm and theoretical guarantees
- Unsupervised adaptation for deep neural networks using Alternating Direction Method of Multipliers
- Unsupervised adaptation of deep neural networks for sound source localization using entropy minimization
- Unsupervised feature extraction for hyperspectral images using combined low rank representation and locally linear embedding
- Unsupervised image segmentation using convolutional autoencoder with total variation regularization as preprocessing
- Unsupervised latent behavior manifold learning from acoustic features: Audio2behavior
- Unsupervised learning of asymmetric high-order autoregressive stochastic volatility model
- Unsupervised speaker adaptation of batch normalized acoustic models for robust ASR
- Unsupervised utterance-wise beamformer estimation with speech recognition-level criterion
- Uplink and downlink user pairing in full-duplex multi-user systems: Complexity and algorithms
- Use of affect based interaction classification for continuous emotion tracking
- User assisted separation of repeating patterns in time and frequency using magnitude projections
- Using optimal transport for estimating inharmonic pitch signals
- Using regional saliency for speech emotion recognition
- Variational inference for nonparametric subspace dictionary learning with hierarchical beta process
- Variational manifold learning for speaker recognition
- Vehicle tracking in Wide area motion imagery: A facility location motivated combinatorial approach
- Very deep convolutional networks for end-to-end speech recognition
- Very deep convolutional neural networks for raw waveforms
- Very low bitrate spatial audio coding with dimensionality reduction
- Vid2speech: Speech reconstruction from silent video
- Visual features for context-aware speech recognition
- Visually informed multi-pitch analysis of string ensembles
- Voice-transformation-based data augmentation for prosodic classification
- Wavelet based head movement artifact removal from electrooculography signals
- Wavelet-based single image super-resolution with an overall enhancement procedure
- Weak interference detection with signal cancellation in satellite communications
- Weak law of large numbers for stationary graph processes
- Weakly supervised spoken term discovery using cross-lingual side information
- Weakly-supervised audio event detection using event-specific Gaussian filters and fully convolutional networks
- Wearable motion sensor based phasic analysis of tennis serve for performance feedback
- When sparsity meets low-rankness: Transform learning with non-local low-rank constraint for image restoration
- Word level lyrics-audio synchronization using separated vocals
- X-ray Computed Tomography simultaneous image reconstruction and contour detection using a hierarchical Markovian model
- Xampling-enabled coexistence in spectrally crowded environments
- e-vectors: JFA and i-vectors revisited
- eAMR: Wideband speech over legacy narrowband networks
- i-Vector/PLDA speaker recognition using support vectors with discriminant analysis
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.