ICASSP 2016 Accepted Papers
The full list of 1,322 papers accepted at ICASSP 2016 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- 1-Bit compressed sensing of positive semi-definite matrices via rank-1 measurement matrices
- 2.5D higher order ambisonics for a sound field described by angular spectrum coefficients
- 3D DOA estimation of multiple sound sources based on spatially constrained beamforming driven by intensity vectors
- 3D acoustic source localization in the spherical harmonic domain based on optimized grid search
- 3D mesh steganalysis using local shape features
- 3D panorama reconstruction based on sitemap joining
- 3D pseudolinear Kalman filter with own-ship path optimization for AOA target tracking
- 3D video frame interpolation via adaptive hybrid motion estimation and compensation
- A Bayesian framework for the multifractal analysis of images using data augmentation and a whittle approximation
- A GMM-based stair quality model for human perceived JPEG images
- A Gaussian mixture regression approach toward modeling the affective dynamics between acoustically-derived vocal arousal score (VC-AS) and internal brain fMRI bold signal response
- A KL divergence and DNN approach to cross-lingual TTS
- A MIL-based interactive approach for hotspot segmentation from bone scintigraphy
- A Metropolis-within-Gibbs sampler to infer task-based functional brain connectivity
- A behavior-based evaluation of product quality
- A benchmark for robustness analysis of visual tracking algorithms
- A comparative study of multi-channel processing methods for noisy automatic speech recognition in urban environments
- A comparative study of recurrent neural network models for lexical domain classification
- A comparative study of robustness of deep learning approaches for VAD
- A comparison between deep neural nets and kernel acoustic models for speech recognition
- A comparison of ASR and human errors for transcription of non-native spontaneous speech
- A cost-effective minutiae disk code for fingerprint recognition and its implementation
- A data set providing synthetic and real-world fisheye video sequences
- A deep auto-encoder based low-dimensional feature extraction from FFT spectral envelopes for statistical parametric speech synthesis
- A deep bidirectional long short-term memory based multi-scale approach for music dynamic emotion prediction
- A deep scattering spectrum - Deep Siamese network pipeline for unsupervised acoustic modeling
- A deterministic plus noise model of excitation signal using principal component analysis for parametric speech synthesis
- A diffusion kernel LMS algorithm for nonlinear adaptive networks
- A distributed algorithm for robust LCMV beamforming
- A divide-and-conquer dictionary learning algorithm and its performance analysis
- A dynamic Bayesian network approach for device-free radio vision: Modeling, learning and inference for body motion recognition
- A fast 3D face reconstruction method from a single image using adjustable model
- A fast direct source localization approach for acoustic sensor array
- A fast dual iterative algorithm for convexly constrained spline smoothing
- A fast method of fog and haze removal
- A frame rate up-conversion method with quadruple motion vector post-processing
- A framework for globally optimal energy-efficient resource allocation in wireless networks
- A full training framework of cross-stream dependence modelling for HMM-based singing voice synthesis
- A general baseband volterra model for dual-band predistortion
- A general framework for reconstruction and classification from compressive measurements with side information
- A generalized Bayesian model for tracking long metrical cycles in acoustic music signals
- A generalized LDPC framework for robust and sublinear compressive sensing
- A generative-discriminative hybrid approach to multi-channel noise reduction for robust automatic speech recognition
- A generator of memory-based, runtime-reconfigurable 2N3M5K FFT engines
- A glottal chink model for the synthesis of voiced fricatives
- A hierarchical algorithm for causality discovery among atrial fibrillation electrograms
- A hierarchical framework for language identification
- A highly parallel coding unit size selection for HEVC
- A joint approach to vector road map registration and vehicle tracking for wide area motion imagery
- A joint design approach for spectrum sharing between radar and communication systems
- A joint learning approach for cross domain age estimation
- A largest matching area approach to image denoising
- A lattice algorithm for optimal phase unwrapping in noise
- A least square approach for distributed sensor fusion in bandwidth-constrained sensor networks
- A linear operator for the computation of soundfield maps
- A linear sensor array with self-bending sensitivity
- A low complexity weighted least squares narrowband DOA estimator for arbitrary array geometries
- A low-cost solution to 3D pinna modeling for HRTF prediction
- A machine learning approach for computationally and energy efficient speech enhancement in binaural hearing aids
- A map-based NMF approach to hyperspectral image unmixing using a linear-quadratic mixture model
- A maximum likelihood-based unscented Kalman filter for multipath mitigation in a multi-correlator based GNSS receiver
- A method for predicting the intelligibility of noisy and non-linearly enhanced binaural speech
- A method to reconstruct coverage loss maps based on matrix completion and adaptive sampling
- A multi-scale approach to extract meaningful annotations from document images
- A multimodal analysis of synchrony during dyadic interaction using a metric based on sequential pattern mining
- A multimodal mixture-of-experts model for dynamic emotion prediction in movies
- A new approach for heart rate monitoring using photoplethysmography signals contaminated by motion artifacts
- A new array geometry for DOA estimation with enhanced degrees of freedom
- A new haze image database with detailed air quality information and a novel no-reference image quality assessment method for haze images
- A new low-rank solution result for a semidefinite program problem subclass with applications to transmit beamforming optimization
- A new time-frequency approach for underdetermined convolutive blind speech separation
- A new uncertainty decoding scheme for DNN-HMM hybrid systems with multichannel speech enhancement
- A normalized spatial spectrum for DOA estimation with uniform linear arrays in the presence of unknown mutual coupling
- A novel DNN-HMM-based approach for extracting single loads from aggregate power signals
- A novel array processing method for precise depth detection of ultrasound point scatter
- A novel color space based on RGB color barycenter
- A novel feedforward noise shaping for word-length reduction
- A novel generalized assignment framework for the classification of hyperspectral image
- A novel image classifier based on Gaussian mixture language model
- A novel sub-Nyquist Fourier transform estimator based on alias-free hybrid stratified sampling
- A novel time-frequency feature extraction algorithm based on dictionary learning
- A novel video-based smoke detection method based on color invariants
- A parameter-free Cauchy-Schwartz information measure for independent component analysis
- A partial least squares based ranker for fast and accurate age estimation
- A partitioned approach to signal separation with microphone ad hoc arrays
- A penalty-BSUM approach for rate optimization in full-duplex MIMO relay networks with relay processing delay
- A phonetically aware system for speech activity detection
- A polynomial optimization approach for robust beamforming design in a device-to-device two-hop one-way relay network
- A practical clock synchronization algorithm for UWB positioning systems
- A precision-improved processing architecture of physical computing for energy-efficient SIFT feature extraction
- A rate-splitting approach to robust multiuser MISO transmission
- A real-time example-based single-image super-resolution algorithm via cross-scale high-frequency components self-learning
- A recursive predictive risk estimate for proximal algorithms
- A risk-unbiased approach to a new Cramér-Rao bound
- A robust Gaussian approximate filter for nonlinear systems with heavy tailed measurement noises
- A robust speech rate estimation based on the activation profile from the selected acoustic unit dictionary
- A score-informed shift-invariant extension of complex matrix factorization for improving the separation of overlapped partials in music recordings
- A segment-sliding reconstruction scheme for pulsed radar echoes with sub-Nyquist sampling
- A semi-global matching method for large-scale light field images
- A semidefinite relaxation approach to the geolocation of two unknown co-channel emitters by a cluster of formation-flying satellites using both TDOA and FDOA measurements
- A short-graph fourier transform via personalized pagerank vectors
- A single-channel noise cancelation filter in the short-time-fourier-transform domain
- A source/filter model with adaptive constraints for NMF-based speech separation
- A sparse regression based approach for cuff-less blood pressure measurement
- A sparse-graph-coded filter bank approach to minimum-rate spectrum-blind sampling
- A speaker adaptation technique for Gaussian process regression based speech synthesis using feature space transform
- A speech enhancement system using binaural hearing aids and an external microphone
- A study of different weighting schemes for spoken language understanding based on convolutional neural networks
- A study of rank-constrained multilingual DNNS for low-resource ASR
- A subjective listening test of six different artificial bandwidth extension approaches in English, Chinese, German, and Korean
- A swiss army knife for finite rate of innovation sampling theory
- A thin-slice perception of emotion? An information theoretic-based framework to identify locally emotion-rich behavior segments for global affect recognition
- A topography structure used in audio steganography
- A transfer learning method for PLDA-based speaker verification
- A unified approach to the design of IIR and FIR notch filters
- A unified framework for atlas-based segmentation with forward deformation and label refinement
- A unified sparse signal decomposition and reconstruction framework for elimination of muscle artifacts from ECG signal
- A weakly-supervised discriminative model for audio-to-score alignment
- A weighted STOI intelligibility metric based on mutual information
- A weighted atomic norm approach to spectral super-resolution with probabilistic priors
- AAC encoding detection and bitrate estimation using a convolutional neural network
- ALADDIN: A locality aligned deep model for instance search
- Abnormal event detection based on sparse reconstruction in crowded scenes
- Abnormal sound event detection using temporal trajectories mixtures
- Accelerated spectral clustering using graph filtering of random signals
- Accelerating multi-user large vocabulary continuous speech recognition on heterogeneous CPU-GPU platforms
- Accelerating stochastic computation for binary classification applications
- Accurate asymptotic analysis for John's test in multichannel signal detection
- Accurate recovery of a specularity from a few samples of the reflectance function
- Achieving global optimality for wirelessly-powered multi-antenna TWRC with lattice codes
- Acoustic data-driven pronunciation lexicon generation for logographic languages
- Acoustic event detection based on non-negative matrix factorization with mixtures of local dictionaries and activation aggregation
- Acoustic scene classification with matrix factorization for unsupervised feature learning
- Acoustic simultaneous localization and mapping (A-SLAM) of a moving microphone array and its surrounding speakers
- Acoustic source separation using the short-time quaternion fourier transforms of particle velocity signals
- Action recognition using interest points capturing differential motion information
- Active eavesdropping via spoofing relay attack
- Active learning for magnetic resonance image quality assessment
- Active learning on weighted graphs using adaptive and non-adaptive approaches
- Active online learning of trusts in social networks
- Adapting ASR for under-resourced languages using mismatched transcriptions
- Adaptive Boolean compressive sensing by using multi-armed bandit
- Adaptive algorithms for hypergraph learning
- Adaptive consensus-based distributed detection in WSN with unreliable links
- Adaptive distributed compressed estimation based on recursive least squares with sensing matrix design
- Adaptive enhancement of luminance and details in images under ambient light
- Adaptive extraction of repeating non-negative temporal patterns for single-channel speech enhancement
- Adaptive learning for stochastic generalized Nash equilibrium problems
- Adaptive margin slack minimization in RKHS for classification
- Adaptive radar detection in the presence of Gaussian clutter with symmetric spectrum
- Adaptive rate control algorithm for SHVC: Application to HD/UHD
- Adaptive regularization for BEM channel estimation in multicarrier systems
- Adaptive reverberation cancelation for multizone soundfield reproduction using sparse methods
- Adaptive sequential optimization with applications to machine learning
- Adaptive sparsity tradeoff for ℓ1-constraint NLMS algorithm
- Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent network
- Advanced b-vector system based deep neural network as classifier for speaker verification
- Adversarial Bandit for online interactive active learning of zero-shot spoken language understanding
- Agreement and disagreement classification of dyadic interactions using vocal and gestural cues
- Algebraic solution for stationary emitter geolocation by a LEO satellite using Doppler frequency measurements
- Algorithm for DNA copy number variation detection with read depth and paramorphism information
- An acoustic keystroke transient canceler for speech communication terminals using a semi-blind adaptive filter model
- An adaptive fixed-point IVA algorithm applied to multi-subject complex-valued FMRI data
- An adaptive multi-level wavelet denoising method for 40-Hz ASSR
- An adaptive resolution rate control method for intra coding in HEVC
- An adaptive robust regression method: Application to galaxy spectrum baseline estimation
- An alternative approach for auditory attention tracking using single-trial EEG
- An alternative proof for the identifiability of independent vector analysis using second order statistics
- An approximate message passing approach for tensor-based seismic data interpolation with randomly missing traces
- An effective color space for face recognition
- An effective performance ranking mechanism to image dehazing methods with psychological inference benchmark
- An efficient anomaly detection approach in surveillance video based on oriented GMM
- An efficient method for polyphonic audio-to-score alignment using onset detection and constant Q transform
- An empirical exploration of CTC acoustic models
- An energy-aware auction for hybrid access in heterogeneous networks under QoS requirements
- An energy-efficient compressive sensing framework incorporating online dictionary learning for long-term wireless health monitoring
- An engineer's guide to particle filtering on matrix Lie groups
- An estimation method of voice timbre evaluation values using feature extraction with Gaussian mixture model based on reference singer
- An expectation-maximization eigenvector clustering approach to direction of arrival estimation of multiple speech sources
- An extensible speaker identification sidekit in Python
- An image smoothing operator for fast and accurate scale space approximation
- An improved DOA estimation algorithm for circular and non-circular signals with high resolution
- An improved anthropometry-based customization method of individual head-related transfer functions
- An improved local binary pattern operator for texture classification
- An information theoretic framework for order of operations forensics
- An introduction to hypergraph signal processing
- An inverse-gamma source variance prior with factorized parameterization for audio source separation
- An investigation into using parallel data for far-field speech recognition
- An iterative hard thresholding approach to ℓ0 sparse Hellinger NMF
- An iterative sure-let approach to sparse reconstruction
- An iteratively reweighted method for recovery of block-sparse signal with unknown block partition
- An online algorithm for throughput maximization of wireless powered communication networks
- An online tensor robust PCA algorithm for sequential 2D data
- An optimization framework for combining multiple graphs
- An unbiased risk estimator for Gaussian mixture noise distributions - Application to speech denoising
- Analog multiple descriptions: A zero-delay source-channel coding approach
- AnalogCast: Full linear coding and pseudo analog transmission for satellite remote-sensing images
- Analysis of DNN approaches to speaker identification
- Analysis of distributed ADMM algorithm for consensus optimization in presence of error
- Analysis of error resiliency of belief propagation in computer vision
- Analysis of natural and synthetic speech using Fujisaki model
- Analysis of p-norm regularized subproblem minimization for sparse photon-limited image recovery
- Analysis of secure communication in millimeter wave networks: Are blockages beneficial?
- Analytical performance assessment of esprit-type algorithms for coexisting circular and strictly non-circular signals
- Annealed learning based block transforms for HEVC video coding
- Anti-occlusion observation model and automatic recovery for multi-view ball tracking in sports analysis
- Applications of 3D spherical transforms to personalization of head-related transfer functions
- Approximate search of audio queries by using DTW with phone time boundary and data augmentation
- Are there approximate fast fourier transforms on graphs?
- Array thinning for antenna selection in millimeter wave MIMO systems
- Asking for a second opinion: Re-querying of noisy multi-class labels
- Aspect Ratio Similarity (ARS) for image retargeting quality assessment
- Asymptotic analysis of downlink MISO systems over Rician fading channels
- Asymptotic closed-loop design of error resilient predictive compression systems
- Asymptotic optimal quantizer design for distributed Bayesian estimation
- Asymptotic perfect secrecy in distributed detection against a global eavesdropper
- Asymptotic performance analysis for 1-bit Bayesian smoothing
- Asynchronous distributed alternating direction method of multipliers: Algorithm and convergence analysis
- Asynchronous local voltage control in power distribution networks
- Asynchronous systems for constraint satisfaction: Filtering and stability
- Atmospheric turbulence mitigation based on turbulence extraction
- Audio enhancing with DNN autoencoder for speaker recognition
- Audio watermarking based on empirical mode decomposition and beat detection
- Audio word similarity for clustering with zero resources based on iterative HMM classification
- Audio-based multimedia event detection using deep recurrent neural networks
- Auditory attention decoding with EEG recordings using noisy acoustic reference signals
- Autocalibration of lidar and optical cameras via edge alignment
- Automatic Chord estimation on seventhsbass Chord vocabulary using deep neural network
- Automatic allocation of NTF components for user-guided audio source separation
- Automatic composition of broadcast news summaries using rank classifiers trained with acoustic and lexical features
- Automatic gain control for parametric array loudspeakers
- Automatic human fall detection in fractional fourier domain for assisted living
- Automatic image region annotation through segmentation based visual semantic analysis and discriminative classification
- Automatic speech recognition for acoustical analysis and assessment of cantonese pathological voice and speech
- Auxiliary beam pair design in mmWave cellular systems with hybrid precoding and limited feedback
- BFGUI: An interactive tool for the synthesis and analysis of microphone array beamformers
- BIAS correction methods for adaptive recursive smoothing with applications in noise PSD estimation
- Bagging regularized common spatial pattern with hybrid motor imagery and myoelectric signal
- Bandlimited field reconstruction from samples obtained on a discrete grid with unknown random locations
- Basis compensation in non-negative matrix factorization model for speech enhancement
- Batch normalized recurrent neural networks
- Bayesian quickest detection with unknown post-change parameter
- Bayesian tuning for support detection and sparse signal estimation via iterative shrinkage-thresholding
- Benchmarking of scoring functions for bias-based fingerprinting code
- Benchmarking state-of-the-art visual saliency models for image quality assessment
- Ber analysis of the box relaxation for BPSK signal recovery
- Bernstein filter: A new solver for mean curvature regularized models
- Better acoustic normalization in subject independent acoustic-to-articulatory inversion: Benefit to recognition
- Beyond L2-loss functions for learning sparse models
- Beyond low rank + sparse: Multi-scale low rank matrix decomposition
- Beyond union of subspaces: Subspace pursuit on Grassmann manifold for data representation
- Bi-directional recurrent neural network with ranking loss for spoken language understanding
- Binary code learning with semantic ranking based supervision
- Binaural sound generation corresponding to omnidirectional video view using angular region-wise source enhancement
- Binaural speaker localization and separation based on a joint ITD/ILD model and head movement tracking
- Bipartite subgraph decomposition for critically sampled wavelet filterbanks on arbitrary graphs
- Bird species recognition using HMM-based unsupervised modelling of individual syllables with incorporated duration modelling
- Bit-depth expansion for noisy contour reduction in natural images
- Blind CFO estimation for multiuser OFDM uplink with large number of receive antennas
- Blind channel estimation in OFDM-based amplify-and-forward two-way relay networks
- Blind deconvolution of sparse but filtered pulses with linear state space models
- Blind estimation of unknown time delay in periodic non-uniform sampling: Application to desynchronized time interleaved-ADCs
- Blind identification of graph filters with multiple sparse inputs
- Blind image quality assessment for multiply distorted images via convolutional neural networks
- Blind mobile sensor calibration using an informed nonnegative matrix factorization with a relaxed rendezvous model
- Blind polychromatic X-ray CT reconstruction from poisson measurements
- Blind separation of underdetermined linear mixtures based on source nonstationarity and AR(1) modeling
- Blind speech separation based on complex spherical k-mode clustering
- Blind sub-Nyquist GNSS signal detection
- Block compressed sensing based distributed resource allocation for M2M communications
- Boosted classification of breast cancer by retrieval of cases having similar disease likelihood
- Boosted multi-scale dictionaries for image compression
- Boosting objectness: Semi-supervised learning for object detection and segmentation in multi-view images
- Bottleneck capacity of random graphs for connectomics
- Bottleneck linear transformation network adaptation for speaker adaptive training-based hybrid DNN-HMM speech recognizer
- Buffer aided distributed space time coding techniques for cooperative DS-CDMA systems
- CNMF-based acoustic features for noise-robust ASR
- CS based processing for high resolution GSM passive bistatic radar
- CS-based device-free localization in the presence of model errors
- CUED-RNNLM - An open-source toolkit for efficient training and evaluation of recurrent neural network language models
- Calibration of the attenuation-rain rate power-law parameters using measurements from commercial microwave networks
- Camera based estimation of respiration rate by analyzing shape and size variation of structured light
- Capacity analysis of WCC-FBMC/OQAM systems
- Capacity maximization for distributed broadband beamforming
- Carrier frequency and bandwidth estimation of cyclostationary multiband signals
- Channel gain prediction for multi-agent networks in the presence of location uncertainty
- Channel learning in indoor localization
- Character proposal network for robust text extraction
- Character-level incremental speech recognition with recurrent neural networks
- Choosing the diagonal loading factor for linear signal estimation using cross validation
- Chroma scaling for high dynamic range video compression
- Chute based automated fish length measurement and water drop detection
- Classification of bisyllabic lexical stress patterns in disordered speech using deep learning
- Classification of breath and snore sounds using audio data recorded with smartphones in the home environment
- Classification of head movement patterns to aid patients undergoing home-based cervical spine rehabilitation
- Classification of human cough signals using spectro-temporal Gabor filterbank features
- Classification of hyperspectral data with ensemble of subspace ICA and edge-preserving filtering
- Classification of medical images using edge-based features and sparse representation
- Classification of respiratory effort and disordered breathing during sleep from audio and pulse oximetry signals
- Classification of voices that elicit soothing effect by applying a voiced vs. unvoiced feature engineering strategy
- Cluster-based dictionary learning and locality-constrained sparse reconstruction for trajectory classification
- Clustering of interictal spikes by dynamic time warping and affinity propagation
- Co-segmentation of multiple images through random walk on graphs
- Codebook enhancement of vlad representation for visual recognition
- Coded excitation ultrasound: Efficient implementation via frequency domain processing
- Coherence regularized dictionary learning
- Column-wise symmetric block partitioned tensor decomposition
- Combining dirty-paper coding and artificial noise for secrecy
- Combining i-vector representation and structured neural networks for rapid adaptation
- Combining multiple kernel models for automatic intelligibility detection of pathological speech
- Combining non-negative matrix factorization and deep neural networks for speech enhancement and automatic speech recognition
- Combining soft decisions of several unreliable experts
- Comix: Joint estimation and lightspeed comparison of mixture models
- Common fate model for unison source separation
- Communication-efficient weighted ADMM for decentralized network optimization
- Community detection game
- Commuting operator of offset linear canonical transform and its applications
- Compact convolutional neural network transfer learning for small-scale image classification
- Compact kernel models for acoustic modeling via random feature selection
- Comparison of different development kits and its suitability in signal processing education
- Comparison of statistical algorithms for power system line outage detection
- Comparison of unsupervised sequence adaptations for deep neural networks
- Compensation of attacks on consensus networks
- Completion of structurally-incomplete matrices with reweighted low-rank and sparsity priors
- Complex NMF under phase constraints based on signal modeling: Application to audio source separation
- Complex ratio masking for joint enhancement of magnitude and phase
- Complexity reduction of SUMIS MIMO soft detection based on box optimization for large systems
- Compressed training adaptive equalization
- Compression and reconstruction methodology for neural signals based on patch ordering inpainting for brain monitoring
- Compression of dynamic 3D point clouds using subdivisional meshes and graph wavelet transforms
- Compressive sensing based target counting and localization exploiting joint sparsity
- Computational agile beam ladar imaging
- Computationally efficient estimation of multi-dimensional spectral lines
- Computed tomography reconstruction based on a hierarchical model and variational Bayesian method
- Conditional MMSE-based single-channel speech enhancement using inter-frame and inter-band correlations
- Confidence assessment for spectral estimation based on estimated covariances
- Connectivity for overlaid wireless networks with outage constraints
- Consensus inference on mobile phone sensors for activity recognition
- Constructive interference exploitation for downlink beamforming based on noise robustness and outage probability
- Content-aware local variability vector for speaker verification with short utterance
- Context adaptive deep neural networks for fast acoustic model adaptation in noisy conditions
- Context adaptive thresholding and entropy coding for very low complexity JPEG transcoding
- Context-dependent point process models for keyword search and detection-based ASR
- Continuous ultrasound based tongue movement video synthesis from speech
- Contour-based 3D tongue motion visualization using ultrasound image sequences
- Convergence analysis for Guassian belief propagation: Dynamic behaviour of marginal covariances
- Convergence-optimized variable node structure for stochastic LDPC decoder
- Convolutional neural network for robust pitch determination
- Convolutional neural network pre-trained with projection matrices on linear discriminant analysis
- Cooperative joint synchronization and localization using time delay measurements
- Cooperative localization based on severely quantized RSS measurements in wireless sensor network
- Coordinated uplink scheduling and beamforming for wireless cellular networks via sum-of-ratio programming and matching
- Coprime array adaptive beamforming based on compressive sensing virtual array signal
- Correlation-statistics-based simulator of perturbed phases triggered by the ionospheric irregularities for HF radar systems
- Coupled dictionary learning for multimodal data: An application to concurrent intracranial and scalp EEG
- Coupled rank-(Lm, Ln, •) block term decomposition by coupled block simultaneous generalized Schur decomposition
- Cross lingual speech emotion recognition using canonical correlation analysis on principal component subspace
- Cross-acoustic transfer learning for sound event classification
- Cross-corpus acoustic emotion recognition from singing and speaking: A multi-task learning approach
- Cross-lingual deep neural network based submodular unbiased data selection for low-resource keyword search
- Cute: A concatenative method for voice conversion using exemplar-based unit selection
- Cyclostationary-based detection of steady-state visually evoked potential signals recorded from EEG
- D-FW: Communication efficient distributed algorithms for high-dimensional sparse optimization
- D3M: Distributed multi-cell multigroup multicasting
- DCT based region log-tiedrank covariance matrices for face recognition
- DEMV-matchmaker: Emotional temporal course representation and deep similarity matching for automatic music video generation
- DNN speaker adaptation using parameterised sigmoid and ReLU hidden activation functions
- DNN-based enhancement of noisy and reverberant speech
- DOA estimation of audio sources in reverberant environments
- DOA estimation of closely-spaced and spectrally-overlapped sources based on time-frequency sparse representation
- DTM: Deformable template matching
- Data selection for noise robust exemplar matching
- Data selection from multiple ASR systems' hypotheses for unsupervised acoustic model training
- Data sketching for large-scale Kalman filtering
- Data-guided random walks for fine-structured object segmentation
- Data-weighted ensemble learning for privacy-preserving distributed learning
- Dealing with uncertain models in wireless communications
- Decentralized coordination of energy resources in electricity distribution networks
- Decoding visemes: Improving machine lip-reading
- Decreasing the measurement time of blood sugar tests using particle filtering
- Deep beamforming networks for multi-channel speech recognition
- Deep belief network-based post-filtering for statistical parametric speech synthesis
- Deep clustering: Discriminative embeddings for segmentation and separation
- Deep complementary bottleneck features for visual speech recognition
- Deep convolutional acoustic word embeddings using word-pair side information
- Deep discriminative manifold learning
- Deep kernel map networks for image annotation
- Deep multi-view representation learning for multi-modal features of the schizophrenia and schizo-affective disorder
- Deep neural network based posteriors for text-dependent speaker verification
- Deep neural network-guided unit selection synthesis
- Deep neural networks for automatic detection of screams and shouted speech in subway trains
- Deep unfolding for multichannel source separation
- Deep unfolding inference for supervised topic model
- Degradedness and stochastic orders of fast fading Gaussian broadcast channels with statistical channel state information at the transmitter
- Delay estimation between EEG and EMG via coherence with time lag
- Delay-Doppler estimation via structured low-rank matrix recovery
- Depth estimation from single images using modified stacked generalization
- Depth fused from intensity range and blur estimation for light-field cameras
- Depth guided image completion for structure and texture synthesis
- Depth map coding based on virtual view quality
- Depth map estimation using census transform for light field cameras
- Depth propagation in 2D-to-3D conversion based on frame clustering
- Depth propagation with tensor voting for 2D-to-3D video conversion
- Depth-aware saliency detection using discriminative saliency fusion
- Design space exploration for hardware-efficient stochastic computing: A case study on discrete cosine transformation
- Detectability prediction of hidden Markov models with cluttered observation sequences
- Detecting double MPEG compression with the same quantiser scale based on MBM feature
- Detecting occlusion from color information to improve visual tracking
- Detecting the instant of emotion change from speech using a martingale framework
- Detection of cyclostationarity in the presence of temporal or spatial structure with applications to cognitive radio
- Detection of drops measured by the time shift technique for spray characterization
- Detection of faint extended sources in hyperspectral data and application to HDF-S MUSE observations
- Detection of overlapping acoustic events using a temporally-constrained probabilistic model
- Detection of pilot contamination attack in T.D.D./S.D.M.A. systems
- Detection of the number of superimposed signals using modified MDL criterion: A random matrix approach
- Detection with phaseless measurements
- Deterministic maximum likelihood method for direction-of-arrival estimation of strictly noncircular signals
- Dictionary learning for Poisson compressed sensing
- Dictionary learning from phaseless measurements
- Diffusion LMS over multitask networks with noisy links
- Diffusion filtering of graph signals and its use in recommendation systems
- Diffusion social learning over weakly-connected graphs
- Diffusion stochastic optimization with non-smooth regularizers
- Diffusive particle filtering for distributed multisensor estimation
- Direction of arrival estimation based on information geometry
- Direction of arrival estimation in MIMO radar systems with nonlinear reflectors
- Direction-of-arrival estimation based on Toeplitz covariance matrix reconstruction
- Direction-of-arrival estimation with espar antennas using Bayesian compressive sensing
- Directional maximum likelihood self-estimation of the path-loss exponent
- Directly modeling voiced and unvoiced components in speech waveforms by neural networks
- Dirichlet process mixture models for time-dependent clustering
- Discontinuous operation for precoded G.fast
- Discourse connective detection in spoken conversations
- Discovering rāga motifs by characterizing communities in networks of melodic patterns
- Discriminant correlation analysis for feature level fusion with application to multimodal biometrics
- Discriminative deep recurrent neural networks for monaural speech separation
- Discriminative feature extraction from X-ray images using deep convolutional neural networks
- Discriminative multi-domain PLDA for speaker verification
- Discriminatively learned filter bank for acoustic features
- Discriminatively trained joint speaker and environment representations for adaptation of deep neural network acoustic models
- Distances between directed networks and applications
- Distributed LMS estimation of scaled and delayed impulse responses
- Distributed MIMO systems: Receiver design and ML detection
- Distributed beamforming in relay networks for energy harvesting multi-group multicast systems
- Distributed beamforming using mobile robots
- Distributed dyadic cyclic descent for non-negative matrix factorization
- Distributed estimation of latent parameters in state space models using separable likelihoods
- Distributed estimation via paid crowd work
- Distributed generalized likelihood ratio tests: Fundamental limits and tradeoffs
- Distributed linear blind source separation over wireless sensor networks with arbitrary connectivity patterns
- Distributed multi-sensor CPHD filter using pairwise gossiping
- Distributed nonconvex optimization over time-varying networks
- Distributed path optimization of multiple UAVs for AOA target localization
- Distributed sparse MVDR beamforming using the bi-alternating direction method of multipliers
- Distributional semantics for understanding spoken meal descriptions
- Distributionally robust chance-constrained minimum variance beamforming
- Divergence estimation based on deep neural networks and its use for language identification
- Document level semantic context for retrieving OOV proper names
- Domain adaptation for speech emotion recognition by sharing priors between related source and target classes
- Domain adaptation using maximum likelihood linear transformation for PLDA-based speaker verification
- Downlink SINR balancing in C-RAN under limited fronthaul capacity
- Dropped pronoun generation for dialogue machine translation
- Dual-microphone voice activity detection based on using optimally weighted maximum a posteriori probabilities
- Dual-stage algorithm to identify channels with poor electrode-to-neuron interface in cochlear implant users
- Dynamic analysis of resting state fMRI data and its applications
- Dynamic relative impulse response estimation using structured sparse Bayesian learning
- EchoSLAM: Simultaneous localization and mapping with acoustic echoes
- Effective utilization of multiple examples in query-by-example spoken term detection
- Effectiveness of fundamental frequency (F0) and strength of excitation (SOE) for spoofed speech detection
- Efficient algorithms for linear polyhedral bandits
- Efficient channel statistics estimation for millimeter-wave MIMO systems
- Efficient deblocking filter implementation on reconfigurable processor
- Efficient estimation of inter-subband speech correlations
- Efficient keypoint detection and description via polynomial regression of scale space
- Efficient near optimal joint modulation classification and detection for MU-MIMO systems
- Efficient neighborhood-based topic modeling for collaborative audio enhancement on massive crowdsourced recordings
- Efficient non-linear feature adaptation using Maxout networks
- Efficient object feature selection for action recognition
- Efficient one-vs-one kernel ridge regression for speech recognition
- Efficient parameter inference in general hidden Markov models using the filter derivatives
- Efficient sensor position selection using graph signal sampling theory
- Efficient stochastic detector for large-scale MIMO
- Efficient subspace detection for high-order MIMO systems
- Efficient target-response interpolation for a graphic equalizer
- Egocentric activity recognition with multimodal fisher vector
- Eigen and multimodal analysis for localizing moving sounding objects
- Emotion classification: How does an automated system compare to Naive human coders?
- Emotion recognition from peripheral physiological signals enhanced by EEG
- Emotion-flow guided music accompaniment generation
- Empirically-estimable multi-class classification bounds
- End-to-end attention-based large vocabulary speech recognition
- End-to-end text-dependent speaker verification
- Energy detection in ISI channels using large-scale receiver arrays
- Energy efficient beamforming for secure communication in cognitive radio networks
- Energy-efficient pilot and data power allocation in massive MIMO communication systems based on MMSE channel estimation
- Enhanced HEVC intra prediction with ordered dither technique
- Enhanced just noticeable difference model with visual regularity consideration
- Enhanced semi-supervised learning for multimodal emotion recognition
- Enhanced vote count circuit based on nor flash memory for fast similarity search
- Epileptiform spike detection via convolutional neural networks
- Equalization matching of speech recordings in real-world environments
- Error performance analysis of the symbol-decision SC polar decoder
- Error-resilient sequential cells with successive time borrowing for stochastic computing
- Estimating direct-to-reverberant ratio mapped from power spectral density using deep neural network
- Estimating ear canal geometry and eardrum reflection coefficient from ear canal input impedance
- Estimating high-dimensional covariance matrices with misses for Kronecker product expansion models
- Estimating orientation in tracking individuals of flying swarms
- Estimating parameters in noisy low frequency exponentially damped sinusoids and exponentials
- Estimation efficiency, accuracy and robustness improvement by exploiting the geometry information in SAR-GMTI system
- Estimation of TDOA for room reflections by iterative weighted l1 constraint
- Estimation of the reliability of multiple rhythm features extraction from a single descriptor
- Eulerian emotion magnification for subtle expression recognition
- Evaluating instrumental measures of speech quality using Bayesian model selection: Correlations can be misleading!
- Evaluation of estimated hammerstein models via normalized projection misalignment of linear and nonlinear subsystems
- Exemplar-based sparse representation of timbre and prosody for voice conversion
- Exemplar-inspired strategies for low-resource spoken keyword search in Swahili
- Experimental study of generalized subspace filters for the cocktail party situation
- Experimental validation of TOA-based methods for microphones array positions calibration
- Exploiting LSTM structure in deep neural networks for speech recognition
- Exploiting correlations among channels in distributed compressive sensing with convolutional deep stacking networks
- Exploiting low-dimensional structures to enhance DNN based acoustic modeling in speech recognition
- Exploiting sparsity for image-based object surface anomaly detection
- Exploiting spectro-temporal structures using NMF for DNN-based supervised speech separation
- Exploratory analysis of speech features related to depression in adults with Aphasia
- Exploring articulatory characteristics of Cantonese dysarthric speech using distinctive features
- Exploring deep learning architectures for automatically grading non-native spontaneous speech
- Exploring multidimensional lstms for large vocabulary ASR
- Exploring persistent local homology in topological data analysis
- Exploring the role of phonetic bottleneck features for speaker and language recognition
- Extension of SeDJoCo and its use in a combination of multicast and coordinated multi-point systems
- Extension of nested arrays with the fourth-order difference co-array enhancement
- Extension of the semi-algebraic framework for approximate CP decompositions via non-symmetric simultaneous matrix diagonalization
- Extensions of semidefinite programming methods for atomic decomposition
- Extensions of the binaural MWF with interference reduction preserving the binaural cues of the interfering source
- Extraction of tongue contour in real-time magnetic resonance imaging sequences
- F0 estimation for noisy speech by exploring temporal harmonic structures in local time frequency spectrum segment
- FPGA based implementation of deep neural networks using on-chip memory only
- Face alignment by deep convolutional network with adaptive learning rate
- Face hallucination via locality-constrained low-rank representation
- Face liveness detection and recognition using shearlet based feature descriptors
- Face recognition with local contourlet combined patterns
- Factored spatial and spectral multichannel raw waveform CLDNNs
- Fall detection in RGB-D videos by combining shape and motion features
- Fast adaptive PARAFAC decomposition algorithm with linear complexity
- Fast alternating projected gradient descent algorithms for recovering spectrally sparse signals
- Fast and easy crowdsourced perceptual audio evaluation
- Fast and efficient rejection of background waveforms in interictal EEG
- Fast and statistically efficient fundamental frequency estimation
- Fast anomaly detection in traffic surveillance video based on robust sparse optical flow
- Fast continuous HRTF acquisition with unconstrained movements of human subjects
- Fast depth image denoising and enhancement using a deep convolutional network
- Fast dynamic MRI using linear dynamical system model
- Fast intra mode decision and block matching for HEVC screen content compression
- Fast keypoint detection in video sequences
- Fast lossless compression of whole slide pathology images using HEVC intra-prediction
- Fast online orthonormal dictionary learning for efficient full waveform inversion
- Fast response aggregation for depth estimation using light field camera
- Fast sparse 2-D DFT computation using sparse-graph alias codes
- Fast variational Bayesian signal recovery in the presence of Poisson-Gaussian noise
- Fast voxel line update for time-space image reconstruction
- Feature adapted convolutional neural networks for downbeat tracking
- Feature mapping, score-, and feature-level fusion for improved normal and whispered speech speaker verification
- Feature-enriched word embeddings for named entity recognition in open-domain conversations
- Feedback of differential precoder for geometrical mean decomposition systems
- Fetal heart rate analysis by hierarchical dirichlet process mixture models
- Filterbank learning using Convolutional Restricted Boltzmann Machine for speech recognition
- Finding the minimum rate of innovation in the presence of noise
- Finding unique dense communities
- Fine-structured object segmentation via edge-guided graph cut with interaction simplification
- Fingerprint recognition with ridge features and minutiae on distortion
- Finite-state channel models for signal transduction in neural systems
- First order echo based room shape recovery using a single mobile device
- Fixed-complexity variants of the effective LLL algorithm with greedy convergence for MIMO detection
- Fixed-point performance analysis of recurrent neural networks
- Flat start training of CD-CTC-SMBR LSTM RNN acoustic models
- Flexibeam: Analytic spatial filtering by beamforming
- Formant shifting for speech intelligibility improvement in car noise environment
- Framewise speech-nonspeech classification by neural networks for voice activity detection with statistical noise suppression
- Frank-Wolfe works for non-Lipschitz continuous gradient objectives: Scalable poisson phase retrieval
- Frequency recognition of steady-state visually evoked potentials using binary subband canonical correlation analysis with reduced dimension of reference signals
- Frequency-based customization of multizone sound system design
- From HMMS to DNNS: Where do the improvements come from?
- From acoustic room reconstruction to slam
- Functional connectivity brain network analysis through network to signal transform based on the resistance distance
- Further results on mainlobe orientation reversal of the first-order steerable differential array due to microphone phase errors
- Fusion of algorithms for multiple measurement vectors
- Fusion of depth, skeleton, and inertial data for human action recognition
- Fuzzy entropy based nonnegative matrix factorization for muscle synergy extraction
- Gain relaxation: A useful technique for signal enhancement with an unaware local noise source targeted at speech recognition
- Gammatone filter based on stochastic computation
- Gating recurrent mixture density networks for acoustic modeling in statistical parametric speech synthesis
- Gauss-Seidel based non-negative matrix factorization for gene expression clustering
- Generalized Laplacian precision matrix estimation for graph signal processing
- Generalized coprime sampling of Toeplitz matrices
- Generalized k-level cutset sampling and reconstruction
- Generalized wave-domain transforms for listening room equalization with azimuthally irregularly spaced loudspeaker arrays
- Generating a morphable model of ears
- Geo-location dependent deep neural network acoustic model for speech recognition
- Geodesic-based pavement shadow removal revisited
- Geometric-guided label propagation for moving object detection
- Geometrical room geometry estimation from room impulse responses
- Geometry and radiometry invariant matched manifold detection and tracking
- Ghosting-free multi-exposure image fusion in gradient domain
- Globally optimized least-squares post-filtering for microphone array speech enhancement
- Gold classification of COPDGene cohort based on deep learning
- Grab-n-Pull: An optimization framework for fairness-achieving networks
- Gradient schemes for robust FFT-based motion estimation
- Gram Schmidt based greedy hybrid precoding for frequency selective millimeter wave MIMO systems
- Graph filter banks with M-channels, maximal decimation, and perfect reconstruction
- Graph signal recovery from incomplete and noisy information using approximate message passing
- Graph-based lifting transform for intra-predicted video coding
- Graph-based representation and coding of 3D images for interactive multiview navigation
- Group diffusion LMS
- Group nonnegative matrix factorisation with speaker and session variability compensation for speaker identification
- Group sparse Bayesian learning via exact and fast marginal likelihood maximization
- Group-blind detection with very large antenna arrays in the presence of pilot contamination
- Groupwise learning for ASR k-best list reranking in spoken language translation
- Hardware implementation of FIR/IIR digital filters using integral stochastic computation
- Harmonic-percussive-residual sound separation using the structure tensor on spectrograms
- Heart-trend: An affordable heart condition monitoring system exploiting morphological pattern
- Heterogeneous domain adaptation with label and structure consistency
- High accuracy indoor localization: A WiFi-based approach
- High diagnostic quality ECG compression and CS signal reconstruction in body sensor networks
- High dynamic range imaging via truncated nuclear norm minimization of low-rank matrix
- High-pitched excitation generation for glottal vocoding in statistical parametric speech synthesis using a deep neural network
- High-resolution sinusoidal modeling of unvoiced speech
- Higher-order listening room compensation with additive compensation signals
- Highway long short-term memory RNNS for distant speech recognition
- Histogram feature deblurring
- Honey chatting: A novel instant messaging system robust to eavesdropping over communication
- How neural network features and depth modify statistical properties of HMM acoustic models
- Hybrid beamforming with two bit RF phase shifters in single group multicasting
- Hybrid music recommender using content-based and social information
- Hyperspectral image classification using set-to-set distance
- IMISOUND: An unsupervised system for sound query by vocal imitation
- IVA for abandoned object detection: Exploiting dependence across color channels
- Identity association using PHD filters in multiple head tracking with depth sensors
- Image colorization based on ADMM with fast singular value thresholding by Chebyshev polynomial approximation
- Image phylogeny tree reconstruction based on region selection
- Image restoration using a stochastic variant of the alternating direction method of multipliers
- Image sentiment analysis using latent correlations among visual, textual, and sentiment views
- Image-assisted geometry simplification for the plenoptic sampling
- Imaging in radio interferometry by iterative subset scanning using a modified AMP algorithm
- Impact of channel access issues and packet losses on distributed outlier detection within wireless sensor networks
- Implementation of F0 transformation for statistical singing voice conversion based on direct waveform modification
- Implementation of the precoder matrix indicator selection using MMSE trace criterion for the downlink transmission in LTE
- Implicit kernel presentation aware object segmentation framework
- Importance sampling of delta-AUC: A basis for active learning for improved keyword search
- Improved DNN-based segmentation for multi-genre broadcast audio
- Improved decoding of analog modulo block codes for noise mitigation
- Improved forgery detection with lateral chromatic aberration
- Improved illumination invariant homomorphic filtering using the dual tree complex wavelet transform
- Improved multi-microphone noise reduction preserving binaural cues
- Improved set-membership partial-update affine projection algorithm
- Improved speaker independent lip reading using speaker adaptive training and deep neural networks
- Improved spoken document summarization with coverage modeling techniques
- Improving SHVC performance with a joint layer coding mode
- Improving adaptive feedback cancellation in hearing aids using an affine combination of filters
- Improving face detection with depth
- Improving non-native mispronunciation detection and enriching diagnostic feedback with DNN-based speech attribute modeling
- Improving resolution in supervised patch-based target detection
- Improving semantic video indexing: Efforts in Waseda TRECVID 2015 SIN system
- Improving speech privacy in personal sound zones
- Incorporating relative transfer function preservation into the binaural multi-channel wiener filter for hearing aids
- Independent versus repeated measurements: A performance quantification via state evolution
- Inferring depolarization of cells from 3D-electrode measurements using a bank of linear state space models
- Information fusion based on kernel entropy component analysis in discriminative canonical correlation space with application to audio emotion recognition
- Information point set registration for shape recognition
- Information theoretic clustering for unsupervised domain-adaptation
- Information theoretic multivariate change detection for multisensory information processing in Internet of Things
- Informed Direction of Arrival estimation using a spherical-head model for Hearing Aid applications
- Infrared small target detection with compressive measurements
- Initial investigation of speech synthesis based on complex-valued neural networks
- Insight into a phase modulation technique for signal decorrelation in multi-channel acoustic echo cancellation
- Instantaneous pitch estimation algorithm based on multirate sampling
- Integrated adaptation with multi-factor joint-learning for far-field speech recognition
- Integrated approach of feature extraction and sound source enhancement based on maximization of mutual information
- Integration of machine learning and human learning for training optimization in robust linear regression
- Integration of orthogonal feature detectors in parameter learning of artificial neural networks to improve robustness and the evaluation on hand-written digit recognition tasks
- Intelligible enhancement of 3D articulation animation by incorporating airflow information
- Intensity-only optical compressive imaging using a multiply scattering material and a double phase retrieval approach
- Inter-speaker variability in forensic voice comparison: A preliminary evaluation
- Interlaced sigma-point information filtering for distributed state estimation of multi-agent systems
- Interpreting the prediction process of a deep network constructed from supervised topic models
- Intrinsic two-dimensional local structures for micro-expression recognition
- Introduction to the special session on Topological Data Analysis, ICASSP 2016
- Intrusive howling detection methods for hearing aid evaluations
- Investigating gated recurrent networks for speech synthesis
- Investigating techniques for low resource conversational speech recognition
- Investigation of speaker embeddings for cross-show speaker diarization
- Investigation on log-linear interpolation of multi-domain neural network language model
- Investigations into vowel and consonant structures in articulatory and auditory spaces using Laplacian eigenmaps
- Investigations on speaker adaptation of LSTM RNN models for speech recognition
- Iterative estimation of phase using complex cepstrum representation
- Iterative linear regression classification for image recognition
- Iterative quadratic relaxation method for optimization of multiple radar waveforms
- Iteratively reweighted tensor SVD for robust multi-dimensional harmonic retrieval
- Joint ML calibration and DOA estimation with separated arrays
- Joint acoustic factor learning for robust deep neural network based automatic speech recognition
- Joint action recognition and summarization by sub-modular inference
- Joint device-to-device transmission activation and transceiver design for sum-rate maximization in MIMO interfering channels
- Joint dictionary training for bandwidth extension of speech signals
- Joint estimation of sound source location and boundary impedance with physics-driven cosparse regularization
- Joint information from nonlinear and linear features for spoofing detection: An i-vector/DNN based approach
- Joint instance and feature importance re-weighting for person reidentification
- Joint maximum likelihood estimation of late reverberant and speech power spectral density in noisy environments
- Joint offloading decision and resource allocation for mobile cloud with computing access point
- Joint sub-band based neighbor embedding for image super-resolution
- Joint transceiver designs for secure communications over MIMO relay
- Joint user association and content placement for Cache-enabled wireless access networks
- Joint-view Kalman-filter recovery of compressed-sensed multiview videos
- Jointly optimal near-end and far-end multi-microphone speech intelligibility enhancement based on mutual information
- Kalman filter for speech enhancement in cocktail party scenarios using a codebook-based approach
- Kalman filters with Bayesian quadratic game fusion in networks
- Keyword search using query expansion for graph-based rescoring of hypothesized detections
- Knowledge-aided hyperparameter-free Bayesian detection in stochastic homogeneous environments
- L1-L1 norms for face super-resolution with mixed Gaussian-impulse noise
- LCMV beamforming with subspace projection for multi-speaker speech enhancement
- LDADEEP+: Latent aspect discovery with deep representations
- Landmark of Mandarin nasal codas and its application in pronunciation error detection
- Language model adaptation for ASR of spoken translations using phrase-based translation models and named entity models
- Language recognition using deep neural networks with very limited training data
- Language-independent acoustic cloning of HTS voices: A preliminary study
- Laplacian deep kernel learning for image annotation
- Large region acoustic source mapping: A generalized sparse constrained deconvolution approach
- Large-scale l0 sparse inverse covariance estimation
- Latent feature representation with 3-D multi-view deep convolutional neural network for bilateral analysis in digital breast tomosynthesis
- Learning compact recurrent neural networks
- Learning compact structural representations for audio events using regressor banks
- Learning cross-lingual information with multilingual BLSTM for speech synthesis of low-resource languages
- Learning data triage: Linear decoding works for compressive MRI
- Learning deep neural network using max-margin minimum classification error
- Learning discriminative and shareable patches for scene classification
- Learning full-range affinity for diffusion-based saliency detection
- Learning in constrained stochastic dynamic potential games
- Learning network structures from firing patterns
- Learning separable fixed-point kernels for deep convolutional neural networks
- Learning structured dictionary based on inter-class similarity and representative margins
- Learning to separate vocals from polyphonic mixtures via ensemble methods and structured output prediction
- Learning-based fully 3D face reconstruction from a single image
- Least squares phase retrieval using feasible point pursuit
- Lightly-supervised utterance-level emotion identification using latent topic modeling of multimodal words
- Linear network operators using node-variant graph filters
- Linearly augmented deep neural network
- Lipreading with long short-term memory
- Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
- Local Q-linear convergence and finite-time active set identification of ADMM on a class of penalized regression problems
- Local fisher discriminant analysis for spoken language identification
- Local likelihood estimation of time-variant Hawkes models
- Localization of sound sources with known statistics in the presence of interferers
- Long short term memory recurrent neural network based encoding method for emotion recognition in video
- Long-CPI multi-channel SAR based ground moving target indication
- Long-term general rank multiuser downlink beamforming with shaping constraints using QOSTBC
- Low bit-rate intra coding scheme based on constrained quantization and median-type filter
- Low complexity tonality control in the Intelligent Gap Filling tool
- Low complexity transform competition for HEVC
- Low level descriptors based DBLSTM bottleneck feature for speech driven talking avatar
- Low rank approximation based hybrid precoding schemes for multi-carrier single-user massive MIMO systems
- Low-complexity beamforming designs of sum secrecy rate maximization for the Gaussian MISO multi-receiver wiretap channel
- Low-complexity recursive convolutional precoding for OFDM-based large-scale antenna systems
- Low-complexity robust multi-cell MISO downlink precoder design
- Low-rank matrices recovery via entropy function
- Low-rank plus diagonal adaptation for deep neural networks
- Low-resolution reconstruction of intensity functions on the sphere for single-particle diffraction imaging
- Lower bounds on the L2-norms of digital resampling filters with zero-valued input samples
- Lucky ranging with towed arrays in underwater environments subject to non-stationary spatial coherence loss
- MIMO radar waveform design for multiple extended target estimation based on greedy SINR maximization
- MMSE denoising of sparse and non-Gaussian AR(1) processes
- MMSE precoder for massive MIMO using 1-bit quantization
- Machine translation based data augmentation for Cantonese keyword spotting
- Magnetic beamforming for wireless power transfer
- Maintaining throughput network connectivity in ad hoc networks
- Manga-specific features and latent style model for manga style analysis
- Manifold denoising based on spectral graph wavelets
- Manifold-based Bayesian inference for semi-supervised source localization
- Masked correlation filters for partially occluded face recognition
- Maximally improper interference in underlay cognitive radio networks
- Maximum likelihood PSD estimation for speech enhancement in reverberant and noisy conditions
- Maximum likelihood and maximum a posteriori direction-of-arrival estimation in the presence of sirp noise
- Maximum likelihood rumor source detection in a star network
- Measure-transformed quasi likelihood ratio test
- Measurement partitioning and observational equivalence in state estimation
- Mediated experts for deep convolutional networks
- Medical image super-resolution with non-local embedding sparse representation and improved IBP
- Membrane shape and boundary conditions estimation using eigenmode decomposition
- Memory reduction techniques for successive cancellation decoding of polar codes
- Memory-restricted multiscale dynamic time warping
- Micro-Doppler extraction from ISAR image
- Microtexture inpainting through Gaussian conditional simulation
- Millimeter wave communications channel estimation via Bayesian group sparse recovery
- Minimum distance criterion for non-negative hyperspectral image deconvolution
- Minimum word error training of long short-term memory recurrent neural network language models for speech recognition
- Mining representative actions for actor identification
- Mitigation of sparsely sampled nonstationary jammers for multi-antenna GNSS receivers
- Mobile beamforming & spatially controlled relay communications
- Modeling audio directional statistics using a complex bingham mixture model for blind source extraction from diffuse noise
- Modeling deep bidirectional relationships for image classification and generation
- Modeling spectral envelopes using deep conditional restricted Boltzmann machines for statistical parametric speech synthesis
- Modelling stress in public speaking: Evolution of stress levels during conference presentations
- Modulation spectrum compensation for HMM-based speech synthesis using line spectral pairs
- Mood state prediction from speech of varying acoustic quality for individuals with bipolar disorder
- Multi-centrality graph spectral decompositions and their application to cyber intrusion detection
- Multi-channel power allocation for device-to-device communication underlaying cellular networks
- Multi-focus image fusion via coupled dictionary training
- Multi-focus pixel-based image fusion in dual domain
- Multi-fold Gabor filter convolution descriptor for face recognition
- Multi-index voting for asymmetric distance computation in a large-scale binary codes
- Multi-kernel based nonlinear models for connectivity identification of brain networks
- Multi-pair two-way AF relaying systems with massive arrays and imperfect CSI
- Multi-pass feature enhancement based on generative-discriminative hybrid approach for noise robust speech recognition
- Multi-processor approximate message passing using lossy compression
- Multi-stream spectral representation for statistical parametric speech synthesis
- Multi-view distributed source coding of binary features for visual sensor networks
- Multiantenna spectrum sensing for improper signals over frequency selective channels
- Multichannel audio declipping
- Multichannel blind source separation based on non-negative tensor factorization in wavenumber domain
- Multichannel identification of room acoustic systems with adaptive filters based on orthonormal basis functions
- Multicore implementation of LDPC decoders based on ADMM algorithm
- Multicriteria optimization for nonunitary joint block diagonalization
- Multilingual data selection for training stacked bottleneck features
- Multilingual region-dependent transforms
- Multimodal Kalman filtering
- Multimodal human action recognition in assistive human-robot interaction
- Multipath radar tracking with large uncertainty in the environment
- Multipath removal by online blind deconvolution in through-the-wall-imaging
- Multiple instance discriminative dictionary learning for action recognition
- Multiple instance learning for model ensemble and meta data transfer
- Multiple scattering effects on the localization of two point scatterers
- Multiple-kernel adaptive segmentation and tracking (MAST) for robust object tracking
- Multiplicative update of AR gains in codebook-driven speech enhancement
- Multiview learning via deep discriminative canonical correlation analysis
- Music emotion recognition with adaptive aggregation of Gaussian process regressors
- Mutual information based radar waveform design for joint radar and cellular communication systems
- NMF-based informed source separation
- NMF-based source separation utilizing prior knowledge on encoding vector
- Network topology adaptation and interference coordination for energy saving in heterogeneous networks
- Neural network based spectral mask estimation for acoustic beamforming
- Neural network shape: Organ shape representation with radial basis function neural networks
- News story clustering with fisher embedding
- No-reference image quality assessment for photographic images of consumer device
- Noise and reverberation effects on depression detection from speech
- Noise estimation for speech reinforcement in the presence of strong echoes
- Noise robust recognition method based on scatterer pattern for radar HRRP data
- Noise robust speech recognition using recent developments in neural networks for computer vision
- Noise suppression method for body-conducted soft speech enhancement based on external noise monitoring
- Non-asymptotic performance bounds of eigenvalue based detection of signals in non-Gaussian noise
- Non-cooperative cross-channel gain estimation using full-duplex amplify-and-forward relaying in cognitive radio networks
- Non-linear regression for bivariate self-similarity identification - application to anomaly detection in Internet traffic based on a joint scaling analysis of packet and byte counts
- Non-monotone quadratic potential games with single quadratic constraints
- Non-negative decomposition of linear relationships: Application to multi-source ocean remote sensing data
- Non-negative intermediate-layer DNN adaptation for a 10-KB speaker adaptation profile
- Non-stationary blind super-resolution
- Non-stationary noise power spectral density estimation based on regional statistics
- Non-verbal speech analysis of interviews with schizophrenic patients
- Nonconvex compressive sensing reconstruction for tensor using structures in modes
- Nonnegative matrix factorization using ADMM: Algorithm and convergence analysis
- Nonnegative matrix factorization-based frequency lowering technology for Mandarin-speaking hearing aid users
- Nonparametric detection of an anomalous disk over a two-dimensional lattice network
- Novel 3D-WPP algorithms for parallel HEVC encoding
- Novel acoustic features for automatic dialog-act tagging
- Novel favorite music classification using EEG-based optimal audio features selected via KDLPCCA
- Novel neural network based fusion for multistream ASR
- Novel quaternion matrix factorisations
- OCR-aided person annotation and label propagation for speaker modeling in TV shows
- Object recognition in art drawings: Transfer of a neural network
- Object saliency using a background prior
- Oligopoly dynamic pricing: A repeated game with incomplete information
- On Renyi's entropy estimation with one-dimensional Gaussian kernels
- On adaptive selection of estimation bandwidth for analysis of locally stationary multivariate processes
- On combining i-vectors and discriminative adaptation methods for unsupervised speaker normalization in DNN acoustic models
- On convexity and identifiability in 1-D Fourier phase retrieval
- On gridless sparse methods for multi-snapshot DOA estimation
- On multiple solutions of the "sequentially drilled" joint congruence transformation (SeDJoCo) problem for semi-blind source separation
- On parameter estimation of symmetric alpha-stable distribution
- On parametric lower bounds for discrete-time filtering
- On pilot-symbol aided channel estimation in FBMC-OQAM
- On privacy preference in collusion-deterrence games for secure multi-party computation
- On projected stochastic gradient descent algorithm with weighted averaging for least squares regression
- On scalable coding of hidden Markov sources
- On simplifying the primal-dual method of multipliers
- On sparse controllability of graph signals
- On spatio-frequential smoothing for joint angles and times of arrival estimation of multipaths
- On target localization with communication costs via tensor completion: A multi-modal approach
- On the LP-convergence of a Girsanov theorem based particle filter
- On the average staleness of global channel state information in wireless networks with random transmit node selection
- On the compression of recurrent neural networks with an application to LVCSR acoustic modeling for embedded speech recognition
- On the decay - and the smoothness behavior of the Fourier transform, and the construction of signals having strong divergent Shannon sampling series
- On the detection of non-stationary signals in the matched signal transform domain
- On the impact of residual CFO in UL MU-MIMO
- On the importance of event detection for ASR
- On the importance of harmonic phase modification for improved speech signal reconstruction
- On the influence of momentum acceleration on online learning
- On the influence of quantization on the identifiability of emotions from voice coding parameters
- On the performance of cloud radio access networks using Matérn hard-core point processes
- On the periodically time-varying bias in adaptive feedback cancellation systems with frequency shifting
- On the separability of signal and interference-plus-noise subspaces in blind pilot decontamination
- On time delay estimation based on multichannel spatiotemporal sparse linear prediction
- On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech recognition
- One plus two may not equal two plus one in a social sensing network with unknown parameters
- One-bit ADCs in wideband massive MIMO systems with OFDM transmission
- Online adaptation of the number of particles of SMC methods
- Online change detection of linear regression models
- Online incremental higher-order partial least squares regression for fast reconstruction of motion trajectories from tensor streams
- Online learning and optimization of Markov jump linear models
- Online least-squares one-class support vector machine for outlier detection in power grid data
- Online low-rank + sparse structure learning for dynamic network tracking
- Online low-rank tensor subspace tracking from incomplete data by CP decomposition using recursive least squares
- Online nonnegative matrix factorization with outliers
- Online speaker diarization using adapted i-vector transforms
- Online speaking rate estimation using recurrent neural networks
- Open-set microphone classification via blind channel analysis
- Opening big in box office? Trailer content can help
- Opinion dynamics in multi-agent systems with binary decision exchanges
- Opportunistic spectrum access with temporal-spatial reuse in cognitive radio networks
- Optimal UAV localisation in vision based navigation systems
- Optimal copula transport for clustering multivariate time series
- Optimal design of constant-modulus channel training sequences
- Optimal linear cooperation for signal classification
- Optimal pilot length for uplink massive MIMO systems with pilot reuse
- Optimal resource block allocation and muting in heterogeneous networks
- Optimal space signalling for intensity modulated MIMO optical wireless communications
- Optimal zero forcing precoder and decoder design for multi-user MIMO FBMC under strong channel selectivity
- Optimizing DTW-based audio-to-MIDI alignment and matching
- Orthogonal sparse eigenvectors: A procrustes problem
- Outlier-robust recovery of low-rank positive semidefinite matrices from magnitude measurements
- Outlying sequence detection in large datasets: Comparison of universal hypothesis testing and clustering
- Overlapping clustering of network data using cut metrics
- PCA using graph total variation
- PLIP based unsharp masking for medical image enhancement
- PROJET - Spatial audio separation using projections
- Pansharpening via coupled triple factorization dictionary learning
- Parallel metropolis chains with cooperative adaptation
- Parallel proximal methods for total variation minimization
- Parallelizing WFST speech decoders
- Parameter estimation of polynomial phase signal based on low-complexity LSU-EKF algorithm in entire identifiable region
- Parametric Frugal sensing of autoregressive power spectra
- Parametric analog mappings for correlated Gaussian sources over AWGN channels
- Partial face recognition: A sparse representation-based approach
- Particle filtering for slice-to-volume motion correction in EPI based functional MRI
- Particle filters with independent resampling
- Particle flow for particle filtering
- Partitioned successive-cancellation list decoding of polar codes
- Pathological speech processing: State-of-the-art, current challenges, and future directions
- Pattern-based 3D model compression
- Perceptual and instrumental evaluation of the perceived level of reverberation
- Perfect error compensation via algorithmic error cancellation
- Performance advantage of quaternion widely linear estimation: An approximate uncorrelating transform approach
- Performance analysis for pilot-based 1-bit channel estimation with unknown quantization threshold
- Performance analysis of EWF codes with intermediate feedback
- Performance analysis of a modified Rao test for adaptive subspace detection
- Performance analysis of joint-sparse recovery from multiple measurement vectors with prior information via convex optimization
- Performance analysis of spectral community detection in realistic graph models
- Performance limits of single-agent and multi-agent sub-gradient stochastic learning
- Persistent homology lower bounds on network distances
- Persistent homology of toroidal sliding window embeddings
- Personalized mispronunciation detection and diagnosis based on unsupervised error pattern discovery
- Personalized speech recognition on mobile devices
- Phaseless super-resolution using masks
- Phoneme-specific speech separation
- Phrase-based rĀga recognition using vector space modeling
- Phylogenetic analysis of near-duplicate images using processing age metrics
- Physical object authentication: Detection-theoretic comparison of natural and artificial randomness
- Physical-model based efficient data representation for many-channel microphone array
- Piecewise sparse signal recovery via piecewise orthogonal matching pursuit
- Pilot aided direction of arrival estimation for mmWave cellular systems
- Pilot-based channel estimation for FBMC/OQAM systems under strong frequency selectivity
- Pinpoint extraction of distant sound source based on DNN mapping from multiple beamforming outputs to prior SNR
- Policy recognition via expectation maximization
- Portfolio optimization with asset selection and risk parity control
- Posterior probabilistic modeling for inter-channel phase and time difference estimation in audio signals
- Practical considerations on the use of preference learning for ranking emotional speech
- Precise phase transition of total variation minimization
- Precise player segmentation in team sports videos using contrast-aware co-segmentation
- Predicting humor response in dialogues from TV sitcoms
- Predicting visual attention using gamma kernels
- Prediction-adaptation-correction recurrent neural networks for low-resource language speech recognition
- Predominant melody extraction from vocal polyphonic music signal by combined spectro-temporal method
- Presentation quality assessment using acoustic information and hand movements
- Principal components analysis-based visual saliency detection
- Printed document authentication using two level or code
- Privacy-preserving energy flow control in smart grids
- Privacy-preserving nonparametric decentralized detection
- Privacy-preserving sound to degrade automatic speaker verification performance
- Progress on phoneme recognition with a continuous-state HMM
- Proportionate affine projection algorithms for block-sparse system identification
- Prosparse denoise: Prony's based sparse pattern recovery in the presence of noise
- Proximity without consensus in online multi-agent optimization
- Pruning subsequence search with attention-based embedding
- Pushing the limit of non-rigid structure-from-motion by shape clustering
- Quadtree decision for depth intra coding in 3D-HEVC by good feature
- Quality-aware adaptive delivery of multi-view video
- Quantification of balance in single limb stance using kinect
- Quantifying cooperation in choir singing: Respiratory and cardiac synchronisation
- Quantization bin matching for cloud storage of JPEG images
- Quantized consensus ADMM for multi-agent distributed optimization
- Quantizer design for exploiting common information in layered coding
- Quasi closed phase analysis of speech signals using time varying weighted linear prediction for accurate formant tracking
- Question detection from acoustic features using recurrent neural network with gated recurrent unit
- Quickest convergence of online algorithms via data selection
- Quickest search over correlated sequences with model uncertainty
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.