ICASSP 2017 Accepted Papers
The full list of 1,320 papers accepted at ICASSP 2017 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- 1+N fusion: Cascaded self-portrait enhancement
- 3D audio-visual speaker tracking with an adaptive particle filter
- 3D colored mesh graph signals multi-layer morphological enhancement
- 3D reconstruction from web harvested images using a forensic quality metric
- 3D tracking swimming fish school with learned kinematic model using LSTM network
- A "polyphase" structure of two-channel spectral graph wavelets and filter banks
- A Bayesian approach to Top-Scoring Pairs classification
- A Bayesian lower bound for parameter estimation of Poisson data including multiple changes
- A Deep Learning approach to modeling competitiveness in spoken conversations
- A Diagonal-Augmented quasi-Newton method with application to factorization machines
- A Maximum Likelihood "identification-correction" scheme of sub-optimal "SeDJoCo" solutions for semi-Blind Source Separation
- A Multi-resolution approach to Common Fate-based audio separation
- A Neural Network approach for mixing language models
- A PLLR and multi-stage Staircase Regression framework for speech-based emotion prediction
- A bayesian multi-frame image super-resolution algorithm using the Gaussian Information Filter
- A blind transform based approach for the detection of isolated astrophysical pulses
- A cache-based bandwidth optimized motion compensation architecture for video decoder
- A case study of machine learning hardware: Real-time source separation using Markov Random Fields via sampling-based inference
- A compact formulation for the l21 mixed-norm minimization problem
- A compact pairwise trajectory representation for action recognition
- A comparative study of acoustic-to-articulatory inversion for neutral and whispered speech
- A comparison between real and complex Schott spherical symmetry test for PolSAR data analysis
- A comparison of Deep Learning methods for environmental sound detection
- A comprehensive performance comparison of RFI mitigation techniques for UWB radar signals
- A comprehensive study of deep bidirectional LSTM RNNS for acoustic modeling in speech recognition
- A constrained adaptive scan order approach to transform coefficient entropy coding
- A convolutional Riemannian texture model with differential entropic active contours for unsupervised pest detection
- A cross-modal adaptation approach for brain decoding
- A data centric approach to utility change detection in online social media
- A data-driven compressive sensing framework tailored for energy-efficient wearable sensing
- A deep learning approach to multiple kernel fusion
- A deep learning approach towards pore extraction for high-resolution fingerprint recognition
- A deep neural network integrated with filterbank learning for speech recognition
- A deep-learning approach to translate between brain structure and functional connectivity
- A differentially private ensemble Kalman Filter for road traffic estimation
- A distributed constrained-form support vector machine
- A distributed robust transmit beamforming design for full-duplex relay-aided wireless communication systems
- A double incremental aggregated gradient method with linear convergence rate for large-scale optimization
- A dual estimation approach for removing the show-through effect in the scanned documents
- A dynamic Bayesian nonparametric model for blind calibration of sensor networks
- A dynamic programming approach for automatic stride detection and segmentation in acoustic emission from the knee
- A fast covariance matrix reconstruction method for two-dimensional direction-of-arrival estimation
- A fast face clustering method for indexing applications on mobile phones
- A fast intra-prediction decision algorithm in inter-frame based on a novel feature of HEVC
- A feature-based linear regression model for predicting perceptual ratings of music by cochlear implant listeners
- A finite rate of innovation multichannel sampling hardware system for multi-pulse signals
- A first attempt at polyphonic sound event detection using connectionist temporal classification
- A first look into a Convolutional Neural Network for speech emotion detection
- A full reference stereoscopic video quality assessment metric
- A generalization of the sparse iterative covariance-based estimator
- A generalized Swendsen-Wang algorithm for Bayesian nonparametric joint segmentation of multiple images
- A generalized log-spectral amplitude estimator for single-channel speech enhancement
- A generalized matrix-decomposition processor for joint MIMO transceiver design
- A geometric learning approach on the space of complex covariance matrices
- A greedy algorithm with learned statistics for sparse signal reconstruction
- A hierarchical Dirichlet process mixture of GID Distributions with feature selection for spatio-temporal video modeling and segmentation
- A joint detection-classification model for audio tagging of weakly labelled data
- A joint learning based Face Super Resolution approach via contextual topological structure
- A k-nearest neighbor multilabel ranking algorithm with application to content-based image retrieval
- A knowledge transfer and boosting approach to the prediction of affect in movies
- A knowledge-driven framework for ECG representation and interpretation for wearable applications
- A lattice method for resolving range ambiguity in dual-frequency RFID tag localisation
- A locality-preserving essence vector modeling framework for spoken document retrieval
- A locally linear embbeding based postfiltering approach for speech enhancement
- A low-complexity algorithm for utility based spectrum coordination in DSL systems
- A low-complexity beamforming method by orthogonal codebooks for millimeterwave links
- A majorization-minimization algorithm with projected gradient updates for time-domain spectrogram factorization
- A minimum variance partially distortionless response filter for single-channel noise reduction
- A mixture model-based real-time audio sources classification method
- A model-free causality measure based on multi-variate delay embedding
- A modulation feature set for robust Automatic Speech Recognition in additive noise and reverberation
- A multiple bandwidth objective speech intelligibility estimator based on articulation index band correlations and attention
- A network of deep neural networks for Distant Speech Recognition
- A neural filter-based scheme for synchronizing chaotic systems
- A neural network alternative to non-negative audio models
- A new chaotic feature for EEG classification based seizure diagnosis
- A new framework for designing incoherent sparsifying dictionaries
- A new generalization of the discrete Teager-Kaiser energy operator - application to biomedical signals
- A new noise annoyance measurement metric for urban noise sensing and evaluation
- A new perceptual assessment methodology for selective HEVC video encryption
- A new two-dimensional Fourier transform algorithm based on image sparsity
- A noise suppression method for body-conducted soft speech based on non-negative tensor factorization of air- and body-conducted signals
- A non-intrusive Short-Time Objective Intelligibility measure
- A nonconvex splitting method for symmetric nonnegative matrix factorization: Convergence analysis and optimality
- A novel LBP-based Color descriptor for face recognition
- A novel dictionary based SRC for face recognition
- A novel ensemble classifier of hyperspectral and LiDAR data using morphological features
- A novel iterative online rating attack based on market self-exciting property
- A novel layerwise pruning method for model reduction of fully connected deep neural networks
- A novel methodology to quantify dense EEG in cognitive tasks
- A novel pitch extraction based on jointly trained deep BLSTM Recurrent Neural Networks with bottleneck features
- A novel re-tracking strategy for monocular SLAM
- A novel sparse model for multi-source localization using distributed microphone array
- A parallelized dynamic programming approach to zero resource spoken term discovery
- A particle filter for sequential infection source estimation
- A performance-based approach to designing the stimulus presentation paradigm for the P300-based BCI by exploiting coding theory
- A practical high-dimensional Sparse Fourier Transform
- A provable nonconvex model for factoring nonnegative matrices
- A pseudo-Voigt component model for high-resolution recovery of constituent spectra in Raman spectroscopy
- A railroad detection algorithm for infrastructure surveillance using enduring airborne systems
- A real-time 3D head mesh modeling and expressive articulatory animation system
- A reassigned based singing voice pitch contour extraction method
- A robust FISTA-like algorithm
- A robust feature descriptor based on multiple gradient-related features
- A scalable convolutional neural network for task-specified scenarios via knowledge distillation
- A self-calibrating bidirectional indoor localization system
- A semi-supervised method for multi-subject FMRI functional alignment
- A simple way to approximate average robust multiuser MISO transmit optimization under covariance-based CSIT
- A software-defined radio implementation of timestamp-free network synchronization
- A sparse CCA algorithm with application to model-order selection for small sample support
- A speech enhancement algorithm by iterating single- and multi-microphone processing and its application to robust ASR
- A statistical approach to semi-supervised speech enhancement with low-order non-negative matrix factorization
- A stochastic maximum-likelihood framework for simplex structured matrix factorization
- A study of speaker verification performance with expressive speech
- A study on data augmentation of reverberant speech for robust speech recognition
- A study on motion mode identification for cyborg roaches
- A subspace approach for shrinkage parameter selection in undersampled configuration for Regularised Tyler Estimators
- A systematic approach to compute perceptual distribution of monosyllables
- A tensor based framework for community detection in dynamic networks
- A time-reversal spatial hardening effect for indoor speed estimation
- A transfer learning and progressive stacking approach to reducing deep model sizes with an application to speech enhancement
- A two-stage algorithm for noisy and reverberant speech enhancement
- A two-stage optimization approach to the asynchronous multi-sensor registration problem
- A unified convergence analysis of the multiplicative update algorithm for nonnegative matrix factorization
- A unified diversity measure for distributed inference
- A wavelet-based approach to monitoring Parkinson's disease symptoms
- A weakly-convex formulation for phaseless imaging
- ADC bit allocation under a power constraint for mmWave massive MIMO communication receivers
- ADMM for harmonic retrieval from one-bit sampling with time-varying thresholds
- AMOS: An automated model order selection algorithm for spectral graph clustering
- ARIMA-GARCH modeling for epileptic seizure prediction
- About zero bitwatermarking error exponents
- Accelerated dual gradient-based methods for total variation image denoising/deblurring problems
- Accelerated sensor position selection using graph localization operator
- Accelerating Deep Convolutional Networks using low-precision and sparsity
- Accelerating the hybrid steepest descent method for affinely constrained convex composite minimization tasks
- Acceleration of Adaptive normalized quasi-Newton algorithm with improved upper bounds of the condition number
- Achievable uplink rates for massive MIMO with coarse quantization
- Acoustic classification using semi-supervised Deep Neural Networks and stochastic entropy-regularization over nearest-neighbor graphs
- Acoustic imaging of sparse Sources with Orthogonal Matching Pursuit and clustering of basis vectors
- Action-vectors: Unsupervised movement modeling for action recognition
- Active learning for low-resource speech recognition: Impact of selection size and language modeling data
- Active learning for sound event classification by clustering unlabeled data
- Active speech control using wave-domain processing with a linear wall of dipole secondary sources
- Adaptation of PLDA for multi-source text-independent speaker verification
- Adapting and controlling DNN-based speech synthesis using input codes
- Adaptive DCTNet for audio signal classification
- Adaptive gain control and time warp for enhanced speech intelligibility under reverberation
- Adaptive matching pursuit for sparse signal recovery
- Adaptive superpixel segmentation aggregating local contour and texture features
- Advances in Empirical Mode Decomposition for computing Instantaneous Amplitudes and Instantaneous Frequencies
- Advances in all-neural speech recognition
- Affect recognition from lip articulations
- Alpha-stable multichannel audio source separation
- Alternating diffusion maps for dementia severity assessment
- Alternative networks for monolingual bottleneck features
- An EM algorithm for joint source separation and diarisation of multichannel convolutive speech mixtures
- An FFT-based synchronization approach to recognize human behaviors using STN-LFP signal
- An FPGA prototype of dual link algorithm for MIMO interference network
- An LSTM-CTC based verification system for proxy-word based OOV keyword search
- An M-channel critically sampled filter bank for graph signals
- An accumulative fusion architecture for discriminating people and vehicles using acoustic and seismic signals
- An accurate perturbation analysis algorithm for music with Toeplitz covariance matrix
- An augmented Lagrangian algorithm for decomposition of symmetric tensors of order-4
- An autoregressive recurrent mixture density network for parametric speech synthesis
- An efficient online Adaptive Sampling strategy for Matrix Completion
- An embedding mechanism for natural steganography after down-sampling
- An empirical evaluation of zero resource acoustic unit discovery
- An engineer's guide to Particle Filtering on the Stiefel manifold
- An evaluation of score-informed methods for estimating fundamental frequency and power from polyphonic audio
- An incremental quasi-Newton method with a local superlinear convergence rate
- An investigation into language model data augmentation for low-resourced STT and KWS
- An investigation into learning effective speaker subspaces for robust unsupervised DNN adaptation
- An iterative auction mechanism for data trading
- An iterative reconstruction algorithm for amplitude sampling
- An online NIPALS algorithm for Partial Least Squares
- An online feature selection architecture for Human Activity Recognition
- Analysis and prediction of heart rate using speech features from natural speech
- Analysis of a covert communication method utilizing non-coherent DPSK masked by pulsed radar interference
- Analysis of keyword spotting performance across IARPA babel languages
- Analytical approach to 2.5D sound field control using a circular double-layer array of fixed-directivity loudspeakers
- Anchor-based group detection in crowd scenes
- Anomaly detection in IP networks based on randomized subspace methods
- Anuran call classification with deep learning
- Aperture Domain Model Image REconstruction (ADMIRE) for improved ultrasound imaging
- Appearance-based gesture recognition in the compressed domain
- Applying compensation techniques on i-vectors extracted from short-test utterances for speaker verification using deep neural network
- Applying the unit circle constraint to the diagonally loaded minimum variance distortionless response beamformer
- Approximate simulation of linear continuous time models driven by asymmetric stable Lévy processes
- Array covariance matrix-based atomic norm minimization for off-grid coherent direction-of-arrival estimation
- Artificial bandwidth extension using the constant Q transform
- Assessment of broadband SNR estimation for hearing aid applications
- Assessment of musical noise using localization of isolated peaks in time-frequency domain
- Assisted dictionary learning for FMRI data analysis
- Asymmetric cross-view dictionary learning for person re-identification
- Asymptotic analysis of a GLR test for detection with large sensor arrays: New results
- Asymptotic analysis of multicell massive MIMO over Rician fading channels
- Asymptotic optimality of consensus-based sequential probability ratio test
- Asymptotic perfect secrecy in distributed estimation for large sensor networks
- Asynchronous online ADMM for consensus problems
- Asynchronous parallel nonconvex large-scale optimization
- Atlas based 3D liver segmentation using adaptive thresholding and superpixel approaches
- Atomic norm minimization for modal analysis with random spatial compression
- Audio Set: An ontology and human-labeled dataset for audio events
- Audio source separation based on convolutive transfer function and frequency-domain lasso optimization
- Audio time stretching with an adaptive multiresolution phase vocoder
- Audio-visual object localization and separation using low-rank and sparsity
- Auto-weighted two-dimensional principal component analysis with robust outliers
- Autoencoders trained with relevant information: Blending Shannon and Wiener's perspectives
- Automated robust Anuran classification by extracting elliptical feature pairs from audio spectrograms
- Automatic assessment of dysarthria severity level using audio descriptors
- Automatic conversion of Pop music into chiptunes for 8-bit pixel art
- Automatic detection of motion artifacts in MR images using CNNS
- Automatic detection of syllable stress using sonority based prominence features for pronunciation evaluation
- Automatic dynamic template tracking of inner lips based on CLNF
- Automatic gain control with integrated signal enhancement for specified target and background-noise levels
- Automatic image cropping with aesthetic map and gradient energy map
- Automatic insect recognition using optical flight dynamics modeled by kernel adaptive ARMA network
- Automatic matching and synchronization of user generated videos from a large scale sport event
- Automatic multi-lingual arousal detection from voice applied to real product testing applications
- Automatic musical key estimation with adaptive mode bias
- Automatic node selection for Deep Neural Networks using Group Lasso regularization
- Automatic parameter tuning for image denoising with learned sparsifying transforms
- Automatic radar waveform recognition based on time-frequency analysis and convolutional neural network
- Automatic segmentation of retinal vasculature
- Automatic shrinkage tuning based on a system-mismatch estimate for sparsity-aware adaptive filtering
- Automatic speech emotion recognition using recurrent neural networks with local attention
- Autoregressive moving average graph filters a stable distributed implementation
- Average SCR loss analysis for polarimetric STAP with Kronecker structured covariance matrix
- Average consensus-based asynchronous tracking
- Axiomatic hierarchical clustering given intervals of metric distances
- BER analysis of regularized least squares for BPSK recovery
- BLSTM-HMM hybrid system combined with sound activity detection network for polyphonic Sound Event Detection
- BSmCCA: A block sparse multiple-set canonical correlation analysis algorithm for multi-subject fMRI data sets
- Bag of Fisher Vectors representation of images by saliency-based spatial partitioning
- Balanced sensor management across multiple time instances via l-1/l-infinity norm minimization
- Balancing exploration and exploitation in reinforcement learning using a value of information criterion
- Barker-Coded node-pore resistive pulse sensing with built-in coincidence correction
- Bayesian Blind Deconvolution with application to acoustic Feedback Path modeling
- Bayesian information criterion for multidimensional sinusoidal order selection
- Bayesian joint-sequence models for grapheme-to-phoneme conversion
- Bayesian learning in a network with multi-hypothesis decision exchanges
- Bayesian multi-antenna sensing in cognitive radio networks using Fractional Bayes Factor
- Bayesian multichannel nonnegative matrix factorization for audio source separation and localization
- Bayesian nonparametric subspace estimation
- Bayesian phonotactic Language Model for Acoustic Unit Discovery
- Bayesian reconstruction of hyperspectral images by using compressed sensing measurements and a local structured prior
- Bayesian-driven criterion to automatically select the regularization parameter in the ℓ1-Potts model
- Beamnet: End-to-end training of a beamformer-supported multi-channel ASR system
- Belief control strategies for interactions over weak graphs
- Bernoulli filter based algorithm for joint target tracking and classification in a cluttered environment
- Binary matrix completion with performance guarantees for single individual haplotyping
- Biobjective transmitter optimization for service integration in MIMO Gaussian broadcast channel
- Biobotic motion and behavior analysis in response to directional neurostimulation
- Biologically inspired speech emotion recognition
- Bivariate probabilistic constrained programming for interference exploitation in the cognitive radio
- Blind bandwidth extension using K-means and Support Vector Regression
- Blind compensation of polynomial mixtures of Gaussian signals with application in nonlinear blind source separation
- Blind estimation of directional properties of room reverberation using a spherical microphone array
- Blind image deblurring based on sparse representation and structural self-similarity
- Blind image deconvolution using Student's-t prior with overlapping group sparsity
- Blind on board wideband antenna RF calibration for multi-antenna satellites
- Blind source separation based on independent low-rank matrix analysis with sparse regularization for time-series activity
- Blood vessels extraction using Fuzzy Mathematical Morphology
- Body structure based triplet Convolutional Neural Network for person re-identification
- Boolean Kalman Filter with correlated observation noise
- Borehole image correspondence and automated alignment
- Brain signal analytics from graph signal processing perspective
- Building recurrent networks by unfolding iterative thresholding for sequential sparse recovery
- CNN architectures for large-scale audio classification
- CNN-LTE: A class of 1-X pooling convolutional neural networks on label tree embeddings for audio scene classification
- Capacity results on the finite state Markov wiretap channel with delayed state feedback
- Change detection between multi-band images using a robust fusion-based approach
- Change detection with unknown post-change parameter using Kiefer-Wolfowitz method
- Channel estimation for crosstalk cancellation in wireless acoustic networks
- Character-level deep conflation for business data analytics
- Character-level language modeling with hierarchical recurrent neural networks
- Classification of Gaussian trajectories with missing data in Boolean gene regulatory networks
- Classification of thyroid nodules in ultrasound images using deep model based transfer learning and hybrid features
- Classification of voice modes using neck-surface accelerometer data
- Clinical decision support system for Parkinson's disease and related movement disorders
- Coalitional game theoretic optimization of electricity cost for communities of smart households
- Codec independent lossy audio compression detection
- Coding of 3D holoscopic image by using spatial correlation of rendered view images
- Coding of fine granular audio signals using High Resolution Envelope Processing (HREP)
- Coherence-adjusted monopole dictionary and convex clustering for 3D localization of mixed near-field and far-field sources
- Collaborative Deep Learning for speech enhancement: A run-time model selection method using autoencoders
- Collaborative method based on the acoustical interaction effects on active noise control systems over distributed networks
- Collaborative voting of 3D features for robust gesture estimation
- Color channel-wise recurrent learning for facial expression recognition
- Color demosaicking via nonlocal tensor representation
- Color image coding based on linear combination of adaptive colorspaces
- Color prediction in image coding using Steered Mixture-of-Experts
- Combination strategy based on relative performance monitoring for multi-stream reverberant speech recognition
- Combinatorial bounds on the α-divergence of univariate mixture models
- Combined Weighted Prediction Error and Minimum Variance Distortionless Response for dereverberation
- Combining belief propagation and successive cancellation list decoding of polar codes on a GPU platform
- Combining unidirectional long short-term memory with convolutional output layer for high-performance speech synthesis
- Comparison of two binaural beamforming approaches for hearing aids
- Complex NMF with the generalized Kullback-Leibler divergence
- Complexity control of HEVC for video conferencing
- Compressed beam-selection in millimeterwave systems with out-of-band partial support information
- Compressed cyclostationary detection for Cognitive Radio
- Compressed sensing MRI using double sparsity with additional training images
- Compressed sensing and optimal denoising of monotone signals
- Compressing higher order ambisonics of a multizone soundfield
- Compressive K-means
- Compressive imaging with iterative forward models
- Compressive information acquisition with hardware impairments and constraints: A case study
- Compressive pulse-Doppler radar sensing via 1-bit sampling with time-varying threshold
- Compressive sensing based ECG monitoring with effective AF detection
- Compressive sensing based spectrum sharing and coexistence for machine-to-machine communications
- Compressive sensing strategy for classification of bearing faults
- Computation and visualization of posterior densities in scalar nonlinear and non-Gaussian Bayesian filtering and smoothing problems
- Computational microscopy: illumination coding and nonlinear optimization enables Gigapixel 3D phase imaging
- Computing the largest eigenvalue distribution for complex Wishart matrices
- Concomitant of ordered multivariate normal distribution with application to parametric inference
- Confidence measures for CTC-based phone synchronous decoding
- Consensus clustering on data fragments
- Constrain the Docile CTUs: An In-Frame complexity allocator for HEVC Intra encoders
- Constructing sub-word units for spoken term detection
- Contextual multi-armed bandit algorithms for personalized learning action selection
- Contour-enhanced resampling of 3D point clouds via graphs
- Convergence analysis of the information matrix in Gaussian Belief Propagation
- Convergence rates of inertial splitting schemes for nonconvex composite optimization
- Convex combination framework for a priori SNR estimation in speech enhancement
- Convolutional Neural Network for speaker change detection in telephone speaker diarization system
- Convolutional approximations to linear dimensionality reduction operators
- Convolutional neural networks for passive monitoring of a shallow water environment using a single sensor
- Convolutional recurrent neural networks for music classification
- Copula application in nonlinear/non-Gaussian Bayesian tracking in the case of correlated sensors
- Correlation-based detection of TCM signals for cognitive radios
- Cost-effective diffusion Kalman filtering with implicit measurement exchanges
- Coupled hidden Markov model for automatic ECG and PCG segmentation
- Cover song identification with 2D Fourier Transform sequences
- Cramér-Rao bounds for the localization of anisotropic sources
- Critical sampling for wavelet filterbanks on arbitrary graphs
- Cross-correlations of zero crossings in jointly Gaussian and stationary processes with zero means
- Cross-modal transfer with neural word vectors for image feature learning
- Cross-modality matching based on Fisher Vector with neural word embeddings and deep image features
- Crowd-ML: A library for privacy-preserving machine learning on smart devices
- Cumulative moving averaged bottleneck speaker vectors for online speaker adaptation of CNN-based acoustic models
- Cyber attacks on estimation sensor networks and iots: Impact, mitigation and implications to unattacked systems
- D2L: Decentralized dictionary learning over dynamic networks
- DFVR: Deformable finger vein recognition
- DNN approach to speaker diarisation using speaker channels
- DNN-based source enhancement self-optimized by reinforcement learning using sound quality measurements
- DNN-based speech mask estimation for eigenvector beamforming
- DOA estimation in structured phase-noisy environments
- DOA estimation with histogram analysis of spatially constrained active intensity vectors
- Data analysis as a web service: A case study using IoT sensor data
- Data-driven fusion of multi-camera video sequences: Application to abandoned object detection
- Data-driven solo voice enhancement for jazz music retrieval
- Decentralized independent vector analysis
- Decoding emotional experiences through physiological signal processing
- Decorrelation for audio object coding
- Deductive refinement of species labelling in weakly labelled birdsong recordings
- Deep Neural Network based learning and transferring mid-level audio features for acoustic scene classification
- Deep attractor network for single-microphone speaker separation
- Deep clustering and conventional networks for music separation: Stronger together
- Deep fusion of heterogeneous sensor data
- Deep learning based automatic volume control and limiter system
- Deep learning on symbolic representations for large-scale heterogeneous time-series event prediction
- Deep long short-term memory adaptive beamforming networks for multichannel robust speech recognition
- Deep mixture density network for statistical model-based feature enhancement
- Deep multi-view models for glitch classification
- Deep multi-view robust representation learning
- Deep neural network based wake-up-word speech recognition with two-stage detection
- Deep neural networks based speaker modeling at different levels of phonetic granularity
- Deep ranking: Triplet MatchNet for music metric learning
- Deep salience map guided arbitrary direction scene text recognition
- Deep-net fusion to classify shots in concert videos
- DeepText: A new approach for text proposal generation and text detection in natural images
- Delay and Doppler processing for multi-target detection with IEEE 802.11 OFDM signaling
- Demixing sparse signals via convex optimization
- Density ridge manifold traversal
- Dereverberation based on bin-wise temporal variations of complex spectrogram
- Design of space-time block coded unique word OFDM systems
- Designing efficient architectures for modeling temporal features with convolutional neural networks
- Designing secure networks with q-composite key predistribution under different link constraints
- Detecting stress and depression in adults with aphasia through speech analysis
- Detection of Visual Evoked Potentials using Ramanujan Periodicity Transform for real time brain computer interfaces
- Detection of anomaly acoustic scenes based on a temporal dissimilarity model
- Detection of impulsive disturbances in archive audio signals
- Detection rate optimization in radar systems with unknown disturbance power
- Detection rate optimization in surveillance radars with two-step sequential detection
- Detection with multimodal dependent data using low-dimensional random projections
- Deterministic annealing based design of error resilient predictive compression systems
- Diagonal microphone placement for the landscape/portrait interchangeable mode of a personal computer
- Dialog context language modeling with recurrent neural networks
- Dictionary-based Equivalent Source Method for Near-Field Acoustic Holography
- Diffusion gradient boosting for networked learning
- Digital predistortion for hybrid precoding architecture in millimeter-wave massive mimo systems
- Direction finding using sparse linear arrays with missing data
- Directional discrete cosine transforms arising from discrete cosine and sine transforms for directional block-wise image representation
- Directional graph weight prediction for image compression
- Dirichlet Mixture Matching Projection for supervised linear dimensionality reduction of proportional data
- Dirichlet process mixture models for clustering i-vector data
- Disc-GLasso: Discriminative graph learning with sparsity regularization
- Discovering dimensions of perceived vocal expression in semi-structured, unscripted oral history accounts
- Discovering sound concepts and acoustic relations in text
- Discriminative autoencoders for speaker verification
- Discriminative feature domains for reverberant acoustic environments
- Discriminative importance weighting of augmented training data for acoustic model training
- Discriminative recurring signal detection and localization
- Disjunctive Normal Shape Boltzmann Machine
- Disparity estimation in stereo videos using spatio-temporal disparity hyperplane models
- Distance metric learning for posteriorgram based keyword search
- Distance-preserving property of random projection for subspaces
- Distributed TV-L1 image fusion using PDMM
- Distributed blind equalization in networked systems
- Distributed decision-making over mobile adaptive networks
- Distributed largest eigenvalue detection
- Distributed max-SINR speech enhancement with ad hoc microphone arrays
- Distributed nonconvex optimization for sparse representation
- Distributed optimization for evolving networks of growing connectivity
- Distributed probabilistic bisection search using social learning
- Distributed recursive least-squares with data-adaptive censoring
- Distributed sensor selection for field estimation
- Distributed sparsified graph filters for denoising and diffusion tasks
- Divide-and-warp temporal alignment of speech signals between speakers: Validation using articulatory data
- DoF analysis in a two-layered heterogeneous wireless interference network
- Domain adaptation of DNN acoustic models using knowledge distillation
- Double Relay Communication Protocol with power control for achieving fairness in cellular systems
- Double-bit quantization and weighting for nearest neighbor search
- Drum extraction in single channel audio signals using multi-layer Non negative Matrix Factor Deconvolution
- Drum transcription from polyphonic music with recurrent neural networks
- Dual-Tree wavelet scattering network with parametric log transformation for object classification
- Dual-fisheye lens stitching for 360-degree imaging
- Duration prediction using multiple Gaussian process experts for GPR-based speech synthesis
- Dynamic Graph Fourier Transform on temporal functional connectivity networks
- Dynamic Probabilistic Linear Discriminant Analysis for video classification
- Dynamic cloud Offloading for View Synthesis
- Dynamic polygon cloud compression
- Dynamic reconstruction of influence graphs with adaptive directed information
- Dynamic tracking attention model for action recognition
- ECG-based biometrics using recurrent neural networks
- EEG channel optimization via sparse common spatial filter
- EEG source imaging assists decoding in a face recognition task
- Edge-preserving filtering by projection onto L0 gradient constraint
- Edited film alignment via selective Hough transform and accurate template matching
- Effect of acoustic conditions on algorithms to detect Parkinson's disease from speech
- Effect of sampling on the estimation of the apparent coefficient of diffusion in MRI
- Effective Fisher vector aggregation for 3D object retrieval
- Effective articulatory modeling for pronunciation error detection of L2 learner without non-native training data
- Effective compressive sensing via reweighted total variation and weighted nuclear norm regularization
- Effective emotion recognition in movie audio tracks
- Effective estimation of the desired-signal subspace and its application to robust adaptive beamforming
- Effective joint training of denoising feature space transforms and Neural Network based acoustic models
- Effective keyword search for low-resourced conversational speech
- Effects of gender information in text-independent and text-dependent speaker verification
- Efficient adaptive filtering in compressive domains for sparse systems and relation to transform-domain adaptive filtering
- Efficient bridging-based destination inference in object tracking
- Efficient hybrid space-ground precoding techniques for multi-beam satellite systems
- Efficient large scale antenna selection by partial switching connectivity
- Efficient methods to train multilingual bottleneck feature extractors for low resource keyword search
- Efficient mode decision for noisy video transcoding
- Efficient multidimensional parameter estimation for joint wideband radar and communication systems based on OFDM
- Efficient multiplier-less structures for Ramanujan filter banks
- Efficient pooling of image based CNN features for action recognition in videos
- Efficient postcoding filter in LU-based beamforming scheme
- Efficient representation of segmentation contours using chain codes
- Efficient single/multiple unimodular waveform design with low weighted correlations
- Eigenvalue decomposition based estimators of carrier frequency offset in multicarrier underwater acoustic communication
- Embedded clustering via robust orthogonal least square discriminant analysis
- Emitter source localization using time-of-arrival measurements from single moving receiver
- Emotion estimation via tensor-based supervised decision-level fusion from multiple Brodmann areas
- Emotion recognition through integrating EEG and peripheral signals
- Encoder-decoder with focus-mechanism for sequence labelling based spoken language understanding
- End-to-end ASR-free keyword search from speech
- End-to-end joint learning of natural language understanding and dialogue manager
- End-to-end speech recognition and keyword search on low-resource languages
- End-to-end spoofing detection with raw waveform CLDNNS
- End-to-end visual speech recognition with LSTMS
- Energy blowup for truncated stable LTI systems
- Energy reduction opportunities in an HEVC real-time encoder
- Energy-efficient design for non-regenerative MIMO relay networks
- Engagement detection for children with Autism Spectrum Disorder
- Enhanced LBP texture features from time frequency representations for acoustic scene classification
- Enhanced canonical correlation analysis with local density for cross-domain visual classification
- Enhanced depth estimation for hand-held light field cameras
- Enhanced indoor localization through crowd sensing
- Enhanced pixel-wise voting for image vanishing point detection in road scenes
- Enhanced single antenna interference cancellation from MMSE third-order complex Volterra filters
- Enhanced ultrasound image reconstruction using a compressive blind deconvolution approach
- Enhancing ICA performance by exploiting sparsity: Application to FMRI analysis
- Enhancing QoS in spatially controlled beamforming networks via distributed stochastic programming
- Enhancing noise and pitch robustness of children's ASR
- Enhancing observability in power distribution grids
- Enhancing retinal vessel segmentation by color fusion
- Enhancing utility and privacy with noisy minimax filters
- Ensemble classification based on Random linear base classifiers
- Ensemble feature selection for domain adaptation in speech emotion recognition
- Environment aware speaker diarization for moving targets using parallel DNN-based recognizers
- Epithelium-stroma classification in histopathological images via convolutional neural networks and self-taught learning
- Estimating sparse signals using integrated wide-band dictionaries
- Estimation accuracy of non-standard maximum likelihood estimators
- Estimation and learning of Dynamic Nonlinear Networks (DyNNets)
- Estimation in autoregressive processes with partial observations
- Estimation of multiple pitches in stereophonic mixtures using a codebook-based approach
- Estimation of vocal tract area function from volumetric Magnetic Resonance Imaging
- Evaluating automatic speech recognition systems in comparison with human perception results using distinctive feature measures
- Evaluation of a complementary hearing aid for spatial sound segregation
- Evaluation of weight sparsity regularizion schemes of deep neural networks applied to functional neuroimaging data
- Event-based consensus for a class of heterogeneous multi-agent systems: An LMI approach
- Event-related synchronisation responses to N-back memory tasks discriminate between healthy ageing, mild cognitive impairment, and mild Alzheimer's disease
- Evolutionary affinity propagation
- Example-based Visual Object Counting for complex background with a local low-rank constraint
- Exemplar selection methods in voice conversion
- Exemplar-based image completion via new quality measure based on phaseless texture features
- Exemplar-embed complex matrix factorization for facial expression recognition
- Expected Likelihood sphericity test distribution for complex angular central Gaussian data
- Experimental demonstration of nullforming from a fully wireless distributed array
- Exploiting different word clusterings for class-based RNN language modeling in speech recognition
- Exploiting mutual coupling by means of analog-digital zero forcing
- Exploiting sequence information for text-dependent Speaker Verification
- Exploiting sequential Low-Rank Factorization for multilingual DNNS
- Exploring universal speech attributes for speaker verification
- Expressive visual text to speech and expression adaptation using deep neural networks
- Extended Kalman filter for extended object tracking
- Extended low-rank plus diagonal adaptation for deep and recurrent neural networks
- Extracting Fourier descriptors from compressive measurements
- Extracting structural spectral features using what-where auto-encoders for statistical parametric speech synthesis
- Extraction of common task signals and spatial maps from group fMRI using a PARAFAC-based tensor decomposition technique
- Extreme image completion
- FRI sampling and time-varying pulses: Some theory and four short stories
- FRIDA: FRI-based DOA estimation for arbitrary array layouts
- Face Album: Towards automatic photo management based on person identity on mobile phones
- Face detection and recognition for home service robots with end-to-end deep neural networks
- Face recognition in real-world images
- Facial attractiveness prediction using psychologically inspired convolutional neural network (PI-CNN)
- Factor analysis methods for joint speaker verification and spoof detection
- Fast HEVC intra coding algorithm based on machine learning and Laplacian Transparent Composite Model
- Fast HRFT measurement system with unconstrained head movements for 3D audio in virtual and augmented reality applications
- Fast Spectral Clustering with efficient large graph construction
- Fast algorithm for statistical phrase/accent command estimation based on generative model incorporating spectral features
- Fast and privacy preserving distributed low-rank regression
- Fast camera self-calibration for synthesizing Free Viewpoint soccer Video
- Fast convolutional sparse coding with separable filters
- Fast exemplar selection algorithm for matrix approximation and representation: A variant oASIS algorithm
- Fast feasibility pursuit for non-convex QCQPS via first-order methods
- Fast harmonic chirp summation
- Fast human segmentation using color and depth
- Fast hyperspectral unmixing in presence of sparse multiple scattering nonlinearities
- Fast implementation for symmetric non-separable transforms based on grids
- Fast interpolation of bandlimited functions
- Fast inverse tone mapping with Reinhard's global operator
- Fast orthogonal approximations of sampled sinusoids and bandlimited signals
- Fast path localization on graphs via multiscale Viterbi decoding
- Fast sparse recovery for any RIP-1 matrix
- Fast tagging of natural sounds using marginal co-regularization
- Faster sequence training
- Faster-than-Nyquist spatiotemporal symbol-level precoding in the downlink of multiuser MISO channels
- Feature encoding in band-limited distributed surveillance systems
- Feature extraction using multimodal convolutional neural networks for visual speech recognition
- Feature mapping for speaker diarization in noisy conditions
- Feature++: Cross dimension feature fusion for road detection
- Feature-based ROI generation for stereo-based pedestrian detection
- Feedback connection for deep neural network-based acoustic modeling
- Fetal heart rate classification by non-parametric Bayesian methods
- Filter design for delay-based anonymous communications
- First-person action recognition through Visual Rhythm texture description
- Fixed-point optimization of deep neural networks with adaptive step size retraining
- Flat focus: depth of field analysis for the FlatCam lensless imaging system
- Flexarray: Random phased array layouts for analytical spatial filtering
- Flexible large-scale fMRI analysis: A survey
- Flipflop correlation tracking with Convolution Kernels Networks
- Flow based botnet detection through semi-supervised active learning
- Forecasting covariance for optimal carry trade portfolio allocations
- Frequency-domain under-modelled blind system identification based on cross power spectrum and sparsity regularization
- Frequency-tuned ACM for biomedical image segmentation
- Frequency-warped time-weighted linear prediction for glottal vocoding
- From biomedical imaging to urban data mining: Theory of signal representations
- From focal stacks to tensor display: A method for light field visualization without multi-view images
- From image quality to patch quality: An Image-Patch Model for No-Reference image quality assessment
- Full-duplex relaying under I/Q imbalance using improper Gaussian signaling
- Full-duplex self-interference mitigation analysis for direct conversion RF nonlinear MIMO channel models with IQ mismatch
- Fully adaptive mode decomposition from time-frequency ridges
- Fully complex deep neural network for phase-incorporating monaural source separation
- Fused estimation of sparse connectivity patterns from rest fMRI
- Fusing shallow and deep learning for bioacoustic bird species classification
- Fusing structure from motion and lidar for dense accurate depth map estimation
- Fusing transcription results from polyphonic and monophonic audio for singing melody transcription in polyphonic music
- Fusion of multiple emotion perspectives: Improving affect recognition through integrating cross-lingual emotion information
- GDspike: An accurate spike estimation algorithm from noisy calcium fluorescence signals
- Game theoretic resource allocation form-dependent channels with application to OFDMA
- General scale interpolation via context-aware autoregressive model and multiplanar constraint
- Generalization of spoofing countermeasures: A case study with ASVspoof 2015 and BTAS 2016 corpora
- Generalized Barankin-type lower bounds for misspecified models
- Generalized Linear Models for count time series
- Generative adversarial network-based postfilter for statistical parametric speech synthesis
- Geometry-adapted Gaussian random field regression
- Global behavior of parallel projection method for certain nonconvex feasibility problems
- Globally optimal beamforming design for downlink CoMP transmission with limited backhaul capacity
- Good features to track for RGBD images
- Gradient magnitude similarity deviation on multiple scales for color image quality assessment
- Gradient-based solution for hybrid precoding in MIMO systems
- Graph Fourier Transform for directed graphs based on Lovász extension of min-cut
- Graph learning under sparsity priors
- Graph regularised tensor factorisation of EEG signals based on network connectivity measures
- Graph-signal reconstruction and blind deconvolution for diffused sparse inputs
- Grasp: A matlab toolbox for graph signal processing
- Greedy alternative for room geometry estimation from acoustic echoes: A subspace-based method
- Greedy search for descriptive spatial face features
- Gridless compressed sensing under shift-invariant sampling
- Group-level support recovery guarantees for group lasso estimator
- Guided deep network for depth map super-resolution: How much can color help?
- HEVC-based motion compensated joint temporal-spatial video denoising
- Hand pose recognition in First Person Vision through graph spectral analysis
- Hardware and software for reproducible research in audio array signal processing
- Hardware-based linear programming decoding via the alternating direction method of multipliers
- Harmonic feature fusion for robust neural network-based acoustic modeling
- Harmonic minimum mean squared error filters for multichannel speech enhancement
- Harnessing neural networks: A random matrix approach
- Hearing in a shoe-box: Binaural source position and wall absorption estimation using virtually supervised learning
- Heartmate: automated integrated anomaly analysis for effective remote cardiac health management
- Heuristic methods for designing unimodular code sequences with performance guarantees
- Hierarchical Structured Dictionary Learning for image categorization
- Hierarchical joint-guided networks for semantic image segmentation
- Hierarchical saliency optimization
- High accuracy event detection for Non-Intrusive Load Monitoring
- High dimensional decomposition of coherent/structured matrices via sequential column/row sampling
- High frequency moments via max-stability
- High level synthesis of Smith-Waterman dataflow implementations
- High precision robust modeling of long room responses using wavelet transform
- High-level synthesis implementation of HEVC 2-D DCT/DST on FPGA
- High-resolution Direction-of-Arrival estimation in SNR and snapshot challenged scenarios using multi-frequency coprime arrays
- Homography-based low rank approximation of light fields for compression
- How little does non-exact recovery help in group testing?
- How should we evaluate supervised hashing?
- Human action recognition using Adaptive Hierarchical Depth Motion Maps and Gabor filter
- Human interaction recognition using low-rank matrix approximation and super descriptor tensor decomposition
- Human recognition from photoplethysmography (PPG) based on non-fiducial features
- Hybrid beamforming design with finite-resolution phase-shifters for frequency selective massive MIMO channels
- Hybrid beamforming for large-scale MIMO systems using uplink-downlink duality
- Hybrid beamforming in uplink massive MIMO systems in the presence of blockers
- Hybrid precoding using long-term channel statistics for massive MIMO systems
- Hyperarticulation detection in repetitive voice queries using pairwise comparison for improved speech recognition
- Hyperspectral image restoration by Hybrid Spatio-Spectral Total Variation
- Hyperspectral unmixing with endmember variability using Partial Membership Latent Dirichlet Allocation
- Hypothesis testing in the presence of maxwell's daemon: signal detection by unlabeled observations
- ICA based single microphone Blind Speech Separation technique using non-linear estimation of speech
- Identifying FMRI dynamic connectivity states using affinity propagation clustering method: Application to schizophrenia
- Identifying a multiple plane plenoptic function from a swiped image
- Identifying correlated components in high-dimensional multivariate Gaussian models
- Identifying directional connections in brain networks via multi-kernel granger models
- Illumination-robust face recognition with Block-based Local Contrast Patterns
- Image classification: A hierarchical dictionary learning approach
- Image co-saliency detection via locally adaptive saliency map fusion
- Image compression with Stochastic Winner-Take-All Auto-Encoder
- Image denoising via collaborative support-agnostic recovery
- Image denoising via group sparsity residual constraint
- Image formation methods in quantitative acoustic microscopy
- Image recognition based on discriminative models using features generated from separable lattice HMMS
- Image reconstruction from partial Fourier measurements via curl constrained sparse gradient estimation
- Image retrieval based on deep Convolutional Neural Networks and binary hashing learning
- Impact of low-precision deep regression networks on single-channel source separation
- Implementation of efficient, low power deep neural networks on next-generation intel client platforms
- Implementation strategies of the seismic Full Waveform Inversion
- Improved Local Spectral Unmixing of hyperspectral data using an algorithmic regularization path for collaborative sparse regression
- Improved cepstra minimum-mean-square-error noise reduction algorithm for robust speech recognition
- Improved eigenvalue shrinkage using weighted Chebyshev polynomial approximation
- Improved template based chord recognition using the CRP feature
- Improving audio-visual speech recognition using deep neural networks with dynamic stream reliability estimates
- Improving latency-controlled BLSTM acoustic models for online speech recognition
- Improving mesh-based motion compensation by using edge adaptive graph-based compensated wavelet lifting for medical data sets
- Improving music source separation based on deep neural networks through data augmentation and network blending
- Improving the perceptual quality of ideal binary masked speech
- Improving the spatial dimensionality of Gauss-Legendre and equiangular sampling schemes on the sphere
- In-situ calibration of accelerometers in body-worn sensors using quiescent gravity
- Inband full-duplex radio access system with self-backhauling: transmit power minimization under QOS requirements
- Incident field recovery for an arbitrary-shaped scatterer
- Incremental adaptation using active learning for acoustic emotion recognition
- Indoor mapping using MIMO radio channel measurements
- Indoor multi-sound source localization based on nonparametric Bayesian clustering
- Induced bias in attenuation measurements taken from commercial microwave links
- Inference Machines for supervised Bluetooth localization
- Inferring emotions from heterogeneous social media data: A Cross-media Auto-Encoder solution
- Inferring latent states in a network influenced by neighbor activities: An undirected generative approach
- Inferring sparse graphs from smooth signals with theoretical guarantees
- Infinite-dimensional SVD for analyzing microphone array
- Infomax-ICA using Hessian-free optimization
- Information diffusion in interconnected heterogeneous networks
- Information geometry metric for random signal detection in large random sensing systems
- Information theoretic structure learning with confidence
- Informed source separation via compressive graph signal sampling
- Infrasonic scene fingerprinting for authenticating speaker location
- Infrastructure-less indoor localization using light fingerprints
- Inpainting-based error concealment for low-delay video communication
- Integrated DNN-based model adaptation technique for noise-robust speech recognition
- Integrating DNN-based and spatial clustering-based mask estimation for robust MVDR beamforming
- Integration of multiple genomic imaging data for the study of schizophrenia using joint nonnegative matrix factorization
- Intelligent compressive data gathering using data ferries for wireless sensor networks
- Inter dataset variability modeling for speaker recognition
- Inter-block dependencies consideration for intra coding in H.264/AVC and HEVC standards
- Interaural time delay personalisation using incomplete head scans
- Interference alignment on MIMO X channel with synergistic CSIT
- Interference cancellation in two-channel nuclear quadrupole resonance measurements
- Interference reduction in music recordings combining Kernel Additive Modelling and Non-Negative Matrix Factorization
- Interpretable human action recognition in compressed domain
- Interpretable phonological features for clinical applications
- Intra-class covariance adaptation in PLDA back-ends for speaker verification
- Introducing complex functional link polynomial filters
- Investigations on byte-level convolutional neural networks for language modeling in low resource speech recognition
- Iterative beam alignment algorithms for TDD MIMO systems
- Iterative block tensor singular value thresholding for extraction of lowrank component of image data
- Iterative diffusion-based anomaly detection
- Jamming Massive MIMO using Massive MIMO: Asymptotic separability results
- Jamming resistant receivers for massive MIMO
- Jazz: A companion to music for frequency estimation with missing data
- Jeffrey's divergence between moving-average and autoregressive models
- Joint Bayesian Gaussian Discriminant Analysis for speaker verification
- Joint CTC-attention based end-to-end speech recognition using multi-task learning
- Joint Near-End Listening Enhancement and far-end noise reduction
- Joint alpha-fairness based DSM and user encoding ordering for zero-forcing nonlinear precoding in G. fast downstream transmission
- Joint analog and digital self-interference cancellation and full-duplex system performance
- Joint channel and carrier frequency estimation for M-ary CPM over frequency-selective channel using PAM decomposition
- Joint modeling of articulatory and acoustic spaces for continuous speech recognition tasks
- Joint optimisation of tandem systems using Gaussian mixture density neural network discriminative sequence training
- Joint parameter and state estimation for wave-based imaging and inversion
- Joint power and subcarrier allocation for multicarrier full-duplex systems
- Joint transmit beamforming optimization and uplink/downlink user selection in a full-duplex multi-user MIMO system
- Jointly optimized transform domain temporal prediction and sub-pixel interpolation
- Kalman filter based system identification exploiting the decorrelation effects of linear prediction
- Kernel least mean square based on conjugate gradient
- Kernel principal component analysis of the ear morphology
- Kernel weighted Fisher sparse analysis on multiple maps for audio event recognition
- Key frames extraction using graph modularity clustering for efficient video summarization
- Knowledge distillation across ensembles of multilingual models for low-resource languages
- Knowledge distillation for small-footprint highway networks
- LBP edge-mapped descriptor using MGM interest points for face recognition
- LDA-based context dependent recurrent neural network language model using document-based topic distribution of words
- LDPC code design for Gaussian multiple-access channels using dynamic EXIT chart analysis
- LPCV: Learning projections from corresponding views for person re-identification
- Laplace gradient based Discriminative and Contrast Invertible descriptor
- Laplace mixtures models for efficient compressed sensing with side information
- Large scale 2D spectral compressed sensing in continuous domain
- Large-scale audio event discovery in one million YouTube videos
- Large-scale nonconvex stochastic optimization by Doubly Stochastic Successive Convex approximation
- Largest center-specific margin for dimension reduction
- Late reverberant power spectral density estimation based on an eigenvalue decomposition
- Latent tree approximation in linear model
- Learning Grassmann manifolds for object state discovery
- Learning a hierarchical spatio-temporal model for human activity recognition
- Learning and free energies for vector approximate message passing
- Learning and inferring human actions with temporal pyramid features based on conditional random fields
- Learning by networked agents under partial information
- Learning complex-valued latent filters with absolute cosine similarity
- Learning concepts through conversations in spoken dialogue systems
- Learning conditional independence structure for high-dimensional uncorrelated vector processes
- Learning cross-lingual knowledge with multilingual BLSTM for emphasis detection with limited training data
- Learning deep vector regression model for no-reference image quality assessment
- Learning dictionary for efficient signal compression
- Learning discriminative features from electroencephalography recordings by encoding similarity constraints
- Learning environmental sounds with end-to-end convolutional neural network
- Learning online alignments with continuous rewards policy gradient
- Learning representations of emotional speech with deep convolutional generative adversarial networks
- Learning rotation invariance in deep hierarchies using circular symmetric filters
- Learning sparse graphs under smoothness prior
- Learning spectrum opportunities in non-stationary radio environments
- Learning time varying graphs
- Learning to invert: Signal recovery via Deep Convolutional Networks
- Learning utterance-level representations for speech emotion and age/gender recognition using deep neural networks
- Least 1-norm pole-zero modeling with sparse deconvolution for speech analysis
- Leveraging manifold learning for extractive broadcast news summarization
- Line detection in speckle images using Radon transform and ℓ1 regularization
- Linear Discriminant Analysis with few training data
- Linear demixed domain multichannel nonnegative matrix factorization for speech enhancement
- Linear systems approach to identifying performance bounds in indirect imaging
- Listening-area-informed sound field reproduction based on circular harmonic expansion
- Local detection and estimation of multiple objects from images with overlapping observation areas
- Local trilateral upsampling for thermal image
- Locality Sensitive Hashing based deepmatching for optical flow estimation
- Localization of multiple sources using time-difference of arrival measurements
- Locally linear embedded sparse coding for image representation
- Location-aware network operation for cloud radio access network
- LogNet: Energy-efficient neural networks using logarithmic computation
- Lombard speech synthesis using long short-term memory recurrent neural networks
- Long-term non-contact tracking of caged rodents
- Low Dimensional Deep Features for facial landmark alignment
- Low angle direction of arrival estimation by time reversal
- Low light image enhancement based on two-step noise suppression
- Low rank phase retrieval
- Low-complexity optimization for two-dimensional direction-of-arrival estimation via decoupled atomic norm minimization
- Low-latency real-time blind source separation for hearing aids based on time-domain implementation of online independent vector analysis with truncation of non-causal components
- Low-rank and sparse soft targets to learn better DNN acoustic models
- Low-rank physical model recovery from low-rank signal approximation
- Low-resource grapheme-to-phoneme conversion using recurrent neural networks
- Lyric recognition in monophonic singing using pitch-dependent DNN
- Machine learning based non-intrusive quality estimation with an augmented feature set
- Making and gaming in signal processing classes
- Malware classification with LSTM and GRU language models and a character-level CNN
- Maritime anomaly detection in ferry tracks
- Massive MIMO processing at the semiconductor edge: Exploiting the system and circuit margins for power savings
- Massive device activity detection by approximate message passing
- Matched subspace detection using compressively sampled data
- Matrix completion based MIMO radars with clutter and interference mitigation via transmit precoding
- Matrix completion of noisy graph signals via proximal gradient minimization
- Maximum secrecy rate in inhomogeneous poisson networks
- Measurement of 2D vibration modes using amplification of high speed video in the presence of noise
- Measurement of sound fields using moving microphones
- Measuring, modelling and predicting perceived reverberation
- Meeting different QoS requirements of vehicular networks: A D2D-based approach
- Melody extraction and detection through LSTM-RNN with harmonic sum loss
- Memory visualization for gated recurrent neural networks in speech recognition
- Minimum Bayes risk training of CTC acoustic models in maximum a posteriori based decoding framework
- Minimum entropy pursuit: Noise analysis
- Minimum mean square deviation in ZA-NLMS algorithm
- Minimum number of possibly non-contiguous samples to distinguish two periods
- Minimum precision requirements for the SVM-SGD learning algorithm
- Minimum probability-of-error perturbation precoding for the one-bit massive MIMO downlink
- Mismatched sparse denoiser requires overestimating the support length
- Mixture source identification in non-stationary data streams with applications in compression
- Mobile phone clustering from acquired speech recordings using deep Gaussian supervector and spectral clustering
- Model based binaural enhancement of voiced and unvoiced speech
- Model order selection for sampling FRI signals
- Modeling Sallen-Key audio filters in the Wave Digital domain
- Modeling interest-based social networks: Superimposing Erdős-Rényi graphs over random intersection graphs
- Modification on LSA speech enhancement for speech recognition
- Modified nonnegative matrix factorization for endmember spectra extraction from highly mixed hyperspectral images combined with multispectral data
- Monte Carlo exploration for active binaural localization
- Mood detection from daily conversational speech using denoising autoencoder and LSTM
- Morph-to-word transduction for accurate and efficient automatic speech recognition and keyword search
- Motion clustering with hybrid-sample-based foreground segmentation for moving cameras
- Motion compensated frame rate up-conversion using 3D frequency selective extrapolation and a multi-layer consistency check
- Motion informed audio source separation
- Moving target localization in multistatic sonar using time delays, Doppler shifts and arrival angles
- Multi-accent speech recognition with hierarchical grapheme based models
- Multi-armed bandits in multi-agent networks
- Multi-channel noise reduction for hands-free voice communication on mobile phones
- Multi-channel signal enhancement with speech and noise covariance estimates computed by a probabilistic localization model
- Multi-pitch estimation using semidefinite programming
- Multi-pitch streaming of interwoven streams
- Multi-rate polar codes for solid state drives
- Multi-scale higher order singular value decomposition (MS-HoSVD) for resting-state FMRI compression and analysis
- Multi-scale spot segmentation with selection of image scales
- Multi-speaker conversations, cross-talk, and diarization for speaker recognition
- Multi-speaker voice activity detection by an improved multiplicative non-negative independent component analysis with sparseness constraints
- Multi-task deep neural network with shared hidden layers: Breaking down the wall between emotion representations
- Multi-task learning for face identification and attribute estimation
- Multi-task learning of structured output layer bidirectional LSTMS for speech synthesis
- Multi-view representation learning via gcca for multimodal analysis of Parkinson's disease
- Multichannel audio source separation: Variational inference of time-frequency sources from time-domain observations
- Multicore distributed dictionary learning: A microarray gene expression biclustering case study
- Multilayer sensor network for information privacy
- Multimodal sparse Bayesian dictionary learning applied to multimodal data classification
- Multiple illumination phaseless super-resolution (MIPS) with applications to phaseless DoA estimation and diffraction imaging
- Multiple parallel branch with folding architecture for multichannel filtered-x least mean square algorithm
- Multiple particle filtering for inference in the presence of state correlation of unknown mixing parameters
- Multiple sound source localization based on TDOA clustering and multi-path matching pursuit
- Multiple source localization using Estimation Consistency in the Time-Frequency domain
- Multiple subspace matching pursuit for spectrum sensing
- Multiple wavelength sensing array design
- Multiple-input multiple-output (MIMO) MRI: An efficient pulse design algorithm to combine parallel excitation and parallel imaging
- Multiprocessor approximate message passing with column-wise partitioning
- Multisensor detection of improper signals in improper noise
- Multistream quickest change detection: Asymptotic optimality under a sparse signal
- Multisymbol with memory noncoherent detection of CPFSK
- Multivariate Linear Time-Frequency modeling and adaptive robust target detection in highly textured monovariate SAR image
- Multivariate Scale mixtures for joint sparse regularization in multi-task learning
- Multivariate scale-free dynamics: Testing fractal connectivity
- Music staging AI
- NAPLib: An open source toolbox for real-time and offline Neural Acoustic Processing
- NIQSV: A no reference image quality assessment metric for 3D synthesized views
- Near-optimal sample complexity bounds for circulant binary embedding
- Nesterov-based parallel algorithm for large-scale nonnegative tensor factorization
- Network architectures for multilingual speech representation learning
- Network discovery using content and homophily
- Network topology inference from non-stationary graph signals
- Network-based genome wide study of hippocampal imaging phenotype in Alzheimer's Disease to identify functional interaction modules
- Neural decoding systems using Markov Decision Processes
- New analysis of radar micro-Doppler gait signatures for rehabilitation and assisted living
- New asymptotic properties for the robust ANMF
- New residue arithmetic based Barrett algorithms: Modular polynomial computations
- Node embedding for network community discovery
- Noise detection in smartphone phonocardiogram
- Noise enhanced distributed Bayesian estimation
- Noisy objective functions based on the f-divergence
- Non-blind image deconvolution using deep dual-pathway rectifier neural network
- Non-convex consensus ADMM for satellite precoder design
- Non-convex shredded signal reconstruction via sparsity enhancement
- Non-invasive gearbox fault diagnosis using scattering transform of acoustic emission
- Non-iterative impulse response shortening method for system latency reduction
- Non-negative matrix factorization of signals with overlapping events for event detection applications
- Non-orthogonal constrained independent vector analysis: Application to data fusion
- Non-parallel voice conversion using i-vector PLDA: towards unifying speaker verification and transformation
- Non-parametric analog Joint Source Channel Coding for amplify-and-forward two-hop networks
- Non-parametric spectrum cartography using adaptive radial basis functions
- Non-separable quadruple lifting structure for four-dimensional integer Wavelet Transform with reduced rounding noise
- Noncontact respiration monitoring of multiple closely positioned patients using ultra-wideband array radar with adaptive beamforming technique
- Nonparametric learning for Hidden Markov Models with preferential attachment dynamics
- Normal-to-shouted speech spectral mapping for speaker recognition under vocal effort mismatch
- Novel Amplitude Scaling method for bilinear frequency Warping-based Voice Conversion
- Novel medical video compression methods over lossless HEVC coder
- Novelty detection for predicting falls risk using smartphone gait data
- Null-steering beamformer for acoustic feedback cancellation in a multi-microphone earpiece optimizing the maximum stable gain
- Numerical filtering of linear state-space models with Markov switching
- ORGB: Offset correction in RGB color space for illumination-robust image processing
- Object detection refinement using Markov random field based pruning and learning based rescoring
- Objective assessment of pathological speech using distribution regression
- Objective characterization of audio signal quality: Applications to music collection description
- Omnidirectional bats, point-to-plane distances, and the price of uniqueness
- On DNN posterior probability combination in multi-stream speech recognition for reverberant environments
- On TOA estimation of vibration signals for localizing impacts on solid surfaces
- On classification of distorted images with deep convolutional neural networks
- On classification of environmental acoustic data using crowds
- On cognitive radio systems with directional antennas and imperfect spectrum sensing
- On methods for privacy-preserving energy disaggregation
- On mitigation of pilot spoofing attack
- On mutual coupling for ULAs: Estimating AoAs in the presence of more coupling parameters
- On random weights for texture generation in one layer CNNS
- On relationships between amplitude and phase of short-time Fourier transform
- On saturation of the Cramér Rao Bound for Sparse Bayesian Learning
- On spatial dependency in molecular distributed detection
- On spectrogram local maxima
- On the bias of pseudolinear estimators for time-of-arrival based localization
- On the impact of non-modal phonation on phonological features
- On the information rate of speech communication
- On the robustness of constrained convolutional neural networks to JPEG post-compression for image resampling detection
- On the role of head motion in affective expression
- On the security of block scrambling-based ETC systems against jigsaw puzzle solver attacks
- On time-frequency mask estimation for MVDR beamforming with application in robust speech recognition
- One-bit sparse array DOA estimation
- Online Empirical Mode Decomposition
- Online action detection and forecast via Multitask deep Recurrent Neural Networks
- Online environmental adaptation of CNN-based acoustic models using spatial diffuseness features
- Online learning of time-frequency patterns
- Online secondary path modelling in wave-domain active noise control
- Optical Tomography based on a nonlinear model that handles multiple scattering
- Optical-flow features empirical mode decomposition for motion anomaly detection
- Optimal achievable rate trade-off in cooperative cognitive radio systems
- Optimal biased estimation using Lehmann-unbiasedness
- Optimal low-rank Dynamic Mode Decomposition
- Optimal sparse L1-norm principal-component analysis
- Optimal transmit strategy for MIMO channels with joint sum and per-antenna power constraints
- Optimization of compound regularization parameters based on Stein's unbiased risk estimate
- Optimization over directed graphs: Linear convergence rate
- Optimized compressive sensing-based direction-of-arrival estimation in massive MIMO
- Optimizing Non Constant Luminance into Constant Luminance for High Dynamic Range Video Distribution
- Optimizing neural-network supported acoustic beamforming by algorithmic differentiation
- Optimizing speaker-specific filter banks for speaker verification
- Optimum array configurations of maximum output SNR for quiescent beamforming
- Orthogonal precoding for sidelobe suppression in DFT-based systems using block reflectors
- Overlapping sound event detection with supervised Nonnegative Matrix Factorization
- P-leader multifractal analysis for text type identification
- POKEMON: A non-linear beamforming algorithm for 1-bit massive MIMO
- PPG-based heart rate estimation using Wiener filter, phase vocoder and Viterbi decoding
- Pairwise learning using multi-lingual bottleneck features for low-resource query-by-example spoken term detection
- Parallel phonetically aware DNNs and LSTM-RNNS for frame-by-frame discriminative modeling of spoken language identification
- Parallelized Stochastic Gradient Markov Chain Monte Carlo algorithms for non-negative matrix factorization
- Parameter-free Plug-and-Play ADMM for image restoration
- Parameter-free automated extraction of neuronal signals from calcium imaging data
- Parametric estimation of spectrum driven by an exogenous signal
- Parametrized design of the generalized sequential probability ratio test
- Parsimonious Online Learning with Kernels via sparse projections in function space
- Part-level fully convolutional networks for pedestrian detection
- Partial image blur detection and segmentation from a single snapshot
- Particle PHD filter based multi-target tracking using discriminative group-structured dictionary learning
- Particle flow SMC delta-GLMB filter
- Particle flow for sequential Monte Carlo implementation of probability hypothesis density
- Partitioned Hierarchical alternating least squares algorithm for CP tensor decomposition
- Partitioned inverse image reconstruction for millimeter-wave SAR imaging
- Patch-based multiple view image denoising with occlusion handling
- Patch-based segmentation of overlapping cervical cells using active contour with local edge information
- Pattern recognition of functional brain networks
- Peak load minimization in load coupled interference networks
- Penalty dual decomposition method with application in signal processing
- Perceptual evaluation of a multiband acoustic crosstalk canceler using a linear loudspeaker array
- Performance analysis for time-of-arrival estimation with oversampled low-complexity 1-bit a/d conversion
- Performance analysis of (TDD) massive MIMO with Kalman channel prediction
- Performance analysis of an AoA estimator in the presence of more mutual coupling parameters
- Performance analysis of coarray-based MUSIC and the Cramér-Rao bound
- Performance bounds for Poisson compressed sensing using Variance Stabilization Transforms
- Performance of time delay estimation in a cognitive radar
- Performance trade-off in an adaptive IEEE 802.11AD waveform design for a joint automotive radar and communication system
- Permutation invariant training of deep models for speaker-independent multi-talker speech separation
- Personalized acoustic modeling by weakly supervised multi-task deep learning using acoustic tokens discovered from unlabeled data
- Personalized video emotion tagging through a topic model
- Personalized video preference estimation based on early fusion using multiple users' viewing behavior
- Perturbation analysis of Joint Eigenvalue Decomposition Algorithms
- Phase Congruency for image understanding with applications in computational seismic interpretation
- Phase estimation in single-channel speech enhancement using phase invariance constraints
- Phase reconstruction method based on time-frequency domain harmonic structure for speech enhancement
- Phase retrieval from STFT measurements via non-convex optimization
- Phase retrieval with a multivariate Von Mises prior: From a Bayesian formulation to a lifting solution
- Phase unmixing: Multichannel source separation with magnitude constraints
- Phase-dependent anisotropic Gaussian model for audio source separation
- Phaseless super-resolution in the continuous domain
- Phonological content impact on wrongful convictions in Forensic Voice Comparison context
- Pickup position and plucking point estimation on an electric guitar
- Pilot precoding and combining in multiuser MIMO networks
- Pitch contour tracking in music using Harmonic Locked Loops
- Pitch-based non-intrusive objective intelligibility prediction
- Polarimetric radar crosstalk removal during sparse image formation
- Polarization spectrogram of bivariate signals
- Polyphonic piano note transcription with non-negative matrix factorization of differential spectrogram
- Portable modeling of virtual physics for audio and haptic interaction design
- Pose-based composition improvement for portrait photographs
- Post-ICA phase de-noising for resting-state complex-valued FMRI data
- Power-law stochastic neighbor embedding
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.