ICASSP 2018 Accepted Papers
The full list of 1,393 papers accepted at ICASSP 2018 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- (Almost) Zero-Shot Cross-Lingual Spoken Language Understanding
- 2.5D Multizone Reproduction Using Weighted Mode Matching
- 2D Vector Map Reversible Data Hiding with Topological Relation Preservation
- 3-D CNN Models for Far-Field Multi-Channel Speech Recognition
- 3D Exterior Soundfield Reproduction Using a Planar Loudspeaker Array
- 3D Image Reconstruction from Multi-Focus Microscope: Axial Super-Resolution and Multiple-Frame Processing
- 3D Mouth Tracking from a Compact Microphone Array Co-Located with a camera
- 3D-Hog Embedding Frameworks for Single and Multi-Viewpoints Action Recognition Based on Human Silhouettes
- A 203 FPS VLSI Architecture of Improved Dense Trajectories for Real-Time Human Action Recognition
- A 320M Pixel/S Vlsi Architecture Design of Weighted Mode Filter for 4K Ultra-Hd Depth Upsampling
- A Bayesian Framework to Optimize Double Band Spectra Spatial Filters for Motor Imagery Classification
- A Bayesian Hierarchical Model for Speech Enhancement
- A Bayesian Model for Joint Unmixing and Robust Classification of Hyperspectral Images
- A Capitalist Scheme for Energy Management in Inferential Sensor Networks
- A Captcha Design Based on Visual Reasoning
- A Casa Approach to Deep Learning Based Speaker-Independent Co-Channel Speech Separation
- A Cascaded Framework for Model-Based 3D Face Reconstruction
- A Comparison of Recent Waveform Generation and Acoustic Modeling Methods for Neural-Network-Based Speech Synthesis
- A Complete End-to-End Speaker Verification System Using Deep Neural Networks: From Raw Signals to Verification Result
- A Compressive Sensing-Based Active User and Symbol Detection Technique for Massive Machine-Type Communications
- A Constant Step Stochastic Douglas-Rachford Algorithm with Application to non Separable Regularizations
- A Conversational Neural Language Model for Speech Recognition in Digital Assistants
- A Corrective Learning Approach for Text-Independent Speaker Verification
- A Coupled Compressive Sensing Scheme for Unsourced Multiple Access
- A Deep Dictionary Model for Image Super-Resolution
- A Deep Encoder-Decoder Networks for Joint Deblurring and Super-Resolution
- A Deep Learning Based Alternative to Beamforming Ultrasound Images
- A Deep Learning Based No-Reference Image Quality Assessment Model for Single-Image Super-Resolution
- A Deep Neural Network Approach for Time-Of- Arrival Estimation in Multipath Channels
- A Deep Neural Network Based Method of Source Localization in a Shallow Water Environment
- A Deep Reinforcement Learning Framework for Identifying Funny Scenes in Movies
- A Deeper Look at Gaussian Mixture Model Based Anti-Spoofing Systems
- A Deeply-Recursive Convolutional Network For Crowd Counting
- A Dimension-Independent Discriminant Between Distributions
- A Discrete Signal Processing Framework for Set Functions
- A Discriminatively Learned Feature Embedding Based on Multi-Loss Fusion For Person Search
- A Distance-Based Formulation for Sampling Signals on Graphs
- A Diverse Large-Scale Dataset for Evaluating Rebroadcast Attacks
- A Dynamic Latent Variable Model for Source Separation
- A Family of Matrices for Generating Hermite-Gaussian-Like DFT Eigenvectors
- A Fast and Memory-Efficient Algorithm for Robust PCA (MEROP)
- A Feature Fusion Method Based on Extreme Learning Machine for Speech Emotion Recognition
- A Flexible Dirty Model Dictionary Learning Approach for Classification
- A Fully Convolutional Tri-Branch Network (FCTN) for Domain Adaptation
- A Generalized Uncorrelated Ridge Regression with Nonnegative Labels for Unsupervised Feature Selection
- A Generative Adversarial Network Based Framework for Unsupervised Visual Surface Inspection
- A Generative Auditory Model Embedded Neural Network for Speech Processing
- A Graph-CNN for 3D Point Cloud Classification
- A Greedy Pursuit Algorithm for Separating Signals from Nonlinear Compressive Observations
- A Hybrid Approach to Combining Conventional and Deep Learning Techniques for Single-Channel Speech Enhancement and Recognition
- A Hybrid Dictionary Approach for Distributed Kernel Adaptive Filtering in Diffusion Networks
- A Hybrid Neural Network Based on the Duplex Model of Pitch Perception for Singing Melody Extraction
- A Joint Detection and Reconstruction Method for Blind Graph Signal Recovery
- A Joint Multi-Task Learning Framework for Spoken Language Understanding
- A Joint Perspective of Periodically Excited Efficient NLMS Algorithm and Inverse Cyclic Convolution
- A Joint Separation-Classification Model for Sound Event Detection of Weakly Labelled Data
- A Joint Source Channel Arithmetic Map Decoder Using Probabilistic Relations Among Intra Modes in Predictive Video Compression
- A Joint Target Localization and Classification Framework for Sensor Networks
- A Large-Scale Study of Language Models for Chord Prediction
- A Learning Algorithm with Compression-Based Regularization
- A Light-Weight Multimodal Framework for Improved Environmental Audio Tagging
- A Low Power Hardware Implementation of Multi-Object DPM Detector for Autonomous Driving
- A Low-Complexity Video Encoder for Equirectangular Projected 360 Video Content
- A Matrix Completion Approach for Wall-Clutter Mitigation in Compressive Radar Imaging of Indoor Targets
- A Modified Signal Phase Unwrapping Algorithm for Range Estimation
- A Motion Aided Merge Mode For Hevc
- A Multi-Camera Deep Neural Network for Detecting Elevated Alertness in Drivers
- A Multi-Resolution Approach to Complexity Reduction in Tomographic Reconstruction
- A Multi-Seed 3D Local Graph Matching Model for Tracking of Densely Packed Cells
- A Natural Shape-Preserving Stereoscopic Image Stitching
- A New Proximal Method for Joint Image Restoration and Edge Detection with the Mumford-Shah Model
- A Nonconvex Variational Approach for Robust Graphical Lasso
- A Nonlinear 3D Geometric Tongue Model
- A Novel Crowd-Resilient Visual Localization Algorithm Via Robust Pca Background Extraction
- A Novel Ego-Noise Suppression Algorithm for Acoustic Signal Enhancement in Autonomous Systems
- A Novel Image-Specific Transfer Approach for Prostate Segmentation in MR Images
- A Novel Joint Radar and Communication System Based on Randomized Partition of Antenna Array
- A Novel LSTM-Based Speech Preprocessor for Speaker Diarization in Realistic Mismatch Conditions
- A Novel Learnable Dictionary Encoding Layer for End-to-End Language Identification
- A Novel Method for Human Bias Correction of Continuous- Time Annotations
- A Novel Selective Active Noise Control Algorithm to Overcome Practical Implementation Issue
- A Novel Semantic Attribute-Based Feature for Image Caption Generation
- A Novel Thresholding Technique for the Denoising of Multicomponent Signals
- A Parallel Best-Response Algorithm with Exact Line Search for Nonconvex Sparsity-Regularized Rank Minimization
- A Parallel Fusion Approach to Piano Music Transcription Based on Convolutional Neural Network
- A Parametric Approach for Classification of Distortions in Pathological Voices
- A Penalized Method for the Predictive Limit of Learning
- A Practical Guide to Multi-Image Alignment
- A Pragmatic Authentication System Using Electroencephalography Signals
- A Priori SNR Estimation Using Discriminative Non-Negative Matrix Factorization
- A Pruned Rnnlm Lattice-Rescoring Algorithm for Automatic Speech Recognition
- A Quaternion Kernel Minimum Error Entropy Adaptive Filter
- A Random Matrix and Concentration Inequalities Framework for Neural Networks Analysis
- A Refined Analysis of the Gap Between Expected Rate for Partial Csit and the Massive Mimo Rate Limit
- A Reliable Video Storage Architecture in Hybrid SLC/MLC Nand Flash
- A Revisit of Action Detection Using Improved Trajectories
- A Robust Change Detector for Highly Heterogeneous Multivariate Images
- A Robust Event-Triggered Consensus Strategy for Linear Multi-Agent Systems with Uncertain Network Topology
- A Robust Hierarchical Qp Setting for Screen Content Coding
- A Robust Machine Learning Method for Cell-Load Approximation in Wireless Networks
- A Rotation-Invariant Convolutional Neural Network for Image Enhancement Forensics
- A Second-Order Variational Framework for Joint Depth Map Estimation and Image Dehazing
- A Shuffled-Based Iterative Demodulation and Decoding Scheme for Ldpc Coded Flash Memory
- A Simple Cepstral Domain DNN Approach to Artificial Speech Bandwidth Extension
- A Simple and Effective Framework for a Priori SNR Estimation
- A Single-Channel Noise Reduction Filtering/Smoothing Technique in the Time Domain
- A Sparse Coding Framework for Gaze Prediction in Egocentric Video
- A Statistical Signal Processing Approach to Clustering over Compressed Data
- A Stem Reu Site on the Integrated Design of Sensor Devices and Signal Processing Algorithms
- A Study of All-Convolutional Encoders for Connectionist Temporal Classification
- A Study of Noise PSD Estimators for Single Channel Speech Enhancement
- A Study of Training Targets for Deep Neural Network-Based Speech Enhancement Using Noise Prediction
- A Supervised Air-Tissue Boundary Segmentation Technique in Real-Time Magnetic Resonance Imaging Video Using a Novel Measure of Contrast and Dynamic Programming
- A Supervised Approach to Global Signal-to-Noise Ratio Estimation for Whispered and Pathological Voices
- A Supervised Stdp-Based Training Algorithm for Living Neural Networks
- A Tensor Decomposition Technique for Source Localization from Multimodal Data
- A Time-Restricted Self-Attention Layer for ASR
- A Time-Weighted Method for Predicting the Intelligibility of Speech in the Presence of Interfering Sounds
- A Toeplitz-Tyler Estimation of the Model Order in Large Dimensional Regime
- A Triplet-Loss Embedded Deep Regressor Network for Estimating Blood Pressure Changes Using Prosodic Features
- A Two-Layer Reinforcement Learning Solution for Energy Harvesting Data Dissemination Scenarios
- A Unified Approach to Generating Sound Zones Using Variable Span Linear Filters
- A Unified Estimator for Source Positioning and DOA Estimation Using AOA
- A Wavenet for Speech Denoising
- A Weighted Least Squares Beam Shaping Technique for Sound Field Control
- ADA-PT: An Adaptive Parameter Tuning Strategy Based on the Weighted Stein Unbiased Risk Estimator
- ASR Performance Prediction on Unseen Broadcast Programs Using Convolutional Neural Networks
- Accelerated Image Reconstruction for Nonlinear Diffractive Imaging
- Accelerating Recurrent Neural Network Language Model Based Online Speech Recognition System
- Accent Conversion Using Phonetic Posteriorgrams
- Accounting for Room Acoustics in Audio-Visual Multi-Speaker Tracking
- Achievable Rate Maximization by Passive Intelligent Mirrors
- Achieving Accompanying Beampattern Peak for High-Speed Users Via Frequency Diverse Array
- Acoustic Analysis and Assessment of the Knee in Osteoarthritis During Walking
- Acoustic Feature Learning Using Cross-Domain Articulatory Measurements
- Acoustic Modeling of Speech Waveform Based on Multi-Resolution, Neural Network Signal Processing
- Acoustic Reflector Localization and Classification
- Acoustic Scene Classification Using Discrete Random Hashing for Laplacian Kernel Machines
- Acoustic-to-Word Attention-Based Model Complemented with Character-Level CTC-Based Model
- Active Anomaly Detection in Heterogeneous Processes
- Active Camera Relocalization with RGBD Camera from a Single 2D Image
- Active Covariance Estimation by Random Sub-Sampling of Variables
- Active Occlusion Cancellation with Hear-Through Equalization for Headphones
- Adaptation of an Expressive Single Speaker Deep Neural Network Speech Synthesis System
- Adaptive Bayesian Channel Gain Cartography
- Adaptive Clustering Algorithm for Cooperative Spectrum Sensing in Mobile Environments
- Adaptive Coding of Non-Negative Factorization Parameters with Application to Informed Source Separation
- Adaptive Noise Canceller with Snr Estimate Switchover for Stepsize Control
- Adaptive Parameters Adjustment for Group Reweighted Zero-Attracting LMS
- Adaptive Permutation Invariant Training with Auxiliary Information for Monaural Multi-Talker Speech Recognition
- Adaptive STFT with Chirp-Modulated Gaussian Window
- Adaptive Sparse Array Reconfiguration based on Machine Learning Algorithms
- Adaptive Travel Time Tomography with Local Sparsity
- Adaptive Visual Target Tracking Based on Label Consistent K-Svd Sparse Coding and Kernel Particle Filter
- Advanced LSTM: A Study About Better Time Dependency Modeling in Emotion Recognition
- Advancing Acoustic-to-Word CTC Model
- Advancing Connectionist Temporal Classification with Attention Modeling
- Adversarial Advantage Actor-Critic Model for Task-Completion Dialogue Policy Learning
- Adversarial Learning of Raw Speech Features for Domain Invariant Speech Recognition
- Adversarial Multi-Agent Target Tracking with Inexact Online Gradient Descent
- Adversarial Multilingual Training for Low-Resource Speech Recognition
- Adversarial Semi-Supervised Audio Source Separation Applied to Singing Voice Extraction
- Adversarial Teacher-Student Learning for Unsupervised Domain Adaptation
- Affine-Projection Least-Mean-Magnitude-Phase Algorithms Using a Posteriori Updates
- Alpha-Stable Low-Rank Plus Residual Decomposition for Speech Enhancement
- Alternating Minimization Approach for Identification of Piecewise Continuous Hammerstein Systems
- Alternative Objective Functions for Deep Clustering
- Altitude Measurement of Low-Angle Target Under Complex Terrain Environment for Meter-Wave Radar
- An Adaptive Combination Rule for Diffusion LMS Based on Consensus Propagation
- An Algorithm for Multi Subject Fmri Analysis Based on the SVD and Penalized Rank-1 Matrix Approximation
- An Analysis of Incorporating an External Language Model into a Sequence-to-Sequence Model
- An Analytical Method to Determine Minimum Per-Layer Precision of Deep Neural Networks
- An Architecture for Self -Aware IOT Applications
- An Attenuation Adapted Pulse Compression Technique to Enhance the Bandwidth and the Resolution Using Ultrafast Ultrasound Imaging
- An Efficient Deep Convolutional Laplacian Pyramid Architecture for Cs Reconstruction At Low Sampling Ratios
- An Efficient Residual Echo Suppression for Multi-Channel Acoustic Echo Cancellation Based on the Frequency-Domain Adaptive Kalman Filter
- An Efficient Target Localization Estimator from Bistatic Range and Tdoa Measurements in Multistatic Radar
- An End-To-End Siamese Convolutional Neural Network for Loop Closure Detection in Visual Slam System
- An End-to-End Approach to Joint Social Signal Detection and Automatic Speech Recognition
- An End-to-End Language-Tracking Speech Recognizer for Mixed-Language Speech
- An Ensemble Framework of Voice-Based Emotion Recognition System for Films and TV Programs
- An Ensemble Learning Approach to Detect Epileptic Seizures from Long Intracranial EEG Recordings
- An Ensemble Learning Method Based on Random Subspace Sampling for Palmprint Identification
- An Event-Triggered Average Consensus Algorithm with Performance Guarantees for Distributed Sensor Networks
- An Experimental Analysis of the Power Consumption of Convolutional Neural Networks for Keyword Spotting
- An Improved Doa Estimator Based on Partial Relaxation Approach
- An Improved Initialization for Low-Rank Matrix Completion Based on Rank-L Updates
- An Improved Iterative Algorithm for Band-Limited Signal Extrapolation on the Sphere
- An Investigation of Noise Shaping with Perceptual Weighting for Wavenet-Based Speech Generation
- An Investigation of Subband Wavenet Vocoder Covering Entire Audible Frequency Range with Limited Acoustic Features
- An Investigation of a Knowledge Distillation Method for CTC Acoustic Models
- An Iterative Approach for Shadow Removal in Document Images
- An Open-Source Speaker Gender Detection Framework for Monitoring Gender Equality
- An Unsupervised Anomalous Event Detection Framework with Class Aware Source Separation
- An Upper-Bound on the Required Size of a Neural Network Classifier
- An 𝓁0 Solution to Sparse Approximation Problems with Continuous Dictionaries
- Analysis and Optimization of Aperture Design in Computational Imaging
- Analysis of Multilingual Blstm Acoustic Model on Low and High Resource Languages
- Anatomy-Guided Inverse-Gradient Susceptibility Artifact Correction Method for High-Resolution FMRI
- Angle Dependent Match Filter Design for Circulating Code
- Anisotropic Total Variation Regularized Low-Rank Tensor Completion Based On Tensor Nuclear Norm for Color Image Inpainting
- Anscombe Meets Hough: Noise Variance Stablization Via Parametric Model Estimation
- Antenna Selection for Large-Scale Mimo Systems with Low-Resolution Adcs
- Aphash: Anchor-Based Probability Hashing for Image Retrieval
- Application of Progressive Neural Networks for Multi-Stream Wfst Combination in One-Pass Decoding
- Applying Multitask Learning to Acoustic-Phonemic Model for Mispronunciation Detection and Diagnosis in L2 English Speech
- Approximate Belief Propagation Decoder for Polar Codes
- Articulatory Information and Multiview Features for Large Vocabulary Continuous Speech Recognition
- Assessing Cross-Dependencies Using Bivariate Multifractal Analysis
- Asymmetric Dct-Jnd for Luminance Adaptation Effects: an Application To Perceptual Video Coding in Mv-Hevc
- Asymptotic Signal Detection Rates with 1-Bit Array Measurements
- Asynchronous Blind Network Division Multiple Access
- Attention-Based Dialog State Tracking for Conversational Interview Coaching
- Attention-Based End-to-End Speech Recognition on Voice Search
- Attention-Based LSTM for Psychological Stress Detection from Spoken Language Using Distant Supervision
- Attention-Based Models for Text-Dependent Speaker Verification
- Attitude Classification in Adjacency Pairs of a Human-Agent Interaction with Hidden Conditional Random Fields
- Audio Set Classification with Attention Model: A Probabilistic Perspective
- Audio Source Separation with Magnitude Priors: The Beads Model
- Audio Style Transfer
- Audio-Visual Conversation Analysis by Smart Posterboard and Humanoid Robot
- Audio-Visual Person Recognition in Multimedia Data From the Iarpa Janus Program
- Augmented Data and Improved Noise Residual-Based CNN for Printer Source Identification
- Augmented Latent Dirichlet Allocation (Lda) Topic Model with Gaussian Mixture Topics
- Augmenting Classrooms with AI for Personalized Education
- Autoencoder Based Image Compression: Can the Learning be Quantization Independent?
- Autoencoder Inspired Unsupervised Feature Selection
- Automated Detection of High FDG Uptake Regions in CT Images
- Automatic Bird Vocalization Identification Based on Fusion of Spectral Pattern and Texture Features
- Automatic Conflict Detection in Police Body-Worn Audio
- Automatic Motion Artifact Detection for Whole-Body Magnetic Resonance Imaging
- Automatic Music Transcription Leveraging Generalized Cepstral Features and Deep Learning
- Automatic Segmentation and Cardiopathy Classification in Cardiac Mri Images Based on Deep Neural Networks
- Automatic Shrinkage Tuning Robust to Input Correlation for Sparsity-Aware Adaptive Filtering
- Automatic Speech Assessment for Aphasic Patients Based on Syllable-Level Embedding and Supra-Segmental Duration Features
- Automatic Temporal Segmentation of Hand Movements for Hand Positions Recognition in French Cued Speech
- Automatically Linking Digital Signal Processing Assessment Questions to Key Engineering Learning Outcomes
- B-Spline Pdf: A Generalization of Histograms to Continuous Density Models for Generative Audio Networks
- BSS Eval or Peass? Predicting the Perception of Singing-Voice Separation
- Bandlimited Spatiotemporal Field Sampling with Location and Time Unaware Mobile Sensors
- Bayesian Anisotropic Gaussian Model for Audio Source Separation
- Bayesian Generative Model Based on Color Histogram of Oriented Phase and Histogram of Oriented Optical Flow for Rare Event Detection in Crowded Scenes
- Bayesian Inference for Multi-Line Spectra in Linear Sensor Array
- Bayesian Models for Unit Discovery on a Very Low Resource Language
- Bayesian Sparse Signal Detection Exploiting Laplace Prior
- Beamforming Design for Full-Duplex Cellular and Mimo Radar Coexistence: A Rate Maximization Approach
- Being Low-Rank in the Time-Frequency Plane
- Benchmarking Uncertainty Estimates with Deep Reinforcement Learning for Dialogue Policy Optimisation
- Binaural Rendering of Dynamic Head and Sound Source Orientation Using High-Resolution HRTF and Retarded Time
- Binaural Spectral Complexity Reduction of Music Signals for Cochlear Implant Listeners
- Binaural Speech Source Localization Using Template Matching of Interaural Time Difference Patterns
- Bindctnet: A Simple Binary Dct Network for Image Classification
- Birdvox-Full-Night: A Dataset and Benchmark for Avian Flight Call Detection
- Bitwise Neural Networks for Efficient Single-Channel Source Separation
- Bitwise Source Separation on Hashed Spectra: An Efficient Posterior Estimation Scheme Using Partial Rank Order Metrics
- Blind Bandwidth Extension Based on Convolutional and Recurrent Deep Neural Networks
- Blind Calibration for Acoustic Vector Sensor Arrays
- Blind Estimation of the Speech Transmission Index for Speech Quality Prediction
- Blind Image Deblurring Via Reweighted Graph Total Variation
- Blind Image Quality Assessment Based on Visuo-Spatial Series Statistics
- Blind Source Separation Using Mixtures of Alpha-Stable Distributions
- Block-Coordinate Proximal Algorithms for Scale-Free Texture Segmentation
- Boosting Noise Robustness of Acoustic Model via Deep Adversarial Training
- Boundary Objectness Network for Object Detection and Localization
- Breast Density Classification with Deep Convolutional Neural Networks
- Bridgenets: Student-Teacher Transfer Learning Based on Recursive Neural Networks and Its Application to Distant Speech Recognition
- Building Competitive Direct Acoustics-to-Word Models for English Conversational Speech Recognition
- CBLDNN-Based Speaker-Independent Speech Separation Via Generative Adversarial Training
- COMPASS: Coding and Multidirectional Parameterization of Ambisonic Sound Scenes
- CPD Updating Using Low-Rank Weights
- CRoss-lingual and Multilingual Speech Emotion Recognition on English and French
- CTC Loss Function with a Unit-Level Ambiguity Penalty
- Calibrating Cameras in Poor-Conditioned Pitch-Based Sports Games
- Can you Find a Face in a HEVC Bitstream?
- Capturing Shared and Individual Information in fMRI Data
- Cascade: Channel-Aware Structured Cosparse Audio Declipper
- Catseyes: Categorizing Seismic Structures with Tessellated Scattering Wavelet Networks
- Cell Subclass Identification in Single-Cell RNA-Sequencing Data Using Orthogonal Nonnegative Matrix Factorization
- Change-Point Detection of Gaussian Graph Signals with Partial Information
- Channel Dependent Codebook Design in Spatial Modulation
- Channel Dependent Mutual Information in Index Modulations
- Characterizing Performance of Speaker Diarization Systems on Far-Field Speech Using Standard Methods
- Classification of Corals in Reflectance and Fluorescence Images Using Convolutional Neural Network Representations
- Classification vs. Regression in Supervised Learning for Single Channel Speaker Count Estimation
- Classifier Cascade to Aid in Detection of Epileptiform Transients in Interictal EEG
- Classifying Pump-Probe Images of Melanocytic Lesions Using the WEYL Transform
- Cloud Radio Access Network with Optimized Base-Station Caching
- Cluster-Based Point Cloud Coding with Normal Weighted Graph Fourier Transform
- Clustering of Data with Missing Entries
- Clustering-Guided Gp-Ucb for Bayesian Optimization
- Coarray Interpolation-Based Coprime Array Doa Estimation Via Covariance Matrix Reconstruction
- Cognitive Analysis of Working Memory Load from Eeg, by a Deep Recurrent Neural Network
- Coherence Bounds for Sensing Matrices in Spherical Harmonics Expansion
- Coherent Time Reversal Sub-Array Processing for Microwave Breast Imaging
- Color Affine Subspace Pursuit for Color Artifact Removal
- Combining Acoustic Embeddings and Decoding Features for End-of-Utterance Detection in Real-Time Far-Field Speech Recognition Systems
- Combining Multiple Deep Features for Glaucoma Classification
- Combining Range and Direction for Improved Localization
- Common and Individual Feature Extraction Using Tensor Decompositions: a Remedy for the Curse of Dimensionality?
- Community Detection from Low-Rank Excitations of a Graph Filter
- Comparative Evaluations of Various Factored Deep Convolutional Rnn Architectures for Noise Robust Speech Recognition
- Comparing the Influence of Depth and Width of Deep Neural Network Based on Fixed Number of Parameters for Audio Event Detection
- Comparison of Speech Tasks for Automatic Classification of Patients with Amyotrophic Lateral Sclerosis and Healthy Subjects
- Complementary Complex-Valued Spectrum for Real-Valued Data: Real Time Estimation of the Panorama Through Circularity-Preserving Dft
- Complementary Set Variational Autoencoder for Supervised Anomaly Detection
- Complex Evolution Recurrent Neural Networks (ceRNNs)
- Complex-Valued Gaussian Process Latent Variable Model for Phase-Incorporating Speech Enhancement
- Complexity Reduction Algorithm for Optimum Quantizer Design Based on Amplitude Sparseness
- Complexity Reduction of Eigenvalue Decomposition-Based Diffuse Power Spectral Density Estimators Using the Power Method
- Compressed Convex Spectral Embedding for Bird Species Classification
- Compressed Sensing Mask Feature in Time-Frequency Domain for Civil Flight Radar Emitter Recognition
- Compressive Regularized Discriminant Analysis of High-Dimensional Data with Applications to Microarray Studies
- Compressive Sampling of Sound Fields Using Moving Microphones
- Compressive networked storage with lazy-encoding
- Computationally Efficient Iv-Based Bias Reduction for Closed-Form Tdoa Localization
- Concatenative Articulatory Video Synthesis Using Real-Time MRI Data for Spoken Language Training
- Concave Losses for Robust Dictionary Learning
- Concurrent Clutter and Noise Suppression via Low Rank Plus Sparse Optimization for Non-Contrast Ultrasound Flow Doppler Processing in Microvasculature
- Concurrent Target Following with Active Directional Sensors
- Confidence Based Acoustic Event Detection
- Confnet: Predict with Confidence
- Consecutive Independence and Correlation Transform for Multimodal Fusion: Application to Eeg and Fmri Data
- Considerations Regarding Individualization of Head-Related Transfer Functions
- Consistent Run Selection for Independent Component Analysis: Application to Fmri Analysis
- Constant False Alarm Rate for Online one Class Svm Learning
- Constant Modulus Probing Waveform Design for Mimo Radar Via Admm Algorithm
- Constrained Bayesian Active Learning of a Linear Classifier
- Constrained Convolutional-Recurrent Networks to Improve Speech Quality with Low Impact on Recognition Accuracy
- Content Delivery Design for Cache-Aided Cloud Radio Access Network to Achieve Low Latency
- Content-Based Representations of Audio Using Siamese Neural Networks
- Context-Sensitive Deep Learning for Detection of Clustered Micro Calcifications in Mammograms
- Continuous Security in IoT Using Blockchain
- Contourlet Based Natural Scene Statistics Using Student'S T Distribution
- Contrast Enhancement Using Phase Transition Information and Total Variation
- Control of Graph Signals Over Random Time-Varying Graphs
- Convergence Analysis on a Fast Iterative Phase Retrieval Algorithm Without Independence Assumption
- Convergence of Variance-Reduced Learning Under Random Reshuffling
- Convolutional Group-Sparse Coding and Source Localization
- Convolutional Neural Network Approach for Eeg-Based Emotion Recognition Using Brain Connectivity and its Spatial Information
- Convolutional Neural Networks and Multitask Strategies for Semantic Mapping of Natural Language Input to a Structured Database
- Convolutional Sequence to Sequence Model with Non-Sequential Greedy Decoding for Grapheme to Phoneme Conversion
- Convolutional Sparse Representations with Gradient Penalties
- Convolutional-Recurrent Neural Networks for Speech Enhancement
- Cooperative Tracking Using Marginal Diffusion Particle Filters
- Correlated Tensor Factorization for Audio Source Separation
- Correlation-Based Face Detection for Recognizing Faces in Videos
- Correntropy-Based Adaptive Filtering of Noncircular Complex Data
- Cortico-Muscular Coherence Enhancement Via Sparse Signal Representation
- Cover Song Identification Using Song-to-Song Cross-Similarity Matrix with Convolutional Neural Network
- Cramér-Rao Bound for Line Constrained Trajectory Tracking
- Crepe: A Convolutional Representation for Pitch Estimation
- Crime Incidents Embedding Using Restricted Boltzmann Machines
- Critically-Sampled Graph Filter Banks with Spectral Domain Sampling
- Cross-Lingual Phoneme Mapping for Language Robust Contextual Speech Recognition
- Cross-Modal Learning to Rank with Adaptive Listwise Constraint
- Cross-Modal Message Passing for Two-Stream Fusion
- Cross-Modality Distillation: A Case for Conditional Generative Adversarial Networks
- Cross-Validated Bandwidth Selection for Precision Matrix Estimation
- Crowdsourced Pairwise-Comparison for Source Separation Evaluation
- Crowdsourcing Emotional Speech
- Cyborg Speech: Deep Multilingual Speech Synthesis for Generating Segmental Foreign Accent with Natural Prosody
- DNN Based Embeddings for Language Recognition
- DNN Based Speaker Embedding Using Content Information for Text-Dependent Speaker Verification
- DNN-Based Concurrent Speakers Detector and its Application to Speaker Extraction with LCMV Beamforming
- Data Driven Convolutional Sparse Coding for Visual Recognition
- Data Injection Attack on Decentralized Optimization
- Data-Aided Fast Beamforming Selection for 5G
- Data-Driven Multi-Channel Filter Design with Peak-Interference Suppression for Threshold-Based Spike Sorting in High-Density Neural Probes
- Data-Driven Nonparametric Hypothesis Testing
- Decentralized Load Balancing in Mobile Communication Networks
- Deep Attractor Networks for Speaker Re-Identification and Blind Source Separation
- Deep Blind Image Quality Assessment by Learning Sensitivity Map
- Deep CNN Based Feature Extractor for Text-Prompted Speaker Recognition
- Deep Clustering with Gated Convolutional Networks
- Deep Factorization for Speech Signal
- Deep Feature Embedding Learning for Person Re-Identification Using Lifted Structured Loss
- Deep Feed-Forward Sequential Memory Networks for Speech Synthesis
- Deep Geometric Matrix Completion: A New Way for Recommender Systems
- Deep Image Super Resolution via Natural Image Priors
- Deep Layer Prior Optimization for Single Image Rain Streaks Removal
- Deep Learning Based Speech Beamforming
- Deep Learning for Accelerated Ultrasound Imaging
- Deep Learning for Frame Error Probability Prediction in BICM-OFDM Systems
- Deep Learning for Joint Source-Channel Coding of Text
- Deep Learning for Predicting Image Memorability
- Deep Mul Timodal Learning for Emotion Recognition in Spoken Language
- Deep Neural Network Based Discriminative Training for I-Vector/PLDA Speaker Verification
- Deep Residual Learning for Model-Based Iterative CT Reconstruction Using Plug-and-Play Framework
- Deep Residual Learning for Small-Footprint Keyword Spotting
- Deep Stock Representation Learning: From Candlestick Charts to Investment Decisions
- Deep Transfer Learning for EEG-Based Brain Computer Interface
- Deep Uniqueness-Aware Hashing for Fine-Grained Multi-Label Image Retrieval
- Deep Word Embeddings for Visual Speech Recognition
- Deep-FSMN for Large Vocabulary Continuous Speech Recognition
- Deepcasd: An End-to-End Approach for Multi-Spectral Image Super-Resolution
- Deeptongue: Tongue Segmentation Via Resnet
- Defending Against Packet-Size Side-Channel Attacks in Iot Networks
- Deformation Stability of Deep Convolutional Neural Networks on Sobolev Spaces
- Delta-Sigma Modulators for Constant Envelop Transmissions with Guaranteed Stability
- Demixing and Blind Deconvolution of Graph-Diffused Sparse Signals
- Demystifying Deep Learning: a Geometric Approach to Iterative Projections
- Densely Connected Progressive Learning for LSTM-Based Speech Enhancement
- Depression Speaks: Automatic Discrimination between Depressed and Non-Depressed Speakers Based on Nonverbal Speech Features
- Depth Super-Resolution Using Joint Adaptive Weighted Least Squares And Patching Gradient
- Depth Super-Resolution with Deep Edge-Inference Network and Edge-Guided Depth Filling
- Dereverberation and Beamforming in Far-Field Speaker Recognition
- Design of Optimal Entropy-Constrained Unrestricted Polar Quantizer for Bivariate Circularly Symmetric Sources
- Designing Signals with Good Correlation and Distribution Properties
- Detection of Cyclostationarity Using Generalized Coherence
- Determined Blind Source Separation via Proximal Splitting Algorithm
- Developing Far-Field Speaker System Via Teacher-Student Learning
- Developing a Geometric Deformable Model for Radar Shape Inversion
- Diabetic Retinopathy Detection Based on Deep Convolutional Neural Networks
- Dictionary Learning Algorithm for Multi-Subject Fmri Analysis Via Temporal and Spatial Concatenation
- Dictionary Learning for Gaussian Kernel Adaptive Filtering with Variablekernel Center and Width
- Dictionary Learning for High Dimensional Graph Signals
- Differentially Private Distributed Principal Component Analysis
- Digital-Analog Superposition Coding for Ofdm Channels with Application To Video Transmission
- Digitalseal: a Transaction Authentication Tool for Online and Offline Transactions
- Digraph Fourier Transform via Spectral Dispersion Minimization
- Direct Ensemble Estimation of Density Functionals
- Direct, Near Real Time Animation of a 3D Tongue Model Using Non-Invasive Ultrasound Images
- Directivity Synthesis with Multipoles Comprising a Cluster of Focused Sources Using a Linear Loudspeaker Array
- Directly Solving the Original Ratiocut Problem for Effective Data Clustering
- Discovering Correspondence Among Image Sets with Projection View Preservation For 3D Object Detection in Point Clouds
- Discriminative Clustering with Cardinality Constraints
- Discriminative Probabilistic Framework for Generalized Multi-Instance Learning
- Distributed Analytical Graph Identification
- Distributed Approximate Message Passing with Summation Propagation
- Distributed Censoring with Energy Constraint in Wireless Sensor Networks
- Distributed Coupled Learning Over Adaptive Networks
- Distributed Diffusion Adaptation Over Graph Signals
- Distributed Estimation Under Network Model Uncertainty
- Distributed Large Neural Network with Centralized Equivalence
- Distributed Maximum Likelihood Using Dynamic Average Consensus
- Distributed Model Construction in Radio Interferometric Calibration
- Distributed Optimal Consensus-Based Kalman Filtering and its Relation to Map Estimation
- Distributed Solution of Large-Scale Linear Systems Via Accelerated Projection-Based Consensus
- Distributed Splitting-Over-Features Sparse Bayesian Learning with Alternating Direction Method of Multipliers
- Distributed Submodular Maximization for Large Vocabulary Continuous Speech Recognition
- Distributed Tdoa-Based Indoor Source Localisation
- Dithered Beamforming for Channel Estimation in Mmwave-Based Massive Mimo
- Dnn-Based Ar-Wiener Filtering for Speech Enhancement
- Dnn-Based Voice Activity Detection Using Auxiliary Speech Models in Noisy Environments
- Dnn-Based Wireless Positioning in an Outdoor Environment
- Doa Estimation in Heteroscedastic Noise with Sparse Bayesian Learning
- Document Quality Estimation Using Spatial Frequency Response
- Domain Adversarial Training for Accented Speech Recognition
- Domain Independent Key Term Extraction from Spoken Content Based on Context and Term Location Information in the Utterances
- Domain and Speaker Adaptation for Cortana Speech Recognition
- Dpca: Dimensionality Reduction for Discriminative Analytics of Multiple Large-Scale Datasets
- Driver Estimation in Non-Linear Autoregressive Models
- Dropout Approaches for LSTM Based Speech Recognition Systems
- Dual Frequency- and Block-Permutation Alignment for Deep Learning Based Block-Online Blind Source Separation
- Dual-Channel Modulation Energy Metric for Direct-to-Reverberation Ratio Estimation
- Dynamic Frame Skipping for Fast Speech Recognition in Recurrent Neural Network Based Acoustic Models
- Dynamic Matrix Recovery from Partially Observed and Erroneous Measurements
- Dynamic Multi-Rater Gaussian Mixture Regression Incorporating Temporal Dependencies of Emotion Uncertainty Using Kalman Filters
- EAR-EEG for Detecting Inter-Brain Synchronisation in Continuous Cooperative Multi-Person Scenarios
- EEG-Based Auditory Attention Decoding Using Steerable Binaural Superdirective Beamformer
- Eadnet: Efficient Architecture for Decomposed Convolutional Neural Networks
- Ecg Delineation for Qt Interval Analysis Using an Unsupervised Learning Method
- Edge-Aware Context Encoder for Image Inpainting
- Edge-Based Loss Function for Single Image Super-Resolution
- Eeg-Based Video Identification Using Graph Signal Modeling and Graph Convolutional Neural Network
- Effective Attention Mechanism in Dynamic Models for Speech Emotion Recognition
- Effective Cover Song Identification Based on Skipping Bigrams
- Effective Noise Removal and Unified Model of Hybrid Feature Space Optimization for Automated Cardiac Anomaly Detection Using Phonocardiogarm Signals
- Efficacy of Multiuser Massive Miso Wireless Energy Transfer Under iq Imbalance and Channel Estimation Errors Over Rician Fading
- Efficient Circulant Matrix Construction and Implementation for Compressed Sensing
- Efficient Constrained Tensor Factorization by Alternating Optimization with Primal-Dual Splitting
- Efficient Convolutional Dictionary Learning Using Partial Update Fast Iterative Shrinkage-Thresholding Algorithm
- Efficient Deep Convolutional Neural Networks Accelerator without Multiplication and Retraining
- Efficient Estimation of Scatter Matrix with Convex Structure Under $T$ -Distribution
- Efficient Integration of Fixed Beamformers and Speech Separation Networks for Multi-Channel Far-Field Speech Separation
- Efficient Learning of Articulatory Models Based on Multi-Label Training and Label Correction for Pronunciation Learning
- Efficient Model-Free Learning to Overcome Hardware Nonidealities in Analog-to-Information Converters
- Efficient Non-Convex Graph Clustering for Big Data
- Efficient Sampling on HEALPix Grid
- Efficient Super-Wide Bandwidth Extension Using Linear Prediction Based Analysis-Synthesis
- Efficient Worker Assignment in Crowdsourced Data Labeling Using Graph Signal Processing
- Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention
- Em-Based Semi-Blind Mimo-Ofdm Channel Estimation
- Emg Acquisition and Hand Pose Classification for Bionic Hands from Randomly-Placed Sensors
- Emphatic Speech Generation with Conditioned Input Layer and Bidirectional LSTMS for Expressive Speech Synthesis
- Emphatic Speech Prosody Prediction with Deep Lstm Networks
- Enabling Early Audio Event Detection with Neural Networks
- End-To-End Low-Resource Lip-Reading with Maxout Cnn and Lstm
- End-To-End Optimized Speech Coding with Deep Neural Networks
- End-to-End Audiovisual Speech Recognition
- End-to-End Automatic Speech Translation of Audiobooks
- End-to-End Continuous Emotion Recognition from Video Using 3D Convlstm Networks
- End-to-End DNN Based Speaker Recognition Inspired by I-Vector and PLDA
- End-to-End Dynamic Query Memory Network for Entity-Value Independent Task-Oriented Dialog
- End-to-End Hierarchical Language Identification System
- End-to-End Multi-Speaker Speech Recognition
- End-to-End Neural Network Based Automated Speech Scoring
- End-to-End Sound Source Enhancement Using Deep Neural Network in the Modified Discrete Cosine Transform Domain
- End-to-End Speech Emotion Recognition Using Deep Neural Networks
- End-to-end Multimodal Speech Recognition
- Endmembers as Directional Data for Robust Material Variability Retrieval in Hyperspectral Image Unmixing
- Energy Efficiency in MIMO Interference Channels: Social Optimality and Max-Min Fairness
- Energy Efficient Consensus Over Directed Graphs
- Energy-Efficient Speaker Identification with Low-Precision Networks
- Enhancement and Analysis of Conversational Speech: JSALT 2017
- Entropy Based Pruning of Backoff Maxent Language Models with Contextual Features
- Envelope Estimation by Tangentially Constrained Spline
- Epileptic State Segmentation with Temporal-Constrained Clustering
- Essence Vector-Based Query Modeling for Spoken Document Retrieval
- Estimation of Source Panning Parameters and Segmentation of Stereophonic Mixtures
- Estimation of Time-Varying Room Impulse Responses of Multiple Sound Sources from Observed Mixture and Isolated Source Signals
- Estimation of the Sound Field at Arbitrary Positions in Distributed Microphone Networks Based on Distributed Ray Space Transform
- Evaluating Models of Dynamic Functional Connectivity Using Predictive Classification Accuracy
- Evaluation of the Penalized Inequality Constrained Minimum Variance Beamformer for Hearing Aids
- Event-Triggered Particle Filtering Via Diffusion Strategies for Distributed Estimation in Autonomous Systems
- Eventness: Object Detection on Spectrograms for Temporal Localization of Audio Events
- Evolutionary Spectra Based on the Multitaper Method with Application To Stationarity Test
- Exploitation of Semantic Keywords for Malicious Event Classification
- Exploiting Convolutional Neural Networks for Phonotactic Based Dialect Identification
- Exploiting Explicit Memory Inclusion for Artificial Bandwidth Extension
- Exploring Ctc-Network Derived Features with Conventional Hybrid System
- Exploring Hashing and Cryptonet Based Approaches for Privacy-Preserving Speech Emotion Recognition
- Exploring Motor Imagery Eeg Patterns for Stroke Patients with Deep Neural Networks
- Exploring Practical Aspects of Neural Mask-Based Beamforming for Far-Field Speech Recognition
- Exploring Sequential Characteristics in Speaker Bottleneck Feature for Text-Dependent Speaker Verification
- Exploring Speech Enhancement with Generative Adversarial Networks for Robust Speech Recognition
- Exploring the Non-Local Similarity Present in Variational Mode Functions for Effective ECG Denoising
- Exploring the Use of Group Delay for Generalised VTS Based Noise Compensation
- Exponentially Consistent K-Means Clustering Algorithm Based on Kolmogrov-Smirnov Test
- Extendable Neural Matrix Completion
- Extended Pipeline for Content-Based Feature Engineering in Music Genre Recognition
- Extension and Evaluation of a Spectro-Temporal Modulation Method to Improve Acoustic Feedback Performance in Hearing Aids
- Extension of Decoding Problem of HMM Based on LP-Norm
- Extracting Domain Invariant Features by Unsupervised Learning for Robust Automatic Speech Recognition
- F0 Estimation for DNN-Based Ultrasound Silent Speech Interfaces
- FDD Massive MIMO Channel Spatial Covariance Conversion Using Projection Methods
- Face Hallucination Based on Key Parts Enhancement
- Facial Feature-Integrated Inter-Camera Human Tracking
- Factorized Hidden Variability Learning for Adaptation of Short Duration Language Identification Models
- Fairness in Multiterminal Data Compression: A Splitting Method for the Egalitarian Solution
- Far-Field Audio-Visual Scene Perception of Multi-Party Human-Robot Interaction for Children and Adults
- Fast 3D-Hevc Depth Maps Intra-Frame Prediction Using Data Mining
- Fast Adaptation on Deepmixture Generative Network Based Acoustic Modeling
- Fast And Robust Recursive Filter for Image Denoising
- Fast Decentralized Learning Via Hybrid Consensus Admm
- Fast Detection of Abnormal Events in Videos with Binary Features
- Fast Dictionary-Based Approach for Mass Spectrometry Data Analysis
- Fast Distributed Subspace Projection via Graph Filters
- Fast Oov Words Incorporation Using Structured Word Embeddings for Neural Network Language Model
- Fast Projection onto the 𝓁∞, 1-Mixed Norm Ball Using Steffensen Root Search
- Fast Projection-Based Solvers for the Non-Convex Quadratically Constrained Feasibility Problem
- Fast Robust Tracking Via Double Correlation Filter Formulation
- Fast Texture Intra Size Coding Based On Big Data Clustering for 3D-Hevc
- Fast Variational Level Set Based Image Segmentation via Two-Scale Filtering Model
- Fast Vehicle Detection with Lateral Convolutional Neural Network
- Fast and Adaptive Blind Audio Source Separation Using Recursive Levenberg-Marquardt Synchrosqueezing
- Fast-Convergence Singular Value Decomposition for Tracking Time-Varying Channels in Massive Mimo Systems
- Faster ICA Under Orthogonal Constraint
- Faster and Still Safe: Combining Screening Techniques and Structured Dictionaries to Accelerate the Lasso
- Faster-Than-Nyquist Signaling with Differential Encoding and Non Coherent Detection
- Fault Detection Using Attention Models Based on Visual Saliency
- Feature Based Adaptation for Speaking Style Synthesis
- Feature Design Using Audio Decomposition for Intelligent Control of the Dynamic Range Compressor
- Feature LMS Algorithms
- Feature Matching Based on Top K Rank Similarity
- Fftnet: A Real-Time Speaker-Dependent Neural Vocoder
- Filter-and-Convolve: A Cnn Based Multichannel Complex Concatenation Acoustic Model
- Fine-Grained Wound Tissue Analysis Using Deep Neural Network
- Finite Sample Performance of Linear Least Squares Estimators Under Sub-Gaussian Martingale Difference Noise
- Finite-Alphabet Noma for Two-User Uplink Channel
- First-Order Bifurcation Detection for Dynamic Complex Networks
- First-Order Difference Energy Regularization for Enhancing Reconstruction Performance in Compressive Sensing of Foot-Gait Signals
- First-Order Perturbation Analysis of Secsi With Generalized Unfoldings
- Flexible Multi-Group Single-Carrier Modulation: Optimal Subcarrier Grouping and Rate Maximization
- Flipping Large Classes on a Shoestring Budget
- Focal Kl-Divergence Based Dilated Convolutional Neural Networks for Co-Channel Speaker Identification
- Fooling End-To-End Speaker Verification With Adversarial Examples
- Foreground Harmonic Noise Reduction for Robust Audio Fingerprinting
- Forward Attention in Sequence- To-Sequence Acoustic Modeling for Speech Synthesis
- Forward Vehicle Collision Warning Based on Quick Camera Calibration
- Fps-Sft: A Multi-Dimensional Sparse Fourier Transform Based on the Fourier Projection-Slice Theorem
- Frame-Subsampled, Drift-Resilient Video Object Tracking
- Frame-by-Frame Closed-Form Update for Mask-Based Adaptive MVDR Beamforming
- Framework for Evaluation of Sound Event Detection in Web Videos
- Frontal Face Generation from Multiple Pose-Variant Faces with CGAN in Real-World Surveillance Scene
- Full-Info Training for Deep Speaker Feature Learning
- Full-Reference Quality Assessment of Contrast Changed Images Based on Local Linear Model
- Fully Automatic Segmentation of the Right Ventricle Via Multi-Task Deep Neural Networks
- Functional Connectivity States of the Brain Using Restricted Boltzmann Machines
- Fusion of Multiple Multiband Images with Complementary Spatial and Spectral Resolutions
- GLRT Particle Filter for Tracking Nlos Target in Around-the-Corner Radar
- GM-PHD Filter Based Online Multiple Human Tracking Using Deep Discriminative Correlation Matching
- GMM-Based Iterative Entropy Coding for Spectral Envelopes of Speech and Audio
- GSC-Based Binaural Speaker Separation Preserving Spatial Cues
- Gamification of DSP: Electronic vs Pen-and-Paper
- Gated Residual Networks with Dilated Convolutions for Supervised Speech Separation
- Generalised Discriminative Transform via Curriculum Learning for Speaker Recognition
- Generalised Sidelobe Canceller for Noise Reduction in Hearing Devices Using an External Microphone
- Generalization of Deep Neural Networks for Chest Pathology Classification in X-Rays Using Generative Adversarial Networks
- Generalized End-to-End Loss for Speaker Verification
- Generalized Linear Mixing Model Accounting for Endmember Variability
- Generalized Tensor Contractions for an Improved Receiver Design in MIMO-OFDM Systems
- Generalized Uncertainty Principles for the Two-Sided Quaternion Linear Canonical Transform
- Generating Sound Words from Audio Signals of Acoustic Events with Sequence-to-Sequence Model
- Generative Adversarial Networks Based Data Augmentation for Noise Robust Speech Recognition
- Generative Adversarial Source Separation
- Generative Model and Associated Metric for Coordinated-Motion Target Groups
- Generative Scatternet Hybrid Deep Learning (G-Shdl) Network with Structural Priors for Semantic Image Segmentation
- Geographic Language Models for Automatic Speech Recognition
- Geolocation of Unknown Emitters Using Tdoa of Path Rays Through the Ionosphere by Multiple Coordinated Distant Receivers
- Geometric Information Based Monaural Speech Separation Using Deep Neural Network
- Geometric Transformation Invariant Image Quality Assessment Using Convolutional Neural Networks
- Global Optimality in Inductive Matrix Completion
- Globally Optimal Energy Efficiency Maximization for Capacity-Limited Fronthaul Crans with Dynamic Power Amplifiers' Efficiency
- Graph Error Effect in Graph Signal Processing
- Graph Learning Based on Total Variation Minimization
- Graph Regularized Tensor Factorization for Single-Trial EEG Analysis
- Graph Sampling with and Without Input Priors
- Graph Signal Processing of Human Brain Imaging Data
- Graph-based Transforms for Predictive Light Field Compression based on Super-Pixels
- Grassmann Singular Spectrum Analysis for Bioacoustics Classification
- Greedy Algorithm with Approximation Ratio for Sampling Noisy Graph Signals
- Greedy Pursuits Based Gradual Weighting Strategy for Weighted $\ell_{1}$-Minimization
- Grid-Free Direction-of-Arrival Estimation with Compressed Sensing and Arbitrary Antenna Arrays
- Gridless Two-Dimensional Doa Estimation With L-Shaped Array Based on the Cross-Covariance Matrix
- Group Sparsity Residual with Non-Local Samples for Image Denoising
- Guided Image Filtering with Arbitrary Window Function
- HNSR: Highway Networks Based Deep Convolutional Neural Networks Model for Single Image Super-Resolution
- Hand-Raising Gesture Detection in Real Classroom
- Hand: Header-Assisted Network Decoding
- Hands-on in Signal Processing Education at Technische Universitat Darmstadt
- Hard Shadows Removal Using an Approximate Illumination Invariant
- Harnessing Bandit Online Learning to Low-Latency Fog Computing
- Hi, Bcd! Hybrid Inexact Block Coordinate Descent for Hyperspectral Super-Resolution
- Hierarchical Attention and Context Modeling for Group Activity Recognition
- Hierarchical Heavy Hitter Detection Under Unknown Models
- Hierarchical Segmentation Based Point Cloud Attribute Compression
- High Accuracy Acoustic Estimation of Multiple Targets
- High Efficiency Compression for Object Detection
- High Order Recurrent Neural Networks for Acoustic Modelling
- High-Accuracy Stochastic Computing-Based FIR Filter Design
- High-Order Tensor Completion for Data Recovery via Sparse Tensor-Train Optimization
- High-Quality Nonparallel Voice Conversion Based on Cycle-Consistent Adversarial Network
- High-Speed Light Field Image Formation Analysis Using Wavefield Modeling with Flexible Sampling
- High-Speed Optical Camera Communication Using an Optimally Modulated Signal
- Higher Order Exponential Splittings for the Fast Non-Linear Fourier Transform of the Korteweg-De Vries Equation
- Hough Transform Guided Deep Feature Extraction for Dense Building Detection in Remote Sensing Images
- How Sampling Rate Affects Cross-Domain Transfer Learning for Video Description
- How are the Centered Kernel Principal Components Relevant to Regression Task? -An Exact Analysis
- How to Interconnect for Massive Mimo Self-Calibration?
- How to Mobilize Mmwave: A Joint Beam and Channel Tracking Approach
- Human Motion Classification with Micro-Doppler Radar and Bayesian-Optimized Convolutional Neural Networks
- Human and Machine Speaker Recognition Based on Short Trivial Events
- Human and Machine Type Communications Can Coexist in Uplink Massive Mimo Systems
- Human-Like Emotion Recognition: Multi-Label Learning from Noisy Labeled Audio-Visual Expressive Speech
- Human-Machine Inference Networks for Smart Decision Making: Opportunities and Challenges
- Hybrid Lstm-Fsmn Networks for Acoustic Modeling
- Hybridnet for Depth Estimation and Semantic Segmentation
- Hypercomplex Tensor Completion with Cayley-Dickson Singular Value Decomposition
- Hyperspectral Super-Resolution Via Coupled Tensor Factorization: Identifiability and Algorithms
- ILAPF: Incremental Learning Assisted Particle Filtering
- IVA-Based Spatio-Temporal Dynamic Connectivity Analysis in Large-Scale FMRI Data
- Identification of Bilinear Forms with the Kalman Filter
- Identification of Multiple-Input Multiple-Output Channels Under Linear Side Constraints
- Identifying Susceptible Agents in Time Varying Opinion Dynamics Through Compressive Measurements
- Identifying Undirected Network Structure via Semidefinite Relaxation
- Image Alignment via Multi-Model Geometric Fitting and Hierarchical Homography Estimation
- Image Augmentation Using Radial Transform for Training Deep Neural Networks
- Image Fusion Using Belief Propagation
- Image Fusion: an Introduction to Multispectral Signal Processing
- Image Quality Assessment Based Label Smoothing in Deep Neural Network Learning
- Image Recognition Based on Separable Lattice Hmms Using a Deep Neural Network for Output Probability Distributions
- Image Reconstruction for Quanta Image Sensors Using Deep Neural Networks
- Image Representation Using Supervised and Unsupervised Learning Methods on Complex Domain
- Image Restoration with Deep Generative Models
- Image-Based PM2.5 Estimation and its Application on Depth Estimation
- Impact of Microphone Array Configurations on Robust Indirect 3d Acoustic Source Localization
- Importance Sampling Estimator of Outage Probability under Generalized Selection Combining Model
- Improved Algorithms for Differentially Private Orthogonal Tensor Decomposition
- Improved Audio-Visual Laughter Detection Via Multi-Scale Multi-Resolution Image Texture Features and Classifier Fusion
- Improved Detection of Semi-Percussive Onsets in Audio Using Temporal Reassignment
- Improved Noise Characterization for Relative Impulse Response Estimation
- Improved Steady State Analysis of the Recursive Least Squares Algorithm
- Improved Tdnns Using Deep Kernels and Frequency Dependent Grid-RNNS
- Improved Weighted Instrumental Variable Estimator for Doppler-Bearing Source Localization in Heavy Noise
- Improving Accuracy of Nonparametric Transfer Learning Via Vector Segmentation
- Improving Consensus-Based Distributed Camera Calibration Via Edge Pruning and Graph Traversal Initialization
- Improving Convolutional Neural Networks Via Compacting Features
- Improving Disparity Map Estimation for Multi-View Noisy Images
- Improving End-of-Turn Detection in Spoken Dialogues by Detecting Speaker Intentions as a Secondary Task
- Improving End-to-End Speech Recognition with Policy Learning
- Improving Mandarin Tone Mispronunciation Detection for Non-Native Learners with Soft-Target Tone Labels and BLSTM-Based Deep Models
- Improving Multichannel Speech Recognition with Generalized Cross Correlation Inputs and Multitask Learning
- Improving Multikernel Adaptive Filtering with Selective Bias
- Improving Sar Automatic Target Recognition Using Simulated Images Under Deep Residual Refinements
- Improving Semi-Supervised Classification for Low-Resource Speech Interaction Applications
- Improving the Capacity of Very Deep Networks with Maxout Units
- Improving the Performance of Online Neural Transducer Models
- Incorporating ASR Errors with Attention-Based, Jointly Trained RNN for Intent Detection and Slot Filling
- Incorporating Scalability in Unsupervised Spatio- Temporal Feature Learning
- Independent Low-Rank Matrix Analysis Based on Multivariate Complex Exponential Power Distribution
- Indian Buffet Process Deep Generative Models for Semi-Supervised Classification
- Individual Difference of Ultrasonic Transducers for Parametric Array Loudspeaker
- Individual Ship Detection Using Underwater Acoustics
- Inexact Proximal Operators for 𝓁p-Quasinorm Minimization
- Influence of the Number of Loudspeakers on the Timbre in Mixed-Order Ambisonics Reprodution
- Information Fusion Using Particles Intersection
- Insense: Incoherent Sensor Selection for Sparse Signals
- Insights in-to-End Learning Scheme for Language Identification
- Instlistener: An Expressive Parameter Estimation System Imitating Human Performances of Monophonic Musical Instruments
- Integrating Perceivers Neural-Perceptual Responses Using a Deep Voting Fusion Network for Automatic Vocal Emotion Decoding
- Intelligent Signal Processing Mechanisms for Nuanced Anomaly Detection in Action Audio-Visual Data Streams
- Interference Reduction on Full-Length Live Recordings
- Interpretable Clustering Ensembles Using Binary Matrix Factorization
- Interpreting DNN Output Layer Activations: A Strategy to Cope with Unseen Data in Speech Recognition
- Invariances and Data Augmentation for Supervised Music Transcription
- Inverse Atmoshperic Scattering Modeling with Convolutional Neural Networks for Single Image Dehazing
- Investigating Label Noise Sensitivity of Convolutional Neural Networks for Fine Grained Audio Signal Labelling
- Investigating the Effect of Sound-Event Loudness on Crowdsourced Audio Annotations
- Investigation in Spatial-Temporal Domain for Face Spoof Detection
- Investigations on End- to-End Audiovisual Fusion
- Invisible Geo-Location Signature in A Single Image
- Iterative Deep Neural Networks for Speaker-Independent Binaural Blind Speech Separation
- JND-Based Perceptual Video Coding for 4: 4: 4 Screen Content Data in HEVC
- Joint Adaptive Impulse Response Estimation and Inverse Filtering for Enhancing In-Car Audio
- Joint Audio-Video Driven Facial Animation
- Joint Estimation of the Room Geometry and Modes with Compressed Sensing
- Joint Gender-, Tone-, Vowel- Classification Via Novel Hierarchical Classification for Annotation of Monosyllabic Mandarin Word Tokens
- Joint I-Vector with End-to-End System for Short Duration Text-Independent Speaker Verification
- Joint Independent Subspace Analysis by Coupled Block Decomposition: Non-Identifiable Cases
- Joint Late Reverberation and Noise Power Spectral Density Estimation in a Spatially Homogeneous Noise Field
- Joint License Plate Super-Resolution and Recognition in One Multi-Task Gan Framework
- Joint List Polar Decoder with Successive Cancellation and Sphere Decoding
- Joint Mobile Sink Scheduling and Data Aggregation in Asynchronous Wireless Sensor Networks Using Q-Learning
- Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition
- Joint Probabilistic Forecasts of Temperature and Solar Irradiance
- Joint Screening Tests for Lasso
- Joint Separation and Dereverberation of Reverberant Mixtures with Determined Multichannel Non-Negative Matrix Factorization
- Joint Source Localization and Dereverberation by Sound Field Interpolation Using Sparse Regularization
- Joint Source and Sensor Placement for Sound Field Control Based on Empirical Interpolation Method
- Joint Space-(Slow) Time Transmission with Unimodular Waveforms and Receive Adaptive Filter Design for Radar
- Joint Speaker Diarization and Recognition Using Convolutional and Recurrent Neural Networks
- Joint Time Synchronization and Localization for Target Sensors Using a Single Mobile Anchor with Position Uncertainties
- Joint Topology and Radio Resource Optimization for Device-to-Device Based Mobile Social Networks
- Joint Verification-Identification in end-to-end Multi-Scale CNN Framework for Topic Identification
- Jointly Tracking and Separating Speech Sources Using Multiple Features and the Generalized Labeled Multi-Bernoulli Framework
- Kalman Filtering and Clustering in Sensor Networks
- Kernel-Induced Sampling Theorem for Translation-Invariant Reproducing Kernel Hilbert Spaces with Uniform Sampling
- Knowledge Transfer from Weakly Labeled Audio Using Convolutional Neural Network for Sound Events and Scenes
- Knowledge Transfer in Permutation Invariant Training for Single-Channel Multi-Talker Speech Recognition
- L1 Patch-Based Image Partitioning into Homogeneous Textured Regions
- Language Model Domain Adaptation Via Recurrent Neural Networks with Domain-Shared and Domain-Specific Representations
- Language Transfer of Audio Word2Vec: Learning Audio Segment Representations Without Target Language Data
- Language and Noise Transfer in Speech Enhancement Generative Adversarial Network
- Large-Scale High-Dimensional Clustering with Fast Sketching
- Large-Scale Regularized Sumcor GCCA via Penalty-Dual Decomposition
- Large-Scale Weakly Supervised Audio Classification Using Gated Convolutional Neural Network
- Late Reverberation Suppression Using Recurrent Neural Networks with Long Short-Term Memory
- Learned Convolutional Sparse Coding
- Learned Forensic Source Similarity for Unknown Camera Models
- Learning Deep Representations Using Convolutional Auto-Encoders with Symmetric Skip Connections
- Learning Explicit Shape and Motion Evolution Maps for Skeleton-Based Human Action Recognition
- Learning Filterbanks from Raw Speech for Phone Recognition
- Learning Gaussian Graphical Models Using Discriminated Hub Graphical Lasso
- Learning Hard Alignments with Variational Inference
- Learning In-Place Residual Homogeneity for Image Detail Enhancement
- Learning Lexical Coherence Representation Using LSTM Forget Gate for Children with Autism Spectrum Disorder During Story-Telling
- Learning Neural Trans-Dimensional Random Field Language Models with Noise-Contrastive Estimation
- Learning Statistically Accurate Resource Allocations in Non-Stationary Wireless Systems
- Learning Temporal Relationships Between Financial Signals
- Learning an Inverse Tone Mapping Network with a Generative Adversarial Regularizer
- Learning on a Budget for User Authentication on Mobile Devices
- Learning-Based Acoustic Source-Microphone Distance Estimation Using the Coherent-to-Diffuse Power Ratio
- Learning-Based Complexity Reduction and Scaling for HEVC Encoders
- Learning-Based Design of Measurement Matrix with Inter-Column Correlation for Compressive Sensing
- Lensless 3D Imaging Using Mask-Based Cameras
- Leveraging LSTM Models for Overlap Detection in Multi-Party Meetings
- Lexico-Acoustic Neural-Based Models for Dialog Act Classification
- Limited-Memory BFGS Optimization of Recurrent Neural Network Language Models for Speech Recognition
- Limiting Numerical Precision of Neural Networks to Achieve Real-Time Voice Activity Detection
- Linear Classification in Speech-Based Objective Differential Diagnosis of Parkinsonism
- Linear Networks Based Speaker Adaptation for Speech Synthesis
- Linear Quantization by Effective-Resistance Sampling
- Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 Workshop
- Lip2Audspec: Speech Reconstruction from Silent Lip Movements Video
- Listening to Each Speaker One by One with Recurrent Selective Hearing Networks
- Lo-Regularized Hybrid Gradient Sparsity Priors for Robust Single-Image Blind Deblurring
- Locality-Preserving Complex-Valued Gaussian Process Latent Variable Model for Robust Face Recognition
- Localization-Free Power Cartography
- Locally Optimal Invariant Detector for Testing Equality of Two Power Spectral Densities
- Loudspeaker and Listening Position Estimation Using Smart Speakers
- Low Complexity Heart Rate Measurement from Wearable Wrist-Type Photoplethysmographic Sensors Robust to Motion Artifacts
- Low Complexity Implementation of Carrier and Symbol Timing Synchronization for a Fully Digital Downhole Telemetry System
- Low Complexity Joint RDO of Prediction Units Couples for HEVC Intra Coding
- Low Rank Fourier Ptychography
- Low Resolution Face Recognition and Reconstruction Via Deep Canonical Correlation Analysis
- Low-Complexity Secure Watermark Encryption for Compressed Sensing-Based Privacy Preserving
- Low-Complexity Weighted Mrt Multicast Beamforming in Massive Mimo Cellular Networks
- Low-Energy Graph Fourier Basis Functions Span Salient Objects
- Low-Overhead Receiver-Side Channel Tracking for Mmwave Mimo
- Low-Rank Matrix Recovery from One-Bit Comparison Information
- Low-Rank Optimization for Data Shuffling in Wireless Distributed Computing
- Low-Rank and Joint-Sparse Signal Recovery for Spatially and Temporally Correlated Data Using Sparse Bayesian Learning
- MIMO Radar Target Detection Using Low-Complexity Receiver
- MIMO Transmit Beampattern Matching Under Waveform Constraints
- MMSE Adaptive Waveform Design for a MIMO Active Sensing System Tracking Multiple Moving Targets
- Machine Assisted Human Decision Making
- Machine Load Estimation Via Stacked Autoencoder Regression
- Man-Made Object Recognition from Underwater Optical Images Using Deep Learning and Transfer Learning
- Manifold-Based Analysis of Natural Stochastic Textures with Application in Texture Synthesis
- Manifold-Based Inference for a Supervised Gaussian Process Classifier
- Marginal Bayesian Bhattacharyya Bounds for Discrete-Time Filtering
- Mask Weighted Stft Ratios for Relative Transfer Function Estimation and ITS Application to Robust ASR
- Matching Projection Decoding Method for Ambisonics System
- Matching Pursuit Based Convolutional Sparse Coding
- Matrix Completion as Graph Bandlimited Reconstruction
- Maximal Figure-of-Merit Embedding for Multi-Label Audio Classification
- Maximum-A-Posteriori Signal Recovery with Prior Information: Applications to Compressive Sensing
- Maximum-Likelihood Online Speaker Diarization in Noisy Meetings Based on Categorical Mixture Model and Probabilistic Spatial Dictionary
- Measuring Uncertainty in Deep Regression Models: The Case of Age Estimation from Speech
- Measuring the Effect of Linguistic Resources on Prosody Modeling for Speech Synthesis
- Meeting Recognition with Asynchronous Distributed Microphone Array Using Block-Wise Refinement of Mask-Based MVDR Beamformer
- Method of Estimating Direction of Arrival of Sound Source for Monaural Hearing Based on Temporal Modulation Perception
- Mgn: Multi-Glimpse Network for Action Recognition
- Min-Max Latency Optimization for Multiuser Computation Offloading in Fog-Radio Access Networks
- Minimum Spanning Distance for Image Segmentation
- Minimum Word Error Rate Training for Attention-Based Sequence-to-Sequence Models
- Mitigation of Nonlinear Distortion in Sound Zone Control by Constraining Individual Loudspeaker Driver Amplitudes
- Mmse-Based Autocorrelation Sampling for Comprime Arrays
- Mobile Bayesian Spectrum Learning for Heterogeneous Networks
- Modal Decomposition of Musical Instrument Sound Via Alternating Direction Method of Multipliers
- Modality-Specific Structure Preserving Hashing for Cross-Modal Retrieval
- Mode Domain Spatial Active Noise Control Using Sparse Signal Representation
- Model-Based Free-Breathing Cardiac MRI Reconstruction Using Deep Learned & Storm Priors: MODL-STORM
- Model-Based Noise PSD Estimation from Speech in Non-Stationary Noise
- Modeling Non-Linguistic Contextual Signals in LSTM Language Models Via Domain Adaptation
- Modeling and Detection of Evolving Threats Using Random Finite Set Statistics
- Modeling the Acquisition of Intonation: A First Step
- Modeling-By-Generation-Structured Noise Compensation Algorithm for Glottal Vocoding Speech Synthesis System
- Modelling Jitter in Wireless Channel Created by Processor-Memory Activity
- Monaural Singing Voice Separation with Skip-Filtering Connections and Recurrent Inference of Time-Frequency Mask
- Monaural Speech Enhancement Using Deep Neural Networks by Maximizing a Short-Time Objective Intelligibility Measure
- Monophone-Based Background Modeling for Two-Stage On-Device Wake Word Detection
- Motor Imagery for Eeg Biometrics Using Convolutional Neural Network
- Mrt-Based Joint Unicast and Multigroup Multicast Transmission in Massive Mimo Systems
- Mse-Optimal 1-Bit Precoding for Multiuser Mimo Via Branch and Bound
- Multi Scale Feedback Connection for Noise Robust Acoustic Modeling
- Multi Task Learning with Positive and Unlabeled Data and its Application to Mental State Prediction
- Multi-Armed Bandits for Human-Machine Decision Making
- Multi-Channel Deep Clustering: Discriminative Spectral and Spatial Embeddings for Speaker-Independent Speech Separation
- Multi-Dialect Speech Recognition with a Single Sequence-to-Sequence Model
- Multi-Exposure Image Fusion Based on Exposure Compensation
- Multi-Kernel Regression for Graph Signal Processing
- Multi-Kernel, Deep Neural Network and Hybrid Models for Privacy Preserving Machine Learning
- Multi-Microphone Neural Speech Separation for Far-Field Multi-Talker Speech Recognition
- Multi-Scale Object Detection with Feature Fusion and Region Objectness Network
- Multi-Scale Recurrent Neural Network for Sound Event Detection
- Multi-Scenario Deep Learning for Multi-Speaker Source Separation
- Multi-Segment Reconstruction Using Invariant Features
- Multi-Task Autoencoder for Noise-Robust Speech Recognition
- Multi-View Audio-Articulatory Features for Phonetic Recognition on RTMRI-TIMIT Database
- Multi-View Source Localization Based on Power Ratios
- Multichannel Kalman Filtering for Speech Ehnancement
- Multichannel Speaker Activity Detection for Meetings
- Multichannel Speech Separation with Recurrent Neural Networks from High-Order Ambisonics Recordings
- Multilayer Adaptation Based Complex Echo Cancellation and Voice Enhancement
- Multilingual Adaptation of RNN Based ASR Systems
- Multilingual Speech Recognition with a Single End-to-End Model
- Multimodal Bag-of-Words for Cross Domains Sentiment Analysis
- Multimodal Signal Processing and Learning Aspects of Human-Robot Interaction for an Assistive Bathing Robot
- Multipitch Estimation Using Block Sparse Bayesian Learning and Intra-Block Clustering
- Multiple Feature Fusion for Automatic Emotion Recognition Using EEG Signals
- Multiple Jpeg Compression Detection Through Task-Driven Non-Negative Matrix Factorization
- Multiple Peer-to-Peer Bidirectional Cooperative Communications Using Massive MIMO Relays
- Multiple-Input Neural Network-Based Residual Echo Suppression
- Multiple-Model and Reduced-Order Kalman Filtering for Pathological Hand Tremor Extraction
- Multisource Mint Using Convolutive Transfer Function
- Multistream Diarization Fusion Using the Minimum Variance Bayesian Information Criterion
- Music Chord Recognition Based on Midi-Trained Deep Feature and BLSTM-CRF Hybird Decoding
- Music Structure Boundary Detection and Labelling by a Deconvolution of Path-Enhanced Self-Similarity Matrix
- Mutual-Information-Private Online Gradient Descent Algorithm
- Narrowband Channel Estimation for Hybrid Beamforming Millimeter Wave Communication Systems with One-Bit Quantization
- Nasal Speech Sounds Detection Using Connectionist Temporal Classification
- Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions
- Nearest-Instance-Centroid-Estimation Linear Discriminant Analysis (Nice Lda)
- Negative Binomial Optimization for Biomedical Structural Variant Signal Reconstruction
- Neural Adaptive Image Denoiser
- Neural Confnet Classification: Fully Neural Network Based Spoken Utterance Classification Using Word Confusion Networks
- Neural Network Based Time-Frequency Masking and Steering Vector Estimation for Two-Channel Mvdr Beamforming
- Neural Network Language Modeling with Letter-Based Features and Importance Sampling
- Neural Sequential Malware Detection with Parameters
- New Multi-Carrier Demodulation Method Applied to Gearbox Vibration Analysis
- No Need for a Lexicon? Evaluating the Value of the Pronunciation Lexica in End-to-End Models
- No-Reference Hdr Image Quality Assessment Method Based on Tensor Space
- No-Reference Weighting Factor Selection for Bimodal Tomography
- Noise Robust Speech Recognition on Aurora4 by Humans and Machines
- Non-Asymptotic Guarantees for Correlation-Aware Support Detection
- Non-Euclidean Vector Product for Neural Networks
- Non-Iterative Missing Samples Recovery of ECG Signals by Lmmse Estimation for an Autoregressive Cyclostationary Model
- Non-Native Children Speech Recognition Through Transfer Learning
- Non-Negative Online Estimation for Hawkes Process Networks
- Non-Parallel Voice Conversion Using Variational Autoencoders Conditioned by Phonetic Posteriorgrams and D-Vectors
- Non-Zero Diffusion Particle Flow SMC-PHD Filter for Audio-Visual Multi-Speaker Tracking
- Noncircularity-Based Localization for Mixed Near-Field and Far-Field Sources with Unknown Mutual Coupling
- Nonconvex Sparse Logistic Regression via Proximal Gradient Descent
- Nonlinear Acoustic Echo Cancellation Using Elitist Resampling Particle Filter
- Nonlinear Speech Enhancement Under Speech PSD Uncertainty
- Nonnegative Matrix Factorization with Transform Learning
- Nonnegative Tensor Factorization for Source Separation of Loops in Audio
- Normalization of Partly Overlapping Audio Recordings from the Same Event Based on Relative Signal Powers
- Novel Algorithms for Exact and Efficient L1-NORM-BASED Tucker2 Decomposition
- Novel Bayesian Cluster Enumeration Criterion for Cluster Analysis with Finite Sample Penalty Term
- Novel Realizations of Speech-Driven Head Movements with Generative Adversarial Networks
- ON the Use of Wavenet as a Statistical Vocoder
- Object-Oriented Anomaly Detection in Surveillance Videos
- Oct Volumetric Data Restoration via Primal-Dual Plug-and-Play Method
- Octagonal-Axis Raster Pattern for Improved Test Zone Search Motion Estimation
- On Adversarial Training and Loss Functions for Speech Enhancement
- On Approximation of Bandlimited Functions with Compressed Sensing
- On Compressive Sensing of Sparse Covariance Matrices Using Deterministic Sensing Matrices
- On Consistency and Asymptotic Uniqueness in Quasi-Maximum Likelihood Blind Separation of Temporally-Diverse Sources
- On Error Resilient Design of Predictive Scalable Coding Systems
- On Information Coupling in Cooperative Network Synchronization
- On Maximum Likelihood Angle of Arrival Estimation Using Orthogonal Projections
- On Modular Training of Neural Acoustics-to-Word Model for LVCSR
- On SDW-MWF and Variable Span Linear Filter with Application to Speech Recognition in Noisy Environments
- On Selecting Antenna Placements in Indoor Radio Environments
- On Sequential Random Distortion Testing of Non-Stationary Processes
- On Spatial Features for Supervised Speech Separation and its Application to Beamforming and Robust ASR
- On Speech Enhancement Using Microphone Arrays in the Presence of Co-Directional Interference
- On Teaching Signals & Systems in a Project-Based Learning Environment
- On Using Backpropagation for Speech Texture Generation and Voice Conversion
- On the Analysis of Training Data for Wavenet-Based Speech Synthesis
- On the Comparison of Two Room Compensation / Dereverberation Methods Employing Active Acoustic Boundary Absorption
- On the Computability of System Approximations Under Causality Constraints
- On the Design of Robust Steerable Frequency-Invariant Beampatterns with Concentric Circular Microphone Arrays
- On the Equivalence of $f$-Divergence Balls and Density Bands in Robust Detection
- On the Geometry of Mixtures of Prescribed Distributions
- On the High-Snr Receiver Operating Characteristic of Glrt for The Conditional Signal Model
- On the Importance of Analytic Phase of Speech Signals in Spoken Language Recognition
- On the Modulus of Continuity for Noisy Positive Super-Resolution
- On the Performance Analysis of Wifi Based Localization
- On the Sample Complexity of Graphical Model Selection from Non-Stationary Samples
- On the Supermodularity of Active Graph-Based Semi-Supervised Learning with Stieltjes Matrix Regularization
- On the Use of Grapheme Models for Searching in Large Spoken Archives
- On-Talk and Off-Talk Detection: A Discrete Wavelet Transform Analysis of Electroencephalogram
- One-Bit Massive Mimo Precoding via a Minimum Symbol-Error Probability Design
- Online Direction of Arrival Estimation Based on Deep Learning
- Online Education Evaluation for Signal Processing Course Through Student Learning Pathways
- Online Multi-Kernel Learning with Orthogonal Random Features
- Open Set Recognition by Regularising Classifier with Fake Data Generated by Generative Adversarial Networks
- Opportunistic Sensing with MIC Arrays on Smart Speakers for Distal Interaction and Exercise Tracking
- Opportunistic Synchronisation of Multi-Static Staring Array Radars via Track-Before-Detect
- Optimal Algorithms and CRB for Reciprocity Calibration in Massive Mimo
- Optimal Crowdsourced Classification with a Reject Option in the Presence of Spammers
- Optimal Online Cyberbullying Detection
- Optimal Pooling of Covariance Matrix Estimates Across Multiple Classes
- Optimal Power Control Law for Equal-Rate DS-CDMA Networks Governed by a Successive Soft Interference Cancellation Scheme
- Optimal Power and Bit Allocation for Graph Signal Interpolation
- Optimal Selection of Subset of Images with Highest Intra-Class Similarity For 3D Scene Reconstruction
- Optimal Spectral Estimation and System Trade-Off in Long-Distance Frequency-Modulated Continuous-Wave Lidar
- Optimal Stopping Times for Estimating Bernoulli Parameters with Applications to Active Imaging
- Optimal Tone Reservation for Peak to Average Power Control of Cdma Systems
- Optimization of Speaker-Aware Multichannel Speech Extraction with ASR Criterion
- Optimized Sparse Array Design Based on the Sum Coarray
- Optimized Transmission for Consensus in Wireless Sensor Networks
- Optimizing Multilingual Knowledge Transfer for Time-Delay Neural Networks with Low-Rank Factorization
- Optimum Configurations of Sparse Subarray Beamformers
- Optimum Exact Histogram Specification
- Optimum Sparse Array Design for Maximizing Signal- to-Noise Ratio in Presence of Local Scatterings
- Optimum Sparse Array Design for Multiple Beamformers with Common Receiver
- Orthogonality-Regularized Masked NMF for Learning on Weakly Labeled Audio Data
- Orthogonally Regularized Deep Networks for Image Super-Resolution
- Out-of-Vocabulary Word Recovery using FST-Based Subword Unit Clustering in a Hybrid ASR System
- Outlier Removal for Enhancing Kernel-Based Classifier Via the Discriminant Information
- Overlapping Animal Sound Classification Using Sparse Representation
- PIPA: A New Proximal Interior Point Algorithm for Large-Scale Convex Optimization
- PVDC: A Binary Descriptor Using Pore-Valley Disk Code Structure for High-Resolution Partial Fingerprint Recognition
- Papr Minimization Through Spatio-Temporal Symbol-Level Precoding for the Non-Linear Multi-User MISO Channel
- Parallel Beamforming Design in Full Duplex Systems with Per-Antenna Power Constraints
- Parallel Stochastic Successive Convex Approximation Method for Large-Scale Dictionary Learning
- Parallel Vector Field Regularized Non-Negative Matrix Factorization for Image Representation
- Parallel-Data-Free Dictionary Learning for Voice Conversion Using Non-Negative Tucker Decomposition
- Parameter Estimation of Heavy-Tailed Random Walk Model from Incomplete Data
- Parameter Selection Strategy for Sparsity Enforcing Prior Models
- Parametric Approximation of Piano Sound Based on Kautz Model with Sparse Linear Prediction
- Particle Filtering and Inference for Limit Order Books in High Frequency Finance
- Particle Flow Particle Filter for Gaussian Mixture Noise Models
- Partitioning Relational Matrices of Similarities or Dissimilarities Using the Value of Information
- Pattern Localization in Time Series Through Signal-To-Model Alignment in Latent Space
- Perceptual Loss for Superpixel-Level Multispectral and Panchromatic Image Classification
- Perceptually Guided Speech Enhancement Using Deep Neural Networks
- Perceptually Motivated Analysis of Numerically Simulated Head-Related Transfer Functions Generated By Various 3D Surface Scanning Systems
- Performance of Interleaved Training for Single-User Hybrid Massive Antenna Downlink
- Performance of Mask Based Statistical Beamforming in a Smart Home Scenario
- Permissible Support Patterns for Identifying the Spreading Function of Time-Varying Channels
- Permutation Invariant Training for Speaker-Independent Multi-Pitch Tracking
- Permutation-Free Cgmm: Complex Gaussian Mixture Model with Inverse Wishart Mixture Model Based Spatial Prior for Permutation-Free Source Separation and Source Counting
- Phase Corrected Total Variation for Audio Signals
- Phase Retrieval via Smoothing Projected Gradient Method
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.