ICASSP 2019 Accepted Papers
The full list of 1,730 papers accepted at ICASSP 2019 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- Learning from Multiview Correlations in Open-domain Videos
- Learning from the Best: A Teacher-student Multilingual Framework for Low-resource Languages
- Learning the Spiral Sharing Network with Minimum Salient Region Regression for Saliency Detection
- Learning to Dequantize Speech Signals by Primal-dual Networks: an Approach for Acoustic Sensor Networks
- Learning to Detect Dysarthria from Raw Speech
- Learning to Fuse Latent Representations for Multimodal Data
- Learning to Match Transient Sound Events Using Attentional Similarity for Few-shot Sound Recognition
- Learning to Rank: A Progressive Neural Network Learning Approach
- Learning-Based Pricing for Privacy-Preserving Job Offloading in Mobile Edge Computing
- Lessons from Building Acoustic Models with a Million Hours of Speech
- Leveraging Image-to-image Translation Generative Adversarial Networks for Face Aging
- Leveraging Weakly Supervised Data to Improve End-to-end Speech-to-text Translation
- Leveraging mmWave Imaging and Communications for Simultaneous Localization and Mapping
- Light Field Denoising Using 4D Anisotropic Diffusion
- Light Field Image Compression Using Depth-based CNN in Intra Prediction
- Linear Prediction-based Part-defined Auto-encoder Used for Speech Enhancement
- Linearized Kernel Representation Learning from Video Tensors by Exploiting Manifold Geometry for Gesture Recognition
- Local Convergence of the Heavy Ball Method in Iterative Hard Thresholding for Low-rank Matrix Completion
- Local Phase U-net for Fundus Image Segmentation
- Localization of an Unknown Number of Speakers in Adverse Acoustic Conditions Using Reliability Information and Diarization
- Localized Random Sampling for Robust Compressive Beam Alignment
- Long Term Background Reference Based Satellite Video Coding
- Long Text Analysis Using Sliced Recurrent Neural Networks with Breaking Point Information Enrichment
- Look, Listen, and Learn More: Design Choices for Deep Audio Embeddings
- Lora Digital Receiver Analysis and Implementation
- Loss and Double-edge-triggered Detector for Robust Small-footprint Keyword Spotting
- Low Bit-rate Speech Coding with VQ-VAE and a WaveNet Decoder
- Low Frequency Crosstalk Cancellation and Its Relationship to Amplitude Panning
- Low Power Pilot Aided Sub-sample Based Channel Estimation for Mmwave Cellular Systems
- Low Power Ultrasonic Gesture Recognition for Mobile Handsets
- Low-Complexity Compressive Analysis in Sub-Eigenspace for ECG Telemonitoring System
- Low-complexity Detection and Performance Analysis for Decode-and-forward Relay Networks
- Low-complexity Recurrent Neural Network-based Polar Decoder with Weight Quantization Mechanism
- Low-cost Measurement of Industrial Shock Signals via Deep Learning Calibration
- Low-latency Deep Clustering for Speech Separation
- Low-latency Speaker-independent Continuous Speech Separation
- Low-pass Filtering as Bayesian Inference
- Low-power Continuous Heart and Respiration Rates Monitoring on Wearable Devices
- Low-power Programmable Processor for Fast Fourier Transform Based on Transport Triggered Architecture
- Low-rank Embedding of Kernels in Convolutional Neural Networks under Random Shuffling
- Low-rank Estimation Based Evolutionary Clustering for Community Detection in Temporal Networks
- Low-rank Matrix Approximation Based on Intermingled Randomized Decomposition
- Low-rankness of Complex-valued Spectrogram and Its Application to Phase-aware Audio Processing
- Low-resolution Visual Recognition via Deep Feature Distillation
- Lung Nodule Detection with a 3D ConvNet via IoU Self-normalization and Maxout Unit
- M-vectors: Sub-band Based Energy Modulation Features for Multi-stream Automatic Speech Recognition
- MIMO Radar Transmit Beampattern Synthesis via Waveform Design for Target Localization
- MSE Based Precoding Schemes for Partially Correlated Transmissions in Interference Channels
- Machine Learning for Condition Monitoring and Innovation
- Magnetic Resonance Fingerprinting Using a Residual Convolutional Neural Network
- Majorization-minimization Algorithms for Convolutive NMF with the Beta-divergence
- Making Decisions with Shuffled Bits
- Mask-based MVDR Beamformer for Noisy Multisource Environments: Introduction of Time-varying Spatial Covariance Model
- Massive MIMO Multicast Beamforming via Accelerated Random Coordinate Descent
- Massive Mimo Channel Estimation with 1-Bit Spatial Sigma-delta ADCS
- Material Identification Using RF Sensors and Convolutional Neural Networks
- Material Segmentation in Hyperspectral Images with a Spatio-spectral Texture Descriptor
- Matrix Completion with Variational Graph Autoencoders: Application in Hyperlocal Air Quality Inference
- Maximally Separated Averages Prediction for High Fidelity Reversible Data Hiding
- Maximally Smooth Dirichlet Interpolation from Complete and Incomplete Sample Points on the Unit Circle
- Maximum-entropy Scattering Models for Financial Time Series
- Measuring the Spherical-harmonic Representation of a Sound Field Using a Cylindrical Array
- Measuring the Task Induced Oscillatory Brain Activity Using Tensor Decomposition
- Median Activation Functions for Graph Neural Networks
- Memorization Capacity of Deep Neural Networks under Parameter Quantization
- Methodical Design and Trimming of Deep Learning Networks: Enhancing External BP Learning with Internal Omnipresent-supervision Training Paradigm
- Mid-depth Based Block Structure Determination for AV1
- Mid-level Chord Transition Features for Musical Style Analysis
- Minimax Magnitude Response Approximation of Pole-radius Constrained IIR Digital Filters
- Minimum-volume Rank-deficient Nonnegative Matrix Factorizations
- Mirage: 2D Source Localization Using Microphone Pair Augmentation with Echoes
- Missing Data in Traffic Estimation: A Variational Autoencoder Imputation Method
- Misspecified CRB on Parameter Estimation for a Coupled Mixture of Polynomial Phase and Sinusoidal FM Signals
- Mitigating the Impact of Speech Recognition Errors on Spoken Question Answering by Adversarial Domain Adaptation
- Modality Attention for End-to-end Audio-visual Speech Recognition
- Model Change Detection with Application to Machine Learning
- Model Selection for Nonnegative Matrix Factorization by Support Union Recovery
- Modeling Characteristics of Real Loudspeakers Using Various Acoustic Models: Modal-domain Approaches
- Modeling Melodic Feature Dependency with Modularized Variational Auto-encoder
- Modeling Nonlinear Audio Effects with End-to-end Deep Neural Networks
- Modeling and Estimation of Interactions of Yule-Simon Processes
- Modelling Sample Informativeness for Deep Affective Computing
- Models of Visually Grounded Speech Signal Pay Attention to Nouns: A Bilingual Experiment on English and Japanese
- Monitoring of Trees' Health Condition Using a UAV Equipped with Low-cost Digital Camera
- Motion Artefact Removal in Functional Near-infrared Spectroscopy Signals Based on Robust Estimation
- Motion-adapted Three-dimensional Frequency Selective Extrapolation
- Multi Label Restricted Boltzmann Machine for Non-intrusive Load Monitoring
- Multi-attention Network for Thoracic Disease Classification and Localization
- Multi-band PIT and Model Integration for Improved Multi-channel Speech Separation
- Multi-channel Itakura Saito Distance Minimization with Deep Neural Network
- Multi-channel Time Encoding for Improved Reconstruction of Bandlimited Signals
- Multi-channel Wind Noise Reduction Using the Corcos Model
- Multi-classification of Breast Cancer Histology Images by Using Gravitation Loss
- Multi-feature Fusion Based on Supervised Multi-view Multi-label Canonical Correlation Projection
- Multi-frame Super-resolution for Time-of-flight Imaging
- Multi-geometry Spatial Acoustic Modeling for Distant Speech Recognition
- Multi-level Supervised Network for Person Re-identification
- Multi-modal Blind Source Separation with Microphones and Blinkies
- Multi-modal Image Stitching with Nonlinear Optimization
- Multi-objective Optimization Training of PLDA for Speaker Verification
- Multi-scale Dense Network for Single-image Super-resolution
- Multi-scale Spatial-temporal Network for Person Re-identification
- Multi-scale Vehicle Re-identification Using Self-adapting Label Smoothing Regularization
- Multi-speaker Emotional Acoustic Modeling for CNN-based Speech Synthesis
- Multi-speaker Sequence-to-sequence Speech Synthesis for Data Augmentation in Acoustic-to-word Speech Recognition
- Multi-spectral Image Denoising with Shared Dictionaries and Low-rank Representation
- Multi-step Self-attention Network for Cross-modal Retrieval Based on a Limited Text Space
- Multi-target Motion Parameter Estimation Exploiting Collaborative UAV Network
- Multi-task Adaptive Matching Pursuit for Sparse Signal Recovery Exploiting Signal Structures
- Multi-teacher Knowledge Distillation for Compressed Video Action Recognition on Deep Neural Networks
- Multi-user Communication in Difficult Interference
- Multi-view Networks for Multi-channel Audio Classification
- Multicarrier Radar-communications Waveform Design for RF Convergence and Coexistence
- Multicast Beamforming Using Semidefinite Relaxation and Bounded Perturbation Resilience
- Multichannel Quaternion Least Mean Square Algorithm
- Multichannel Sparse Blind Deconvolution on the Sphere
- Multimodal Grounding for Sequence-to-sequence Speech Recognition
- Multimodal One-shot Learning of Speech and Images
- Multimodal Retinal Image Registration and Fusion Based on Sparse Regularization via a Generalized Minimax-concave Penalty
- Multimodal Speaker Adaptation of Acoustic Model and Language Model for Asr Using Speaker Face Embedding
- Multipath-enabled Private Audio with Noise
- Multiple Agents Representation Using Motion Fields
- Multiple Linear Regression for High Efficiency Video Intra Coding
- Multiple Sound Source Localization with Rigid Spherical Microphone Arrays via Residual Energy Test
- Multiple Subspace Alignment Improves Domain Adaptation
- Multiple Temporal Scales Based Speaker Embeddings Learning for Text-dependent Speaker Recognition
- Multiple-graph Recurrent Graph Convolutional Neural Network Architectures for Predicting Disease Outcomes
- Multiresolution Time-of-arrival Estimation from Multiband Radio Channel Measurements
- Multiscale Directional Fusion for Depth Map Super Resolution with Denoising
- Multiscale Structure Tensor Total Variation for Image Recovery
- Multisource Remote Sensing Data Classification Using Deep Hierarchical Random Walk Networks
- Multisource Surveillance Video Coding by Exploiting 3D and 2D Knolwedge
- Multitask Learning for Frame-level Instrument Recognition
- Multiview Canonical Correlation Analysis over Graphs
- Muse-ing on the Impact of Utterance Ordering on Crowdsourced Emotion Annotations
- Music Boundary Detection Based on a Hybrid Deep Model of Novelty, Homogeneity, Repetition and Duration
- Mvdr Robust Adaptive Beamforming Design with Direction of Arrival and Generalized Similarity Constraints
- NN-based Ordinal Regression for Assessing Fluency of ESL Speech
- Native Language and Stimuli Signal Prediction from EEG
- Near-infrared Image Guided Neural Networks for Color Image Denoising
- Near-optimal Coded Apertures for Imaging via Nazarov's Theorem
- Negative Correlation, Non-linear Filtering, and Discovering of Repetitiveness for Cache Timing Channel Detection
- Network Adaptation Strategies for Learning New Classes without Forgetting the Original Ones
- Neural Approaches to Automated Speech Scoring of Monologue and Dialogue Responses
- Neural CRF Transducers for Sequence Labeling
- Neural Codes to Factor Language in Multilingual Speech Recognition
- Neural Music Synthesis for Flexible Timbre Control
- Neural Networks Sequential Training Using Variational Gaussian Particle Filter
- Neural Source-filter-based Waveform Model for Statistical Parametric Speech Synthesis
- Neural Variational Identification and Filtering for Stochastic Non-linear Dynamical Systems with Application to Non-intrusive Load Monitoring
- Neuromorphic Vision Sensing for CNN-based Action Recognition
- Node-asynchronous Implementation of Rational Filters on Graphs
- Noise-tolerant Audio-visual Online Person Verification Using an Attention-based Neural Network Fusion
- Noisy 1-Bit Compressed Sensing with Heterogeneous Side-information
- Non-coherent Sensor Fusion via Entropy Regularized Optimal Mass Transport
- Non-harmonic Analysis Based Instantaneous Heart Rate Estimation from Photoplethysmography
- Non-intrusive Speech Quality Assessment Using Neural Networks
- Non-intrusive Speech Quality Assessment for Super-wideband Speech Communication Networks
- Non-local Self-attention Structure for Function Approximation in Deep Reinforcement Learning
- Non-negative Matrix Factorization Using Bregman Monotone Operator Splitting
- Nonlinear Acceleration of Constrained Optimization Algorithms
- Nonlinear Multi-scale Super-resolution Using Deep Learning
- Nonlinear Prediction of Multidimensional Signals via Deep Regression with Applications to Image Coding
- Nonlinear State Estimation Using Particle Filters on the Stiefel Manifold
- Nonnegative Low-rank Sparse Component Analysis
- Nose, Eyes and Ears: Head Pose Estimation by Locating Facial Keypoints
- Novel Detection Methods for Zero-padded Single Carrier Spatial Modulation in Doubly Selective Channels
- Novel Lower Bound on the Performance of a Partial Zero Forcing Receiver in a Mimo Cellular Network
- Novel Metric Learning for Non-parallel Voice Conversion
- OMP and Continuous Dictionaries: Is k-step Recovery Possible?
- Object Counting in Video Surveillance Using Multi-scale Density Map Regression
- Object Detection in Curved Space for 360-Degree Camera
- Object and Text-guided Semantics for CNN-based Activity Recognition
- Objective Assessment of Spatial Audio Quality Using Directional Loudness Maps
- Objective Assessment of Vocal Tremor
- Objective Comparison of Speech Enhancement Algorithms with Hearing Loss Simulation
- Objective Measures of Plosive Nasalization in Hypernasal Speech
- Obtaining Narrow Transition Region in STFT Domain Processing Using Subband Filters
- Occupancy Pattern Recognition with Infrared Array Sensors: A Bayesian Approach to Multi-body Tracking
- On Achievable Rates for Massive Mimo System with Imperfect Channel Covariance Information
- On Evaluating CNN Representations for Low Resource Medical Image Classification
- On Massive MIMO Cellular Systems Resilience to Radar Interference
- On Modified Squared Givens Rotations for Sphere Decoder Preprocessing
- On Nonparametric Identification of Wiener Systems with Deterministic Inputs
- On Optimal Beam Steering Directions in Millimeter Wave Systems
- On Radar Privacy in Shared Spectrum Scenarios
- On Reducing the Effect of Speaker Overlap for Chime-5
- On Role and Location of Normalization before Model-based Data Augmentation in Residual Blocks for Classification Tasks
- On Self-assessment of Proficiency of Autonomous Systems
- On Training Targets and Objective Functions for Deep-learning-based Audio-visual Speech Enhancement
- On Using 2D Sequence-to-sequence Models for Speech Recognition
- On the Accuracy Limit of Time-delay Estimation with a Band-limited Signal
- On the Adversarial Robustness of Subspace Learning
- On the Computability of the Secret Key Capacity under Rate Constraints
- On the Design of Flexible Kronecker Product Beamformers with Linear Microphone Arrays
- On the Equivalence of Semidifinite Relaxations for MIMO Detection with General Constellations
- On the Fourier Representation of Computable Continuous Signals
- On the Move: Localization with Kinetic Euclidean Distance Matrices
- On the Performance of DIBR Methods When Using Depth Maps from State-of-the-art Stereo Matching Algorithms
- On the Sensitivity of Spectral Initialization for Noisy Phase Retrieval
- On the Transferability of Adversarial Examples against CNN-based Image Forensics
- On the Usefulness of Statistical Normalisation of Bottleneck Features for Speech Recognition
- One-bit Unlimited Sampling
- One-dimensional Edge-preserving Spline Smoothing for Estimation of Piecewise Smooth Functions
- Online Deep Attractor Network for Real-time Single-channel Speech Separation
- Online Estimation and Smoothing of a Target Trajectory in Mixed Stationary/moving Conditions
- Online Learning for Computation Peer Offloading with Semi-bandit Feedback
- Online Learning with Self-tuned Gaussian Kernels: Good Kernel-initialization by Multiscale Screening
- Online Radio Map Update Based on a Marginalized Particle Gaussian Process
- Online Singing Voice Separation Using a Recurrent One-dimensional U-NET Trained with Deep Feature Losses
- Online Single Person Tracking for Unmanned Aerial Vehicles: Benchmark and New Baseline
- Online Variational Bayesian Subspace Filtering
- Optimal Feature Selection for Blind Super-resolution Image Quality Evaluation
- Optimal ROC Curves from Score Variable Threshold Tests
- Optimal Sensor Placement for Signal Extraction
- Optimal Trilateration Is an Eigenvalue Problem
- Optimization of Speaker Extraction Neural Network with Magnitude and Temporal Spectrum Approximation Loss
- Optimization of a Moving Colored Coded Aperture in Compressive Spectral Imaging
- Optimized Color-guided Filter for Depth Image Denoising
- Optimized Quantization in Distributed Graph Signal Processing
- Optimizing QoE of Multiple Users over DASH: A Meta-learning Approach
- Optimum Sampling for Packet Assisted Round Trip Time Measurement
- Outphasing Elements for Hybrid Analogue Digital Beamforming and Single-RF MIMO
- Overlap-add Windows with Maximum Energy Concentration for Speech and Audio Processing
- PGR-Net: A Parallel Network Based on Group and Regression for Age Estimation
- PPSAN: Perceptual-aware 3D Point Cloud Segmentation via Adversarial Learning
- Pairwise Approximate K-SVD
- Parallel Coordinate Descent Algorithms for Sparse Phase Retrieval
- Parameter Uncertainty for End-to-end Speech Recognition
- Parametric Cepstral Mean Normalization for Robust Speech Recognition
- Parametric Hear through Equalization for Augmented Reality Audio
- Particle Filtering: the First 25 Years and beyond
- Passive Detection and Discrimination of Body Movements in the sub-THz Band: A Case Study
- Pathological Speech Intelligibility Assessment Based on the Short-time Objective Intelligibility Measure
- Peak Detection and Baseline Correction Using a Convolutional Neural Network
- Perceptual Audio Coding with Adaptive Non-uniform Time/frequency Tilings Using Subband Merging and Time Domain Aliasing Reduction
- Perceptual Quality Preserving Image Super-resolution via Channel Attention
- Perceptual Soundfield Reconstruction in Three Dimensions via Sound Field Extrapolation
- Perceptually Enhanced Single Frequency Filtering for Dysarthric Speech Detection and Intelligibility Assessment
- Perceptually-motivated Environment-specific Speech Enhancement
- Perfect Match: Improved Cross-modal Embeddings for Audio-visual Synchronisation
- Performance Advantages of Deep Neural Networks for Angle of Arrival Estimation
- Performance Analysis of Convex Data Detection in MIMO
- Performance Analysis of Discrete-valued Vector Reconstruction Based on Box-constrained Sum of L1 Regularizers
- Performance Analysis of One-bit Group-sparse Signal Reconstruction
- Performance Bound for Blind Extraction of Non-gaussian Complex-valued Vector Component from Gaussian Background
- Performance Enhancement of the Measure-transformed Music Algorithm via Mse Based Optimization
- Performance of Jensen Shannon Divergence in Incipient Fault Detection and Estimation
- Perturbed Projected Gradient Descent Converges to Approximate Second-order Points for Bound Constrained Nonconvex Problems
- PhaST: Model-free Phaseless Subspace Tracking
- Phase-aware Harmonic/percussive Source Separation via Convex Optimization
- Phase-only Robust Minimum Dispersion Beamforming
- Phoebe: Pronunciation-aware Contextualization for End-to-end Speech Recognition
- Phoneme Dependent Speaker Embedding and Model Factorization for Multi-speaker Speech Synthesis and Adaptation
- Phoneme Level Language Models for Sequence Based Low Resource ASR
- Phoneme Specific Modelling and Scoring Techniques for Anti Spoofing System
- Phonemic-level Duration Control Using Attention Alignment for Natural Speech Synthesis
- Phonespoof: A New Dataset for Spoofing Attack Detection in Telephone Channel
- Phonetic Analysis of Dysarthric Speech Tempo and Applications to Robust Personalised Dysarthric Speech Recognition
- Phylogenetic Analysis of Software Using Cache Miss Statistics
- Piano Sustain-pedal Detection Using Convolutional Neural Networks
- Pigment Unmixing of Hyperspectral Images of Paintings Using Deep Neural Networks
- Pixel Level Data Augmentation for Semantic Image Segmentation Using Generative Adversarial Networks
- Pixel-level Texture Segmentation Based AV1 Video Compression
- Pliable Data Shuffling for On-device Distributed Learning
- Point Cloud Segmentation Using Hierarchical Tree for Architectural Models
- Polynomial Networks Representation of Nonlinear Mixtures with Application in Underdetermined Blind Source Separation
- Polyphonic Music Transcription with Semantic Segmentation
- Polyphonic Sound Event Detection Using Convolutional Bidirectional Lstm and Synthetic Data-based Transfer Learning
- Post-stitching Depth Adjustment for Stereoscopic Panorama
- Postfiltering Using an Adversarial Denoising Autoencoder with Noise-aware Training
- Potential Games for Distributed Parameter Estimation in Networks with Ambiguous Measurements
- Power Minimization in Multi-tier Networks with Flexible Duplexing
- Power Network Parameter Correction via Sparse Unsupervised Regression
- Power System State Forecasting via Deep Recurrent Neural Networks
- Power-efficient Beam Pattern Synthesis via Sequential Outer Approximation Procedure
- Practical Concentric Open Sphere Cardioid Microphone Array Design for Higher Order Sound Field Capture
- Pre-training of Speaker Embeddings for Low-latency Speaker Change Detection in Broadcast News
- Precoding Design for the MIMO-RoC Downlink
- Predicting Tongue Motion in Unlabeled Ultrasound Videos Using Convolutional Lstm Neural Networks
- Predicting Video-frames Using Encoder-convlstm Combination
- Predicting the Precision of Elevation Localization Based on Head Related Transfer Functions
- Predicting the Secret Parameters of a Chaotic Random Number Generator from Time Series
- Prediction of Multi-target Dynamics Using Discrete Descriptors: an Interactive Approach
- Prediction-correction for Nonsmooth Time-varying Optimization via Forward-backward Envelopes
- Prediction-error-ordering for High-fidelity Reversible Data Hiding
- Prewarping Siamese Network: Learning Local Representations for Online Signature Verification
- Price-aware Renewable Energy Management with Transmission Losses
- Privacy-aware Feature Extraction for Gender Discrimination versus Speaker Identification
- Privacy-cost Trade-off in a Smart Meter System with a Renewable Energy Source and a Rechargeable Battery
- Privacy-preserving Online Human Behaviour Anomaly Detection Based on Body Movements and Objects Positions
- Privacy-preserving Paralinguistic Tasks
- Progressive Filtering for Feature Matching
- Promising Accurate Prefix Boosting for Sequence-to-sequence ASR
- Proper Guidance Image Generation Based on Saliency Factor for Better Transmission Refinement in Image Dehazing
- Properties and Limits of the Minimum-norm Differential Beamformers with Circular Microphone Arrays
- Provable Memory-efficient Online Robust Matrix Completion
- Provably Accelerated Randomized Gossip Algorithms
- Proximal Deep Recurrent Neural Network for Monaural Singing Voice Separation
- Prune Your Neurons Blindly: Neural Network Compression through Structured Class-blind Pruning
- Pruning SIFT & SURF for Efficient Clustering of Near-duplicate Images
- PyHTK: Python Library and ASR Pipelines for HTK
- Quadratic Envelope Regularization for Structured Low Rank Approximation
- Quality Control of Voice Recordings in Remote Parkinson's Disease Monitoring Using the Infinite Hidden Markov Model
- Quantized Event-triggered Sampled-data Average Consensus with Guaranteed Rate of Convergence
- Quantized Gaussian Embedding Steganography
- Quasi Black Hole Effect of Gradient Descent in Large Dimension: Consequence on Neural Network Learning
- Quasi-fully Convolutional Neural Network with Variational Inference for Speech Synthesis
- Quaternion Convolutional Neural Networks for Detection and Localization of 3D Sound Events
- Quaternion Convolutional Neural Networks for Heterogeneous Image Processing
- Quaternion-Valued Adaptive Filtering via Nesterov's Extrapolation
- Question Answering for Spoken Lecture Processing
- Quickest Detection of Deviations from Periodic Statistical Behavior
- Quickest Detection of Time-varying False Data Injection Attacks in Dynamic Smart Grids
- RF-based Analytics Generated by Tag-to-tag Networks
- RHFCN: : Fully CNN-based Steganalysis of MP3 with Rich High-pass Filtering
- RTF-steered Binaural MVDR Beamforming Incorporating an External Microphone for Dynamic Acoustic Scenarios
- Radar Stationary and Moving Indoor Target Localization with Low-rank and Sparse Regularizations
- Radial Loss for Learning Fine-grained Video Similarity Metric
- Rain Streak Removal via Multi-scale Mixture Exponential Power Model
- Random Forest Oriented Fast QTBT Frame Partitioning
- Random Infinite Tree and Dependent Poisson Diffusion Process for Nonparametric Bayesian Modeling in Multiple Object Tracking
- Random Sampling for Distributed Coded Matrix Multiplication
- Randomized Tensor Ring Decomposition and Its Application to Large-scale Data Reconstruction
- Randomly Weighted CNNs for (Music) Audio Classification
- Real-time Object Detection via Pruning and a Concatenated Multi-feature Assisted Region Proposal Network
- Real-time Passive Acoustic 3D Tracking of Deep Diving Cetacean by Small Non-uniform Mobile Surface Antenna
- Real-time Prediction for Fine-grained Air Quality Monitoring System with Asynchronous Sensing
- Real-time Speech Enhancement Using an Efficient Convolutional Recurrent Network for Dual-microphone Mobile Phones in Close-talk Scenarios
- Real-time Tracker with Fast Recovery from Target Loss
- Receiver Design for Doppler Positioning with Leo Satellites
- Recognition of Online Handwriting with Variability on Smart Devices
- Reconfigurable Multitask Audio Dynamics Processing Scheme
- Reconstruction-cognizant Graph Sampling Using Gershgorin Disc Alignment
- Recurrent 3D Convolutional Network for Rodent Behavior Recognition
- Recurrent Deep Divergence-based Clustering for Simultaneous Feature Learning and Clustering of Variable Length Time Series
- Recurrent Neural Network Language Model Training Using Natural Gradient
- Recurrent Neural Networks with Stochastic Layers for Acoustic Novelty Detection
- Reduced Complexity Image Clustering Based on Camera Fingerprints
- Reduced-complexity Deep Neural Network-aided Channel Code Decoder: A Case Study for BCH Decoder
- Reducing the Search Space for Hyperparameter Optimization Using Group Sparsity
- Referential Vowel Duration Ratio as a Feature for Automatic Assessment of L2 Word Prosody
- Reflection Symmetry Detection by Embedding Symmetry in a Graph
- Reflection Tomographic Imaging of Highly Scattering Objects Using Incremental Frequency Inversion
- Regular Sampling of Tensor Signals: Theory and Application to FMRI
- Regularized Fourier Ptychography Using an Online Plug-and-play Algorithm
- Reinforcement Learning Based Speech Enhancement for Robust Speech Recognition
- Reinforcement Learning with Safe Exploration for Network Security
- Reinforcing Self-expressive Representation with Constraint Propagation for Face Clustering in Movies
- Relationships between Deep Learning and Linear Adaptive Systems
- Reliability of the Most Common Objective Metrics for Light Field Quality Assessment
- Remote State Preparation for Multiple Parties
- Replay Attack Detection Using Magnitude and Phase Information with Attention-based Adaptive Filters
- Representation Learning Using Convolution Neural Network for Acoustic-to-articulatory Inversion
- Representation Mixing for TTS Synthesis
- Residual Integration Neural Network
- Resource Optimization in Quantum Access Networks
- Rethinking Super-resolution: the Bandwidth Selection Problem
- Rethinking Teaching Practices for Signal Processing Education
- Retrieving Speech Samples with Similar Emotional Content Using a Triplet Loss Function
- Revisiting Hidden Markov Models for Speech Emotion Recognition
- Revisiting and Improving Semi-supervised Learning: A Large Dimensional Approach
- Robust Approximate Message Passing for Nonzero-mean Sensing Matrices
- Robust Audio-visual Speech Recognition Using Bimodal Dfsmn with Multi-condition Training and Dropout Regularization
- Robust Bayesian Beamforming for Sources at Different Distances with Applications in Urban Monitoring
- Robust Beamspace Design for Direct Localization
- Robust Capon Beamforming via ADMM
- Robust Common Spatial Patterns Estimation Using Dynamic Time Warping to Improve BCI Systems
- Robust Detection for Cluster Analysis
- Robust Dictionary Learning Using α-Divergence
- Robust Freeway Accident Detection: A Two-Stage Approach
- Robust Full-sphere Binaural Sound Source Localization Using Interaural and Spectral Cues
- Robust Graph Signal Sampling
- Robust Gridless Sound Field Decomposition Based on Structured Reciprocity Gap Functional in Spherical Harmonic Domain
- Robust Least Mean Squares Estimation of Graph Signals
- Robust Linear Discriminant Analysis Using Tyler's Estimator: Asymptotic Performance Characterization
- Robust Low-tubal-rank Tensor Completion
- Robust M-estimation Based Matrix Completion
- Robust Molecular Dynamics Simulations Using Coded FFT Algorithm
- Robust Recognition of Reverberant and Noisy Speech Using Coherence-based Processing
- Robust Room Equalization Using Sparse Sound-field Reconstruction
- Robust Secure Precoding and Antenna Selection: A Probabilistic Optimization Approach for Interference Exploitation
- Robust Self-calibration of Constant Offset Time-difference-of-arrival
- Robust Sparse Multichannel Active Noise Control
- Robust Speech Activity Detection in Movie Audio: Data Resources and Experimental Evaluation
- Robust Subspace Clustering by Learning an Optimal Structured Bipartite Graph via Low-rank Representation
- Robust Unsupervised Flexible Auto-weighted Local-coordinate Concept Factorization for Image Clustering
- Robust View Synthesis in Wide-baseline Complex Geometric Environments
- Robust Visual Tracking via Adaptive Occlusion Detection
- Robust and Fine-grained Prosody Control of End-to-end Speech Synthesis
- Rodent Sleep Assessment with a Trainable Video-based Approach
- Role Specific Lattice Rescoring for Speaker Role Recognition from Speech Recognition Outputs
- SAMIR: Sparsity Amplified Iteratively-reweighted Beamforming for High-rsolution Ultrasound Imaging
- SDR - Half-baked or Well Done?
- SLiQA-I: Towards Cold-start Development of End-to-end Spoken Language Interface for Question Answering
- SNIPER: Few-shot Learning for Anomaly Detection to Minimize False-negative Rate with Ensured True-positive Rate
- SPFEMD: Super-pixel Based Finger Earth Mover's Distance for Hand Gesture Recognition
- STFT Spectral Loss for Training a Neural Speech Waveform Model
- SURE-TISTA: A Signal Recovery Network for Compressed Sensing
- SVD-PHAT: A Fast Sound Source Localization Method
- SVM-based Seal Imprint Verification Using Edge Difference
- Safety in the Face of Unknown Unknowns: Algorithm Fusion in Data-driven Engineering Systems
- Saliency Aware: Weakly Supervised Object Localization
- Saliency Map on Cnns for Protein Secondary Structure Prediction
- Saliency Prediction for Omnidirectional Images Considering Optimization on Sphere Domain
- Salient Object Detection on Hyperspectral Images Using Features Learned from Unsupervised Segmentation Task
- Sample Complexity of Joint Structure Learning
- Sample Space-time Covariance Matrix Estimation
- Sampling Schemes for Accurate Reconstruction and Computation of Performance Parameters of Antenna Radiation Pattern
- Scalable Gaussian Process Using Inexact Admm for Big Data
- Scalable MCMC in Degree Corrected Stochastic Block Model
- Scalable Mutual Information Estimation Using Dependence Graphs
- Scaling up MIMO Radar for Target Detection
- Scanet: Spatial-channel Attention Network for 3D Object Detection
- Scattering Multi-connectivity Estimation for Indoor mmWave Small Cells under Limited Training Steps
- Scene Privacy Protection
- Scene-dependent Anomalous Acoustic-event Detection Based on Conditional Wavenet and I-vector
- Score-based Learning for Relevance Prediction in Image Similarity Search
- Score-specific Non-maximum Suppression and Coexistence Prior for Multi-scale Face Detection
- Second Order Sequential Best Rotation Algorithm with Householder Reduction for Polynomial Matrix Eigenvalue Decomposition
- Secure Analytics and Resilient Inference for the Internet of Things
- Secure MIMO Interference Channel with Confidential Messages and Delayed CSIT
- Securing Smartphone Handwritten Pin Codes with Recurrent Neural Networks
- Seeing through Sounds: Predicting Visual Semantic Segmentation Results from Multichannel Audio Signals
- Segment-level Training of ANNs Based on Acoustic Confidence Measures for Hybrid HMM/ANN Speech Recognition
- Segmentation, Classification, and Visualization of Orca Calls Using Deep Learning
- Seizure Detection Using Least Eeg Channels by Deep Convolutional Neural Network
- Selecting Optimal Proposal Number for Image-based Object Detection
- Selective Jpeg2000 Encryption of Iris Data: Protecting Sample Data vs. Normalised Texture
- Selective Virtual Sensing Technique for Multi-channel Feedforward Active Noise Control Systems
- Self-attention Aligner: A Latency-control End-to-end Model for ASR Using Self-attention Network and Chunk-hopping
- Self-attention Based Model for Punctuation Prediction Using Word and Speech Embeddings
- Self-attention Based Prosodic Boundary Prediction for Chinese Speech Synthesis
- Self-attention Networks for Connectionist Temporal Classification in Speech Recognition
- Self-supervised Audio-visual Co-segmentation
- Sell-corpus: an Open Source Multiple Accented Chinese-english Speech Corpus for L2 English Learning Assessment
- Semantic Query-by-example Speech Search Using Visual Grounding
- Semantic Super-resolution for Extremely Low-resolution Vehicle License Plate
- Semi-supervised Acoustic Event Detection Based on Tri-training
- Semi-supervised Depth Estimation from a Single Image Based on Confidence Learning
- Semi-supervised End-to-end Speech Recognition Using Text-to-speech and Autoencoders
- Semi-supervised Learning with Generative Adversarial Networks for Arabic Dialect Identification
- Semi-supervised Monaural Singing Voice Separation with a Masking Network Trained on Synthetic Mixtures
- Semi-supervised Multichannel Speech Enhancement with Variational Autoencoders and Non-negative Matrix Factorization
- Semi-supervised Multiclass Clustering Based on Signed Total Variation
- Semi-supervised Nuisance-attribute Networks for Domain Adaptation
- Semi-supervised Training for End-to-end Models via Weak Distillation
- Semi-supervised Training for Improving Data Efficiency in End-to-end Speech Synthesis
- Semi-supervised Transfer Learning for Convolutional Neural Networks for Glaucoma Detection
- Semi-supervised Triplet Loss Based Learning of Ambient Audio Embeddings
- Semi-supervised and Population Based Training for Voice Commands Recognition
- Sensor-Assisted Global Motion Estimation for Efficient UAV Video Coding
- Sentiment Aware Fake News Detection on Online Social Networks
- Separable Simplex-structured Matrix Factorization: Robustness of Combinatorial Approaches
- Seq2Seq Attentional Siamese Neural Networks for Text-dependent Speaker Verification
- Sequence Noise Injected Training for End-to-end Speech Recognition
- Sequence-level Knowledge Distillation for Model Compression of Attention-based Sequence-to-sequence Speech Recognition
- Sequence-to-sequence Modelling of F0 for Speech Emotion Conversion
- Sequential Matching Model for End-to-end Multi-turn Response Selection
- Sequential Structured Dictionary Learning for Block Sparse Representations
- Sergan: Speech Enhancement Using Relativistic Generative Adversarial Networks with Gradient Penalty
- Shadow Removal Detection and Localization for Forensics Analysis
- Sharpening Sparse Regularizers
- Sharpening of Angular Spectra Based on a Directional Re-assignment Approach for Ambisonic Sound-field Visualisation
- Shift-invariant Subspace Tracking with Missing Data
- Ship Wake Detection in X-band SAR Images Using Sparse GMC Regularization
- Short-segment Heart Sound Classification Using an Ensemble of Deep Convolutional Neural Networks
- Shot Type Feasibility in Autonomous UAV Cinematography
- Sign Language Detection "in the Wild" with Recurrent Neural Networks
- SignProx: One-bit Proximal Algorithm for Nonconvex Stochastic Optimization
- Signals and Systems: Casting It as an Action-adventure Rather than a Horror Genre
- Similarity Learning for Authorship Verification in Social Media
- Similarity Metric Based on Siamese Neural Networks for Voice Casting
- Similarity Search-based Blind Source Separation
- Simple Cooperative Transmission Schemes for Underlay Spectrum Sharing Using Symbol-level Precoding and Load-controlled Arrays
- Simulation or Real-time?
- Simultaneous Blind Deconvolution and Phase Retrieval with Tensor Iterative Hard Thresholding
- Simultaneous DFT and IDFT through Widely Linear CLMS
- Simultaneous Optimization of Forgetting Factor and Time-frequency Mask for Block Online Multi-channel Speech Enhancement
- Singing Voice Separation: A Study on Training Data
- Singing Voice Synthesis Based on Generative Adversarial Networks
- Single Image Interpolation Exploiting Semi-local Similarity
- Single-channel Speech Extraction Using Speaker Inventory and Attention Network
- Skin Lesion Classification Using Hybrid Deep Neural Networks
- Sleep Gesture Detection in Classroom Monitor System
- Small Array Reproduction Method for Ambisonic Encodings Using Headtracking
- Smart DSP for a Smarter Power Grid: Teaching Power System Analysis through Signal Processing
- Smooth Signal Recovery on Product Graphs
- Solving Complex Quadratic Equations with Full-rank Random Gaussian Matrices
- Solving Continuous-domain Problems Exactly with Multiresolution B-splines
- Solving Memory Access Conflicts in LTE-4G Standard
- Solving Quadratic Equations via Amplitude-based Nonconvex Optimization
- Sound Event Detection Using Graph Laplacian Regularization Based on Event Co-occurrence
- Sound Event Detection with Sequentially Labelled Data Based on Connectionist Temporal Classification and Unsupervised Clustering
- Sound Event Envelope Estimation in Polyphonic Mixtures
- Sound Source Localization in a Reverberant Room Using Harmonic Based Music
- Sound-based Transportation Mode Recognition with Smartphones
- Soundfield Reconstruction in Reverberant Environments Using Higher-order Microphones and Impulse Response Measurements
- Space Alternating Variational Estimation and Kronecker Structured Dictionary Learning
- Space Warping Based Dimensionality Reduction of Higher Order Ambisonics Signals
- Sparse Bayesian Learning for Robust PCA
- Sparse Blind Demixing for Low-latency Signal Recovery in Massive Iot Connectivity
- Sparse Fractal Array Design with Increased Degrees of Freedom
- Sparse Gaussian Process Audio Source Separation Using Spectrum Priors in the Time-domain
- Sparse Learning of Parsimonious Reproducing Kernel Hilbert Space Models
- Sparse Recovery and Non-stationary Blind Demodulation
- Sparse Recovery over Nonlinear Dictionaries
- Sparse Signal Recovery Using MPDR Estimation
- Sparse Subspace Clustering for Evolving Data Streams
- Sparsity-based Blind Deconvolution of Neural Activation Signal in FMRI
- Spatial Audio Coding without Recourse to Background Signal Compression
- Spatial Constraint on Multi-channel Deep Clustering
- Spatial and Channel Attention Based Convolutional Neural Networks for Modeling Noisy Speech
- Spatial-fourier Retrieval of Head-related Impulse Responses from Fast Continuous-azimuth Recordings in the Time-domain
- Spatially Adaptive Losses for Video Super-resolution with GANs
- Spatio-spectral Modulation Using a Binary Photomask for Compressive Chromotomography
- Speaker Agnostic Foreground Speech Detection from Audio Recordings in Workplace Settings from Wearable Recorders
- Speaker Change Detection Using Fundamental Frequency with Application to Multi-talker Segmentation
- Speaker Characterization Using TDNN-LSTM Based Speaker Embedding
- Speaker Diarisation Using 2D Self-attentive Combination of Embeddings
- Speaker Recognition for Multi-speaker Conversations Using X-vectors
- Speaker Verification Using End-to-end Adversarial Language Adaptation
- Speaker-dependent Wavenet-based Delay-free Adpcm Speech Coding
- Speaker-independent Classification of Phonetic Segments from Raw Ultrasound in Child Speech
- Spectral Efficiency of Noncooperative Uplink Massive MIMO Systems with Joint Decoding
- Spectral Graph Wavelet Transform as Feature Extractor for Machine Learning in Neuroimaging
- Spectral Method for Multiplexed Phase Retrieval and Application in Optical Imaging in Complex Media
- Spectral Partitioning of Time-varying Networks with Unobserved Edges
- Spectrum-adapted Polynomial Approximation for Matrix Functions
- Speech Artifact Removal from Eeg Recordings of Spoken Word Production with Tensor Decomposition
- Speech Augmentation Using Wavenet in Speech Recognition
- Speech Denoising by Parametric Resynthesis
- Speech Emotion Recognition Using Capsule Networks
- Speech Emotion Recognition Using Deep Neural Network Considering Verbal and Nonverbal Speech Sounds
- Speech Emotion Recognition Using Multi-hop Attention Mechanism
- Speech Enhancement with Variational Autoencoders and Alpha-stable Distributions
- Speech Landmark Bigrams for Depression Detection from Naturalistic Smartphone Speech
- Speech Markers for Clinical Assessment of Cocaine Users
- Speech Recognition with No Speech or with Noisy Speech
- Speech Super Resolution Generative Adversarial Network
- Speech Waveform Reconstruction Using Convolutional Neural Networks with Noise and Periodic Inputs
- Speech as a Biomarker for Obstructive Sleep Apnea Detection
- Spherical Clustering of Users Navigating 360° Content
- Spike Detection in Axonal-synaptic Channels with Multiple Synapses
- Spoofing Attack Detection by Anomaly Detection
- Squared-Loss Mutual Information via High-Dimension Coherence Matrix Estimation
- Static and Dynamic State Predictions for Acoustic Model Combination
- Statistical Persistent Homology of Brain Signals
- Statistical Rank Selection for Incomplete Low-rank Matrices
- Stereo Source Separation in the Frequency Domain: Solving the Permutation Problem by a Sliding K-means Method
- Stochastic Adaptive Neural Architecture Search for Keyword Spotting
- Stochastic Data-driven Hardware Resilience to Efficiently Train Inference Models for Stochastic Hardware Implementations
- Stochastic Gradient Descent for Spectral Embedding with Implicit Orthogonality Constraint
- Stochastic Markov Recurrent Neural Network for Source Separation
- Stochastic Ml Simplex-structured Matrix Factorization under the Dirichlet Mixture Model
- Stream Attention-based Multi-array End-to-end Speech Recognition
- Streaming End-to-end Speech Recognition for Mobile Devices
- Strip the Stripes: Artifact Detection and Removal for Scanning Electron Microscopy Imaging
- Structural Recurrent Neural Network for Traffic Speed Prediction
- SubSpectralNet - Using Sub-spectrogram Based Convolutional Neural Networks for Acoustic Scene Classification
- Subband Optimization and Filtering Technique for Practical Personal Audio Systems
- Subband Temporal Envelope Features and Data Augmentation for End-to-end Recognition of Distant Conversational Speech
- Subword Regularization and Beam Search Decoding for End-to-end Automatic Speech Recognition
- Sum Throughput Maximization for Multi-tag MISO Backscattering
- Super-gaussianity of Speech Spectral Coefficients as a Potential Biomarker for Dysarthric Speech Detection
- Super-resolution DOA Estimation for Arbitrary Array Geometries Using a Single Noisy Snapshot
- Super-resolution Results for a 1D Inverse Scattering Problem
- Super-resolution Using Flow Estimation in Contrast Enhanced Ultrasound Imaging
- Supervised Kernel Change Point Detection with Partial Annotations
- Supervised Speech Enhancement with Real Spectrum Approximation
- Support Tensor Machine for Financial Forecasting
- Surgical Activities Recognition Using Multi-scale Recurrent Networks
- System and VLSI Implementation of Phase-based View Synthesis
- TCN: Transferable Coupled Network for Cross-Resolution Face Recognition*
- TCNN: Temporal Convolutional Neural Network for Real-time Speech Enhancement in the Time Domain
- TS-MC: Two Stage Matrix Completion Algorithm for Wireless Sensor Networks
- TV-DCT: Method to Impute Gene Expression Data Using DCT Based Sparsity and Total Variation Denoising
- Target Localization and Mutual Information Improvement for Cooperative MIMO Radar and MIMO Communication Systems
- Target and Non-target Speaker Discrimination by Humans and Machines
- Task-Based Quantization for Massive MIMO Channel Estimation
- Teach an All-rounder with Experts in Different Domains
- Teacher-student Deep Clustering for Low-delay Single Channel Speech Separation
- Teacher-student Training for Acoustic Event Detection Using Audioset
- Teaching Practical DSP with Off-the-shelf Hardware and Free Software
- Teaching Signal Processing Concepts to Digital Natives
- Team Policy Learning for Multi-agent Reinforcement Learning
- Temporal Salience Based Human Action Recognition
- Tensor Matched Kronecker-structured Subspace Detection for Missing Information
- Tensor Robust PCA on Graphs
- Tensor Super-resolution for Seismic Data
- Tensor-Train Discriminant Analysis
- Tensor-based Estimation of mmWave MIMO Channels with Carrier Frequency Offset
- Tensor-ring Nuclear Norm Minimization and Application for Visual : Data Completion
- The CORAL+ Algorithm for Unsupervised Domain Adaptation of PLDA
- The Design of Personal Audio Systems for Speech Transmission Using Analytical and Measured Responses
- The Direction Cosine Matrix Algorithm in Fixed-point: Implementation and Analysis
- The Discrete Cosine Transform on Triangles
- The Effect of Spatio-temporal Inconsistency on the Subjective Quality Evaluation of Omnidirectional Videos
- The Generalization Effect for Multilingual Speech Emotion Recognition across Heterogeneous Languages
- The Geometry of Equality-constrained Global Consensus Problems
- The Good, the Bad, Algorithmic Noise Tolerance (Ant), the Ugly
- The Impact of Stalling on the Perceptual Quality of HTTP-based Omnidirectional Video Streaming
- The Leap Speaker Recognition System for NIST SRE 2018 Challenge
- The Limitation and Practical Acceleration of Stochastic Gradient Algorithms in Inverse Problems
- The Matched Reassignment Applied to Echolocation Data
- The Phasebook: Building Complex Masks via Discrete Representations for Source Separation
- The Pytorch-kaldi Speech Recognition Toolkit
- The Speechtransformer for Large-scale Mandarin Chinese Speech Recognition
- The Universal Manifold Embedding for Estimating Rigid Transformations of Point Clouds
- Tied Normal Variance-Mean Mixtures for Linear Score Calibration
- Time Difference of Arrival Estimation of Speech Signals Using Deep Neural Networks with Integrated Time-frequency Masking
- Time Domain Spherical Harmonic Analysis for Adaptive Noise Cancellation over a Spatial Region
- Time Series Prediction for Kernel-based Adaptive Filters Using Variable Bandwidth, Adaptive Learning-rate, and Dimensionality Reduction
- Time Signal Classification Using Random Convolutional Features
- Time-based Sampling and Reconstruction of Non-bandlimited Signals
- Time-frequency-bin-wise Switching of Minimum Variance Distortionless Response Beamformer for Underdetermined Situations
- Time-frequency-masking-based Determined BSS with Application to Sparse IVA
- Time-varying Graph Learning Based on Sparseness of Temporal Variation
- Timescalenet : A Multiresolution Approach for Raw Audio Recognition
- To Reverse the Gradient or Not: an Empirical Comparison of Adversarial and Multi-task Learning in Speech Recognition
- Toa Source Node Self-positioning with Unknown Clock Skew in Wireless Sensor Networks
- Toeplitz Matrix Completion for Direction Finding Using a Modified Nested Linear Array
- Token-wise Training for Attention Based End-to-end Speech Recognition
- Topic Detection in Conversational Telephone Speech Using CNN with Multi-stream Inputs
- Total-variation-regularized Tensor Ring Completion for Remote Sensing Image Reconstruction
- Toward Robust Interpretable Human Movement Pattern Analysis in a Workplace Setting
- Toward Subjective Violence Detection in Videos
- Toward the Quantum Internet: A Directional-dependent Noise Model for Quantum Signal Processing
- Towards Audio to Scene Image Synthesis Using Generative Adversarial Network
- Towards Automatic Methods to Detect Errors in Transcriptions of Speech Recordings
- Towards Better Confidence Estimation for Neural Models
- Towards Code-switching ASR for End-to-end CTC Models
- Towards Cross-modality Topic Modelling via Deep Topical Correlation Analysis
- Towards Disease-specific Speech Markers for Differential Diagnosis in Parkinsonism
- Towards End-to-end Speech-to-text Translation with Two-pass Decoding
- Towards Generating Ambisonics Using Audio-visual Cue for Virtual Reality
- Towards Learned Color Representations for Image Splicing Detection
- Towards Perceptually Optimized Sound Zones: A Proof-of-concept Study
- Towards Unsupervised Single-channel Blind Source Separation Using Adversarial Pair Unmix-and-remix
- Towards Unsupervised Speech-to-text Translation
- Towards Visually Grounded Sub-word Speech Unit Discovery
- Tracking Dynamic Systems in α-Stable Environments
- Tracking Multiple Image Sharing on Social Networks
- Tracking a Cluster of Space Debris in Low Orbit by Filtering on Lie Groups
- Trainable Adaptive Window Switching for Speech Enhancement
- Trainable Time Warping: Aligning Time-series in the Continuous-time Domain
- Training Dynamic Exponential Family Models with Causal and Lateral Dependencies for Generalized Neuromorphic Computing
- Training Multi-task Adversarial Network for Extracting Noise-robust Speaker Embedding
- Training Neural Audio Classifiers with Few Data
- Transdrums: A Drum Pattern Transfer System Preserving Global Pattern Structure
- Transfer Learning Using Raw Waveform Sincnet for Robust Speaker Diarization
- Transfer Learning of Language-independent End-to-end ASR with Language Model Fusion
- Transfer and Collaborative Learning Method for Personalized Noninvasive Blood Glucose Measurement Modeling
- Transferability of Neural Network Approaches for Low-rate Energy Disaggregation
- Transferable Positive/negative Speech Emotion Recognition via Class-wise Adversarial Domain Adaptation
- Transferring Piano Performance Control across Environments
- Transform Coefficient Coding for Screen Content in Versatile Video Coding (VVC)
- Transform Domain Based Medical Image Super-resolution via Deep Multi-scale Network
- Transmission Line Cochlear Model Based AM-FM Features for Replay Attack Detection
- Transmit Beampattern Design for MIMO Radar with One-bit DACs
- Triggered Attention for End-to-end Speech Recognition
- Trigonometric Interpolation Beamforming for a Circular Microphone Array
- Tropical Modeling of Weighted Transducer Algorithms on Graphs
- Truly Unsupervised Acoustic Word Embeddings Using Weak Top-down Constraints in Encoder-decoder Models
- Tuning Frequency Dependency in Music Classification
- Tuplemax Loss for Language Identification
- Turning a Vulnerability into an Asset: Accelerating Facial Identification with Morphing
- Two-B-real Net: Two-branch Network for Real-time Salient Object Detection
- Two-stream Multi-focus Image Fusion Based on the Latent Decision Map
- Type and Leak Your Ethnicity on Smartphones
- UTD-CRSS Systems for 2018 NIST Speaker Recognition Evaluation
- Understanding Deep Neural Networks through Input Uncertainties
- Unified Framework for Minimax MIMO Transmit Beampattern Matching under Waveform Constraints
- Unifying Isolated and Overlapping Audio Event Detection with Multi-label Multi-task Convolutional Recurrent Neural Networks
- Unifying Probabilistic Models for Time-frequency Analysis
- Universal Acoustic Modeling Using Neural Mixture Models
- Universal Adversarial Attacks on Text Classifiers
- Unmixing Dynamic Pet Images: Combining Spatial Heterogeneity and Non-gaussian Noise
- Unrolled Projected Gradient Descent for Multi-spectral Image Fusion
- Unsupervised Deep Clustering for Source Separation: Direct Learning from Mixtures Using Spatial Information
- Unsupervised Feature Ranking and Selection Based on Autoencoders
- Unsupervised Feature Selection Based on Reconstruction Error Minimization
- Unsupervised Learning of Deep Features for Music Segmentation
- Unsupervised Melody Style Conversion
- Unsupervised Person Re-identification Using Reliable and Soft Labels
- Unsupervised Polyglot Text-to-speech
- Unsupervised Training of a Deep Clustering Model for Multichannel Blind Source Separation
- Unsupervised User Clustering in Non-orthogonal Multiple Access
- Updates in Bayesian Filtering by Continuous Projections on a Manifold of Densities
- Uplink Multi-user MIMO Detection via Parallel Access
- User Constrained Thumbnail Generation Using Adaptive Convolutions
- Using 3D Residual Network for Spatio-temporal Analysis of Remote Sensing Data
- Using Deep-Q Network to Select Candidates from N-best Speech Recognition Hypotheses for Enhancing Dialogue State Tracking
- Using Extreme Gradient Boosting to Detect Glottal Closure Instants in Speech Signal
- Using RFID Technology to Introduce Properties of LMS
- Using Recurrences in Time and Frequency within U-net Architecture for Speech Enhancement
- Utterance-level Aggregation for Speaker Recognition in the Wild
- Utterance-level End-to-end Language Identification Using Attention-based CNN-BLSTM
- Variance Preserving Initialization for Training Deep Neuromorphic Photonic Networks with Sinusoidal Activations
- Variational and Hierarchical Recurrent Autoencoder
- Vehicle Pose Estimation Using Mask Matching
- Video Quality Assessment for Encrypted HTTP Adaptive Streaming: Attention-based Hybrid RNN-HMM Model
- Video-based, Occlusion-robust Multi-view Stereo Using Inner-boundary Depths of Textureless Areas
- View-invariant Action Recognition from RGB Data via 3D Pose Estimation
- Visual Relationship Recognition via Language and Position Guided Attention
- Vocal Melody Extraction via DNN-based Pitch Estimation and Salience-based Pitch Refinement
- Voice Conversion with Cyclic Recurrent Neural Network and Fine-tuned Wavenet Vocoder
- Voice Trigger Detection from Lvcsr Hypothesis Lattices Using Bidirectional Lattice Recurrent Neural Networks
- Wav2Letter++: A Fast Open-source Speech Recognition System
- Wav2Pix: Speech-conditioned Face Generation Using Generative Adversarial Networks
- Waveform Generation for Text-to-speech Synthesis Using Pitch-synchronous Multi-scale Generative Adversarial Networks
- Waveform Modeling by Adaptive Weighted Hermite Functions
- Waveglow: A Flow-based Generative Network for Speech Synthesis
- Wavelength-resolved Neutron Tomography for Crystalline Materials
- Wavenilm: A Causal Neural Network for Power Disaggregation from the Complex Power Signal
- Weakly Standard Interference Mappings: Existence of Fixed Points and Applications to Power Control in Wireless Networks
- Weakly Supervised Instance Segmentation Using Hybrid Networks
- West: Word Encoded Sequence Transducers
- When CTC Training Meets Acoustic Landmarks
- When Can a System of Subnetworks Be Registered Uniquely?
- When Not to Classify: Detection of Reverse Engineering Attacks on DNN Image Classifiers
- Who Do I Sound like? Showcasing Speaker Recognition Technology by Youtube Voice Search
- Why Do Neural Dialog Systems Generate Short and Meaningless Replies? a Comparison between Dialog and Translation
- Widely Linear Kernels for Complex-valued Kernel Activation Functions
- Windowed Attention Mechanisms for Speech Recognition
- Word Characters and Phone Pronunciation Embedding for ASR Confidence Classifier
- Word and Class Common Space Embedding for Code-switch Language Modelling
- Workload-aware Automatic Parallelization for Multi-GPU DNN Training
- Zero Resource Speaking Rate Estimation from Change Point Detection of Syllable-like Units
- Zero-mean Convolutional Network with Data Augmentation for Sound Level Invariant Singing Voice Separation
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.