ICASSP 2022 Accepted Papers
The full list of 1,864 papers accepted at ICASSP 2022 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- Lattention: Lattice-Attention in ASR Rescoring
- Lattice Rescoring Based on Large Ensemble of Complementary Neural Language Models
- LatticeBART: Lattice-to-Lattice Pre-Training for Speech Recognition
- Learnable Hypergraph Laplacian for Hypergraph Learning
- Learnable Nonlinear Compression for Robust Speaker Verification
- Learnable Wavelet Packet Transform for Data-Adapted Spectrograms
- Learning Acoustic Frame Labeling for Phoneme Segmentation with Regularized Attention Mechanism
- Learning Adjustable Image Rescaling with Joint Optimization of Perception and Distortion
- Learning Approach For Fast Approximate Matrix Factorizations
- Learning Common Dependency Structure for Unsupervised Cross-Domain Ner
- Learning Continuous Representation of Audio for Arbitrary Scale Super Resolution
- Learning Correlation for Online Multiple Object Tracking
- Learning Decoupling Features Through Orthogonality Regularization
- Learning Deep Pathological Features for WSI-Level Cervical Cancer Grading
- Learning Domain-Invariant Transformation for Speaker Verification
- Learning Expanding Graphs for Signal Interpolation
- Learning Filterbanks for End-to-End Acoustic Beamforming
- Learning Gaussian Graphical Models with Differing Pairwise Sample Sizes
- Learning Monocular 3D Human Pose Estimation With Skeletal Interpolation
- Learning Monocular Mesh Recovery of Multiple Body Parts Via Synthesis
- Learning Multiple Explainable and Generalizable Cues for Face Anti-Spoofing
- Learning Music Audio Representations Via Weak Language Supervision
- Learning Music Sequence Representation From Text Supervision
- Learning Semantic-Aligned Feature Representation for Text-Based Person Search
- Learning Sound Localization Better from Semantically Similar Samples
- Learning Sparse Graphs with a Core-Periphery Structure
- Learning Structured Sparsity For Time-Frequency Reconstruction
- Learning Subject-Invariant Representations from Speech-Evoked EEG Using Variational Autoencoders
- Learning Task-Specific Representation for Video Anomaly Detection with Spatial-Temporal Attention
- Learning To Integrate Vision Data Into Road Network Data
- Learning to Enhance or Not: Neural Network-Based Switching of Enhanced and Observed Signals for Overlapping Speech Recognition
- Learning to Fuse Heterogeneous Features for Low-Light Image Enhancement
- Learning to Predict Speech in Silent Videos Via Audiovisual Analogy
- Learning to Sample for Sparse Signals
- Learning-Aided Initialization for Variational Bayesian DOA Estimation
- Learning-Based Personal Speech Enhancement for Teleconferencing by Exploiting Spatial-Spectral Features
- Learning-Based Resource Allocation with Dynamic Data Rate Constraints
- Learnings from Federated Learning in The Real World
- Leveraging Bilinear Attention to Improve Spoken Language Understanding
- Leveraging Local Temporal Information for Multimodal Scene Classification
- Leveraging Sparse Coding for EEG Based Emotion Recognition in Shooting
- LightPose: A Lightweight and Efficient Model with Transformer for Human Pose Estimation
- Linear-Time Sampling on Signed Graphs Via Gershgorin Disc Perfect Alignment
- Lipreading Model Based On Whole-Part Collaborative Learning
- Listen, Know and Spell: Knowledge-Infused Subword Modeling for Improving ASR Performance of OOV Named Entities
- LiteHAR: Lightweight Human Activity Recognition from WIFI Signals with Random Convolution Kernels
- LocUNet: Fast Urban Positioning Using Radio Maps and Deep Learning
- Local Context Interaction-Aware Glyph-Vectors for Chinese Sequence Tagging
- Local Information Modeling with Self-Attention for Speaker Verification
- Local and Global Alignments for Generalizable Sensor-Based Human Activity Recognition
- Local-Global Feature Aggregation for Light Field Image Super-Resolution
- Localization based Sequential Grouping for Continuous Speech Separation
- Localizing More Sources than Sensors in Presence of Coherent Sources
- Locate This, Not that: Class-Conditioned Sound Event DOA Estimation
- Location-Based Training for Multi-Channel Talker-Independent Speaker Separation
- Look, Listen and Pay More Attention: Fusing Multi-Modal Information for Video Violence Detection
- Low Complex Accurate Multi-Source RTF Estimation
- Low Complexity Equalization for Afdm In Doubly Dispersive Channels
- Low Precision Local Learning for Hardware-Friendly Neuromorphic Visual Recognition
- Low Resources Online Single-Microphone Speech Enhancement with Harmonic Emphasis
- Low-Complexity Attention Modelling via Graph Tensor Networks
- Low-Complexity Multi-Model CNN in-Loop Filter for AVS3
- Low-Latency Human-Computer Auditory Interface Based on Real-Time Vision Analysis
- Low-Light Image Enhancement via Feature Restoration
- Low-Rank Phase Retrieval with Structured Tensor Models
- M2Met: The Icassp 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
- MA-NET: Multi-Scale Attention-Aware Network for Optical Flow Estimation
- MAG+: An Extended Multimodal Adaptation Gate for Multimodal Sentiment Analysis
- MAKD: MULTIPLE Auxiliary Knowledge Distillation
- MANNER: Multi-View Attention Network For Noise Erasure
- MBA-RainGAN: A Multi-Branch Attention Generative Adversarial Network for Mixture of Rain Removal
- MBNet: A Multi-Resolution Branch Network for Semantic Segmentation Of Ultra-High Resolution Images
- MEJIGCLU: More Effective Jigsaw Clustering For Unsupervised Visual Representation Learning
- MFA: TDNN with Multi-Scale Frequency-Channel Attention for Text-Independent Speaker Verification with Short Utterances
- MLP-SVNET: A Multi-Layer Perceptrons Based Network for Speaker Verification
- MM-DFN: Multimodal Dynamic Fusion Network for Emotion Recognition in Conversations
- MOS Predictor for Synthetic Speech with I-Vector Inputs
- MRI Recovery with a Self-Calibrated Denoiser
- MS-ROCANet: Multi-Scale Residual Orthogonal-Channel Attention Network for Scene Text Detection
- MSDTRON: A High-Capability Multi-Speaker Speech Synthesis System for Diverse Data Using Characteristic Information
- MTAF: Shopping Guide Micro-Videos Popularity Prediction Using Multimodal and Temporal Attention Fusion Approach
- Magic Dust for Cross-Lingual Adaptation of Monolingual Wav2vec-2.0
- Making The Unknown More Certain: A Stacked Ensemble Classifier for Open Gesture Recognition with a Social Robot
- Manifold Learning-Supported Estimation of Relative Transfer Functions For Spatial Filtering
- Mannet: A Large-Scale Manipulated Image Detection Dataset And Baseline Evaluations
- Map: Multispectral Adversarial Patch to Attack Person Detection
- Mask-Based Attention Parallel Network for in-the-Wild Facial Expression Recognition
- Masked Acoustic Unit for Mispronunciation Detection and Correction
- Massive Unsourced Random Access Based on Bilinear Vector Approximate Message Passing
- Massively Multilingual ASR: A Lifelong Learning Solution
- Matching Point Sets with Quantum Circuit Learning
- Material-Guided Siamese Fusion Network for Hyperspectral Object Tracking
- Matrix Decomposition on Graphs: A Simplified Functional View
- Maximizing Audio Event Detection Model Performance on Small Datasets Through Knowledge Transfer, Data Augmentation, and Pretraining: an Ablation Study
- Maximum Batch Frobenius Norm for Multi-Domain Text Classification
- Melons: Generating Melody With Long-Term Structure Using Transformers And Structure Graph
- Memobert: Pre-Training Model with Prompt-Based Learning for Multimodal Emotion Recognition
- Memory in Echo State Networks and the Controllability Matrix Rank
- Memory-Based Message Passing: Decoupling the Message for Propagation from Discrimination
- Message Passing-Based Cooperative Localization with Embedded Particle Flow
- Meta Talk: Learning To Data-Efficiently Generate Audio-Driven Lip-Synchronized Talking Face With High Definition
- MetricGAN-U: Unsupervised Speech Enhancement/ Dereverberation Based Only on Noisy/ Reverberated Speech
- Metricbert: Text Representation Learning Via Self-Supervised Triplet Training
- Mimo Detection by Variational Posterior Inference
- Minimizing Residuals for Native-Nonnative Voice Conversion in a Sparse, Anchor-Based Representation of Speech
- Minimum Word Error Training For Non-Autoregressive Transformer-Based Code-Switching ASR
- Mining Hard Samples Locally And Globally For Improved Speech Separation
- Mismatched Supervised Learning
- Mitigating Closed-Model Adversarial Examples with Bayesian Neural Modeling for Enhanced End-to-End Speech Recognition
- Mixed In Time And Modality: Curse Or Blessingƒ Cross-Instance Data Augmentation for Weakly Supervised Multimodal Temporal Fusion
- Mixed Knowledge Relation Transformer for Image Captioning
- Mixed Precision DNN Quantization for Overlapped Speech Separation and Recognition
- Mixed Transformer U-Net for Medical Image Segmentation
- Mixer-TTS: Non-Autoregressive, Fast and Compact Text-to-Speech Model Conditioned on Language Model Embeddings
- Mixture Model Auto-Encoders: Deep Clustering Through Dictionary Learning
- Mmlatch: Bottom-Up Top-Down Fusion For Multimodal Sentiment Analysis
- Model Selection via Misspecified Cramér-Rao Bound Minimization
- Model-Based Approach for Measuring the Fairness in ASR
- Model-Based Online Learning for Resource Sharing in Joint Radar-Communication Systems
- Model-Based Reconstruction for Collimated Beam Ultrasound Systems
- Modeling Beats and Downbeats with a Time-Frequency Transformer
- Modeling Human Memory in Multi-Object Tracking with Transformers
- Modeling Intention, Emotion and External World in Dialogue Systems
- Modeling The Detection Capability Of High-Speed Spiking Cameras
- Modeling of Pre-Trained Neural Network Embeddings Learned From Raw Waveform for COVID-19 Infection Detection
- Modernn: Towards Fine-Grained Motion Details for Spatiotemporal Predictive Learning
- Modulo Event-Driven Sampling: System Identification and Hardware Experiments
- Monocular Vehicle 3D Bounding Box Estimation Using Homograhy and Geometry in Traffic Scene
- Monotonic Generalized Nash Games with Application to the Management of Energy-Aware Aloha Networks
- Motif-Topology and Reward-Learning Improved Spiking Neural Network for Efficient Multi-Sensory Integration
- Multi-ACCDOA: Localizing And Detecting Overlapping Sounds From The Same Class With Auxiliary Duplicating Permutation Invariant Training
- Multi-Channel Attentive Graph Convolutional Network with Sentiment Fusion for Multimodal Sentiment Analysis
- Multi-Channel End-To-End Neural Diarization with Distributed Microphones
- Multi-Channel Multi-Speaker ASR Using 3D Spatial Feature
- Multi-Channel Narrow-Band Deep Speech Separation with Full-Band Permutation Invariant Training
- Multi-Channel Speaker Diarization Using Spatial Features for Meetings
- Multi-Channel Speaker Verification with Conv-Tasnet Based Beamformer
- Multi-Channel Speech Denoising for Machine Ears
- Multi-Domain Unpaired Ultrasound Image Artifact Removal Using a Single Convolutional Neural Network
- Multi-Domain Unsupervised Image-to-Image Translation with Appearance Adaptive Convolution
- Multi-Feature Integration for Speaker Embedding Extraction
- Multi-Focus Guided Semantic Aggregation for Video Object Detection
- Multi-Frame Full-Rank Spatial Covariance Analysis for Underdetermined BSS in Reverberant Environments
- Multi-Frame Super-Resolution With Raw Images Via Modified Deformable Convolution
- Multi-Head Relu Implicit Neural Representation Networks
- Multi-Hierarchy Proxy Structure for Deep Metric Learning
- Multi-Level Contrastive Learning for Cross-Lingual Alignment
- Multi-Level Relation Aware Network for Person Re-Identification
- Multi-Level Spatial-Temporal Adaptation Network for Motor Imagery Classification
- Multi-Lingual Multi-Task Speech Emotion Recognition Using wav2vec 2.0
- Multi-Modal Acoustic-Articulatory Feature Fusion For Dysarthric Speech Recognition
- Multi-Modal Emotion Recognition with Self-Guided Modality Calibration
- Multi-Modal Learning with Text Merging for TEXTVQA
- Multi-Modal Pre-Training for Automated Speech Recognition
- Multi-Modal Recurrent Fusion for Indoor Localization
- Multi-Pose Virtual Try-On Via Self-Adaptive Feature Filtering
- Multi-Query Multi-Head Attention Pooling and Inter-Topk Penalty for Speaker Verification
- Multi-Relation Message Passing for Multi-Label Text Classification
- Multi-Role Event Argument Extraction as Machine Reading Comprehension with Argument Match Optimization
- Multi-Sample Subband Wavernn Via Multivariate Gaussian
- Multi-Scale Refinement Network Based Acoustic Echo Cancellation
- Multi-Scale Reinforcement Learning Strategy for Object Detection
- Multi-Scale Speaker Embedding-Based Graph Attention Networks For Speaker Diarisation
- Multi-Scale Temporal Frequency Convolutional Network With Axial Attention for Speech Enhancement
- Multi-Scale Temporal Frequency Convolutional Network with Axial Attention for Multi-Channel Speech Enhancement
- Multi-Speaker Pitch Tracking via Embodied Self-Supervised Learning
- Multi-Stage Graph Representation Learning for Dialogue-Level Speech Emotion Recognition
- Multi-Stage and Multi-Loss Training for Fullband Non-Personalized and Personalized Speech Enhancement
- Multi-Task Deep Residual Echo Suppression with Echo-Aware Loss
- Multi-Task Gaussian Process Regression for the Detection of Sleep Cycles in Premature Infants
- Multi-Task Learning Improves Synthetic Speech Detection
- Multi-Task Learning Improves the Brain Stoke Lesion Segmentation
- Multi-Task RNN-T with Semantic Decoder for Streamable Spoken Language Understanding
- Multi-Task Voice Activated Framework Using Self-Supervised Learning
- Multi-Task fMRI Data Fusion Using IVA and PARAFAC2
- Multi-Turn Incomplete Utterance Restoration As Object Detection
- Multi-Turn RNN-T for Streaming Recognition of Multi-Party Speech
- Multi-View And Multi-Modal Event Detection Utilizing Transformer-Based Multi-Sensor Fusion
- Multi-View Data Representation Via Deep Autoencoder-Like Nonnegative Matrix Factorization
- Multi-View Information Bottleneck Without Variational Approximation
- Multi-View Learning Based on Non-Redundant Fusion for Icu Patient Mortality Prediction
- Multi-View Self-Attention Based Transformer for Speaker Recognition
- Multiband Image Fusion with Controllable Error Guarantees
- Multichannel Noise Reduction Using Dilated Multichannel U-Net and Pre-Trained Single-Channel Network
- Multichannel Speech Enhancement Without Beamforming
- Multilingual Second-Pass Rescoring for Automatic Speech Recognition Systems
- Multilingual Text-To-Speech Training Using Cross Language Voice Conversion And Self-Supervised Learning Of Speech Representations
- Multimodal Depression Classification using Articulatory Coordination Features and Hierarchical Attention Based text Embeddings
- Multimodal Emotion Recognition with Surgical and Fabric Masks
- Multimodal Evaluation Method for Sound Event Detection
- Multimodal Graph Signal Denoising Via Twofold Graph Smoothness Regularization with Deep Algorithm Unrolling
- Multimodal Sentiment Analysis on Unaligned Sequences Via Holographic Embedding
- Multimodal Transformer with Learnable Frontend and Self Attention for Emotion Recognition
- Multiple Instance Learning with Task-Specific Multi-Level Features for Weakly Annotated Histopathological Image Classification
- Multiple Kernel K-Means Clustering with Simultaneous Spectral Rotation
- Multiple Offsets Multilateration: A New Paradigm for Sensor Network Calibration with Unsynchronized Reference Nodes
- Multiple Patch-Aware Network for Faster Real-World Image Dehazing
- Multiple Temporal Context Embedding Networks for Unsupervised time Series Anomaly Detection
- Multiplication-Avoiding Variant of Power Iteration with Applications
- Multiscale Attention Aggregation Network for 2D Vessel Segmentation
- Multiscale Crowd Counting and Localization By Multitask Point Supervision
- Multistream Neural Architectures for Cued Speech Recognition Using a Pre-Trained Visual Feature Extractor and Constrained CTC Decoding
- Multisv: Dataset for Far-Field Multi-Channel Speaker Verification
- Multitask Gaussian Process With Hierarchical Latent Interactions
- Multitask Sparse Neural Network for Hyperspectral Image Denoising
- Multivariate Multiscale Cosine Similarity Entropy
- Multiview Long-Short Spatial Contrastive Learning For 3D Medical Image Analysis
- Music Enhancement via Image Translation and Vocoding
- Music Identification Using Brain Responses to Initial Snippets
- Music Phrase Inpainting Using Long-Term Representation and Contrastive Loss
- Music Source Separation With Deep Equilibrium Models
- Musicyolo: A Sight-Singing Onset/Offset Detection Framework Based on Object Detection Instead of Spectrum Frames
- NEX+: Novel View Synthesis with Neural Regularisation Over Multi-Plane Images
- NFT-K: Non-Fungible Tangent Kernels
- NN3A: Neural Network Supported Acoustic Echo Cancellation, Noise Suppression and Automatic Gain Control for Real-Time Communications
- NVC-Net: End-To-End Adversarial Voice Conversion
- Natural-Looking Adversarial Examples from Freehand Sketches
- Navigating Audio-Visual Event Detection Across Mismatched Modalities
- Nearest Subspace Search in The Signed Cumulative Distribution Transform Space For 1d Signal Classification
- Neartracker: Acoustic 2-D Target Tracking with Nearby Reflector in Siso System
- Neighbor-Augmented Transformer-Based Embedding for Retrieval
- Netrca: An Effective Network Fault Cause Localization Algorithm
- Neufa: Neural Network Based End-to-End Forced Alignment with Bidirectional Attention Mechanism
- Neural Architecture Search for Speech Emotion Recognition
- Neural Audio-To-Score Music Transcription For Unconstrained Polyphony Using Compact Output Representations
- Neural Cascade Architecture for Joint Acoustic Echo and Noise Suppression
- Neural Collapse in Deep Homogeneous Classifiers and The Role of Weight Decay
- Neural Grapheme-To-Phoneme Conversion with Pre-Trained Grapheme Models
- Neural HMMS Are All You Need (For High-Quality Attention-Free TTS)
- Neural Network-Based Compression Framework for DOA Estimation Exploiting Distributed Array
- Neural Speech Synthesis on a Shoestring: Improving the Efficiency of Lpcnet
- Neural-FST Class Language Model for End-to-End Speech Recognition
- New Improved Criterion for Model Selection in Sparse High-Dimensional Linear Regression Models
- News Recommendation Via Multi-Interest News Sequence Modelling
- No More Than 6ft Apart: Robust K-Means via Radius Upper Bounds
- No-Reference Quality Assessment of Variable Frame-Rate Videos Using Temporal Bandpass Statistics
- Node Slicing Broad Learning System for Text Classification
- Node-Screening Tests For The L0-Penalized Least-Squares Problem
- Noise Suppression for Improved Few-Shot Learning
- Noise-Robust Speech Recognition With 10 Minutes Unparalleled In-Domain Data
- Non-Autoregressive ASR with Self-Conditioned Folded Encoders
- Non-Autoregressive End-To-End Automatic Speech Recognition Incorporating Downstream Natural Language Processing
- Non-Autoregressive Transformer with Unified Bidirectional Decoder for Automatic Speech Recognition
- Non-Invasive Blood Pressure Monitoring with Multi-Modal In-Ear Sensing
- Non-Rigid Transformation Based Adversarial Attack Against 3d Object Tracking
- Nonlinear Signal Decomposition Based on Block Sparse Approximation
- Nonverbal Sound Detection for Disordered Speech
- Not All Features are Equal: Selection of Robust Features for Speech Emotion Recognition in Noisy Environments
- Novel Class Discovery: A Dependency Approach
- Novel Instance Mining with Pseudo-Margin Evaluation for Few-Shot Object Detection
- OPTE: Online Per-Title Encoding for Live Video Streaming
- ORCA-PARTY: An Automatic Killer Whale Sound Type Separation Toolkit Using Deep Learning
- OT Cleaner: Label Correction as Optimal Transport
- Object Detection and Tracking in Ultrasound Scans Using an Optical Flow and Semantic Segmentation Framework Based on Convolutional Neural Networks
- Object-Oriented Backdoor Attack Against Image Captioning
- Occluded Person Re-Identification Via Relational Adaptive Feature Correction Learning
- Off-The-Grid Covariance-Based Super-Resolution Fluctuation Microscopy
- Off-the-Shelf Deep Integration For Residual-Echo Suppression
- Omni-Sparsity DNN: Fast Sparsity Optimization for On-Device Streaming E2E ASR Via Supernet
- On Adversarial Robustness Of Large-Scale Audio Visual Learning
- On Continuous-Domain Inverse Problems with Sparse Superpositions of Decaying Sinusoids as Solutions
- On Federated Learning with Energy Harvesting Clients
- On Identifiable Polytope Characterization for Polytopic Matrix Factorization
- On Language Model Integration for RNN Transducer Based Speech Recognition
- On Loss Functions and Evaluation Metrics for Music Source Separation
- On Mini-Batch Training with Varying Length Time Series
- On Spectral and Temporal Sparsification of Speech Signals for the Improvement of Speech Perception in CI Listeners
- On Submodular Set Cover Problems for Near-Optimal Online Kernel Basis Selection
- On Synchronization of Wireless Acoustic Sensor Networks in the Presence of Time-Varying Sampling Rate Offsets and Speaker Changes
- On The Convergence of ADAM-Type Algorithms for Solving Structured Single Node and Decentralized Min-Max Saddle Point Games
- On The Effectiveness of Active Learning by Uncertainty Sampling in Classification of High-Dimensional Gaussian Mixture Data
- On The Impact of Normalization Strategies in Unsupervised Adversarial Domain Adaptation for Acoustic Scene Classification
- On The Observability in Visual Slam Networks
- On The Relaxation of Orthogonal Tensor Rank and Its Nonconvex Riemannian Optimization for Tensor Completion
- On the Acquisition of Stationary Signals Using Uniform ADCS
- On the False Alarm Probability of the Normalized Matched Filter for Off-Grid Target Detection
- On the Importance of Different Frequency Bins for Speaker Verification
- On the Interplay between Sparsity, Naturalness, Intelligibility, and Prosody in Speech Synthesis
- On the Potential of Spatially-Spread Orthogonal Time Frequency Space Modulation for ISAC Transmissions
- On the Prediction of the Frequency Response of a Wooden Plate from Its Mechanical Parameters
- On the Stability of Low Pass Graph Filter with a Large Number of Edge Rewires
- On the Use of Component Structural Characteristics for Voxel Segmentation in Semicon 3D Images
- On the Use of Geodesic Triangles between Gaussian Distributions for Classification Problems
- One Model to Enhance Them All: Array Geometry Agnostic Multi-Channel Personalized Speech Enhancement
- One TTS Alignment to Rule Them All
- One-Shot Voice Conversion For Style Transfer Based On Speaker Adaptation
- Online Continual Learning Using Enhanced Random Vector Functional Link Networks
- Online Detection of Scalp-Invisible Mesial-Temporal Brain Interictal Epileptiform Discharges from EEG
- Online Ecg Biometrics Via Hadamard Code
- Online Learning for Latent Yule-Simon Processes
- Online Learning with Probabilistic Feedback
- OpenFEAT: Improving Speaker Identification by Open-Set Few-Shot Embedding Adaptation with Transformer
- Operator Formulation for Linear Transformations and Signal Estimation in the Joint Spatial-Slepian Domain
- Optimal Combination Policies for Adaptive Social Learning
- Optimal Qos-Aware Network Slicing for Service-Oriented Networks with Flexible Routing
- Optimal Resource Allocation and Beamforming for Two-User Miso WPCNS for a Non-Linear Circuit-Based EH Model : (Invited Paper)
- Optimization Guarantees for ISTA and ADMM Based Unfolded Networks
- Optimization of Compressive Light Field Display in Dual-Guided Learning
- Optimization of a Fixed Virtual Sensing Feedback ANC Controller For In-Ear Headphones with Multiple Loudspeakers
- Optimize Wav2vec2s Architecture for Small Training Set Through Analyzing its Pre-Trained Models Attention Pattern
- Optimizing Alignment of Speech and Language Latent Spaces for End-To-End Speech Recognition and Understanding
- Optimizing Latent Space Directions for Gan-Based Local Image Editing
- Optimizing The Consumption Of Spiking Neural Networks With Activity Regularization
- Optm3sec: Optimizing Multicast Irs-Aided Multiantenna Dfrc Secrecy Channel With Multiple Eavesdroppers
- Orthogonal Nonnegative Matrix Tri-Factorization for Community Detection in Multiplex Networks
- Out-Of-Distribution As A Target Class in Semi-Supervised Learning
- Over-Parameterized Network Solves Phase Retrieval Effectively
- Over-the-Air Personalized Federated Learning
- PAMA-TTS: Progression-Aware Monotonic Attention for Stable SEQ2SEQ TTS with Accurate Phoneme Duration Control
- PDD-Net: A Precise Defect Detection Network Based on Point Set Representation
- PEAR: Photographic Embedding for Aesthetic Rating
- PGTRNET: Two-Phase Weakly Supervised Object Detection with Pseudo Ground Truth Refinement
- PMP-NET: Rethinking Visual Context for Scene Graph Generation
- POPO: Pessimistic Offline Policy Optimization
- PU-Refiner: A Geometry Refiner with Adversarial Learning for Point Cloud Upsampling
- PVAE-TTS: Adaptive Text-to-Speech via Progressive Style Adaptation
- PYXIS: An Open-Source Performance Dataset Of Sparse Accelerators
- Pair-Level Supervised Contrastive Learning for Natural Language Inference
- Panchromatic Imagery Copy-Paste Localization Through Data-Driven Sensor Attribution
- Parallel Composition of Weighted Finite-State Transducers
- Parameter Estimation in Sparse Inverse Problems Using Bernoulli-Gaussian Prior
- Parameter-Free Style Projection for Arbitrary Image Style Transfer
- Parametric Modeling of Human Wrist for Bioimpedance-Based Physiological Sensing
- Parametric Models for Doa Trajectory Localization
- Part-of-Speech Models Compression Methods for on-Device Grapheme-to-Phoneme Conversion
- Partial Arithmetic Consensus based Distributed Intensity Particle Flow SMC-PHD Filter for Multi-Target Tracking
- Partial Variable Training for Efficient on-Device Federated Learning
- Partially Fake Audio Detection by Self-Attention-Based Fake Span Discovery
- Partially Relaxed Orthogonal Least Squares Weighted Subspace Fitting Direction-of-Arrival Estimation
- Pas-Mef: Multi-Exposure Image Fusion Based On Principal Component Analysis, Adaptive Well-Exposedness And Saliency Map
- Passtrans: An Improved Password Reuse Model Based on Transformer
- Patch Steganalysis: A Sampling Based Defense Against Adversarial Steganography
- Path Signatures for Non-Intrusive Load Monitoring
- Peer Collaborative Learning for Polyphonic Sound Event Detection
- Perfect Reconstruction of Classes of Non-Bandlimited Signals from Projections with Unknown Angles
- Performance Optimization for Wireless Semantic Communications over Energy Harvesting Networks
- Performance-Efficiency Trade-Offs in Unsupervised Pre-Training for Speech Recognition
- Personalized Automatic Speech Recognition Trained on Small Disordered Speech Datasets
- Personalized Pagerank Graph Attention Networks
- Personalized speech enhancement: new models and Comprehensive evaluation
- Phase Continuity: Learning Derivatives of Phase Spectrum for Speech Enhancement
- Phase Control of Parametric Array Loudspeaker by Optimizing Sideband Weights
- Phase Shifted Bedrosian Filterbank: An Interpretable Audio Front-End for Time-Domain Audio Source Separation
- Phase-Only Reconfigurable Sparse Array Beamforming Using Deep Learning
- Phone-Informed Refinement of Synthesized Mel Spectrogram for Data Augmentation in Speech Recognition
- Phone-to-Audio Alignment without Text: A Semi-Supervised Approach
- Phoneme Mispronunciation Detection By Jointly Learning To Align
- Phonology Recognition in American Sign Language
- Phonotactic Language Recognition Using A Universal Phoneme Recognizer and A Transformer Architecture
- Photon-Limited Deblurring Using Algorithm Unrolling
- Physical Layer Anonymous Communications: An Anonymity Entropy Oriented Precoding Design (Invited Paper)
- Picknet: Real-Time Channel Selection for Ad Hoc Microphone Arrays
- Pixel-Level and Affinity-Level Knowledge Distillation for Unsupervised Segmentation of Covid-19 Lesions
- Pixinwav: Residual Steganography for Hiding Pixels in Audio
- Plug-and-Play and Relay Regularizations on Noisy Low Rank Tensor Completion for Snapshot Multispectral Image Restoration
- Point Cloud Attribute Compression Via Chroma Subsampling
- Point Cloud Denoising Using Normal Vector-Based Graph Wavelet Shrinkage
- Point-Mass Filter with Decomposition of Transient Density
- Polyphone Disambiguation and Accent Prediction Using Pre-Trained Language Models in Japanese TTS Front-End
- Polyphonic Audio Event Detection: Multi-Label or Multi-Class Multi-Task Classification Problem?
- Position-Invariant Adversarial Attacks on Neural Modulation Recognition
- PostGAN: A GAN-Based Post-Processor to Enhance the Quality of Coded Speech
- Power Allocation for Wireless Federated Learning Using Graph Neural Networks
- Power-Efficient Hybrid MIMO Receiver with Task-Specific Beamforming using Low-Resolution ADCs
- Predicting Flat-Fading Channels via Meta-Learned Closed-Form Linear Filters and Equilibrium Propagation
- Predicting Human Motion Using Key Subsequences
- Predicting the Generalization Gap in Deep Models using Anchoring
- Preliminary Results on the Generation of Artificial Handwriting Data Using a Decomposition-Recombination Strategy
- Preserving Trajectory Privacy in Driving Data Release
- Prime Knowledge with Local Pattern Consistency for Knowledge Distillation
- Prior-Bert and Multi-Task Learning for Target-Aspect-Sentiment Joint Detection
- Privacy Attacks for Automatic Speech Recognition Acoustic Models in A Federated Learning Framework
- Privacy Protection In Learning Fair Representations
- Privacy Sensitive Speech Analysis Using Federated Learning to Assess Depression
- Privacy-Aware Communication over a Wiretap Channel with Generative Networks
- Privacy-Enhancing Appliance Filtering For Smart Meters
- Privacy-Preserving Action Recognition
- Privacy-Preserving Distributed Expectation Maximization for Gaussian Mixture Model Using Subspace Perturbation
- Privacy-Preserving Federated Multi-Task Linear Regression: A One-Shot Linear Mixing Approach Inspired By Graph Regularization
- Private Learning Via Knowledge Transfer with High-Dimensional Targets
- Probabilistic Fine-Grained Urban Flow Inference with Normalizing Flows
- Probably Pleasant? A Neural-Probabilistic Approach to Automatic Masker Selection for Urban Soundscape Augmentation
- Progressive Continual Learning for Spoken Keyword Spotting
- Progressive Image Super-Resolution via Neural Differential Equation
- Progressive Multi-Stage Neural Audio Coding with Guided References
- Progressive Teacher-Student Training Framework for Music Tagging
- Progressive-Granularity Retrieval Via Hierarchical Feature Alignment for Person Re-Identification
- Prosodyspeech: Towards Advanced Prosody Model for Neural Text-to-Speech
- Prosospeech: Enhancing Prosody with Quantized Vector Pre-Training in Text-To-Speech
- Prototype Learning for Interpretable Respiratory Sound Analysis
- Prototype-Based Inter-Camera Learning for Person Re-Identification
- Provable Sample Complexity Guarantees For Learning Of Continuous-Action Graphical Games With Nonparametric Utilities
- Provable Second-Order Riemannian Gauss-Newton Method for Low-Rank Tensor Estimation ‖
- Proximal-Based Adaptive Simulated Annealing for Global Optimization
- Pseudo Strong Labels for Large Scale Weakly Supervised Audio Tagging
- Pseudo-Interacting Guided Network for Few-Shot Segmentation
- Pseudo-Label Transfer from Frame-Level to Note-Level in a Teacher-Student Framework for Singing Transcription from Polyphonic Music
- Pseudo-Labeling for Massively Multilingual Speech Recognition
- Punctuation Prediction for Streaming On-Device Speech Recognition
- Pyramid Fusion Attention Network For Single Image Super-Resolution
- QA4QG: Using Question Answering to Constrain Multi-Hop Question Generation
- Qrelation: an Agent Relation-Based Approach for Multi-Agent Reinforcement Learning Value Function Factorization
- Quantifying Discriminability between NMF Bases
- Quantization-Aware Precoding For Mu-Mimo With Limited-Capacity Fronthaul
- Quantized Winograd Acceleration for CONV1D Equipped ASR Models on Mobile Devices
- Quantum Federated Learning with Quantum Data
- Quantum Long Short-Term Memory
- Quickest Detection of Composite and Non-Stationary Changes with Application to Pandemic Monitoring
- RCANet: Row-Column Attention Network for Semantic Segmentation
- RIS-Aided Monostatic Mimo Radar with Co-Located Antennas
- RTSNet: Deep Learning Aided Kalman Smoothing
- Randomized Smoothing Under Attack: How Good is it in Practice?
- Rangeinet: Fast Lidar Point Cloud Temporal Interpolation
- Rank-Based Loss For Learning Hierarchical Representations
- Rate Coding Or Direct Coding: Which One Is Better For Accurate, Robust, And Energy-Efficient Spiking Neural Networks?
- Rate Control for Learned Video Compression
- Rational Arrays for DOA Estimation
- Raw Plenoptic Video Coding Under Hexagonal Lattice Resolution of Motion Vectors
- Raw Source and Filter Modelling for Dysarthric Speech Recognition
- RawNeXt: Speaker Verification System For Variable-Duration Utterances With Deep Layer Aggregation And Extended Dynamic Scaling Policies
- Rawboost: A Raw Data Boosting and Augmentation Method Applied to Automatic Speaker Verification Anti-Spoofing
- Real Additive Margin Softmax for Speaker Verification
- Real-M: Towards Speech Separation on Real Mixtures
- Real-Time Fall Detection Using Mmwave Radar
- Real-World Adversarial Examples Via Makeup
- Real-World On-Board Uav Audio Data Set For Propeller Anomalies
- Realistic Monocular-To-3d Virtual Try-On Via Multi-Scale Characteristics Capture
- Recognition Of Silently Spoken Word From Eeg Signals Using Dense Attention Network (DAN)
- Recovery of Graph Signals From Sign Measurements
- Recovery of Noisy Pooled Tests via Learned Factor Graphs with Application to COVID-19 Testing
- Recurrent Design of Probing Waveform for Sparse Bayesian Learning Based DOA Estimation
- Referee: Towards Reference-Free Cross-Speaker Style Transfer with Low-Quality Data for Expressive Speech Synthesis
- Reference Microphone Selection and Low-Rank Approximation Based Multichannel Wiener Filter with Application to Speech Recognition
- Reformulating Speaker Diarization As Community Detection With Emphasis On Topological Structure
- Region-to-Region Kernel Interpolation of Acoustic Transfer Function with Directional Weighting
- Regression Assisted Matrix Completion for Reconstructing a Propagation Field with Application to Source Localization
- Regularization Using Denoising: Exact and Robust Signal Recovery
- Regularized Latent Space Exploration for Discriminative Face Super-Resolution
- Relation Discovery in Nonlinearly Related Large-Scale Settings
- Relative Viewpoint Estimation Based on Structured 3d Representation Alignment
- Remix-Cycle-Consistent Learning on Adversarially Learned Separator for Accurate and Stable Unsupervised Speech Separation
- Repeat after Me: Self-Supervised Learning of Acoustic-to-Articulatory Mapping by Vocal Imitation
- Repetition Assessment for Speech and Language Disorders: A Study of the Logopenic Variant of Primary Progressive Aphasia
- Representation Learning Through Cross-Modal Conditional Teacher-Student Training For Speech Emotion Recognition
- RescoreBERT: Discriminative Speech Recognition Rescoring With Bert
- Residual Recovery Algorithm for Modulo Sampling
- Residual-Guided Personalized Speech Synthesis based on Face Image
- Restless Multi-Armed Bandits under Exogenous Global Markov Process
- Rethinking Computer-Aided Pelvis Segmentation
- Rethinking Two-B-Real Net for Real-Time Salient Object Detection
- Retrieval Bias Aware Ensemble Model for Conditional Sentence Generation
- Retrieval Enhanced Segment Generation Neural Network for Task-Oriented Dialogue Systems
- Retrieving Speaker Information from Personalized Acoustic Models for Speech Recognition
- Robust Adaptive Beamforming Based on Power Method Processing and Spatial Spectrum Matching
- Robust Adaptive Beamforming Maximizing the Worst-Case SINR Over Distributional Uncertainty Sets for Random INC Matrix And Signal Steering Vector
- Robust Adaptive Noise Canceller Algorithm with Snr-Based Stepsize Control and Noise-Path Gain Compensation
- Robust Bayesian Reconstruction of Multispectral Single-Photon 3D Lidar Data with Non-Uniform Background
- Robust Classification with Flexible Discriminant Analysis in Heterogeneous Data
- Robust Collaborative Learning for Sequence Modelling
- Robust Disentangled Variational Speech Representation Learning for Zero-Shot Voice Conversion
- Robust High-Order Tensor Recovery Via Nonconvex Low-Rank Approximation
- Robust Nonparametric Distribution Forecast with Backtest-Based Bootstrap and Adaptive Residual Selection
- Robust Parameter Estimation Based on the K-Divergence
- Robust Pressure Matching with ATF Perturbation Constraints for Sound Field Control
- Robust Self-Supervised Speaker Representation Learning Via Instance Mix Regularization
- Robust Signal Processing Over Simplicial Complexes
- Robust Speaker Verification Using Population-Based Data Augmentation
- Robust Speaker Verification with Joint Self-Supervised and Supervised Learning
- Robust Thermal Infrared Pedestrian Detection By Associating Visible Pedestrian Knowledge
- Robust Unstructured Knowledge Access in Conversational Dialogue with ASR Errors
- Robust Video Hashing Based on Local Fluctuation Preserving for Tracking Deep Fake Videos
- Robust and Efficient Uncertainty Aware Biosignal Classification via Early Exit Ensembles
- Run-and-Back Stitch Search: Novel Block Synchronous Decoding For Streaming Encoder-Decoder ASR
- S-DCCRN: Super Wide Band DCCRN with Learnable Complex Feature for Speech Enhancement
- S2 Reducer: High-Performance Sparse Communication to Accelerate Distributed Deep Learning
- S3PRL-VC: Open-Source Voice Conversion Framework with Self-Supervised Speech Representations
- S3T: Self-Supervised Pre-Training with Swin Transformer For Music Classification
- SA-SDR: A Novel Loss Function for Separation of Meeting Style Data
- SADN: Learned Light Field Image Compression with Spatial-Angular Decorrelation
- SAGA: Self-Augmentation with Guided Attention for Representation Learning
- SALSA-Lite: A Fast and Effective Feature for Polyphonic Sound Event Localization and Detection with Microphone Arrays
- SDETR: Attention-Guided Salient Object Detection with Transformer
- SDNET: Lightweight Facial Expression Recognition For Sample Disequilibrium
- SDR - Medium Rare with Fast Computations
- SERAB: A Multi-Lingual Benchmark for Speech Emotion Recognition
- SIG-VC: A Speaker Information Guided Zero-Shot Voice Conversion System for Both Human Beings and Machines
- SLUE: New Benchmark Tasks For Spoken Language Understanding Evaluation on Natural Speech
- SODA: Self-Organizing Data Augmentation in Deep Neural Networks Application to Biomedical Image Segmentation Tasks
- SP Attack: Single-Perspective Attack for Generating Adversarial Omnidirectional Images
- SQAPP: No-Reference Speech Quality Assessment Via Pairwise Preference
- SRP-DNN: Learning Direct-Path Phase Difference for Multiple Moving Sound Source Localization
- SRU++: Pioneering Fast Recurrence with Attention for Speech Recognition
- SYNT++: Utilizing Imperfect Synthetic Data to Improve Speech Recognition
- Safari from Visual Signals: Recovering Volumetric 3d Shapes
- Safeguarding UAV Networks through Integrated Sensing, Jamming, and Communications
- Sain: Similarity-Aware Video Frame Interpolation
- Sampling Set Selection for Graph Signals under Arbitrary Signal Priors
- Sar-Shipnet: Sar-Ship Detection Neural Network via Bidirectional Coordinate Attention and Multi-Resolution Feature Fusion
- Scalable Data Association and Multi-Target Tracking Under a Poisson Mixture Measurement Process
- Scalable Neural Architectures for End-to-End Environmental Sound Classification
- Scalable Ridge Leverage Score Sampling for the Nyström Method
- Scattering Statistics of Generalized Spatial Poisson Point Processes
- Score Difficulty Analysis for Piano Performance Education based on Fingering
- Screen & Relax: Accelerating The Resolution Of Elastic-Net By Safe Identification of The Solution Support
- SecMPNN: 3-Party Privacy-Preserving Molecular Structure Properties Inference
- Seed: Sound Event Early Detection Via Evidential Uncertainty
- SegNet-Based Deep Representation Learning for Dysphagia Classification
- Seismic Fault Identification Using Graph High-Frequency Components as Input to Graph Convolutional Network
- Selective Multi-Task Learning For Speech Emotion Recognition Using Corpora Of Different Styles
- Selective Mutual Learning: An Efficient Approach for Single Channel Speech Separation
- Selective Scale Cascade Attention Network for Breast Cancer Histopathology Image Classification
- Self Supervised Representation Learning with Deep Clustering for Acoustic Unit Discovery from Raw Speech
- Self-Attention for Incomplete Utterance Rewriting
- Self-Critical Sequence Training for Automatic Speech Recognition
- Self-Ensemble Variance Regularization for Domain Adaptation
- Self-Knowledge Distillation based Self-Supervised Learning for Covid-19 Detection from Chest X-Ray Images
- Self-Knowledge Distillation via Feature Enhancement for Speaker Verification
- Self-Learned Video Super-Resolution with Augmented Spatial and Temporal Context
- Self-Supervised Acoustic Anomaly Detection Via Contrastive Learning
- Self-Supervised Contrastive Learning for Cross-Domain Hyperspectral Image Representation
- Self-Supervised Learning Method Using Multiple Sampling Strategies for General-Purpose Audio Representation
- Self-Supervised Learning for Sentiment Analysis via Image-Text Matching
- Self-Supervised Learning on A Lightweight Low-Light Image Enhancement Model with Curve Refinement
- Self-Supervised Representation Learning for Unsupervised Anomalous Sound Detection Under Domain Shift
- Self-Supervised Speaker Recognition Training using Human-Machine Dialogues
- Self-Supervised Speaker Recognition with Loss-Gated Learning
- Self-Supervised Speaker Verification with Simple Siamese Network and Self-Supervised Regularization
- Semantic Association Network for Video Corpus Moment Retrieval
- Semantically Proportional Patchmix for Few-Shot Learning
- Semi-Supervised 360° Depth Estimation from Multiple Fisheye Cameras with Pixel-Level Selective Loss
- Semi-Supervised Gaussian Mixture Variational Autoencoder for Pulse Shape Discrimination
- Semi-Supervised Source Localization With Residual Physical Learning
- Semi-Supervised Standardized Detection of Periodic Signals with Application to Exoplanet Detection
- Semidefinite Relaxation Method for Moving Object Localization Using a Stationary Transmitter at Unknown Position
- Sensing-Assisted Beam Tracking in V2I Networks: Extended Target Case
- Sensors to Sign Language: A Natural Approach to Equitable Communication
- Sentiment-Aware Automatic Speech Recognition Pre-Training for Enhanced Speech Emotion Recognition
- Sentiment-Aware Distillation for Bitcoin Trend Forecasting Under Partial Observability
- Sequence Transduction with Graph-Based Supervision
- Sequential MCMC Methods for Audio Signal Enhancement
- Shared Transformer Encoder with Mask-Based 3d Model Estimation for Container Mass Estimation
- Short-and-Sparse Deconvolution Via Rank-One Constrained Optimization (Roco)
- Signal Compression via Neural Implicit Representations
- Signal Processing On Cell Complexes
- Signal Recovery from Inconsistent Nonlinear Observations
- Simple Attention Module Based Speaker Verification with Iterative Noisy Label Detection
- Simpler is Better: Spectral Regularization and Up-Sampling Techniques for Variational Autoencoders
- Simplicial Convolutional Neural Networks
- Simulation-and-Mining: Towards Accurate Source-Free Unsupervised Domain Adaptive Object Detection
- Simultaneous Nonlocal Low-Rank And Deep Priors For Poisson Denoising
- Single Image De-Raining with High-Low Frequency Guidance
- Single-Shot Balanced Detector for Geospatial Object Detection
- Sketch Storytelling
- Sketched RT3D: How to Reconstruct Billions of Photons Per Second
- Skim: Skipping Memory Lstm for Low-Latency Real-Time Continuous Speech Separation
- SleepGAN: Towards Personalized Sleep Therapy Music
- Slim: Explicit Slot-Intent Mapping with Bert for Joint Multi-Intent Detection and Slot Filling
- Social Welfare Maximization in Cross-Silo Federated Learning
- Solving The Long-Tailed Problem Via Intra- And Inter-Category Balance
- Sound Event Detection Guided by Semantic Contexts of Scenes
- Source Separation By Steering Pretrained Music Models
- Spain-Net: Spatially-Informed Stereophonic Music Source Separation
- Sparse Adversarial Attack For Video Via Gradient-Based Keyframe Selection
- Sparse Array Source Enumeration Via Coarray Subspace Optimization
- Sparse Modeling of The Early Part of Noisy Room Impulse Responses with Sparse Bayesian Learning
- Sparse Multi-Reference Alignment: Sample Complexity and Computational Hardness
- Sparse Recovery of Acoustic Waves
- Sparse Self-Attention for Semi-Supervised Sound Event Detection
- Sparse Subspace Tracking in High Dimensions
- Sparse-Group Log-Sum Penalized Graphical Model Learning For Time Series
- SparseBFA: Attacking Sparse Deep Neural Networks with the Worst-Case Bit Flips on Coordinates
- Sparsity Improves Unsupervised Attribute Discovery in Stylegan
- Sparsity-Based Sound Field Separation in the Spherical Harmonics Domain
- Spatial Active Noise Control Based on Individual Kernel Interpolation of Primary and Secondary Sound Fields
- Spatial Active Noise Control with the Remote Microphone Technique: an Approach with a Moving Higher Order Microphone
- Spatial Data Augmentation with Simulated Room Impulse Responses for Sound Event Localization and Detection
- Spatial Mixup: Directional Loudness Modification as Data Augmentation for Sound Event Localization and Detection
- Spatial Processing Front-End for Distant ASR Exploiting Self-Attention Channel Combinator
- Spatial-Context-Aware Deep Neural Network for Multi-Class Image Classification
- Spatial-Temporal Graph Convolution Network for Multichannel Speech Enhancement
- Spatio-Temporal Attention Graph Convolution Network for Functional Connectome Classification
- Spatio-Temporal Graph Complementary Scattering Networks
- Spatio-Temporal Graph Convolutional Networks for Continuous Sign Language Recognition
- Spatio-Temporal Motion Aggregation Network for Video Action Detection
- Spatio-Temporal PRRS Epidemic Forecasting via Factorized Deep Generative Modeling
- Speaker Embedding Conversion for Backward and Cross-Channel Compatibility
- Speaker Generation
- Speaker Identity Preservation in Dysarthric Speech Reconstruction by Adversarial Speaker Adaptation
- Speaker Normalization for Self-Supervised Speech Emotion Recognition
- Speaker Reinforcement Using Target Source Extraction for Robust Automatic Speech Recognition
- Speaker-Targeted Audio-Visual Speech Recognition Using a Hybrid CTC/Attention Model with Interference Loss
- Specialised Video Quality Model For Enhanced User Generated Content (UGC) With Special Effects
- Spectral Permutation Test on Persistence Diagrams
- Spectral-Spatial Symmetrical Aggregation Cross-Linking Multi-Modal Data Fusion Network
- Speech Denoising in the Waveform Domain With Self-Attention
- Speech Emotion Recognition Using Self-Supervised Features
- Speech Emotion Recognition with Co-Attention Based Multi-Level Acoustic Information
- Speech Emotion Recognition with Global-Aware Fusion on Multi-Scale Feature Representation
- Speech Enhancement for Low Bit Rate Speech Codec
- Speech Enhancement with Neural Homomorphic Synthesis
- Speech Pattern Based Black-Box Model Watermarking for Automatic Speech Recognition
- Speech Recognition Using Biologically-Inspired Neural Networks
- Speech Recovery For Real-World Self-Powered Intermittent Devices
- Speech Tasks Relevant to Sleepiness Determined With Deep Transfer Learning
- SpeechSplit2.0: Unsupervised Speech Disentanglement for Voice Conversion without Tuning Autoencoder Bottlenecks
- Speechmoe2: Mixture-of-Experts Model with Improved Routing
- Spell My Name: Keyword Boosted Speech Recognition
- Spherical Convolutional Recurrent Neural Network for Real-Time Sound Source Tracking
- Spoken Language Recognition with Cluster-Based Modeling
- Stability Analysis of Unfolded WMMSE for Power Allocation
- Stability of Neural Networks on Manifolds to Relative Perturbations
- Stable and Transferable Wireless Resource Allocation Policies Via Manifold Neural Networks
- Stacked Multi-Scale Attention Network for Image Colorization
- Statistical Pyramid Dense Time Delay Neural Network for Speaker Verification
- Statistical, Spectral and Graph Representations for Video-Based Facial Expression Recognition in Children
- Stealthy Backdoor Attack with Adversarial Training
- Stgat-Mad : Spatial-Temporal Graph Attention Network For Multivariate Time Series Anomaly Detection
- Stpointgcn: Spatial Temporal Graph Convolutional Network for Multiple People Recognition Using Millimeter-Wave Radar
- Streaming Transformer Transducer based Speech Recognition Using Non-Causal Convolution
- Streaming on-Device Detection of Device Directed Speech from Voice and Touch-Based Invocation
- Structural Prior Models for 3-D Deep Vessel Segmentation
- Study of Positional Encoding Approaches for Audio Spectrogram Transformers
- Study of the Null Directions on The Performance of Differential Beamformers
- Study on Time-of-Flight Estimation in Ultrasonic Well Logging Tool: Model-Driven Transfer Learning
- Studying Three Families of Divergences to Compare Wide-Sense Stationary Gaussian Arma Processes
- Stylegan-Induced Data-Driven Regularization for Inverse Problems
- Subgraph Representation Learning with Hard Negative Samples for Inductive Link Prediction
- Subjective And Objective Quality Assessment Of Mobile Gaming Video
- Subspace Clustering Using Unsupervised Data Augmentation
- Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
- Super-Resolution of Satellite Images by two-Dimensional RRDB and Edge-Enhancement Generative Adversarial Network
- Superresolution and Segmentation of OCT Scans Using Multi-Stage Adversarial Guided Attention Training
- Supervised Attention in Sequence-to-Sequence Models for Speech Recognition
- Supervised Learning Based Sparse Channel Estimation For RIS Aided Communications
- Supervised Training of Siamese Spiking Neural Networks with Earth Mover's Distance
- Supervised and Self-Supervised Pretraining Based Covid-19 Detection Using Acoustic Breathing/Cough/Speech Signals
- Symbol-Level Online Channel Tracking for Deep Receivers
- Synergistic Network Learning and Label Correction for Noise-Robust Image Classification
- Synpose: A Large-Scale and Densely Annotated Synthetic Dataset for Human Pose Estimation in Classroom
- Syntax-Based Graph Matching for Knowledge Base Question Answering
- Synthesis of Adversarial Samples in Two-Stage Classifiers
- Synthesizing Dysarthric Speech Using Multi-Speaker Tts For Dysarthric Speech Recognition
- T-NGA: Temporal Network Grafting Algorithm for Learning to Process Spiking Audio Sensor Events
- T-SVD Based Broadband Non-Synchronous Measurements
- TCRNet: Make Transformer, CNN and RNN Complement Each Other
- TEA-PSE: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System for ICASSP 2022 DNS Challenge
- TED Talk Teaser Generation with Pre-Trained Models
- TFPSNet: Time-Frequency Domain Path Scanning Network for Speech Separation
- TH-Net: A Method Of Single 3d Object Tracking Based On Transformers And Hausdorff Distance
- TINYS2I: A Small-Footprint Utterance Classification Model with Contextual Support for On-Device SLU
- TNTC: Two-Stream Network with Transformer-Based Complementarity for Gait-Based Emotion Recognition
- TP-VIT: A Two-Pathway Vision Transformer for Video Action Recognition
- TPARN: Triple-Path Attentive Recurrent Network for Time-Domain Multichannel Speech Enhancement
- Tackling Data Scarcity in Speech Translation Using Zero-Shot Multilingual Machine Translation Techniques
- Tackling the Score Shift in Cross-Lingual Speaker Verification by Exploiting Language Information
- TalkingFlow: Talking Facial Landmark Generation with Multi-Scale Normalizing Flow Network
- Target-Aware Auto-Augmentation for Unsupervised Domain Adaptive Object Detection
- TargetDrop: A Targeted Regularization Method for Convolutional Neural Networks
- Teaching CNNs to Mimic Human Visual Cognitive Process & Regularise Texture-Shape Bias
- Tempo: Improving Training Performance in Cross-Silo Federated Learning
- Temporal Contrastive-Loss for Audio Event Detection
- Temporal Cross-Graph Network for Brain Functional Activity Prediction
- Temporal Dynamic Convolutional Neural Network for Text-Independent Speaker Verification and Phonemic Analysis
- Temporal Early Exiting for Streaming Speech Commands Recognition
- Temporal Knowledge Distillation for on-device Audio Classification
- Tensor-Based Orthogonal Matching Pursuit with Phase Rotation for Channel Estimation In Hybrid Beamforming Mimo-Ofdm Systems
- Terahertz Image Restoration Benchmarking Dataset
- Test-Time Detection of Backdoor Triggers for Poisoned Deep Neural Networks
- Text Adaptive Detection for Customizable Keyword Spotting
- Text-Free Non-Parallel Many-To-Many Voice Conversion Using Normalising Flow
- Text-Image De-Contextualization Detection Using Vision-Language Models
- Text2Poster: Laying Out Stylized Texts on Retrieved Images
- Text2video: Text-Driven Talking-Head Video Synthesis with Personalized Phoneme - Pose Dictionary
- Texture Information Boosts Video Quality Assessment
- The CUHK-Tencent Speaker Diarization System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
- The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks
- The Coral++ Algorithm for Unsupervised Domain Adaptation of Speaker Recognition
- The DKU Audio-Visual Wake Word Spotting System for the 2021 MISP Challenge
- The Data/Identity Tradeoff with Censored Sensors
- The Dawn of Quantum Natural Language Processing
- The First Multimodal Information Based Speech Processing (Misp) Challenge: Data, Tasks, Baselines And Results
- The Impact of JPEG Compression on Prior Image Noise
- The Impact of Removing Head Movements on Audio-Visual Speech Enhancement
- The Mirrornet : Learning Audio Synthesizer Controls Inspired by Sensorimotor Interaction
- The PCG-AIID System for L3DAS22 Challenge: MIMO and MISO Convolutional Recurrent Network for Multi Channel Speech Enhancement and Speech Recognition
- The Prototype Co-Prime Array with a Robust Difference Co-Array
- The Representation Jensen-Rényi Divergence
- The Royalflush System of Speech Recognition for M2met Challenge
- The Second Dicova Challenge: Dataset and Performance Analysis for Diagnosis of Covid-19 Using Acoustics
- The Sjtu System For Multimodal Information Based Speech Processing Challenge 2021
- The USTC-Ximalaya System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription (M2met) Challenge
- The Vicomtech Audio Deepfake Detection System Based on Wav2vec2 for the 2022 ADD Challenge
- The Volcspeech System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
- The impact of cross language on acoustic-to-articulatory inversion and its influence on articulatory speech synthesis
- Thin Slices of Depression: Improving Depression Detection Performance Through Data Segmentation
- Threshold Independent Evaluation of Sound Event Detection Scores
- Tie Your Embeddings Down: Cross-Modal Latent Spaces for End-to-end Spoken Language Understanding
- Tight Integration Of Neural- And Clustering-Based Diarization Through Deep Unfolding Of Infinite Gaussian Mixture Model
- Time Domain Adversarial Voice Conversion for ADD 2022
- Time Domain Radial Filter Design for Spherical Waves
- Time-Balanced Focal Loss for Audio Event Detection
- Time-Domain Acoustic Contrast Control with A Spatial Uniformity Constraint for Personal Audio Systems
- Time-Domain Audio-Visual Speech Separation on Low Quality Videos
- Time-Frequency Attention for Monaural Speech Enhancement
- Time-Frequency and Geometric Analysis of Task-Dependent Learning in Raw Waveform Based Acoustic Models
- TitaNet: Neural Model for Speaker Representation with 1D Depth-Wise Separable Convolutions and Global Context
- To Catch A Chorus, Verse, Intro, or Anything Else: Analyzing a Song with Structural Functions
- Tonet: Tone-Octave Network for Singing Melody Extraction from Polyphonic Music
- Topological Correlation of Brain Signals
- Torchaudio: Building Blocks for Audio and Speech Processing
- Toward Degradation-Robust Voice Conversion
- Toward mmWave-Based Sound Enhancement and Separation
- Towards A Common Speech Analysis Engine
- Towards Accurate Cross-Domain in-Bed Human Pose Estimation
- Towards Automatic Transcription of Polyphonic Electric Guitar Music: A New Dataset and a Multi-Loss Transformer Model
- Towards Better Meta-Initialization with Task Augmentation for Kindergarten-Aged Speech Recognition
- Towards Closed-Loop Speech Synthesis from Stereotactic EEG: A Unit Selection Approach
- Towards Controllable and Physical Interpretable Underwater Scene Simulation
- Towards End-to-End Integration of Dialog History for Improved Spoken Language Understanding
- Towards Expressive Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis
- Towards Fast And Convenient End-To-End HRTF Personalization
- Towards Faster Continuous Multi-Channel HRTF Measurements Based On Learning System Models
- Towards Identity Preserving Normal to Dysarthric Voice Conversion
- Towards Interpretability of Speech Pause in Dementia Detection Using Adversarial Learning
- Towards Interpreting Deep Learning Models to Understand Loss of Speech Intelligibility in Speech Disorders Step 2: Contribution of the Emergence of Phonetic Traits
- Towards Joint Frame-Level and MOS Quality Predictions with Low-Complexity Objective Models
- Towards Learning Universal Audio Representations
- Towards Lifelong Learning of Multilingual Text-to-Speech Synthesis
- Towards Lightweight Applications: Asymmetric Enroll-Verify Structure for Speaker Verification
- Towards Low-Distortion Multi-Channel Speech Enhancement: The ESPNET-Se Submission to the L3DAS22 Challenge
- Towards Measuring Fairness in Speech Recognition: Casual Conversations Dataset Transcriptions
- Towards Practical and Efficient Long Video Summary
- Towards Reducing the Need for Speech Training Data to Build Spoken Language Understanding Systems
- Towards Robust Speech-to-Text Adversarial Attack
- Towards Robust Visual Transformer Networks via K-Sparse Attention
- Towards Speaker Age Estimation With Label Distribution Learning
- Towards Transferable Speech Emotion Representation: On Loss Functions for Cross-Lingual Latent Representations
- Towards Using Clothes Style Transfer for Scenario-Aware Person Video Generation
- Towards end-to-end Speaker Diarization with Generalized Neural Speaker Clustering
- Tracking the Dimensions of Latent Spaces of Gaussian Process Latent Variable Models
- Training Privacy-Preserving Video Analytics Pipelines by Suppressing Features That Reveal Information About Private Attributes
- Training Robust Zero-Shot Voice Conversion Models with Self-Supervised Features
- Training Stable Graph Neural Networks Through Constrained Learning
- Training Strategies for Automatic Song Writing: A Unified Framework Perspective
- Training Strategies for Improved Lip-Reading
- Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers Using End-to-End Speaker-Attributed ASR
- Transducer-Based Streaming Deliberation for Cascaded Encoders
- Transductive Clip with Class-Conditional Contrastive Learning
- Transformer-Based Domain Adaptation for Event Data Classification
- Transformer-Based Estimation of Spoken Sentences Using Electrocorticography
- Transformer-Based Multi-Aspect Multi-Granularity Non-Native English Speaker Pronunciation Assessment
- Transformer-Based Person Search Model with Symmetric Online Instance Matching
- Transformer-Based Streaming ASR with Cumulative Attention
- Transformer-S2A: Robust and Efficient Speech-to-Animation
- Transient Analysis of Clustered Multitask Diffusion RLS Algorithm
- Transient Detection with Unknown Statistics Via Source Coding
- Transmit Beamforming with Fixed Covariance for Integrated MIMO Radar and Multiuser Communications
- Transtl: Spatial-Temporal Localization Transformer for Multi-Label Video Classification
- TriBYOL: Triplet BYOL for Self-Supervised Representation Learning
- Tts4pretrain 2.0: Advancing the use of Text and Speech in ASR Pretraining with Consistency and Contrastive Losses
- Tunet: A Block-Online Bandwidth Extension Model Based On Transformers And Self-Supervised Pretraining
- Turn-to-Diarize: Online Speaker Diarization Constrained by Transformer Transducer Speaker Turn Detection
- Two Strategies Toward Lightweight Image Super-Resolution
- Two-Path GMM-ResNet and GMM-SENet for ASV Spoofing Detection
- Two-Snapshot DOA Estimation Via Hankel-Structured Matrix Completion
- Type-Aware Medical Visual Question Answering
- U-GAT-VC: Unsupervised Generative Attentional Networks for Non-Parallel Voice Conversion
- UNET-TTS: Improving Unseen Speaker and Style Transfer in One-Shot Voice Cloning
- Ubilung: Multi-Modal Passive-Based Lung Health Assessment
- Ubiquitous Physiological Prediction of SUD Patients' Wellness State Using Memory-Based Convolutional Models
- Uformer: A Unet Based Dilated Complex & Real Dual-Path Conformer Network for Simultaneous Speech Enhancement and Dereverberation
- Uncertainty Estimation with a VAE-Classifier Hybrid Model
- Uncertainty in Data-Driven Kalman Filtering for Partially Known State-Space Models
- Underdetermined Two-Dimensional Localization for Wideband Sources Based on Distributed Sensor Array Networks
- Underwater Image Enhancement Via Learning Water Type Desensitized Representations
- Underwater Small Target Detection Based on Deformable Convolutional Pyramid
- Underwater Stereo Matching Via Unsupervised Appearance And Feature Adaptation Networks
- Unfolding Model-Based Beamforming for High Quality Ultrasound Imaging
- Unified Matrix Coding for NN Originated MIP in H.266/VVC
- Unified Multimodal Punctuation Restoration Framework for Mixed-Modality Corpus
- Unified Speculation, Detection, and Verification Keyword Spotting
- Unimodular Waveform Design with Low Correlation Levels: A Fast Algorithm Development to Support Large-Scale Code Lengths
- Unispeech-Sat: Universal Speech Representation Learning With Speaker Aware Pre-Training
- Universal Efficient Variable-Rate Neural Image Compression
- Universal Paralinguistic Speech Representations Using self-Supervised Conformers
- Unlimited Sampling with Local Averages
- Unlimited Sampling with Sparse Outliers: Experiments with Impulsive and Jump or Reset Noise
- Unrolling Particles: Unsupervised Learning of Sampling Distributions
- Unsupervised Anomaly Detection for Container Cloud Via BILSTM-Based Variational Auto-Encoder
- Unsupervised Audio-Caption Aligning Learns Correspondences Between Individual Sound Events and Textual Phrases
- Unsupervised Clustering and Analysis of Contraction-Dependent Fetal Heart Rate Segments
- Unsupervised Contrastive Hashing for Cross-Modal Retrieval in Remote Sensing
- Unsupervised Data Selection for Speech Recognition with Contrastive Loss Ratios
- Unsupervised Deep Learning Network for Deformable Fundus Image Registration
- Unsupervised Hierarchical Translation-Based Model for Multi-Modal Medical Image Registration
- Unsupervised Model Adaptation for End-to-End ASR
- Unsupervised Speech Enhancement with Speech Recognition Embedding and Disentanglement Losses
- Unsupervised Word-Level Prosody Tagging for Controllable Speech Synthesis
- Unsupervised and Untrained Underwater Image Restoration Based on Physical Image Formation Model
- Upmixing Via Style Transfer: A Variational Autoencoder for Disentangling Spatial Images And Musical Content
- Urban Sound & Sight: Dataset And Benchmark For Audio-Visual Urban Scene Understanding
- User Scheduling Using Graph Neural Networks for Reconfigurable Intelligent Surface Assisted Multiuser Downlink Communications
- Using Acoustic Deep Neural Network Embeddings to Detect Multiple Sclerosis From Speech
- Using Multiple Reference Audios and Style Embedding Constraints for Speech Synthesis
- Using Spectral Sequence-to-Sequence Autoencoders to Assess Mild Cognitive Impairment
- Using a Single Input to Forecast Human Action Keystates in Everyday Pick and Place Actions
- Usted: Improving ASR with a Unified Speech and Text Encoder-Decoder
- VADOI: Voice-Activity-Detection Overlapping Inference for End-To-End Long-Form Speech Recognition
- VCD: View-Constraint Disentanglement for Action Recognition
- VCVTS: Multi-Speaker Video-to-Speech Synthesis Via Cross-Modal Knowledge Transfer from Voice Conversion
- VISinger: Variational Inference with Adversarial Learning for End-to-End Singing Voice Synthesis
- VQA-BC: Robust Visual Question Answering Via Bidirectional Chaining
- VR-FAM: Variance-Reduced Encoder with Nonlinear Transformation for Facial Attribute Manipulation
- VSEGAN: Visual Speech Enhancement Generative Adversarial Network
- VU-BERT: A Unified Framework for Visual Dialog
- VarArray: Array-Geometry-Agnostic Continuous Speech Separation
- Variable Span Trade-Off Filter for Sound Zone Control with Kernel Interpolation Weighting
- Variance Reduction-Boosted Byzantine Robustness in Decentralized Stochastic Optimization
- Varianceflow: High-Quality and Controllable Text-to-Speech using Variance Information via Normalizing Flow
- Variational Bayesian Framework for Advanced Image Generation with Domain-Related Variables
- Variational Bayesian Graph Convolutional Network for Robust Collaborative Filtering
- Variational Bayesian Tensor Networks with Structured Posteriors
- Video Anomaly Detection via Prediction Network with Enhanced Spatio-Temporal Memory Exchange
- Video Frame Interpolation via Local Lightweight Bidirectional Encoding with Channel Attention Cascade
- Violinist Identification Using Note-Level Timbre Feature Distributions
- Vision Transformer Equipped With Neural Resizer On Facial Expression Recognition Task
- Vision Transformer-Based Retina Vessel Segmentation with Deep Adaptive Gamma Correction
- Visual Representation Learning with Self-Supervised Attention for Low-Label High-Data Regime
- Visualtts: TTS with Accurate Lip-Speech Synchronization for Automatic Voice Over
- Vocalsound: A Dataset for Improving Human Vocal Sounds Recognition
- Vocbench: A Neural Vocoder Benchmark for Speech Synthesis
- Voice Filter: Few-Shot Text-to-Speech Speaker Adaptation Using Voice Conversion as a Post-Processing Module
- W-ART: Action Relation Transformer for Weakly-Supervised Temporal Action Localization
- WENETSPEECH: A 10000+ Hours Multi-Domain Mandarin Corpus for Speech Recognition
- WLS Design of Arma Graph Filters Using Iterative Second-Order Cone Programming
- Wasserstein Cross-Lingual Alignment For Named Entity Recognition
- Wassertrain: An Adversarial Training Framework Against Wasserstein Adversarial Attacks
- Watermarking Images in Self-Supervised Latent Spaces
- Wav2CLIP: Learning Robust Audio Representations from Clip
- Wav2vec-Switch: Contrastive Learning from Original-Noisy Speech Pairs for Robust Speech Recognition
- Wave-Domain Approach for Cancelling Noise Entering Open Windows
- Wavebender GAN: An Architecture for Phonetically Meaningful Speech Manipulation
- Waveform Optimization for Wireless Power Transfer with Power Amplifier and Energy Harvester Non-linearities
- Wavelet-Based Unsupervised Label-to-Image Translation
- Weak Target Detection in Massive MIMO Radar via an Improved Reinforcement Learning Approach
- Weakly Supervised Point Cloud Upsampling VIA Optimal Transport
- Wearable Seld Dataset: Dataset For Sound Event Localization And Detection Using Wearable Devices Around Head
- Weighted Graph Embedded Low-Rank Projection Learning for Feature Extraction
- Weighted Wavelet-Based Spectral-Spatial Transforms For CFA-Sampled Raw Camera Image Compression Considering Image Features
- What Is The Patient Looking At? Robust Gaze-Scene Intersection Under Free-Viewing Conditions
- When BERT Meets Quantum Temporal Convolution Learning for Text Classification in Heterogeneous Computing
- When Does Backdoor Attack Succeed in Image Reconstruction? A Study of Heuristics vs. Bi-Level Solution
- Wide-Sense Stationarity and Spectral Estimation for Generalized Graph Signal
- Wikitag: Wikipedia-Based Knowledge Embeddings Towards Improved Acoustic Event Classification
- Win The Lottery Ticket Via Fourier Analysis: Frequencies Guided Network Pruning
- Wishart Localization Prior On Spatial Covariance Matrix In Ambisonic Source Separation Using Non-Negative Tensor Factorization
- Wlinker: Modeling Relational Triplet Extraction As Word Linking
- Word Order does not Matter for Speech Recognition
- WordMarkov: A New Password Probability Model of Semantics
- Zero-Shot Cross-Lingual Transfer Using Multi-Stream Encoder and Efficient Speaker Representation
- Zeroth-Order Randomized Subspace Newton Methods
- nnSpeech: Speaker-Guided Conditional Variational Autoencoder for Zero-Shot Multi-speaker text-to-speech
- r-G2P: Evaluating and Enhancing Robustness of Grapheme to Phoneme Conversion by Controlled Noise Introducing and Contextual Information Incorporation
- r-Local Unlabeled Sensing: Improved Algorithm and Applications
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.