ICASSP 2023 Accepted Papers
The full list of 2,718 papers accepted at ICASSP 2023 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- Recurrent Fine-Grained Self-Attention Network for Video Crowd Counting
- Recursive Estimation of User Intent From Noninvasive Electroencephalography Using Discriminative Models
- Recursive Joint Attention for Audio-Visual Fusion in Regression Based Emotion Recognition
- Recursive/Iterative Unique Projection-Aggregation Decoding of Reed-Muller Codes
- Reducing Language Confusion for Code-Switching Speech Recognition with Token-Level Language Diarization
- Reducing the Communication and Computational Cost of Random Fourier Features Kernel LMS in Diffusion Networks
- Reducing the Computational Complexity of Learning with Random Convolutional Features
- Reducing the GAP Between Streaming and Non-Streaming Transducer-Based ASR by Adaptive Two-Stage Knowledge Distillation
- Refined Pseudo Labeling for Source-Free Domain Adaptive Object Detection
- Region-Awared Transformer with Asymmetric Loss in Multi-Label Classification
- Regression to Classification: Waveform Encoding for Neural Field-Based Audio Signal Representation
- Regularized Deep Generative Model Learning for Real-Time Massive MIMO Channel Tracking
- Regularized EM Algorithm
- Regularized Neural Detection for Millimeter Wave Massive Mimo Communication Systems with One-Bit Adcs
- Relapse Detection in Patients with Psychotic Disorders Using Unsupervised Learning on Smartwatch Signals
- Relapse Prediction from Long-Term Wearable Data Using Self-Supervised Learning and Survival Analysis
- Relate Auditory Speech To Eeg By Shallow-Deep Attention-Based Network
- Relating EEG Recordings to Speech Using Envelope Tracking and The Speech-FFR
- Relational Representation Learning for Zero-Shot Relation Extraction with Instance Prompting and Prototype Rectification
- Relative Dynamic Time Warping Comparison for Pronunciation Errors
- Relevance Propagation through Deep Conditional Random Fields
- Reliability Estimation for Synthetic Speech Detection
- Reliable Beamforming at Terahertz Bands: Are Causal Representations the Way Forward?
- Reliable Cluster-Based Framework for Open Set Domain Adaptation
- Removing Radio Frequency Interference From Auroral Kilometric Radiation With Stacked Autoencoders
- Repackagingaugment: Overcoming Prediction Error Amplification in Weight-Averaged Speech Recognition Models Subject to Self-Training
- Repetition Counting from Compressed Videos Using Sparse Residual Similarity
- Representation Learning of Clinical Multivariate Time Series with Random Filter Banks
- Representation of Vocal Tract Length Transformation Based on Group Theory
- Residual Hybrid Attention Network for Compression Artifact Reduction
- Residual Squeeze-and-Excitation U-Shaped Network for Minutia Extraction in Contactless Fingerprint Images
- Resolving Doppler Ambiguity Via Spread Phase Alignment in FDA-MIMO Radar
- Resource Allocation for UAV-Enabled Integrated Sensing and Communication (ISAC) via Multi-Objective Optimization
- Resource-Efficient Transfer Learning from Speech Foundation Model Using Hierarchical Feature Fusion
- Restoration of Time-Varying Graph Signals using Deep Algorithm Unrolling
- Rethink Long-Tailed Recognition with Vision Transforms
- Rethink Pair-Wise Self-Supervised Cross-Modal Retrieval From A Contrastive Learning Perspective
- Rethinking Implicit Neural Representations For Vision Learners
- Rethinking Learning-Based Method for Lossless Genome Compression
- Rethinking Random Walk in Graph Representation Learning
- Rethinking Rule-Based Approaches in Session-Based Recommendation
- Rethinking the Reasonability of the Test Set for Simultaneous Machine Translation
- Retiformer: Retinex-Based Enhancement In Transformer For Low-Light Image
- Retinal Biomarkers for Detecting Diabetic Retinopaty Using Smartphone-Based Deep Learning Frameworks
- Retrieval-Based Natural 3D Human Motion Generation
- Reverberation as Supervision For Speech Separation
- Revisit Out-Of-Vocabulary Problem For Slot Filling: A Unified Contrastive Framework With Multi-Level Data Augmentations
- Revisit Sampling Theory of Bandlimited Graph Signals: One Bridge Between GSP and DSP
- Rigid-Body Sound Synthesis with Differentiable Modal Resonators
- Ripple Sparse Self-Attention for Monaural Speech Enhancement
- Robust Acoustic And Semantic Contextual Biasing In Neural Transducers For Speech Recognition
- Robust Adaptive Beamforming with Proximal Method
- Robust Angle Estimation for Hybrid mmWave Systems
- Robust Audio-Visual ASR with Unified Cross-Modal Attention
- Robust Autoencoders for Collective Corruption Removal
- Robust Binary Component Decompositions
- Robust Binaural Sound Localisation with Temporal Attention
- Robust Content-Variant Reference Image Quality Assessment Via Similar Patch Matching
- Robust Data-Driven Accelerated Mirror Descent
- Robust Data2VEC: Noise-Robust Speech Representation Learning for ASR by Combining Regression and Improved Contrastive Learning
- Robust Dominant Periodicity Detection for Time Series with Missing Data
- Robust Fir Filters for Wireless Low-Frequency Sound Zones
- Robust GMM Parameter Estimation via the K-BM Algorithm
- Robust Hyperspectral Anomaly Detection with Simultaneous Mixed Noise Removal via Constrained Convex Optimization
- Robust Hypothesis Testing With Moment Constrained Uncertainty Sets
- Robust Iterative Solution for Linear Array-Based 3-D Localization by Message Passing
- Robust Knowledge Distillation from RNN-T Models with Noisy Training Labels Using Full-Sum Loss
- Robust Log-Based Anomaly Detection with Hierarchical Contrastive Learning
- Robust M-Estimation Based Distributed Expectation Maximization Algorithm with Robust Aggregation
- Robust Monocular Localization of Drones by Adapting Domain Maps to Depth Prediction Inaccuracies
- Robust Multi-Object Tracking With Spatial Uncertainty
- Robust Network Topologies for Distributed Learning
- Robust Online Multiband Drift Estimation in Electrophysiology Data
- Robust Self-Guided Deep Image Prior
- Robust Spatiotemporal Fusion of Satellite Images via Convex Optimization
- Robust Subspace Tracking with Contamination Mitigation via α-Divergence
- Robust Time Series Recovery and Classification Using Test-Time Noise Simulator Networks
- Robust Video Anomaly Detection Framework via Prior Knowledge and Multi-Path Frame Prediction
- Robust Video Object Segmentation with Restricted Attention
- Robust Watermarking Scheme in Encrypted Domain Based on Integer Lifting Wavelet Transform and Compressed Sensing
- Robust and Globally Sparse Pca via Majorization-Minimization and Variable Splitting
- Robust and Parallelizable Tensor Completion Based on Tensor Factorization and Maximum Correntropy Criterion
- Robust multi-modal speech emotion recognition with ASR error adaptation
- Robustdistiller: Compressing Universal Speech Representations for Enhanced Environment Robustness
- Robustness and Convergence of Mirror Descent for Blind Deconvolution
- Robustness of Deep Equilibrium Architectures to Changes in the Measurement Model
- Robustness-Preserving Lifelong Learning Via Dataset Condensation
- Role of Bias Terms in Dot-Product Attention
- Role of Lexical Boundary Information in Chunk-Level Segmentation for Speech Emotion Recognition
- Room Impulse Response Reconstruction Based on Spatio-Temporal-Spectral Features Learned from a Spherical Microphone Array Measurement
- Row Conditional-TGAN for Generating Synthetic Relational Databases
- Rumor Detection Via Assessing the Spreading Propensity of Users
- Runtime Prediction of Machine Learning Algorithms in Automl Systems
- RØROS: Building a Responsive Online Recommender System via Meta-Gradients Updating
- S-Feature Pyramid Network and Attention Model for Drone Detection
- S3I-PointHop: SO(3)-Invariant PointHop for 3D Point Cloud Classification
- SADE: A Self-Adaptive Expert for Multi-Dataset Question Answering
- SADI: A Self-Adaptive Decomposed Interpretable Framework for Electric Load Forecasting Under Extreme Events
- SAMO: Speaker Attractor Multi-Center One-Class Learning For Voice Anti-Spoofing
- SAN: A Robust End-to-End ASR Model Architecture
- SAR Image Despeckling with Residual-in-Residual Dense Generative Adversarial Network
- SARdBScene: Dataset and Resnet Baseline for Audio Scene Source Counting and Analysis
- SC-Net: Salient Point and Curvature Based Adversarial Point Cloud Generation Network
- SCA: Streaming Cross-Attention Alignment For Echo Cancellation
- SCSGNet: Spatial-Correlated and Shape-Guided Network for Breast Mass Segmentation
- SD-PINN: Physics Informed Neural Networks for Spatially Dependent PDES
- SDG-L: A Semiparametric Deep Gaussian Process based Framework for Battery Capacity Prediction
- SDRNet: Shape Decoupled Regression Network for 3d face Reconstruction
- SDTN: Speaker Dynamics Tracking Network for Emotion Recognition in Conversation
- SENER: Sentiment Element Named Entity Recognition for Aspect-Based Sentiment Analysis
- SEPDIFF: Speech Separation Based on Denoising Diffusion Model
- SFEMGN: Image Denoising with Shallow Feature Enhancement Network and Multi-Scale ConvGRU
- SFR: Semantic-Aware Feature Rendering of Point Cloud
- SG-VAD: Stochastic Gates Based Speech Activity Detection
- SIAST: A Slot Imbalance-Aware Self-Training Scheme for Semi-Supervised Slot Filling
- SIGVIC: Spatial Importance Guided Variable-Rate Image Compression
- SINCO: A Novel Structural Regularizer for Image Compression Using Implicit Neural Representations
- SL-MoE: A Two-Stage Mixture-of-Experts Sequence Learning Framework for Forecasting Rapid Intensification of Tropical Cyclone
- SLBERT: A Novel Pre-Training Framework for Joint Speech and Language Modeling
- SLICER: Learning Universal Audio Representations Using Low-Resource Self-Supervised Pre-Training
- SMCL: Saliency Masked Contrastive Learning for Long-Tailed Visual Recognition
- SMUG: Towards Robust Mri Reconstruction by Smoothed Unrolling
- SPADE: Self-Supervised Pretraining for Acoustic Disentanglement
- SPASHT: Semantic and Pragmatic Speech Features for Automatic Assessment of Autism
- SPECTRANET-SO(3): Learning Satellite Orientation from Optical Spectra by Implicitly Modeling Mutually Exclusive Probability Distributions on The Rotation Manifold
- SQA: Strong Guidance Query with Self-Selected Attention for Human-Object Interaction Detection
- SQuId: Measuring Speech Naturalness in Many Languages
- SR-init: An Interpretable Layer Pruning Method
- SRTNET: Time Domain Speech Enhancement via Stochastic Refinement
- SS-ADMM: Stationary and Sparse Granger Causal Discovery for Cortico-Muscular Coupling
- SSGD: A Smartphone Screen Glass Dataset for Defect Detection
- SSI-Net: A Multi-Stage Speech Signal Improvement System for ICASSP 2023 SSI Challenge
- SSVMR: Saliency-Based Self-Training for Video-Music Retrieval
- ST-MVDNet++: Improve Vehicle Detection with Lidar-Radar Geometrical Augmentation via Self-Training
- ST360IQ: No-Reference Omnidirectional Image Quality Assessment With Spherical Vision Transformers
- STACKMAPS: A Visualization Technique for Diabetic Retinopathy Grading
- STYX: Adaptive Poisoning Attacks Against Byzantine-Robust Defenses in Federated Learning
- SUVR: A Search-Based Approach to Unsupervised Visual Representation Learning
- SVMV: Spatiotemporal Variance-Supervised Motion Volume for Video Frame Interpolation
- SW-WAVENET: Learning Representation from Spectrogram and Wavegram Using Wavenet for Anomalous Sound Detection
- SYNTACC : Synthesizing Multi-Accent Speech By Weight Factorization
- SafeDeep: A Scalable Robustness Verification Framework for Deep Neural Networks
- Saliency-Driven Hierarchical Learned Image Coding for Machines
- Salient Co-Speech Gesture Synthesizing with Discrete Motion Representation
- Sample-Adapt Fusion Network for RGB-D Hand Detection in the Wild
- Sample-Aware Knowledge Distillation for Long-Tailed Learning
- Sample-Efficient Robust MMV Recovery Algorithm
- Sampling Order-Limited Signals on the Sphere
- Sandformer: CNN and Transformer under Gated Fusion for Sand Dust Image Restoration
- Sanet: Spatial Attention Network with Global Average Contrast Learning for Infrared Small Target Detection
- Scalable Multi-Task Semantic Communication System with Feature Importance Ranking
- Scalable Weight Reparametrization for Efficient Transfer Learning
- Scalable and Secure Federated XGBoost
- Scale-Adaptive Tiny Object Detection Enhanced by Across-Scale and Shape-Preserved Semantic Location
- ScaleMix: Intra- And Inter-Layer Multiscale Feature Combination for Change Detection
- Scaling Law Analysis for Covariance Based Activity Detection in Cooperative Multi-Cell Massive Mimo
- Scoreformer: Score Fusion-Based Transformers for Weakly-Supervised Violence Detection
- Search for Efficient Deep Visual-Inertial Odometry Through Neural Architecture Search
- Second-Order Statistic Deviation to Model Anomalies in the Design of Unsupervised Detectors
- Select The Best: Enhancing Graph Representation with Adaptive Negative Sample Selection
- Selecting Language Models Features VIA Software-Hardware Co-Design
- Selective Film Conditioning with CTC-Based ASR Probability for Speech Enhancement
- Self Supervised Bert for Legal Text Classification
- Self-Adaptive Incremental Machine Speech Chain for Lombard TTS with High-Granularity ASR Feedback in Dynamic Noise Condition
- Self-Adaptive Reasoning on Sub-Questions for Multi-Hop Question Answering
- Self-Attention Based Action Segmentation Using Intra-And Inter-Segment Representations
- Self-Attention for Enhanced OAMP Detection in MIMO Systems
- Self-Convolution for Automatic Speech Recognition
- Self-Distillation Hashing for Efficient Hamming Space Retrieval
- Self-Healing Through Error Detection, Attribution, and Retraining
- Self-Paced Partial Domain-Aware Learning for Face Anti-Spoofing
- Self-Remixing: Unsupervised Speech Separation VIA Separation and Remixing
- Self-Similarity is all You Need for Fast and Light-Weight Generic Event Boundary Detection
- Self-Sufficient Framework for Continuous Sign Language Recognition
- Self-Supervised Accent Learning for Under-Resourced Accents Using Native Language Data
- Self-Supervised Adversarial Training for Contrastive Sentence Embedding
- Self-Supervised Audio-Visual Speaker Representation with Co-Meta Learning
- Self-Supervised Audio-Visual Speech Representations Learning by Multimodal Self-Distillation
- Self-Supervised Facial Action Unit Detection with Region and Relation Learning
- Self-Supervised Guided Hypergraph Feature Propagation for Semi-Supervised Classification with Missing Node Features
- Self-Supervised Hierarchical Metrical Structure Modeling
- Self-Supervised Learning for Speech Enhancement Through Synthesis
- Self-Supervised Learning of Audio Representations using Angular Contrastive Loss
- Self-Supervised Learning with Bi-Label Masked Speech Prediction for Streaming Multi-Talker Speech Recognition
- Self-Supervised Learning with Explorative Knowledge Distillation
- Self-Supervised Learning-Based Source Separation for Meeting Data
- Self-Supervised Representations for Singing Voice Conversion
- Self-Supervised Representations in Speech-Based Depression Detection
- Self-Supervised Speech Representation Learning for Keyword-Spotting With Light-Weight Transformers
- Self-Transriber: Few-Shot Lyrics Transcription With Self-Training
- Selinet: A Lightweight Model for Single Channel Speech Separation
- SemGeo: Semantic Keywords for Cross-View Image Geo-Localization
- Semantic Centralized Contrastive Learning for Unsupervised Hashing
- Semantic Memory Guided Image Representation for Polyp Segmentation
- Semantic Preprocessor for Image Compression for Machines
- Semantic Preserving Learning for Task-Oriented Point Cloud Downsampling
- Semantic-Aware Gated Fusion Network For Interactive Colorization
- Semantic-Preserving Augmentation for Robust Image-Text Retrieval
- SemanticAC: Semantics-Assisted Framework for Audio Classification
- Semantically-Informed Deep Neural Networks For Sound Recognition
- Semantics-Aware Gamma Correction for Unsupervised Low-Light Image Enhancement
- Semantics-Disentangled Contrastive Embedding for Generalized Zero-Shot Learning
- Semantics-Guided Object Removal for Facial Images: with Broad Applicability and Robust Style Preservation
- Semi-Federated Learning for Edge Intelligence with Imperfect SIC
- Semi-Supervised Contrastive Learning with Soft Mask Attention for Facial Action Unit Detection
- Semi-Supervised Domain Generalization with Graph-Based Classifier
- Semi-Supervised Graph Ultra-Sparsifier Using Reweighted ℓ1 Optimization
- Semi-Supervised Learning with Per-Class Adaptive Confidence Scores for Acoustic Environment Classification with Imbalanced Data
- Semi-Supervised Local Structured Feature Learning with Dynamic Maximum Entropy Graph
- Semi-Supervised Remote Sensing Image Change Detection Using Mean Teacher Model for Constructing Pseudo-Labels
- Semi-Supervised Semantic Segmentation with Structured Output Space Adaption
- Semi-Supervised Sound Event Detection with Pre-Trained Model
- Semi-Supervised Speech Enhancement Based On Speech Purity
- Semi-Swinderain: Semi-Supervised Image Deraining Network Using SWIN Transformer
- Sensor Selection for Angle of Arrival Estimation Based on the Two-Target Cramér-Rao Bound
- Sequence-Based Device-Free Gesture Recognition Framework for Multi-Channel Acoustic Signals
- Sequential Datum-Wise Joint Feature Selection and Classification in the Presence of External Classifier
- Sequential Invariant Information Bottleneck
- Seri: Sketching-Reasoning-Integrating Progressive Workflow for Empathetic Response Generation
- Shadocnet: Learning Spatial-Aware Tokens in Transformer for Document Shadow Removal
- Shadow Removal of Text Document Images Using Background Estimation and Adaptive Text Enhancement
- Sharing Low Rank Conformer Weights for Tiny Always-On Ambient Speech Recognition Models
- Shift to Your Device: Data Augmentation for Device-Independent Speaker Verification Anti-Spoofing
- Short-Segment Speaker Verification Using ECAPA-TDNN with Multi-Resolution Encoder
- Show Me the Instruments: Musical Instrument Retrieval From Mixture Audio
- Shuffleaugment: A Data Augmentation Method Using Time Shuffling
- Shuffled Autoregression for Motion Interpolation
- Sign Language Recognition via Deformable 3D Convolutions and Modulated Graph Convolutional Networks
- Signal Analysis-Synthesis Using the Quantum Fourier Transform
- Signal Processing And Quantum State Tomography on Noisy Devices
- Signal Processing Grand Challenge 2023 - E-Prevention: Sleep Behavior as an Indicator of Relapses in Psychotic Patients
- Signal Processing On Product Spaces
- Signal Processing with Optical Quadratic Random Sketches
- Signal Reconstruction for FMCW Radar Interference Mitigation Using Deep Unfolding
- Similarity Relation Preserving Cross-Modal Learning for Multispectral Pedestrian Detection Against Adversarial Attacks
- Simple Pooling Front-Ends for Efficient Audio Classification
- Simplicial Vector Autoregressive Model For Streaming Edge Flows
- Simulating Realistic Speech Overlaps Improves Multi-Talker ASR
- Simultaneous Acoustic Echo Sorting and 3-D Room Geometry Inference
- Simultaneous Estimation of Direction of Arrival and Sound Speed Using a Non-Uniform Sensor Array
- Simultaneous Reconstruction and Uncertainty Quantification for Tomography
- Simultaneously Learning Robust Audio Embeddings and Balanced Hash Codes for Query-by-Example
- Sine: Similarity-Regularized Intra-Class Exploitation for Cross-Granularity Few-Shot Learning
- SingNet: a real-time Singing Voice beat and Downbeat Tracking System
- Singing Voice Synthesis Based on a Musical Note Position-Aware Attention Mechanism
- Single Domain Dynamic Generalization for Iris Presentation Attack Detection
- Single-Anchor UWB Localization Using Channel Impulse Response Distributions
- Single-Channel Speech Enhancement with Deep Complex U-Networks and Probabilistic Latent Space Models
- Single-Particle Tracking by Graph Transformer
- Single-Photon Image Super-Resolution via Self-Supervised Learning
- Single-Sample Direction-of-Arrival Estimation for Fast and Robust 3D Localization With Real Measurements from a Massive MIMO System
- Single-Shot Domain Adaptation via Target-Aware Generative Augmentations
- Single-Shot Fractional Fourier Phase Retrieval
- Single-branch Network for Multimodal Training
- Sinusoidal Frequency Estimation by Gradient Descent
- Sketch Less Face Image Retrieval: A New Challenge
- Skillnet-NLG: General-Purpose Natural Language Generation with a Sparsely Activated Approach
- Slot-Triggered Contextual Biasing For Personalized Speech Recognition Using Neural Transducers
- Small-Footprint Slimmable Networks for Keyword Spotting
- Smart Split-Federated Learning over Noisy Channels for Embryo Image Segmentation
- Smoothing Complex-Valued Signals on Graphs with Monte-Carlo
- Smoothing Point Adjustment-Based Evaluation of Time Series Anomaly Detection
- Soft 2D-to-3D Delivery Using Deep Graph Neural Networks for Holographic-Type Communication
- Soft Dynamic Time Warping for Multi-Pitch Estimation and Beyond
- Soft Label Coding for end-to-end Sound Source Localization with ad-hoc Microphone Arrays
- Solving Audio Inverse Problems with a Diffusion Model
- Solving Jigsaw Puzzle of Large Eroded Gaps Using Puzzlet Discriminant Network
- Sora: Scalable Black-Box Reachability Analyser on Neural Networks
- Source Localization for Extremely Large-Scale Antenna Arrays with Spatial Non-Stationarity
- Source-Filter HiFi-GAN: Fast and Pitch Controllable High-Fidelity Neural Vocoder
- Source-Free Unsupervised Domain Adaptation for Question Answering
- Space-Time Graph Neural Networks with Stochastic Graph Perturbations
- Space-Time Variable Density Samplings for Sparse Bandlimited Graph Signals Driven by Diffusion Operators
- Spammer Detection on Short Video Applications: A new Challenge and Baselines
- Sparse Aggregation-Based Channel Estimation For Massive Mimo Systems With Decentralized Baseband Processing
- Sparse Asynchronous Samples from Networks of Tems for Reconstruction of Classes of Non-Bandlimited Signals
- Sparse Bayesian Learning Assisted Decision Fusion in Millimeter Wave Massive MIMO Sensor Networks
- Sparse Bayesian Learning Based Three-Dimensional Imaging for Antenna Array Radar
- Sparse Black-Box Inversion Attack with Limited Information
- Sparse Convolution Based Octree Feature Propagation for Lidar Point Cloud Compression
- Sparse Delay-Doppler Channel Estimation for OTFS Modulation Using 2D-Music
- Sparse Error Correction for Power Network Parameters
- Sparse Graph Learning with Spectrum Prior for Deep Graph Convolutional Networks
- Sparse Mixture Once-for-all Adversarial Training for Efficient in-situ Trade-off between Accuracy and Robustness of DNNs
- Sparse Non-Contact Multiple People Localization and Vital Signs Monitoring Via FMCW Radar
- Sparse Representations with Cone Atoms
- Sparse and Structured Modelling of Underwater Acoustic Channel Impulse Responses
- Sparsity Constraint Implementation for the Joint Eigenvalue Decomposition of Matrices
- Sparsity-Driven Joint Blind Deconvolution-Demodulation with Application to Motor Fault Detection
- Sparsity-Smoothness-Aware Power Spectral Density Estimation with Application to Phased Array Weather Radar
- Spatial Active Noise Control Method Based on Sound Field Interpolation from Reference Microphone Signals
- Spatial Correlation Fusion Network for Few-Shot Segmentation
- Spatial Cross-Attention for Transformer-Based Image Captioning
- Spatial Graph Signal Interpolation with an Application for Merging BCI Datasets with Various Dimensionalities
- Spatial Inference Using Censored Multiple Testing with Fdr Control
- Spatial Similarity Guidance for Few-Shot Segmentation
- Spatial-Domain Object Detection Under Mimo-Fmcw Automotive Radar Interference
- Spatial-Temporal Graph Convolutional Network Boosted Flow-Frame Prediction For Video Anomaly Detection
- Spatially Informed Independent vector analysis for Source Extraction based on the convolutive Transfer Function Model
- Spatially Selective Deep Non-Linear Filters For Speaker Extraction
- Spatio-Temporal Attention in Multi-Granular Brain Chronnectomes For Detection of Autism Spectrum Disorder
- Spatio-Temporal Hybrid Fusion of CAE and SWin Transformers for Lung Cancer Malignancy Prediction
- Spatio-Temporal Structure Consistency for Semi-Supervised Medical Image Classification
- Speaker Change Detection For Transformer Transducer ASR
- Speaker Diaphragm Excursion Prediction: Deep Attention and Online Adaptation
- Speaker Recognition with Two-Step Multi-Modal Deep Cleansing
- Speaker-Aware Hierarchical Transformer For Personality Recognition In Multiparty Dialogues
- Speaker-Independent Acoustic-to-Articulatory Speech Inversion
- Speakeraugment: Data Augmentation for Generalizable Source Separation via Speaker Parameter Manipulation
- Spectral Clustering-Aware Learning of Embeddings for Speaker Diarisation
- Spectral Super-Resolution on the Unit Circle Via Gradient Descent
- Spectro-Temporal Post-Filtering Via Short-Time Target Cancellation for Directional Speech Enhancement in a Dual-Microphone Hearing AID
- Speech Dereverberation with a Reverberation Time Shortening Target
- Speech Emotion Recognition Based on Low-Level Auto-Extracted Time-Frequency Features
- Speech Emotion Recognition Via Two-Stream Pooling Attention With Discriminative Channel Weighting
- Speech Emotion Recognition via Heterogeneous Feature Learning
- Speech Enhancement with Intelligent Neural Homomorphic Synthesis
- Speech Intelligibility Classifiers from 550k Disordered Speech Samples
- Speech MOS Multi-Task Learning and Rater Bias Correction
- Speech Modeling with a Hierarchical Transformer Dynamical VAE
- Speech Privacy Leakage from Shared Gradients in Distributed Learning
- Speech Reconstruction from Silent Tongue and Lip Articulation by Pseudo Target Generation and Domain Adversarial Training
- Speech Separation with Large-Scale Self-Supervised Learning
- Speech Signal Improvement Using Causal Generative Diffusion Models
- Speech Summarization of Long Spoken Document: Improving Memory Efficiency of Speech/Text Encoders
- Speech and Noise Dual-Stream Spectrogram Refine Network With Speech Distortion Loss For Robust Speech Recognition
- Speech-Based Emotion Recognition with Self-Supervised Models Using Attentive Channel-Wise Correlations and Label Smoothing
- Speech-Text Based Multi-Modal Training with Bidirectional Attention for Improved Speech Recognition
- Speechlmscore: Evaluating Speech Generation Using Speech Language Model
- Spherical Sector Harmonics Based Soundfield Radial Extrapolation And Robustness Analysis
- Spherical Vector Quantization for Spatial Direction Coding
- Spice+: Evaluation of Automatic Audio Captioning Systems with Pre-Trained Language Models
- Spike-Based Optical Flow Estimation Via Contrastive Learning
- Spoofed Training Data for Speech Spoofing Countermeasure Can Be Efficiently Created Using Neural Vocoders
- Spteae: A Soft Prompt Transfer Model for Zero-Shot Cross-Lingual Event Argument Extraction
- Stabilising and Accelerating Light Gated Recurrent Units for Automatic Speech Recognition
- Stacking-Based Attention Temporal Convolutional Network for Action Segmentation
- Stargan-vc Based Cross-Domain Data Augmentation for Speaker Verification
- Static and Dynamic Source and Filter Cues for Classification of Amyotrophic Lateral Sclerosis Patients and Healthy Subjects
- Static-Scene Constrained Optimization for Matrix/Tensor-Decomposition-free Foreground-Background Separation
- Statistical Analysis of Speech Disorder Specific Features to Characterise Dysarthria Severity Level
- Stay In The Middle: A Semi-Supervised Model for CT Metal Artifact Reduction
- Step restriction for improving adversarial attacks
- Stereoscopic Video Retargeting Based on Camera Motion Classification
- Stochastic Optimization of Vector Quantization Methods in Application to Speech and Image Processing
- Stochastic Super-Resolution For Gaussian Textures
- Strategies for Enhanced Signal Modulation Classifications Under Unknown Symbol Rates and Noise Conditions
- Stream Attention Based U-Net for L3DAS23 Challenge
- StreamSpeech: Low-Latency Neural Architecture for High-Quality on-Device Speech Synthesis
- Streaming Joint Speech Recognition and Disfluency Detection
- Streaming Multi-Channel Speech Separation with Online Time-Domain Generalized Wiener Filter
- Streaming Stroke Classification of Online Handwriting
- Streaming Voice Conversion via Intermediate Bottleneck Features and Non-Streaming Teacher Guidance
- String-Based Molecule Generation Via Multi-Decoder VAE
- Structural Optimization of Factor Graphs for Symbol Detection via Continuous Clustering and Machine Learning
- Structural Reparameterization Lightweight Network for Video Action Recognition
- Structure-Aware Multi-Feature Co-Learning for Dual Branch Face Super Resolution
- Structure-Aware Sparse Bayesian Learning-Based Channel Estimation for Intelligent Reflecting Surface-Aided MIMO
- Structure-Preserving and Redundancy-Free Features Refinement for Generalized Zero-Shot Learning
- Structured Errors-in-Variables Modelling for Cortico-Muscular Coherence Enhancement
- Structured Pruning of Self-Supervised Pre-Trained Models for Speech Recognition and Understanding
- Structured State Space Decoder for Speech Recognition and Synthesis
- Structured-Anchor Projected Clustering for Hyperspectral Images
- Stuart: Individualized Classroom Observation of Students with Automatic Behavior Recognition And Tracking
- Study And Design Of Robust Personal Sound Zones With Vast Using Low Rank Rirs
- Study of Manifold Geometry Using Multiscale Non-Negative Kernel Graphs
- Study on the Fairness of Speaker Verification Systems Across Accent and Gender Groups
- Style Modeling for Multi-Speaker Articulation-to-Speech
- Sub-Band Contrastive Learning-Based Knowledge Distillation For Sound Classification
- Subband Dependency Modeling for Sound Event Detection
- Subgradient Descent Learning with Over-the-Air Computation
- Subject-Specific Adaptation for a Causally-Trained Auditory-Attention Decoding System
- Subspace Hybrid Beamforming for Head-Worn Microphone Arrays
- Subspace Modeling Enabled High-Sensitivity X-Ray Chemical Imaging
- Subspace-Based Detector For Distributed Mmwave Mimo Radar Sensors
- Suffix Retrieval-Augmented Language Modeling
- Summary on the Multimodal Information Based Speech Processing (MISP) 2022 Challenge
- Super Dilated Nested Arrays with Ideal Critical Weights and Increased Degrees of Freedom
- Super-Resolution Harmonic Retrieval of Non-Circular Signals
- Super-Resolution Information Enhancement for Crowd Counting
- Super-Resolution for Macro X-Ray Fluorescence Data Collected from Old Master Paintings
- Supercm: Revisiting Clustering for Semi-Supervised Learning
- Supervised Contrastive Learning as Multi-Objective Optimization for Fine-Tuning Large Pre-Trained Language Models
- Supervised Hierarchical Clustering Using Graph Neural Networks for Speaker Diarization
- Surface-Sampling Based Objective Quality Assessment Metrics for Meshes
- Surrogate Based Post-HOC Calibration for Distributional Shift
- Switching Kronecker Product Linear Filtering for Multispeaker Adaptive Speech Dereverberation
- Symbol Level Precoding in the RF Domain for Low Hardware Complexity RIS-Assisted MU-MISO Systems
- Symbol-Level Precoding is Related to Parameter Estimation from Quantized Data
- SyncNet: Correlating Objective for Time Delay Estimation in Audio Signals
- Syngen: A Syntactic Plug-And-Play Module for Generative Aspect-Based Sentiment Analysis
- Synthesizer Preset Interpolation Using Transformer Auto-Encoders
- Synthesizing Speech from ECoG with a Combination of Transformer-Based Encoder and Neural Vocoder
- Synthetic Pseudo Anomalies for Unsupervised Video Anomaly Detection: A Simple Yet Efficient Framework Based on Masked Autoencoder
- T5-SR: A Unified Seq-to-Seq Decoding Strategy for Semantic Parsing
- T5lephone: Bridging Speech and Text Self-Supervised Models for Spoken Language Understanding Via Phoneme Level T5
- TABLEIE: Capturing the Interactions Among Sub-Tasks in Information Extraction via Double Tables
- TAMformer: Multi-Modal Transformer with Learned Attention Mask for Early Intent Prediction
- TAPE: An End-to-End Timbre-Aware Pitch Estimator
- TAPLoss: A Temporal Acoustic Parameter Loss for Speech Enhancement
- TDMA-Based Multi-User Binary Computation Offloading in the Finite-Block-Length Regime
- TEA-PSE 3.0: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System For ICASSP 2023 Dns-Challenge
- TEFISTA-NET: GTD Parameter Estimation of Low-Frequency Ultra- Wideband Radar via Model-Based Deep Learning
- TF-GRIDNET: Making Time-Frequency Domain Models Great Again for Monaural Speaker Separation
- TFCnet: Time-Frequency Domain Corrector for Speech Separation
- TINYCOD: Tiny and Effective Model for Camouflaged Object Detection
- TOLD: a Novel Two-Stage Overlap-Aware Framework for Speaker Diarization
- TOPO-MLP : A Simplicial Network without Message Passing
- TRICL: Triplet Continual Learning
- TRUSTERA: A Live Conversation Redaction System
- TSPTQ-ViT: Two-Scaled Post-Training Quantization for Vision Transformer
- TSpeech-AI System Description to the 5th Deep Noise Suppression (DNS) Challenge
- TT-Net: Dual-Path Transformer Based Sound Field Translation in the Spherical Harmonic Domain
- Tangent Bundle Filters and Neural Networks: From Manifolds to Cellular Sheaves and Back
- Target Sound Extraction with Variable Cross-Modality Clues
- Target Speaker Extraction with Ultra-Short Reference Speech by VE-VE Framework
- Target Speaker Voice Activity Detection with Transformers and Its Integration with End-To-End Neural Diarization
- Target Velocity Estimation for Quantization-Based Cooperative MIMO Radar and Communications System
- Target-Speaker Voice Activity Detection Via Sequence-to-Sequence Prediction
- Targeted Adversarial Attacks Against Neural Machine Translation
- Tayloraecnet: A Taylor Style Neural Network For Full-Band Echo Cancellation
- TeAw: Text-Aware Few-Shot Remote Sensing Image Scene Classification
- Tell Model Where to Attend: Improving Interpretability of Aspect-Based Sentiment Classification via Small Explanation Annotations
- Tempo vs. Pitch: Understanding Self-Supervised Tempo Estimation
- Temporal Contrastive Learning with Curriculum
- Temporal Modeling Matters: A Novel Temporal Emotional Modeling Approach for Speech Emotion Recognition
- Tensor Completion for Efficient and Accurate Hyperparameter Optimisation in Large-Scale Statistical Learning
- Tensor Decomposition Based Latent Feature Clustering for Hyperspectral Band Selection
- Tensor Low Rank Column-Wise Compressive Sensing for Dynamic Imaging
- Tensor-based Complex-valued Graph Neural Network for Dynamic Coupling Multimodal brain Networks
- Tensorized LSSVMS For Multitask Regression
- Tensorized Neural Layer Decomposition for 2-D DOA Estimation
- Terminology-Aware Medical Dialogue Generation
- Ternary Weight Networks
- Test Your Samples Jointly: Pseudo-Reference for Image Quality Evaluation
- Test-Time Training-Free Domain Adaptation
- Text Classification In The Wild: A Large-Scale Long-Tailed Name Normalization Dataset
- Text is all You Need: Personalizing ASR Models Using Controllable Speech Synthesis
- Text-To-Speech Synthesis Based on Latent Variable Conversion Using Diffusion Probabilistic Model and Variational Autoencoder
- Text-to-ECG: 12-Lead Electrocardiogram Synthesis Conditioned on Clinical Text Reports
- Textless Direct Speech-to-Speech Translation with Discrete Speech Representation
- Textless Speech-to-Music Retrieval Using Emotion Similarity
- Tg-Critic: A Timbre-Guided Model For Reference-Independent Singing Evaluation
- The 2nd Clarity Enhancement Challenge for Hearing Aid Speech Intelligibility Enhancement: Overview and Outcomes
- The Ajmide Topic Segmentation System for the ICASSP 2023 General Meeting Understanding and Generation Challenge
- The DKU Post-Challenge Audio-Visual Wake Word Spotting System for the 2021 MISP Challenge: Deep Analysis
- The Edinburgh International Accents of English Corpus: Towards the Democratization of English ASR
- The First Pathloss Radio Map Prediction Challenge
- The MBSTOI Binaural Intelligibility Metric Using a Close-Talking Microphone Reference
- The Multimodal Information Based Speech Processing (Misp) 2022 Challenge: Audio-Visual Diarization And Recognition
- The NERCSLIP-USTC System for the L3DAS23 Challenge Task2: 3D Sound Event Localization and Detection (SELD)
- The NIO System for Audio-Visual Diarization and Recognition in MISP Challenge 2022
- The NPU-ASLP System for Audio-Visual Speech Recognition in MISP 2022 Challenge
- The NPU-Elevoc Personalized Speech Enhancement System for Icassp2023 DNS Challenge
- The Pipeline System of ASR and NLU with MLM-based data Augmentation Toward Stop Low-Resource Challenge
- The Potential of Neural Speech Synthesis-Based Data Augmentation for Personalized Speech Enhancement
- The R3VIVAL Dataset: Repository of Room Responses and 360 Videos of a Variable Acoustics Lab
- The Role of Initial Entanglement in Adaptive Gibbs State Preparation on Quantum Computers
- The Role of Memory in Social Learning When Sharing Partial Opinions
- The Secret Source : Incorporating Source Features to Improve Acoustic-To-Articulatory Speech Inversion
- The Uniqueness Problem of Physical Law Learning
- The Ustc System for Adress-m Challenge
- The WHU-Alibaba Audio-Visual Speaker Diarization System for the MISP 2022 Challenge
- The XMU System for Audio-Visual Diarization and Recognition in MISP Challenge 2022
- Thermal Infrared Image Inpainting Via Edge-Aware Guidance
- Think Before You Speak: Concept-Guided Explicit Persona Reasoning for Personalized Dialogue Generation
- This Changes to That : Combining Causal and Non-Causal Explanations to Generate Disease Progression in Capsule Endoscopy
- Time-Aware Multiway Adaptive Fusion Network for Temporal Knowledge Graph Question Answering
- Time-Domain Speech Enhancement Assisted by Multi-Resolution Frequency Encoder and Decoder
- Time-Frequency Awareness Network For Human Mesh Recovery From Videos
- Time-Resolved FMRI Shared Response Model Using Gaussian Process Factor Analysis
- Time-Varying Signals Recovery Via Graph Neural Networks
- Time-Weighted Frequency Domain Audio Representation with GMM Estimator for Anomalous Sound Detection
- TinyOOD: Effective out-of-Distribution Detection for TinyML
- To Regularize or Not to Regularize: The Role of Positivity in Sparse Array Interpolation with a Single Snapshot
- To Wake-Up or Not to Wake-Up: Reducing Keyword False Alarm by Successive Refinement
- Token2vec: A Joint Self-Supervised Pre-Training Framework Using Unpaired Speech and Text
- Top-K Visual Tokens Transformer: Selecting Tokens for Visible-Infrared Person Re-Identification
- Topgformer: Topological-Based Graph Transformer for Mapping Brain Structural Connectivity to Functional Connectivity
- Topological Signal Processing Over Weighted Simplicial Complexes
- Topological Slepians: Maximally Localized Representations of Signals Over Simplicial Complexes
- Topology Uncertainty Modeling For Imbalanced Node Classification on Graphs
- Torchaudio-Squim: Reference-Less Speech Quality and Intelligibility Measures in Torchaudio
- Toroidal Probabilistic Spherical Discriminant Analysis
- Toward A Multimodal Approach for Disfluency Detection and Categorization
- Toward Asymptotic Optimality: Sequential Unsupervised Regression of Density Ratio for Early Classification
- Toward Auto-Evaluation With Confidence-Based Category Relation-Aware Regression
- Toward Privacy-Enhancing Ambulatory-Based Well-Being Monitoring: Investigating User Re-Identification Risk in Multimodal Data
- Toward Universal Text-To-Music Retrieval
- Towards A Unified Conformer Structure: from ASR to ASV Task
- Towards Accurate and Real-Time End-of-Speech Estimation
- Towards Adversarially Robust Continual Learning
- Towards Bandwidth Estimation for Graph Signal Reconstruction
- Towards Building Text-to-Speech Systems for the Next Billion Users
- Towards Controllable Audio Texture Morphing
- Towards Dialogue Modeling Beyond Text
- Towards Diverse and Coherent Augmentation for Time-Series Forecasting
- Towards Domain Generalisation in ASR with Elitist Sampling and Ensemble Knowledge Distillation
- Towards Efficient and Optimal Joint Beamforming and Antenna Selection: A Machine Learning Approach
- Towards Explainable Recommendation Via Bert-Guided Explanation Generator
- Towards Hyperbolic Regularizers For Point Cloud Part Segmentation
- Towards Improved Room Impulse Response Estimation for Speech Recognition
- Towards Improved Sonar Performance Using Environment-Informed Sparse Sub-Array Processing
- Towards Interpretable Seizure Detection Using Wearables
- Towards Learning Emotion Information from Short Segments of Speech
- Towards Low-Power Heart Rate Estimation Based on User's Demographics and Activity Level For Wearables
- Towards Making a Trojan-Horse Attack on Text-to-Image Retrieval
- Towards Polymorphic Adversarial Examples Generation for Short Text
- Towards Practical Edge Inference Attacks Against Graph Neural Networks
- Towards Privacy and Utility in Tourette TIC Detection Through Pretraining Based on Publicly Available Video Data of Healthy Subjects
- Towards Real-Time Person Search with Invariant Feature Learning
- Towards Real-Time Single-Channel Speech Separation in Noisy and Reverberant Environments
- Towards Realizing the Value of Labeled Target Samples: A Two-Stage Approach for Semi-Supervised Domain Adaptation
- Towards Reducing Patient Effort for the Automatic Prediction of Speech Intelligibility in Head and Neck Cancers
- Towards Reliable Image Outpainting: Learning Structure-Aware Multimodal Fusion with Depth Guidance
- Towards Robust Audio-Based Vehicle Detection Via Importance-Aware Audio-Visual Learning
- Towards Robust Data-Driven Underwater Acoustic Localization: A Deep CNN Solution with Performance Guarantees for Model Mismatch
- Towards Scale Adaptive Underwater Detection Through Refined Pyramid Grid
- Towards Simultaneous Segmentation Of Liver Tumors And Intrahepatic Vessels Via Cross-Attention Mechanism
- Towards Trustworthy Multi-Label Sewer Defect Classification via Evidential Deep Learning
- Towards Trustworthy Phoneme Boundary Detection with Autoregressive Model and Improved Evaluation Metric
- Towards Zero-Shot Code-Switched Speech Recognition
- Towards Zero-Shot Personalized Table-to-Text Generation with Contrastive Persona Distillation
- Towards a More Stable and General Subgraph Information Bottleneck
- Towards a Robust and Efficient Classifier for Real World Radio Signal Modulation Classification
- Towards a Unified Training for Levenshtein Transformer
- TrOMR:Transformer-Based Polyphonic Optical Music Recognition
- Tracking Objects and Activities with Attention for Temporal Sentence Grounding
- Tracking Targets in Hyper-Scale Cameras Using Movement Predication
- Training Graph Neural Networks on Growing Stochastic Graphs
- Training Large-Vocabulary Neural Language Models by Private Federated Learning for Resource-Constrained Devices
- Training Neural Networks for Sequential Change-Point Detection
- Training Robust Spiking Neural Networks on Neuromorphic Data with Spatiotemporal Fragments
- Training Robust Spiking Neural Networks with Viewpoint Transform and Spatiotemporal Stretching
- Training Set Cleansing of Backdoor Poisoning by Self-Supervised Representation Learning
- Training Sound Event Detection with Soft Labels from Crowdsourced Annotations
- Training Stronger Spiking Neural Networks with Biomimetic Adaptive Internal Association Neurons
- TransLink: Transformer-Based Embedding for Tracklets' Global Link
- Transadapt: A Transformative Framework for Online Test Time Adaptive Semantic Segmentation
- Transaudio: Towards the Transferable Adversarial Audio Attack Via Learning Contextualized Perturbations
- Transceiver Design for MIMO-DFRC Systems
- Transcription Free Filler Word Detection with Neural Semi-CRFs
- Transductive Matrix Completion with Calibration for Multi-Task Learning
- Transferring Quantified Emotion Knowledge for the Detection of Depression in Alzheimer's Disease Using Forestnets
- Transformer-Based Bioacoustic Sound Event Detection on Few-Shot Learning Tasks
- Transformer-Based Deep Hashing Method for Multi-Scale Feature Fusion
- Transformer-Based Multi-Prototype Approach for Diabetic Macular Edema Analysis in OCT Images
- Transformer-based tracking Network for Maneuvering Targets
- Transient Dictionary Learning for Compressed Time-of-Flight Imaging
- Transmit Energy Focusing For Parameter Estimation in Transmit Beamspace Slow-Time MIMO Radar
- Transplayer: Timbre Style Transfer with Flexible Timbre Control
- Transwnet: Integrating Transformers into CNNS via Row and Column Attention for Abdominal Multi-Organ Segmentation
- Tree-Like Interaction Learning for Bundle Recommendation
- TreeXGNN: can gradient-boosted decision trees help boost heterogeneous graph neural networks?
- TriAAN-VC: Triple Adaptive Attention Normalization for Any-to-Any Voice Conversion
- TrimTail: Low-Latency Streaming ASR with Simple But Effective Spectrogram-Level Length Penalty
- Trinet: Stabilizing Self-Supervised Learning From Complete or Slow Collapse
- Trust Your Partner's Friends: Hierarchical Cross-Modal Contrastive Pre-Training for Video-Text Retrieval
- Twitter Stance Detection via Neural Production Systems
- Two-Branch Multi-Scale Deep Neural Network for Generalized Document Recapture Attack Detection
- Two-Phase Prototypical Contrastive Domain Generalization for Cross-Subject EEG-Based Emotion Recognition
- Two-Stage Neural Network for ICASSP 2023 Speech Signal Improvement Challenge
- Two-Stage UNet with Multi-Axis Gated Multilayer Perceptron for Monaural Noisy-Reverberant Speech Enhancement
- Two-Stage Video De-Raining with Spatio-Temporal Fusion and Illumination-Invariant Detail Preservation
- Two-Step Band-Split Neural Network Approach For Full-Band Residual Echo Suppression
- Two-Stream Decoder Feature Normality Estimating Network for Industrial Anomaly Detection
- Two-Stream Joint-Training for Speaker Independent Acoustic-to-Articulatory Inversion
- U-Beat: A Multi-Scale Beat Tracking Model Based on Wave-U-Net
- U-Shiftformer: Brain Tumor Segmentation Using A Shifted Attention Mechanism
- UAV Local Path Planning Based on Improved Proximal Policy Optimization Algorithm
- UAV Remote Sensing Image Dehazing Based on Multi-Dimensional Saliency Awareness Unequal Network
- UCONV-Conformer: High Reduction of Input Sequence Length for End-to-End Speech Recognition
- UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction
- UFO2: A Unified Pre-Training Framework for Online and Offline Speech Recognition
- UML: A Universal Monolingual Output Layer For Multilingual Asr
- UNTAG: Learning Generic Features for Unsupervised Type-Agnostic Deepfake Detection
- UNeXt: a Low-Dose CT denoising UNet model with the modified ConvNeXt block
- UPGLADE: Unplugged Plug-and-Play Audio Declipper Based on Consensus Equilibrium of DNN and Sparse Optimization
- URM4DMU: An User Representation Model for Darknet Markets Users
- UWB Localization-of-Things Via Soft Information: Network Experimentation in Indoor Environment
- UX-Net: Filter-and-Process-Based Improved U-Net for real-time time-domain audio Separation
- Ultimate Negative Sampling for Contrastive Learning
- Ultra Real-Time Portrait Matting via Parallel Semantic Guidance
- Ultrasound Image Quality Control Using Speech-Assisted Switchable CycleGAN
- Unbiased Unsupervised Stimulus Reconstruction for EEG-Based Auditory Attention Decoding
- Uncer2Natural: Uncertainty-Aware Unsupervised Image Denoising
- Uncertainty Estimation in Deep Speech Enhancement Using Complex Gaussian Mixture Models
- Uncertainty-Aware Few-Shot Class-Incremental Learning
- Understandable Relu Neural Network For Signal Classification
- Understanding Shared Speech-Text Representations
- Underwater Image Restoration with Light-Aware Progressive Network
- Unified Keyword Spotting and Audio Tagging on Mobile Devices with Transformers
- Unified Prompt Learning Makes Pre-Trained Language Models Better Few-Shot Learners
- Unifying Speech Enhancement and Separation with Gradient Modulation for End-to-End Noise-Robust Speech Separation
- Unique Bispectrum Inversion for Signals with Finite Spectral/Temporal Support
- Unitary Esprit for Coprime Arrays
- Universal Speaker Recognition Encoders for Different Speech Segments Duration
- Unlimited Sampling Radar: Life Below the Quantization Noise
- Unlimited Sampling in Phase Space
- Unlimited Sampling of FRI Signals Independent of Sampling Rate
- Unobtrusive Respiratory Monitoring System for Intensive Care
- Unrestricted Anchor Graph Based GCN for Incomplete Multi-View Clustering
- Unrolled Fourier Disparity Layer Optimization for Scene Reconstruction from Few-Shots Focal Stacks
- Unsupervised Action Segmentation of Untrimmed Egocentric Videos
- Unsupervised Anomaly Detection and Localization of Machine Audio: A Gan-Based Approach
- Unsupervised Deep Digital Staining for Microscopic Cell Images via Knowledge Distillation
- Unsupervised Domain Adaptation for Preference Learning Based Speech Emotion Recognition
- Unsupervised Domain Adaptation via Subspace Interpolating Deep Dictionary Learning: A Case Study in Machine Inspection
- Unsupervised Extractive Summarization With Heterogeneous Graph Embeddings for Chinese Documents
- Unsupervised Feature Selection with self-Weighted and ℓ2,0-Norm Constraint
- Unsupervised Fine-Tuning Data Selection for ASR Using Self-Supervised Speech Models
- Unsupervised Model-Based Speaker Adaptation of End-To-End Lattice-Free MMI Model for Speech Recognition
- Unsupervised Noise Adaptation Using Data Simulation
- Unsupervised Out-of-Distribution Detection Using Few in-Distribution Samples
- Unsupervised Pre-Training for Data-Efficient Text-to-Speech on Low Resource Languages
- Unsupervised Speaker Verification Using Pre-Trained Model and Label Correction
- Unsupervised Video Anomaly Detection For Stereotypical Behaviours in Autism
- Unsupervised Vocal Dereverberation with Diffusion-Based Generative Models
- Unsupervised Voice Type Discrimination Score Adaptation Using X-Vector Clusters
- Unsupervised Word Segmentation Using Temporal Gradient Pseudo-Labels
- Unsupervised word Segmentation Based on Word Influence
- Untargeted Backdoor Attack Against Object Detection
- Using Adapters to Overcome Catastrophic Forgetting in End-to-End Automatic Speech Recognition
- Using Auxiliary Tasks In Multimodal Fusion of Wav2vec 2.0 And Bert for Multimodal Emotion Recognition
- Using Emotion Embeddings to Transfer Knowledge between Emotions, Languages, and Annotation Formats
- Using Machine Learning to Understand the Relationships Between Audiometric Data, Speech Perception, Temporal Processing, And Cognition
- Using Modified Adult Speech as Data Augmentation for Child Speech Recognition
- Using Received Power in Microphone Arrays to Estimate Direction of Arrival
- Utility Polelocalization by Learning from Ambient Traces on Distributed Acoustic Sensing
- Utilization of Bessel Beams in Wideband Sub Terahertz Communication Systems to Mitigate Beamsplit Effects in the Near-field
- Utilizing Wav2Vec In Database-Independent Voice Disorder Detection
- VAN-ICP: GPU-Accelerated Approximate Nearest Neighbor Search for ICP Registration via Voxel Dilation
- VE-KWS: Visual Modality Enhanced End-to-End Keyword Spotting
- VF-Taco2: Towards Fast and Lightweight Synthesis for Autoregressive Models with Variation Autoencoder and Feature Distillation
- VLKP:Video Instance Segmentation with Visual-Linguistic Knowledge Prompts
- VPPT: Visual Pre-Trained Prompt Tuning Framework for Few-Shot Image Classification
- VQ-CL: Learning Disentangled Speech Representations with Contrastive Learning and Vector Quantization
- Vani: Very-Lightweight Accent-Controllable TTS for Native And Non-Native Speakers With Identity Preservation
- Vararray Meets T-Sot: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition
- Variable Attention Masking for Configurable Transformer Transducer Speech Recognition
- Variable Rate Allocation for Vector-Quantized Autoencoders
- Variational Bayesian Channel Estimation in Wideband Multi-Scale Multi-Lag Channels
- Variational Inference Aided Estimation of Time Varying Channels
- Variational Message Passing-Based Respiratory Motion Estimation and Detection Using Radar Signals
- VarietySound: Timbre-Controllable Video to Sound Generation Via Unsupervised Information Disentanglement
- Various Performance Bounds on the Estimation of Low-Rank Probability Mass Function Tensors from Partial Observations
- Vehicle View Synthesis by Generative Adversarial Network
- ViT-Cat: Parallel Vision Transformers With Cross Attention Fusion for Popularity Prediction in MEC Networks
- Video Captioning via Relation-Aware Graph Learning
- Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-to-Speech
- Vision Transformer with Progressive Tokenization for CT Metal Artifact Reduction
- Vision Transformer-Based Feature Extraction for Generalized Zero-Shot Learning
- Vision, Deduction and Alignment: An Empirical Study on Multi-Modal Knowledge Graph Alignment
- Vision2Touch: Imaging Estimation of Surface Tactile Physical Properties
- Visual Answer Localization with Cross-Modal Mutual Knowledge Transfer
- Visual Graph Reasoning Network
- Visual Information Matters for ASR Error Correction
- Visual Onoma-to-Wave: Environmental Sound Synthesis from Visual Onomatopoeias and Sound-Source Images
- Visual Prompting for Adversarial Robustness
- Visual-Aware Text-to-Speech*
- Vitasd: Robust Vision Transformer Baselines for Autism Spectrum Disorder Facial Diagnosis
- Voice Conversion Using Feature Specific Loss Function Based Self-Attentive Generative Adversarial Network
- Voice-Preserving Zero-Shot Multiple Accent Conversion
- Volume-Regularized Nonnegative Tucker Decomposition with Identifiability Guarantees
- Volumetric 3D Reconstruction with Window-Wise Global Feature Aggregation
- Volumetric Attribute Compression for 3D Point Clouds Using Feedforward Network with Geometric Attention
- W2KPE: Keyphrase Extraction with Word-Word Relation
- WAVELET2VEC: A Filter Bank Masked Autoencoder for EEG-Based Seizure Subtype Classification
- WHC: Weighted Hybrid Criterion for Filter Pruning on Convolutional Neural Networks
- WIFI-Based Robust Child Presence Detection for Smart Cars
- WITT: A Wireless Image Transmission Transformer for Semantic Communications
- WL-MSR: Watch and Listen for Multimodal Subtitle Recognition
- WUDA: Unsupervised Domain Adaptation Based on Weak Source Domain Labels
- Wassertein Gan Synthesis for Time Series with Complex Temporal Dynamics: Frugal Architectures and Arbitrary Sample-Size Generation
- Water Leak Detection and Localization Using Convolutional Autoencoders
- Wav2Seq: Pre-Training Speech-to-Text Encoder-Decoder Models Using Pseudo Languages
- Wav2vec-Based Detection and Severity Level Classification of Dysarthria From Speech
- Wave-U-Net Discriminator: Fast and Lightweight Discriminator for Generative Adversarial Network-Based Speech Synthesis
- Waveform Boundary Detection for Partially Spoofed Audio
- Waveform Design to Improve the Estimation of Target Parameters Using the Fourier Transform Method in a MIMO OFDM DFRC System
- Wavsyncswap: End-To-End Portrait-Customized Audio-Driven Talking Face Generation
- WeSinger 2: Fully Parallel Singing Voice Synthesis via Multi-Singer Conditional Adversarial Training
- Weakly- and Semi-Supervised Object Localization
- Weakly-Supervised Scene-Specific Crowd Counting Using Real-Synthetic Hybrid Data
- Weavspeech: Data Augmentation Strategy For Automatic Speech Recognition Via Semantic-Aware Weaving
- Weight Averaging: A Simple Yet Effective Method to Overcome Catastrophic Forgetting in Automatic Speech Recognition
- Weight-Based Mask For Domain Adaptation
- Weight-Sharing Supernet for Searching Specialized Acoustic Event Classification Networks Across Device Constraints
- Weighted Sampling for Masked Language Modeling
- Wekws: A Production First Small-Footprint End-to-End Keyword Spotting Toolkit
- Wespeaker: A Research and Production Oriented Speaker Embedding Learning Toolkit
- When is Mimo Massive in Radar?
- Whether Contribution of Features Differ Between Video-Mediated and In-Person Meetings in Important Utterance Estimation
- Which Country is This Picture From? New Data and Methods For Dnn-Based Country Recognition
- Wiener Filtering Without Covariance Matrix Inversion
- Windowed Fourier Analysis for Signal Processing on Graph Bundles
- Wireless Deep Speech Semantic Transmission
- Wireless Location Tracking via Complex-Domain Super MDS with Time Series Self-Localization Information
- Wireless Power Transfer Using Chirp Waveforms
- Wireless Sensing for Simultaneous Human Vocal Sound and Heart Sound Recognition
- Wordreg: Mitigating the Gap between Training and Inference with Worst-Case Drop Regularization
- X-SEPFORMER: End-To-End Speaker Extraction Network with Explicit Optimization on Speaker Confusion
- YOLOX-B: A Better Yolox Model for Real-Time Driver Behavior Detection
- Yolo-Based Lightweight Object Detection With Structure Simplification And Attention Enhancement
- Your Camera Improves Your Point Cloud Compression
- ZO-DARTS: Differentiable Architecture Search with Zeroth-Order Approximation
- Zephyr: Zero-Shot Punctuation Restoration
- Zero-Shot Anomalous Sound Detection in Domestic Environments Using Large-Scale Pretrained Audio Pattern Recognition Models
- Zero-Shot Domain Adaptation of Anomalous Samples for Semi-Supervised Anomaly Detection
- Zero-Shot Personalized Lip-To-Speech Synthesis with Face Image Based Voice Control
- Zero-Shot Sound Event Classification Using a Sound Attribute Vector with Global and Local Feature Learning
- Zero-Shot Speech Emotion Recognition Using Generative Learning with Reconstructed Prototypes
- Zone Plate Virtual Lenses for Memory-Constrained NLOS Imaging
- ifUNet++: Iterative Feedback UNet++ for Infrared Small Target Detection
- jaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus
- mmSense: Detecting Concealed Weapons with a Miniature Radar Sensor
- mmWave Wi-Fi Trajectory Estimation with Continuous-Time Neural Dynamic Learning
- ψ-Net: Point Structural Information Network for No-Reference Point Cloud Quality Assessment
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.