← All conferences

ICASSP 2023 Accepted Papers

The full list of 2,718 papers accepted at ICASSP 2023 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. Recurrent Fine-Grained Self-Attention Network for Video Crowd Counting
  2. Recursive Estimation of User Intent From Noninvasive Electroencephalography Using Discriminative Models
  3. Recursive Joint Attention for Audio-Visual Fusion in Regression Based Emotion Recognition
  4. Recursive/Iterative Unique Projection-Aggregation Decoding of Reed-Muller Codes
  5. Reducing Language Confusion for Code-Switching Speech Recognition with Token-Level Language Diarization
  6. Reducing the Communication and Computational Cost of Random Fourier Features Kernel LMS in Diffusion Networks
  7. Reducing the Computational Complexity of Learning with Random Convolutional Features
  8. Reducing the GAP Between Streaming and Non-Streaming Transducer-Based ASR by Adaptive Two-Stage Knowledge Distillation
  9. Refined Pseudo Labeling for Source-Free Domain Adaptive Object Detection
  10. Region-Awared Transformer with Asymmetric Loss in Multi-Label Classification
  11. Regression to Classification: Waveform Encoding for Neural Field-Based Audio Signal Representation
  12. Regularized Deep Generative Model Learning for Real-Time Massive MIMO Channel Tracking
  13. Regularized EM Algorithm
  14. Regularized Neural Detection for Millimeter Wave Massive Mimo Communication Systems with One-Bit Adcs
  15. Relapse Detection in Patients with Psychotic Disorders Using Unsupervised Learning on Smartwatch Signals
  16. Relapse Prediction from Long-Term Wearable Data Using Self-Supervised Learning and Survival Analysis
  17. Relate Auditory Speech To Eeg By Shallow-Deep Attention-Based Network
  18. Relating EEG Recordings to Speech Using Envelope Tracking and The Speech-FFR
  19. Relational Representation Learning for Zero-Shot Relation Extraction with Instance Prompting and Prototype Rectification
  20. Relative Dynamic Time Warping Comparison for Pronunciation Errors
  21. Relevance Propagation through Deep Conditional Random Fields
  22. Reliability Estimation for Synthetic Speech Detection
  23. Reliable Beamforming at Terahertz Bands: Are Causal Representations the Way Forward?
  24. Reliable Cluster-Based Framework for Open Set Domain Adaptation
  25. Removing Radio Frequency Interference From Auroral Kilometric Radiation With Stacked Autoencoders
  26. Repackagingaugment: Overcoming Prediction Error Amplification in Weight-Averaged Speech Recognition Models Subject to Self-Training
  27. Repetition Counting from Compressed Videos Using Sparse Residual Similarity
  28. Representation Learning of Clinical Multivariate Time Series with Random Filter Banks
  29. Representation of Vocal Tract Length Transformation Based on Group Theory
  30. Residual Hybrid Attention Network for Compression Artifact Reduction
  31. Residual Squeeze-and-Excitation U-Shaped Network for Minutia Extraction in Contactless Fingerprint Images
  32. Resolving Doppler Ambiguity Via Spread Phase Alignment in FDA-MIMO Radar
  33. Resource Allocation for UAV-Enabled Integrated Sensing and Communication (ISAC) via Multi-Objective Optimization
  34. Resource-Efficient Transfer Learning from Speech Foundation Model Using Hierarchical Feature Fusion
  35. Restoration of Time-Varying Graph Signals using Deep Algorithm Unrolling
  36. Rethink Long-Tailed Recognition with Vision Transforms
  37. Rethink Pair-Wise Self-Supervised Cross-Modal Retrieval From A Contrastive Learning Perspective
  38. Rethinking Implicit Neural Representations For Vision Learners
  39. Rethinking Learning-Based Method for Lossless Genome Compression
  40. Rethinking Random Walk in Graph Representation Learning
  41. Rethinking Rule-Based Approaches in Session-Based Recommendation
  42. Rethinking the Reasonability of the Test Set for Simultaneous Machine Translation
  43. Retiformer: Retinex-Based Enhancement In Transformer For Low-Light Image
  44. Retinal Biomarkers for Detecting Diabetic Retinopaty Using Smartphone-Based Deep Learning Frameworks
  45. Retrieval-Based Natural 3D Human Motion Generation
  46. Reverberation as Supervision For Speech Separation
  47. Revisit Out-Of-Vocabulary Problem For Slot Filling: A Unified Contrastive Framework With Multi-Level Data Augmentations
  48. Revisit Sampling Theory of Bandlimited Graph Signals: One Bridge Between GSP and DSP
  49. Rigid-Body Sound Synthesis with Differentiable Modal Resonators
  50. Ripple Sparse Self-Attention for Monaural Speech Enhancement
  51. Robust Acoustic And Semantic Contextual Biasing In Neural Transducers For Speech Recognition
  52. Robust Adaptive Beamforming with Proximal Method
  53. Robust Angle Estimation for Hybrid mmWave Systems
  54. Robust Audio-Visual ASR with Unified Cross-Modal Attention
  55. Robust Autoencoders for Collective Corruption Removal
  56. Robust Binary Component Decompositions
  57. Robust Binaural Sound Localisation with Temporal Attention
  58. Robust Content-Variant Reference Image Quality Assessment Via Similar Patch Matching
  59. Robust Data-Driven Accelerated Mirror Descent
  60. Robust Data2VEC: Noise-Robust Speech Representation Learning for ASR by Combining Regression and Improved Contrastive Learning
  61. Robust Dominant Periodicity Detection for Time Series with Missing Data
  62. Robust Fir Filters for Wireless Low-Frequency Sound Zones
  63. Robust GMM Parameter Estimation via the K-BM Algorithm
  64. Robust Hyperspectral Anomaly Detection with Simultaneous Mixed Noise Removal via Constrained Convex Optimization
  65. Robust Hypothesis Testing With Moment Constrained Uncertainty Sets
  66. Robust Iterative Solution for Linear Array-Based 3-D Localization by Message Passing
  67. Robust Knowledge Distillation from RNN-T Models with Noisy Training Labels Using Full-Sum Loss
  68. Robust Log-Based Anomaly Detection with Hierarchical Contrastive Learning
  69. Robust M-Estimation Based Distributed Expectation Maximization Algorithm with Robust Aggregation
  70. Robust Monocular Localization of Drones by Adapting Domain Maps to Depth Prediction Inaccuracies
  71. Robust Multi-Object Tracking With Spatial Uncertainty
  72. Robust Network Topologies for Distributed Learning
  73. Robust Online Multiband Drift Estimation in Electrophysiology Data
  74. Robust Self-Guided Deep Image Prior
  75. Robust Spatiotemporal Fusion of Satellite Images via Convex Optimization
  76. Robust Subspace Tracking with Contamination Mitigation via α-Divergence
  77. Robust Time Series Recovery and Classification Using Test-Time Noise Simulator Networks
  78. Robust Video Anomaly Detection Framework via Prior Knowledge and Multi-Path Frame Prediction
  79. Robust Video Object Segmentation with Restricted Attention
  80. Robust Watermarking Scheme in Encrypted Domain Based on Integer Lifting Wavelet Transform and Compressed Sensing
  81. Robust and Globally Sparse Pca via Majorization-Minimization and Variable Splitting
  82. Robust and Parallelizable Tensor Completion Based on Tensor Factorization and Maximum Correntropy Criterion
  83. Robust multi-modal speech emotion recognition with ASR error adaptation
  84. Robustdistiller: Compressing Universal Speech Representations for Enhanced Environment Robustness
  85. Robustness and Convergence of Mirror Descent for Blind Deconvolution
  86. Robustness of Deep Equilibrium Architectures to Changes in the Measurement Model
  87. Robustness-Preserving Lifelong Learning Via Dataset Condensation
  88. Role of Bias Terms in Dot-Product Attention
  89. Role of Lexical Boundary Information in Chunk-Level Segmentation for Speech Emotion Recognition
  90. Room Impulse Response Reconstruction Based on Spatio-Temporal-Spectral Features Learned from a Spherical Microphone Array Measurement
  91. Row Conditional-TGAN for Generating Synthetic Relational Databases
  92. Rumor Detection Via Assessing the Spreading Propensity of Users
  93. Runtime Prediction of Machine Learning Algorithms in Automl Systems
  94. RØROS: Building a Responsive Online Recommender System via Meta-Gradients Updating
  95. S-Feature Pyramid Network and Attention Model for Drone Detection
  96. S3I-PointHop: SO(3)-Invariant PointHop for 3D Point Cloud Classification
  97. SADE: A Self-Adaptive Expert for Multi-Dataset Question Answering
  98. SADI: A Self-Adaptive Decomposed Interpretable Framework for Electric Load Forecasting Under Extreme Events
  99. SAMO: Speaker Attractor Multi-Center One-Class Learning For Voice Anti-Spoofing
  100. SAN: A Robust End-to-End ASR Model Architecture
  101. SAR Image Despeckling with Residual-in-Residual Dense Generative Adversarial Network
  102. SARdBScene: Dataset and Resnet Baseline for Audio Scene Source Counting and Analysis
  103. SC-Net: Salient Point and Curvature Based Adversarial Point Cloud Generation Network
  104. SCA: Streaming Cross-Attention Alignment For Echo Cancellation
  105. SCSGNet: Spatial-Correlated and Shape-Guided Network for Breast Mass Segmentation
  106. SD-PINN: Physics Informed Neural Networks for Spatially Dependent PDES
  107. SDG-L: A Semiparametric Deep Gaussian Process based Framework for Battery Capacity Prediction
  108. SDRNet: Shape Decoupled Regression Network for 3d face Reconstruction
  109. SDTN: Speaker Dynamics Tracking Network for Emotion Recognition in Conversation
  110. SENER: Sentiment Element Named Entity Recognition for Aspect-Based Sentiment Analysis
  111. SEPDIFF: Speech Separation Based on Denoising Diffusion Model
  112. SFEMGN: Image Denoising with Shallow Feature Enhancement Network and Multi-Scale ConvGRU
  113. SFR: Semantic-Aware Feature Rendering of Point Cloud
  114. SG-VAD: Stochastic Gates Based Speech Activity Detection
  115. SIAST: A Slot Imbalance-Aware Self-Training Scheme for Semi-Supervised Slot Filling
  116. SIGVIC: Spatial Importance Guided Variable-Rate Image Compression
  117. SINCO: A Novel Structural Regularizer for Image Compression Using Implicit Neural Representations
  118. SL-MoE: A Two-Stage Mixture-of-Experts Sequence Learning Framework for Forecasting Rapid Intensification of Tropical Cyclone
  119. SLBERT: A Novel Pre-Training Framework for Joint Speech and Language Modeling
  120. SLICER: Learning Universal Audio Representations Using Low-Resource Self-Supervised Pre-Training
  121. SMCL: Saliency Masked Contrastive Learning for Long-Tailed Visual Recognition
  122. SMUG: Towards Robust Mri Reconstruction by Smoothed Unrolling
  123. SPADE: Self-Supervised Pretraining for Acoustic Disentanglement
  124. SPASHT: Semantic and Pragmatic Speech Features for Automatic Assessment of Autism
  125. SPECTRANET-SO(3): Learning Satellite Orientation from Optical Spectra by Implicitly Modeling Mutually Exclusive Probability Distributions on The Rotation Manifold
  126. SQA: Strong Guidance Query with Self-Selected Attention for Human-Object Interaction Detection
  127. SQuId: Measuring Speech Naturalness in Many Languages
  128. SR-init: An Interpretable Layer Pruning Method
  129. SRTNET: Time Domain Speech Enhancement via Stochastic Refinement
  130. SS-ADMM: Stationary and Sparse Granger Causal Discovery for Cortico-Muscular Coupling
  131. SSGD: A Smartphone Screen Glass Dataset for Defect Detection
  132. SSI-Net: A Multi-Stage Speech Signal Improvement System for ICASSP 2023 SSI Challenge
  133. SSVMR: Saliency-Based Self-Training for Video-Music Retrieval
  134. ST-MVDNet++: Improve Vehicle Detection with Lidar-Radar Geometrical Augmentation via Self-Training
  135. ST360IQ: No-Reference Omnidirectional Image Quality Assessment With Spherical Vision Transformers
  136. STACKMAPS: A Visualization Technique for Diabetic Retinopathy Grading
  137. STYX: Adaptive Poisoning Attacks Against Byzantine-Robust Defenses in Federated Learning
  138. SUVR: A Search-Based Approach to Unsupervised Visual Representation Learning
  139. SVMV: Spatiotemporal Variance-Supervised Motion Volume for Video Frame Interpolation
  140. SW-WAVENET: Learning Representation from Spectrogram and Wavegram Using Wavenet for Anomalous Sound Detection
  141. SYNTACC : Synthesizing Multi-Accent Speech By Weight Factorization
  142. SafeDeep: A Scalable Robustness Verification Framework for Deep Neural Networks
  143. Saliency-Driven Hierarchical Learned Image Coding for Machines
  144. Salient Co-Speech Gesture Synthesizing with Discrete Motion Representation
  145. Sample-Adapt Fusion Network for RGB-D Hand Detection in the Wild
  146. Sample-Aware Knowledge Distillation for Long-Tailed Learning
  147. Sample-Efficient Robust MMV Recovery Algorithm
  148. Sampling Order-Limited Signals on the Sphere
  149. Sandformer: CNN and Transformer under Gated Fusion for Sand Dust Image Restoration
  150. Sanet: Spatial Attention Network with Global Average Contrast Learning for Infrared Small Target Detection
  151. Scalable Multi-Task Semantic Communication System with Feature Importance Ranking
  152. Scalable Weight Reparametrization for Efficient Transfer Learning
  153. Scalable and Secure Federated XGBoost
  154. Scale-Adaptive Tiny Object Detection Enhanced by Across-Scale and Shape-Preserved Semantic Location
  155. ScaleMix: Intra- And Inter-Layer Multiscale Feature Combination for Change Detection
  156. Scaling Law Analysis for Covariance Based Activity Detection in Cooperative Multi-Cell Massive Mimo
  157. Scoreformer: Score Fusion-Based Transformers for Weakly-Supervised Violence Detection
  158. Search for Efficient Deep Visual-Inertial Odometry Through Neural Architecture Search
  159. Second-Order Statistic Deviation to Model Anomalies in the Design of Unsupervised Detectors
  160. Select The Best: Enhancing Graph Representation with Adaptive Negative Sample Selection
  161. Selecting Language Models Features VIA Software-Hardware Co-Design
  162. Selective Film Conditioning with CTC-Based ASR Probability for Speech Enhancement
  163. Self Supervised Bert for Legal Text Classification
  164. Self-Adaptive Incremental Machine Speech Chain for Lombard TTS with High-Granularity ASR Feedback in Dynamic Noise Condition
  165. Self-Adaptive Reasoning on Sub-Questions for Multi-Hop Question Answering
  166. Self-Attention Based Action Segmentation Using Intra-And Inter-Segment Representations
  167. Self-Attention for Enhanced OAMP Detection in MIMO Systems
  168. Self-Convolution for Automatic Speech Recognition
  169. Self-Distillation Hashing for Efficient Hamming Space Retrieval
  170. Self-Healing Through Error Detection, Attribution, and Retraining
  171. Self-Paced Partial Domain-Aware Learning for Face Anti-Spoofing
  172. Self-Remixing: Unsupervised Speech Separation VIA Separation and Remixing
  173. Self-Similarity is all You Need for Fast and Light-Weight Generic Event Boundary Detection
  174. Self-Sufficient Framework for Continuous Sign Language Recognition
  175. Self-Supervised Accent Learning for Under-Resourced Accents Using Native Language Data
  176. Self-Supervised Adversarial Training for Contrastive Sentence Embedding
  177. Self-Supervised Audio-Visual Speaker Representation with Co-Meta Learning
  178. Self-Supervised Audio-Visual Speech Representations Learning by Multimodal Self-Distillation
  179. Self-Supervised Facial Action Unit Detection with Region and Relation Learning
  180. Self-Supervised Guided Hypergraph Feature Propagation for Semi-Supervised Classification with Missing Node Features
  181. Self-Supervised Hierarchical Metrical Structure Modeling
  182. Self-Supervised Learning for Speech Enhancement Through Synthesis
  183. Self-Supervised Learning of Audio Representations using Angular Contrastive Loss
  184. Self-Supervised Learning with Bi-Label Masked Speech Prediction for Streaming Multi-Talker Speech Recognition
  185. Self-Supervised Learning with Explorative Knowledge Distillation
  186. Self-Supervised Learning-Based Source Separation for Meeting Data
  187. Self-Supervised Representations for Singing Voice Conversion
  188. Self-Supervised Representations in Speech-Based Depression Detection
  189. Self-Supervised Speech Representation Learning for Keyword-Spotting With Light-Weight Transformers
  190. Self-Transriber: Few-Shot Lyrics Transcription With Self-Training
  191. Selinet: A Lightweight Model for Single Channel Speech Separation
  192. SemGeo: Semantic Keywords for Cross-View Image Geo-Localization
  193. Semantic Centralized Contrastive Learning for Unsupervised Hashing
  194. Semantic Memory Guided Image Representation for Polyp Segmentation
  195. Semantic Preprocessor for Image Compression for Machines
  196. Semantic Preserving Learning for Task-Oriented Point Cloud Downsampling
  197. Semantic-Aware Gated Fusion Network For Interactive Colorization
  198. Semantic-Preserving Augmentation for Robust Image-Text Retrieval
  199. SemanticAC: Semantics-Assisted Framework for Audio Classification
  200. Semantically-Informed Deep Neural Networks For Sound Recognition
  201. Semantics-Aware Gamma Correction for Unsupervised Low-Light Image Enhancement
  202. Semantics-Disentangled Contrastive Embedding for Generalized Zero-Shot Learning
  203. Semantics-Guided Object Removal for Facial Images: with Broad Applicability and Robust Style Preservation
  204. Semi-Federated Learning for Edge Intelligence with Imperfect SIC
  205. Semi-Supervised Contrastive Learning with Soft Mask Attention for Facial Action Unit Detection
  206. Semi-Supervised Domain Generalization with Graph-Based Classifier
  207. Semi-Supervised Graph Ultra-Sparsifier Using Reweighted ℓ1 Optimization
  208. Semi-Supervised Learning with Per-Class Adaptive Confidence Scores for Acoustic Environment Classification with Imbalanced Data
  209. Semi-Supervised Local Structured Feature Learning with Dynamic Maximum Entropy Graph
  210. Semi-Supervised Remote Sensing Image Change Detection Using Mean Teacher Model for Constructing Pseudo-Labels
  211. Semi-Supervised Semantic Segmentation with Structured Output Space Adaption
  212. Semi-Supervised Sound Event Detection with Pre-Trained Model
  213. Semi-Supervised Speech Enhancement Based On Speech Purity
  214. Semi-Swinderain: Semi-Supervised Image Deraining Network Using SWIN Transformer
  215. Sensor Selection for Angle of Arrival Estimation Based on the Two-Target Cramér-Rao Bound
  216. Sequence-Based Device-Free Gesture Recognition Framework for Multi-Channel Acoustic Signals
  217. Sequential Datum-Wise Joint Feature Selection and Classification in the Presence of External Classifier
  218. Sequential Invariant Information Bottleneck
  219. Seri: Sketching-Reasoning-Integrating Progressive Workflow for Empathetic Response Generation
  220. Shadocnet: Learning Spatial-Aware Tokens in Transformer for Document Shadow Removal
  221. Shadow Removal of Text Document Images Using Background Estimation and Adaptive Text Enhancement
  222. Sharing Low Rank Conformer Weights for Tiny Always-On Ambient Speech Recognition Models
  223. Shift to Your Device: Data Augmentation for Device-Independent Speaker Verification Anti-Spoofing
  224. Short-Segment Speaker Verification Using ECAPA-TDNN with Multi-Resolution Encoder
  225. Show Me the Instruments: Musical Instrument Retrieval From Mixture Audio
  226. Shuffleaugment: A Data Augmentation Method Using Time Shuffling
  227. Shuffled Autoregression for Motion Interpolation
  228. Sign Language Recognition via Deformable 3D Convolutions and Modulated Graph Convolutional Networks
  229. Signal Analysis-Synthesis Using the Quantum Fourier Transform
  230. Signal Processing And Quantum State Tomography on Noisy Devices
  231. Signal Processing Grand Challenge 2023 - E-Prevention: Sleep Behavior as an Indicator of Relapses in Psychotic Patients
  232. Signal Processing On Product Spaces
  233. Signal Processing with Optical Quadratic Random Sketches
  234. Signal Reconstruction for FMCW Radar Interference Mitigation Using Deep Unfolding
  235. Similarity Relation Preserving Cross-Modal Learning for Multispectral Pedestrian Detection Against Adversarial Attacks
  236. Simple Pooling Front-Ends for Efficient Audio Classification
  237. Simplicial Vector Autoregressive Model For Streaming Edge Flows
  238. Simulating Realistic Speech Overlaps Improves Multi-Talker ASR
  239. Simultaneous Acoustic Echo Sorting and 3-D Room Geometry Inference
  240. Simultaneous Estimation of Direction of Arrival and Sound Speed Using a Non-Uniform Sensor Array
  241. Simultaneous Reconstruction and Uncertainty Quantification for Tomography
  242. Simultaneously Learning Robust Audio Embeddings and Balanced Hash Codes for Query-by-Example
  243. Sine: Similarity-Regularized Intra-Class Exploitation for Cross-Granularity Few-Shot Learning
  244. SingNet: a real-time Singing Voice beat and Downbeat Tracking System
  245. Singing Voice Synthesis Based on a Musical Note Position-Aware Attention Mechanism
  246. Single Domain Dynamic Generalization for Iris Presentation Attack Detection
  247. Single-Anchor UWB Localization Using Channel Impulse Response Distributions
  248. Single-Channel Speech Enhancement with Deep Complex U-Networks and Probabilistic Latent Space Models
  249. Single-Particle Tracking by Graph Transformer
  250. Single-Photon Image Super-Resolution via Self-Supervised Learning
  251. Single-Sample Direction-of-Arrival Estimation for Fast and Robust 3D Localization With Real Measurements from a Massive MIMO System
  252. Single-Shot Domain Adaptation via Target-Aware Generative Augmentations
  253. Single-Shot Fractional Fourier Phase Retrieval
  254. Single-branch Network for Multimodal Training
  255. Sinusoidal Frequency Estimation by Gradient Descent
  256. Sketch Less Face Image Retrieval: A New Challenge
  257. Skillnet-NLG: General-Purpose Natural Language Generation with a Sparsely Activated Approach
  258. Slot-Triggered Contextual Biasing For Personalized Speech Recognition Using Neural Transducers
  259. Small-Footprint Slimmable Networks for Keyword Spotting
  260. Smart Split-Federated Learning over Noisy Channels for Embryo Image Segmentation
  261. Smoothing Complex-Valued Signals on Graphs with Monte-Carlo
  262. Smoothing Point Adjustment-Based Evaluation of Time Series Anomaly Detection
  263. Soft 2D-to-3D Delivery Using Deep Graph Neural Networks for Holographic-Type Communication
  264. Soft Dynamic Time Warping for Multi-Pitch Estimation and Beyond
  265. Soft Label Coding for end-to-end Sound Source Localization with ad-hoc Microphone Arrays
  266. Solving Audio Inverse Problems with a Diffusion Model
  267. Solving Jigsaw Puzzle of Large Eroded Gaps Using Puzzlet Discriminant Network
  268. Sora: Scalable Black-Box Reachability Analyser on Neural Networks
  269. Source Localization for Extremely Large-Scale Antenna Arrays with Spatial Non-Stationarity
  270. Source-Filter HiFi-GAN: Fast and Pitch Controllable High-Fidelity Neural Vocoder
  271. Source-Free Unsupervised Domain Adaptation for Question Answering
  272. Space-Time Graph Neural Networks with Stochastic Graph Perturbations
  273. Space-Time Variable Density Samplings for Sparse Bandlimited Graph Signals Driven by Diffusion Operators
  274. Spammer Detection on Short Video Applications: A new Challenge and Baselines
  275. Sparse Aggregation-Based Channel Estimation For Massive Mimo Systems With Decentralized Baseband Processing
  276. Sparse Asynchronous Samples from Networks of Tems for Reconstruction of Classes of Non-Bandlimited Signals
  277. Sparse Bayesian Learning Assisted Decision Fusion in Millimeter Wave Massive MIMO Sensor Networks
  278. Sparse Bayesian Learning Based Three-Dimensional Imaging for Antenna Array Radar
  279. Sparse Black-Box Inversion Attack with Limited Information
  280. Sparse Convolution Based Octree Feature Propagation for Lidar Point Cloud Compression
  281. Sparse Delay-Doppler Channel Estimation for OTFS Modulation Using 2D-Music
  282. Sparse Error Correction for Power Network Parameters
  283. Sparse Graph Learning with Spectrum Prior for Deep Graph Convolutional Networks
  284. Sparse Mixture Once-for-all Adversarial Training for Efficient in-situ Trade-off between Accuracy and Robustness of DNNs
  285. Sparse Non-Contact Multiple People Localization and Vital Signs Monitoring Via FMCW Radar
  286. Sparse Representations with Cone Atoms
  287. Sparse and Structured Modelling of Underwater Acoustic Channel Impulse Responses
  288. Sparsity Constraint Implementation for the Joint Eigenvalue Decomposition of Matrices
  289. Sparsity-Driven Joint Blind Deconvolution-Demodulation with Application to Motor Fault Detection
  290. Sparsity-Smoothness-Aware Power Spectral Density Estimation with Application to Phased Array Weather Radar
  291. Spatial Active Noise Control Method Based on Sound Field Interpolation from Reference Microphone Signals
  292. Spatial Correlation Fusion Network for Few-Shot Segmentation
  293. Spatial Cross-Attention for Transformer-Based Image Captioning
  294. Spatial Graph Signal Interpolation with an Application for Merging BCI Datasets with Various Dimensionalities
  295. Spatial Inference Using Censored Multiple Testing with Fdr Control
  296. Spatial Similarity Guidance for Few-Shot Segmentation
  297. Spatial-Domain Object Detection Under Mimo-Fmcw Automotive Radar Interference
  298. Spatial-Temporal Graph Convolutional Network Boosted Flow-Frame Prediction For Video Anomaly Detection
  299. Spatially Informed Independent vector analysis for Source Extraction based on the convolutive Transfer Function Model
  300. Spatially Selective Deep Non-Linear Filters For Speaker Extraction
  301. Spatio-Temporal Attention in Multi-Granular Brain Chronnectomes For Detection of Autism Spectrum Disorder
  302. Spatio-Temporal Hybrid Fusion of CAE and SWin Transformers for Lung Cancer Malignancy Prediction
  303. Spatio-Temporal Structure Consistency for Semi-Supervised Medical Image Classification
  304. Speaker Change Detection For Transformer Transducer ASR
  305. Speaker Diaphragm Excursion Prediction: Deep Attention and Online Adaptation
  306. Speaker Recognition with Two-Step Multi-Modal Deep Cleansing
  307. Speaker-Aware Hierarchical Transformer For Personality Recognition In Multiparty Dialogues
  308. Speaker-Independent Acoustic-to-Articulatory Speech Inversion
  309. Speakeraugment: Data Augmentation for Generalizable Source Separation via Speaker Parameter Manipulation
  310. Spectral Clustering-Aware Learning of Embeddings for Speaker Diarisation
  311. Spectral Super-Resolution on the Unit Circle Via Gradient Descent
  312. Spectro-Temporal Post-Filtering Via Short-Time Target Cancellation for Directional Speech Enhancement in a Dual-Microphone Hearing AID
  313. Speech Dereverberation with a Reverberation Time Shortening Target
  314. Speech Emotion Recognition Based on Low-Level Auto-Extracted Time-Frequency Features
  315. Speech Emotion Recognition Via Two-Stream Pooling Attention With Discriminative Channel Weighting
  316. Speech Emotion Recognition via Heterogeneous Feature Learning
  317. Speech Enhancement with Intelligent Neural Homomorphic Synthesis
  318. Speech Intelligibility Classifiers from 550k Disordered Speech Samples
  319. Speech MOS Multi-Task Learning and Rater Bias Correction
  320. Speech Modeling with a Hierarchical Transformer Dynamical VAE
  321. Speech Privacy Leakage from Shared Gradients in Distributed Learning
  322. Speech Reconstruction from Silent Tongue and Lip Articulation by Pseudo Target Generation and Domain Adversarial Training
  323. Speech Separation with Large-Scale Self-Supervised Learning
  324. Speech Signal Improvement Using Causal Generative Diffusion Models
  325. Speech Summarization of Long Spoken Document: Improving Memory Efficiency of Speech/Text Encoders
  326. Speech and Noise Dual-Stream Spectrogram Refine Network With Speech Distortion Loss For Robust Speech Recognition
  327. Speech-Based Emotion Recognition with Self-Supervised Models Using Attentive Channel-Wise Correlations and Label Smoothing
  328. Speech-Text Based Multi-Modal Training with Bidirectional Attention for Improved Speech Recognition
  329. Speechlmscore: Evaluating Speech Generation Using Speech Language Model
  330. Spherical Sector Harmonics Based Soundfield Radial Extrapolation And Robustness Analysis
  331. Spherical Vector Quantization for Spatial Direction Coding
  332. Spice+: Evaluation of Automatic Audio Captioning Systems with Pre-Trained Language Models
  333. Spike-Based Optical Flow Estimation Via Contrastive Learning
  334. Spoofed Training Data for Speech Spoofing Countermeasure Can Be Efficiently Created Using Neural Vocoders
  335. Spteae: A Soft Prompt Transfer Model for Zero-Shot Cross-Lingual Event Argument Extraction
  336. Stabilising and Accelerating Light Gated Recurrent Units for Automatic Speech Recognition
  337. Stacking-Based Attention Temporal Convolutional Network for Action Segmentation
  338. Stargan-vc Based Cross-Domain Data Augmentation for Speaker Verification
  339. Static and Dynamic Source and Filter Cues for Classification of Amyotrophic Lateral Sclerosis Patients and Healthy Subjects
  340. Static-Scene Constrained Optimization for Matrix/Tensor-Decomposition-free Foreground-Background Separation
  341. Statistical Analysis of Speech Disorder Specific Features to Characterise Dysarthria Severity Level
  342. Stay In The Middle: A Semi-Supervised Model for CT Metal Artifact Reduction
  343. Step restriction for improving adversarial attacks
  344. Stereoscopic Video Retargeting Based on Camera Motion Classification
  345. Stochastic Optimization of Vector Quantization Methods in Application to Speech and Image Processing
  346. Stochastic Super-Resolution For Gaussian Textures
  347. Strategies for Enhanced Signal Modulation Classifications Under Unknown Symbol Rates and Noise Conditions
  348. Stream Attention Based U-Net for L3DAS23 Challenge
  349. StreamSpeech: Low-Latency Neural Architecture for High-Quality on-Device Speech Synthesis
  350. Streaming Joint Speech Recognition and Disfluency Detection
  351. Streaming Multi-Channel Speech Separation with Online Time-Domain Generalized Wiener Filter
  352. Streaming Stroke Classification of Online Handwriting
  353. Streaming Voice Conversion via Intermediate Bottleneck Features and Non-Streaming Teacher Guidance
  354. String-Based Molecule Generation Via Multi-Decoder VAE
  355. Structural Optimization of Factor Graphs for Symbol Detection via Continuous Clustering and Machine Learning
  356. Structural Reparameterization Lightweight Network for Video Action Recognition
  357. Structure-Aware Multi-Feature Co-Learning for Dual Branch Face Super Resolution
  358. Structure-Aware Sparse Bayesian Learning-Based Channel Estimation for Intelligent Reflecting Surface-Aided MIMO
  359. Structure-Preserving and Redundancy-Free Features Refinement for Generalized Zero-Shot Learning
  360. Structured Errors-in-Variables Modelling for Cortico-Muscular Coherence Enhancement
  361. Structured Pruning of Self-Supervised Pre-Trained Models for Speech Recognition and Understanding
  362. Structured State Space Decoder for Speech Recognition and Synthesis
  363. Structured-Anchor Projected Clustering for Hyperspectral Images
  364. Stuart: Individualized Classroom Observation of Students with Automatic Behavior Recognition And Tracking
  365. Study And Design Of Robust Personal Sound Zones With Vast Using Low Rank Rirs
  366. Study of Manifold Geometry Using Multiscale Non-Negative Kernel Graphs
  367. Study on the Fairness of Speaker Verification Systems Across Accent and Gender Groups
  368. Style Modeling for Multi-Speaker Articulation-to-Speech
  369. Sub-Band Contrastive Learning-Based Knowledge Distillation For Sound Classification
  370. Subband Dependency Modeling for Sound Event Detection
  371. Subgradient Descent Learning with Over-the-Air Computation
  372. Subject-Specific Adaptation for a Causally-Trained Auditory-Attention Decoding System
  373. Subspace Hybrid Beamforming for Head-Worn Microphone Arrays
  374. Subspace Modeling Enabled High-Sensitivity X-Ray Chemical Imaging
  375. Subspace-Based Detector For Distributed Mmwave Mimo Radar Sensors
  376. Suffix Retrieval-Augmented Language Modeling
  377. Summary on the Multimodal Information Based Speech Processing (MISP) 2022 Challenge
  378. Super Dilated Nested Arrays with Ideal Critical Weights and Increased Degrees of Freedom
  379. Super-Resolution Harmonic Retrieval of Non-Circular Signals
  380. Super-Resolution Information Enhancement for Crowd Counting
  381. Super-Resolution for Macro X-Ray Fluorescence Data Collected from Old Master Paintings
  382. Supercm: Revisiting Clustering for Semi-Supervised Learning
  383. Supervised Contrastive Learning as Multi-Objective Optimization for Fine-Tuning Large Pre-Trained Language Models
  384. Supervised Hierarchical Clustering Using Graph Neural Networks for Speaker Diarization
  385. Surface-Sampling Based Objective Quality Assessment Metrics for Meshes
  386. Surrogate Based Post-HOC Calibration for Distributional Shift
  387. Switching Kronecker Product Linear Filtering for Multispeaker Adaptive Speech Dereverberation
  388. Symbol Level Precoding in the RF Domain for Low Hardware Complexity RIS-Assisted MU-MISO Systems
  389. Symbol-Level Precoding is Related to Parameter Estimation from Quantized Data
  390. SyncNet: Correlating Objective for Time Delay Estimation in Audio Signals
  391. Syngen: A Syntactic Plug-And-Play Module for Generative Aspect-Based Sentiment Analysis
  392. Synthesizer Preset Interpolation Using Transformer Auto-Encoders
  393. Synthesizing Speech from ECoG with a Combination of Transformer-Based Encoder and Neural Vocoder
  394. Synthetic Pseudo Anomalies for Unsupervised Video Anomaly Detection: A Simple Yet Efficient Framework Based on Masked Autoencoder
  395. T5-SR: A Unified Seq-to-Seq Decoding Strategy for Semantic Parsing
  396. T5lephone: Bridging Speech and Text Self-Supervised Models for Spoken Language Understanding Via Phoneme Level T5
  397. TABLEIE: Capturing the Interactions Among Sub-Tasks in Information Extraction via Double Tables
  398. TAMformer: Multi-Modal Transformer with Learned Attention Mask for Early Intent Prediction
  399. TAPE: An End-to-End Timbre-Aware Pitch Estimator
  400. TAPLoss: A Temporal Acoustic Parameter Loss for Speech Enhancement
  401. TDMA-Based Multi-User Binary Computation Offloading in the Finite-Block-Length Regime
  402. TEA-PSE 3.0: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System For ICASSP 2023 Dns-Challenge
  403. TEFISTA-NET: GTD Parameter Estimation of Low-Frequency Ultra- Wideband Radar via Model-Based Deep Learning
  404. TF-GRIDNET: Making Time-Frequency Domain Models Great Again for Monaural Speaker Separation
  405. TFCnet: Time-Frequency Domain Corrector for Speech Separation
  406. TINYCOD: Tiny and Effective Model for Camouflaged Object Detection
  407. TOLD: a Novel Two-Stage Overlap-Aware Framework for Speaker Diarization
  408. TOPO-MLP : A Simplicial Network without Message Passing
  409. TRICL: Triplet Continual Learning
  410. TRUSTERA: A Live Conversation Redaction System
  411. TSPTQ-ViT: Two-Scaled Post-Training Quantization for Vision Transformer
  412. TSpeech-AI System Description to the 5th Deep Noise Suppression (DNS) Challenge
  413. TT-Net: Dual-Path Transformer Based Sound Field Translation in the Spherical Harmonic Domain
  414. Tangent Bundle Filters and Neural Networks: From Manifolds to Cellular Sheaves and Back
  415. Target Sound Extraction with Variable Cross-Modality Clues
  416. Target Speaker Extraction with Ultra-Short Reference Speech by VE-VE Framework
  417. Target Speaker Voice Activity Detection with Transformers and Its Integration with End-To-End Neural Diarization
  418. Target Velocity Estimation for Quantization-Based Cooperative MIMO Radar and Communications System
  419. Target-Speaker Voice Activity Detection Via Sequence-to-Sequence Prediction
  420. Targeted Adversarial Attacks Against Neural Machine Translation
  421. Tayloraecnet: A Taylor Style Neural Network For Full-Band Echo Cancellation
  422. TeAw: Text-Aware Few-Shot Remote Sensing Image Scene Classification
  423. Tell Model Where to Attend: Improving Interpretability of Aspect-Based Sentiment Classification via Small Explanation Annotations
  424. Tempo vs. Pitch: Understanding Self-Supervised Tempo Estimation
  425. Temporal Contrastive Learning with Curriculum
  426. Temporal Modeling Matters: A Novel Temporal Emotional Modeling Approach for Speech Emotion Recognition
  427. Tensor Completion for Efficient and Accurate Hyperparameter Optimisation in Large-Scale Statistical Learning
  428. Tensor Decomposition Based Latent Feature Clustering for Hyperspectral Band Selection
  429. Tensor Low Rank Column-Wise Compressive Sensing for Dynamic Imaging
  430. Tensor-based Complex-valued Graph Neural Network for Dynamic Coupling Multimodal brain Networks
  431. Tensorized LSSVMS For Multitask Regression
  432. Tensorized Neural Layer Decomposition for 2-D DOA Estimation
  433. Terminology-Aware Medical Dialogue Generation
  434. Ternary Weight Networks
  435. Test Your Samples Jointly: Pseudo-Reference for Image Quality Evaluation
  436. Test-Time Training-Free Domain Adaptation
  437. Text Classification In The Wild: A Large-Scale Long-Tailed Name Normalization Dataset
  438. Text is all You Need: Personalizing ASR Models Using Controllable Speech Synthesis
  439. Text-To-Speech Synthesis Based on Latent Variable Conversion Using Diffusion Probabilistic Model and Variational Autoencoder
  440. Text-to-ECG: 12-Lead Electrocardiogram Synthesis Conditioned on Clinical Text Reports
  441. Textless Direct Speech-to-Speech Translation with Discrete Speech Representation
  442. Textless Speech-to-Music Retrieval Using Emotion Similarity
  443. Tg-Critic: A Timbre-Guided Model For Reference-Independent Singing Evaluation
  444. The 2nd Clarity Enhancement Challenge for Hearing Aid Speech Intelligibility Enhancement: Overview and Outcomes
  445. The Ajmide Topic Segmentation System for the ICASSP 2023 General Meeting Understanding and Generation Challenge
  446. The DKU Post-Challenge Audio-Visual Wake Word Spotting System for the 2021 MISP Challenge: Deep Analysis
  447. The Edinburgh International Accents of English Corpus: Towards the Democratization of English ASR
  448. The First Pathloss Radio Map Prediction Challenge
  449. The MBSTOI Binaural Intelligibility Metric Using a Close-Talking Microphone Reference
  450. The Multimodal Information Based Speech Processing (Misp) 2022 Challenge: Audio-Visual Diarization And Recognition
  451. The NERCSLIP-USTC System for the L3DAS23 Challenge Task2: 3D Sound Event Localization and Detection (SELD)
  452. The NIO System for Audio-Visual Diarization and Recognition in MISP Challenge 2022
  453. The NPU-ASLP System for Audio-Visual Speech Recognition in MISP 2022 Challenge
  454. The NPU-Elevoc Personalized Speech Enhancement System for Icassp2023 DNS Challenge
  455. The Pipeline System of ASR and NLU with MLM-based data Augmentation Toward Stop Low-Resource Challenge
  456. The Potential of Neural Speech Synthesis-Based Data Augmentation for Personalized Speech Enhancement
  457. The R3VIVAL Dataset: Repository of Room Responses and 360 Videos of a Variable Acoustics Lab
  458. The Role of Initial Entanglement in Adaptive Gibbs State Preparation on Quantum Computers
  459. The Role of Memory in Social Learning When Sharing Partial Opinions
  460. The Secret Source : Incorporating Source Features to Improve Acoustic-To-Articulatory Speech Inversion
  461. The Uniqueness Problem of Physical Law Learning
  462. The Ustc System for Adress-m Challenge
  463. The WHU-Alibaba Audio-Visual Speaker Diarization System for the MISP 2022 Challenge
  464. The XMU System for Audio-Visual Diarization and Recognition in MISP Challenge 2022
  465. Thermal Infrared Image Inpainting Via Edge-Aware Guidance
  466. Think Before You Speak: Concept-Guided Explicit Persona Reasoning for Personalized Dialogue Generation
  467. This Changes to That : Combining Causal and Non-Causal Explanations to Generate Disease Progression in Capsule Endoscopy
  468. Time-Aware Multiway Adaptive Fusion Network for Temporal Knowledge Graph Question Answering
  469. Time-Domain Speech Enhancement Assisted by Multi-Resolution Frequency Encoder and Decoder
  470. Time-Frequency Awareness Network For Human Mesh Recovery From Videos
  471. Time-Resolved FMRI Shared Response Model Using Gaussian Process Factor Analysis
  472. Time-Varying Signals Recovery Via Graph Neural Networks
  473. Time-Weighted Frequency Domain Audio Representation with GMM Estimator for Anomalous Sound Detection
  474. TinyOOD: Effective out-of-Distribution Detection for TinyML
  475. To Regularize or Not to Regularize: The Role of Positivity in Sparse Array Interpolation with a Single Snapshot
  476. To Wake-Up or Not to Wake-Up: Reducing Keyword False Alarm by Successive Refinement
  477. Token2vec: A Joint Self-Supervised Pre-Training Framework Using Unpaired Speech and Text
  478. Top-K Visual Tokens Transformer: Selecting Tokens for Visible-Infrared Person Re-Identification
  479. Topgformer: Topological-Based Graph Transformer for Mapping Brain Structural Connectivity to Functional Connectivity
  480. Topological Signal Processing Over Weighted Simplicial Complexes
  481. Topological Slepians: Maximally Localized Representations of Signals Over Simplicial Complexes
  482. Topology Uncertainty Modeling For Imbalanced Node Classification on Graphs
  483. Torchaudio-Squim: Reference-Less Speech Quality and Intelligibility Measures in Torchaudio
  484. Toroidal Probabilistic Spherical Discriminant Analysis
  485. Toward A Multimodal Approach for Disfluency Detection and Categorization
  486. Toward Asymptotic Optimality: Sequential Unsupervised Regression of Density Ratio for Early Classification
  487. Toward Auto-Evaluation With Confidence-Based Category Relation-Aware Regression
  488. Toward Privacy-Enhancing Ambulatory-Based Well-Being Monitoring: Investigating User Re-Identification Risk in Multimodal Data
  489. Toward Universal Text-To-Music Retrieval
  490. Towards A Unified Conformer Structure: from ASR to ASV Task
  491. Towards Accurate and Real-Time End-of-Speech Estimation
  492. Towards Adversarially Robust Continual Learning
  493. Towards Bandwidth Estimation for Graph Signal Reconstruction
  494. Towards Building Text-to-Speech Systems for the Next Billion Users
  495. Towards Controllable Audio Texture Morphing
  496. Towards Dialogue Modeling Beyond Text
  497. Towards Diverse and Coherent Augmentation for Time-Series Forecasting
  498. Towards Domain Generalisation in ASR with Elitist Sampling and Ensemble Knowledge Distillation
  499. Towards Efficient and Optimal Joint Beamforming and Antenna Selection: A Machine Learning Approach
  500. Towards Explainable Recommendation Via Bert-Guided Explanation Generator
  501. Towards Hyperbolic Regularizers For Point Cloud Part Segmentation
  502. Towards Improved Room Impulse Response Estimation for Speech Recognition
  503. Towards Improved Sonar Performance Using Environment-Informed Sparse Sub-Array Processing
  504. Towards Interpretable Seizure Detection Using Wearables
  505. Towards Learning Emotion Information from Short Segments of Speech
  506. Towards Low-Power Heart Rate Estimation Based on User's Demographics and Activity Level For Wearables
  507. Towards Making a Trojan-Horse Attack on Text-to-Image Retrieval
  508. Towards Polymorphic Adversarial Examples Generation for Short Text
  509. Towards Practical Edge Inference Attacks Against Graph Neural Networks
  510. Towards Privacy and Utility in Tourette TIC Detection Through Pretraining Based on Publicly Available Video Data of Healthy Subjects
  511. Towards Real-Time Person Search with Invariant Feature Learning
  512. Towards Real-Time Single-Channel Speech Separation in Noisy and Reverberant Environments
  513. Towards Realizing the Value of Labeled Target Samples: A Two-Stage Approach for Semi-Supervised Domain Adaptation
  514. Towards Reducing Patient Effort for the Automatic Prediction of Speech Intelligibility in Head and Neck Cancers
  515. Towards Reliable Image Outpainting: Learning Structure-Aware Multimodal Fusion with Depth Guidance
  516. Towards Robust Audio-Based Vehicle Detection Via Importance-Aware Audio-Visual Learning
  517. Towards Robust Data-Driven Underwater Acoustic Localization: A Deep CNN Solution with Performance Guarantees for Model Mismatch
  518. Towards Scale Adaptive Underwater Detection Through Refined Pyramid Grid
  519. Towards Simultaneous Segmentation Of Liver Tumors And Intrahepatic Vessels Via Cross-Attention Mechanism
  520. Towards Trustworthy Multi-Label Sewer Defect Classification via Evidential Deep Learning
  521. Towards Trustworthy Phoneme Boundary Detection with Autoregressive Model and Improved Evaluation Metric
  522. Towards Zero-Shot Code-Switched Speech Recognition
  523. Towards Zero-Shot Personalized Table-to-Text Generation with Contrastive Persona Distillation
  524. Towards a More Stable and General Subgraph Information Bottleneck
  525. Towards a Robust and Efficient Classifier for Real World Radio Signal Modulation Classification
  526. Towards a Unified Training for Levenshtein Transformer
  527. TrOMR:Transformer-Based Polyphonic Optical Music Recognition
  528. Tracking Objects and Activities with Attention for Temporal Sentence Grounding
  529. Tracking Targets in Hyper-Scale Cameras Using Movement Predication
  530. Training Graph Neural Networks on Growing Stochastic Graphs
  531. Training Large-Vocabulary Neural Language Models by Private Federated Learning for Resource-Constrained Devices
  532. Training Neural Networks for Sequential Change-Point Detection
  533. Training Robust Spiking Neural Networks on Neuromorphic Data with Spatiotemporal Fragments
  534. Training Robust Spiking Neural Networks with Viewpoint Transform and Spatiotemporal Stretching
  535. Training Set Cleansing of Backdoor Poisoning by Self-Supervised Representation Learning
  536. Training Sound Event Detection with Soft Labels from Crowdsourced Annotations
  537. Training Stronger Spiking Neural Networks with Biomimetic Adaptive Internal Association Neurons
  538. TransLink: Transformer-Based Embedding for Tracklets' Global Link
  539. Transadapt: A Transformative Framework for Online Test Time Adaptive Semantic Segmentation
  540. Transaudio: Towards the Transferable Adversarial Audio Attack Via Learning Contextualized Perturbations
  541. Transceiver Design for MIMO-DFRC Systems
  542. Transcription Free Filler Word Detection with Neural Semi-CRFs
  543. Transductive Matrix Completion with Calibration for Multi-Task Learning
  544. Transferring Quantified Emotion Knowledge for the Detection of Depression in Alzheimer's Disease Using Forestnets
  545. Transformer-Based Bioacoustic Sound Event Detection on Few-Shot Learning Tasks
  546. Transformer-Based Deep Hashing Method for Multi-Scale Feature Fusion
  547. Transformer-Based Multi-Prototype Approach for Diabetic Macular Edema Analysis in OCT Images
  548. Transformer-based tracking Network for Maneuvering Targets
  549. Transient Dictionary Learning for Compressed Time-of-Flight Imaging
  550. Transmit Energy Focusing For Parameter Estimation in Transmit Beamspace Slow-Time MIMO Radar
  551. Transplayer: Timbre Style Transfer with Flexible Timbre Control
  552. Transwnet: Integrating Transformers into CNNS via Row and Column Attention for Abdominal Multi-Organ Segmentation
  553. Tree-Like Interaction Learning for Bundle Recommendation
  554. TreeXGNN: can gradient-boosted decision trees help boost heterogeneous graph neural networks?
  555. TriAAN-VC: Triple Adaptive Attention Normalization for Any-to-Any Voice Conversion
  556. TrimTail: Low-Latency Streaming ASR with Simple But Effective Spectrogram-Level Length Penalty
  557. Trinet: Stabilizing Self-Supervised Learning From Complete or Slow Collapse
  558. Trust Your Partner's Friends: Hierarchical Cross-Modal Contrastive Pre-Training for Video-Text Retrieval
  559. Twitter Stance Detection via Neural Production Systems
  560. Two-Branch Multi-Scale Deep Neural Network for Generalized Document Recapture Attack Detection
  561. Two-Phase Prototypical Contrastive Domain Generalization for Cross-Subject EEG-Based Emotion Recognition
  562. Two-Stage Neural Network for ICASSP 2023 Speech Signal Improvement Challenge
  563. Two-Stage UNet with Multi-Axis Gated Multilayer Perceptron for Monaural Noisy-Reverberant Speech Enhancement
  564. Two-Stage Video De-Raining with Spatio-Temporal Fusion and Illumination-Invariant Detail Preservation
  565. Two-Step Band-Split Neural Network Approach For Full-Band Residual Echo Suppression
  566. Two-Stream Decoder Feature Normality Estimating Network for Industrial Anomaly Detection
  567. Two-Stream Joint-Training for Speaker Independent Acoustic-to-Articulatory Inversion
  568. U-Beat: A Multi-Scale Beat Tracking Model Based on Wave-U-Net
  569. U-Shiftformer: Brain Tumor Segmentation Using A Shifted Attention Mechanism
  570. UAV Local Path Planning Based on Improved Proximal Policy Optimization Algorithm
  571. UAV Remote Sensing Image Dehazing Based on Multi-Dimensional Saliency Awareness Unequal Network
  572. UCONV-Conformer: High Reduction of Input Sequence Length for End-to-End Speech Recognition
  573. UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction
  574. UFO2: A Unified Pre-Training Framework for Online and Offline Speech Recognition
  575. UML: A Universal Monolingual Output Layer For Multilingual Asr
  576. UNTAG: Learning Generic Features for Unsupervised Type-Agnostic Deepfake Detection
  577. UNeXt: a Low-Dose CT denoising UNet model with the modified ConvNeXt block
  578. UPGLADE: Unplugged Plug-and-Play Audio Declipper Based on Consensus Equilibrium of DNN and Sparse Optimization
  579. URM4DMU: An User Representation Model for Darknet Markets Users
  580. UWB Localization-of-Things Via Soft Information: Network Experimentation in Indoor Environment
  581. UX-Net: Filter-and-Process-Based Improved U-Net for real-time time-domain audio Separation
  582. Ultimate Negative Sampling for Contrastive Learning
  583. Ultra Real-Time Portrait Matting via Parallel Semantic Guidance
  584. Ultrasound Image Quality Control Using Speech-Assisted Switchable CycleGAN
  585. Unbiased Unsupervised Stimulus Reconstruction for EEG-Based Auditory Attention Decoding
  586. Uncer2Natural: Uncertainty-Aware Unsupervised Image Denoising
  587. Uncertainty Estimation in Deep Speech Enhancement Using Complex Gaussian Mixture Models
  588. Uncertainty-Aware Few-Shot Class-Incremental Learning
  589. Understandable Relu Neural Network For Signal Classification
  590. Understanding Shared Speech-Text Representations
  591. Underwater Image Restoration with Light-Aware Progressive Network
  592. Unified Keyword Spotting and Audio Tagging on Mobile Devices with Transformers
  593. Unified Prompt Learning Makes Pre-Trained Language Models Better Few-Shot Learners
  594. Unifying Speech Enhancement and Separation with Gradient Modulation for End-to-End Noise-Robust Speech Separation
  595. Unique Bispectrum Inversion for Signals with Finite Spectral/Temporal Support
  596. Unitary Esprit for Coprime Arrays
  597. Universal Speaker Recognition Encoders for Different Speech Segments Duration
  598. Unlimited Sampling Radar: Life Below the Quantization Noise
  599. Unlimited Sampling in Phase Space
  600. Unlimited Sampling of FRI Signals Independent of Sampling Rate
  601. Unobtrusive Respiratory Monitoring System for Intensive Care
  602. Unrestricted Anchor Graph Based GCN for Incomplete Multi-View Clustering
  603. Unrolled Fourier Disparity Layer Optimization for Scene Reconstruction from Few-Shots Focal Stacks
  604. Unsupervised Action Segmentation of Untrimmed Egocentric Videos
  605. Unsupervised Anomaly Detection and Localization of Machine Audio: A Gan-Based Approach
  606. Unsupervised Deep Digital Staining for Microscopic Cell Images via Knowledge Distillation
  607. Unsupervised Domain Adaptation for Preference Learning Based Speech Emotion Recognition
  608. Unsupervised Domain Adaptation via Subspace Interpolating Deep Dictionary Learning: A Case Study in Machine Inspection
  609. Unsupervised Extractive Summarization With Heterogeneous Graph Embeddings for Chinese Documents
  610. Unsupervised Feature Selection with self-Weighted and ℓ2,0-Norm Constraint
  611. Unsupervised Fine-Tuning Data Selection for ASR Using Self-Supervised Speech Models
  612. Unsupervised Model-Based Speaker Adaptation of End-To-End Lattice-Free MMI Model for Speech Recognition
  613. Unsupervised Noise Adaptation Using Data Simulation
  614. Unsupervised Out-of-Distribution Detection Using Few in-Distribution Samples
  615. Unsupervised Pre-Training for Data-Efficient Text-to-Speech on Low Resource Languages
  616. Unsupervised Speaker Verification Using Pre-Trained Model and Label Correction
  617. Unsupervised Video Anomaly Detection For Stereotypical Behaviours in Autism
  618. Unsupervised Vocal Dereverberation with Diffusion-Based Generative Models
  619. Unsupervised Voice Type Discrimination Score Adaptation Using X-Vector Clusters
  620. Unsupervised Word Segmentation Using Temporal Gradient Pseudo-Labels
  621. Unsupervised word Segmentation Based on Word Influence
  622. Untargeted Backdoor Attack Against Object Detection
  623. Using Adapters to Overcome Catastrophic Forgetting in End-to-End Automatic Speech Recognition
  624. Using Auxiliary Tasks In Multimodal Fusion of Wav2vec 2.0 And Bert for Multimodal Emotion Recognition
  625. Using Emotion Embeddings to Transfer Knowledge between Emotions, Languages, and Annotation Formats
  626. Using Machine Learning to Understand the Relationships Between Audiometric Data, Speech Perception, Temporal Processing, And Cognition
  627. Using Modified Adult Speech as Data Augmentation for Child Speech Recognition
  628. Using Received Power in Microphone Arrays to Estimate Direction of Arrival
  629. Utility Polelocalization by Learning from Ambient Traces on Distributed Acoustic Sensing
  630. Utilization of Bessel Beams in Wideband Sub Terahertz Communication Systems to Mitigate Beamsplit Effects in the Near-field
  631. Utilizing Wav2Vec In Database-Independent Voice Disorder Detection
  632. VAN-ICP: GPU-Accelerated Approximate Nearest Neighbor Search for ICP Registration via Voxel Dilation
  633. VE-KWS: Visual Modality Enhanced End-to-End Keyword Spotting
  634. VF-Taco2: Towards Fast and Lightweight Synthesis for Autoregressive Models with Variation Autoencoder and Feature Distillation
  635. VLKP:Video Instance Segmentation with Visual-Linguistic Knowledge Prompts
  636. VPPT: Visual Pre-Trained Prompt Tuning Framework for Few-Shot Image Classification
  637. VQ-CL: Learning Disentangled Speech Representations with Contrastive Learning and Vector Quantization
  638. Vani: Very-Lightweight Accent-Controllable TTS for Native And Non-Native Speakers With Identity Preservation
  639. Vararray Meets T-Sot: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition
  640. Variable Attention Masking for Configurable Transformer Transducer Speech Recognition
  641. Variable Rate Allocation for Vector-Quantized Autoencoders
  642. Variational Bayesian Channel Estimation in Wideband Multi-Scale Multi-Lag Channels
  643. Variational Inference Aided Estimation of Time Varying Channels
  644. Variational Message Passing-Based Respiratory Motion Estimation and Detection Using Radar Signals
  645. VarietySound: Timbre-Controllable Video to Sound Generation Via Unsupervised Information Disentanglement
  646. Various Performance Bounds on the Estimation of Low-Rank Probability Mass Function Tensors from Partial Observations
  647. Vehicle View Synthesis by Generative Adversarial Network
  648. ViT-Cat: Parallel Vision Transformers With Cross Attention Fusion for Popularity Prediction in MEC Networks
  649. Video Captioning via Relation-Aware Graph Learning
  650. Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-to-Speech
  651. Vision Transformer with Progressive Tokenization for CT Metal Artifact Reduction
  652. Vision Transformer-Based Feature Extraction for Generalized Zero-Shot Learning
  653. Vision, Deduction and Alignment: An Empirical Study on Multi-Modal Knowledge Graph Alignment
  654. Vision2Touch: Imaging Estimation of Surface Tactile Physical Properties
  655. Visual Answer Localization with Cross-Modal Mutual Knowledge Transfer
  656. Visual Graph Reasoning Network
  657. Visual Information Matters for ASR Error Correction
  658. Visual Onoma-to-Wave: Environmental Sound Synthesis from Visual Onomatopoeias and Sound-Source Images
  659. Visual Prompting for Adversarial Robustness
  660. Visual-Aware Text-to-Speech*
  661. Vitasd: Robust Vision Transformer Baselines for Autism Spectrum Disorder Facial Diagnosis
  662. Voice Conversion Using Feature Specific Loss Function Based Self-Attentive Generative Adversarial Network
  663. Voice-Preserving Zero-Shot Multiple Accent Conversion
  664. Volume-Regularized Nonnegative Tucker Decomposition with Identifiability Guarantees
  665. Volumetric 3D Reconstruction with Window-Wise Global Feature Aggregation
  666. Volumetric Attribute Compression for 3D Point Clouds Using Feedforward Network with Geometric Attention
  667. W2KPE: Keyphrase Extraction with Word-Word Relation
  668. WAVELET2VEC: A Filter Bank Masked Autoencoder for EEG-Based Seizure Subtype Classification
  669. WHC: Weighted Hybrid Criterion for Filter Pruning on Convolutional Neural Networks
  670. WIFI-Based Robust Child Presence Detection for Smart Cars
  671. WITT: A Wireless Image Transmission Transformer for Semantic Communications
  672. WL-MSR: Watch and Listen for Multimodal Subtitle Recognition
  673. WUDA: Unsupervised Domain Adaptation Based on Weak Source Domain Labels
  674. Wassertein Gan Synthesis for Time Series with Complex Temporal Dynamics: Frugal Architectures and Arbitrary Sample-Size Generation
  675. Water Leak Detection and Localization Using Convolutional Autoencoders
  676. Wav2Seq: Pre-Training Speech-to-Text Encoder-Decoder Models Using Pseudo Languages
  677. Wav2vec-Based Detection and Severity Level Classification of Dysarthria From Speech
  678. Wave-U-Net Discriminator: Fast and Lightweight Discriminator for Generative Adversarial Network-Based Speech Synthesis
  679. Waveform Boundary Detection for Partially Spoofed Audio
  680. Waveform Design to Improve the Estimation of Target Parameters Using the Fourier Transform Method in a MIMO OFDM DFRC System
  681. Wavsyncswap: End-To-End Portrait-Customized Audio-Driven Talking Face Generation
  682. WeSinger 2: Fully Parallel Singing Voice Synthesis via Multi-Singer Conditional Adversarial Training
  683. Weakly- and Semi-Supervised Object Localization
  684. Weakly-Supervised Scene-Specific Crowd Counting Using Real-Synthetic Hybrid Data
  685. Weavspeech: Data Augmentation Strategy For Automatic Speech Recognition Via Semantic-Aware Weaving
  686. Weight Averaging: A Simple Yet Effective Method to Overcome Catastrophic Forgetting in Automatic Speech Recognition
  687. Weight-Based Mask For Domain Adaptation
  688. Weight-Sharing Supernet for Searching Specialized Acoustic Event Classification Networks Across Device Constraints
  689. Weighted Sampling for Masked Language Modeling
  690. Wekws: A Production First Small-Footprint End-to-End Keyword Spotting Toolkit
  691. Wespeaker: A Research and Production Oriented Speaker Embedding Learning Toolkit
  692. When is Mimo Massive in Radar?
  693. Whether Contribution of Features Differ Between Video-Mediated and In-Person Meetings in Important Utterance Estimation
  694. Which Country is This Picture From? New Data and Methods For Dnn-Based Country Recognition
  695. Wiener Filtering Without Covariance Matrix Inversion
  696. Windowed Fourier Analysis for Signal Processing on Graph Bundles
  697. Wireless Deep Speech Semantic Transmission
  698. Wireless Location Tracking via Complex-Domain Super MDS with Time Series Self-Localization Information
  699. Wireless Power Transfer Using Chirp Waveforms
  700. Wireless Sensing for Simultaneous Human Vocal Sound and Heart Sound Recognition
  701. Wordreg: Mitigating the Gap between Training and Inference with Worst-Case Drop Regularization
  702. X-SEPFORMER: End-To-End Speaker Extraction Network with Explicit Optimization on Speaker Confusion
  703. YOLOX-B: A Better Yolox Model for Real-Time Driver Behavior Detection
  704. Yolo-Based Lightweight Object Detection With Structure Simplification And Attention Enhancement
  705. Your Camera Improves Your Point Cloud Compression
  706. ZO-DARTS: Differentiable Architecture Search with Zeroth-Order Approximation
  707. Zephyr: Zero-Shot Punctuation Restoration
  708. Zero-Shot Anomalous Sound Detection in Domestic Environments Using Large-Scale Pretrained Audio Pattern Recognition Models
  709. Zero-Shot Domain Adaptation of Anomalous Samples for Semi-Supervised Anomaly Detection
  710. Zero-Shot Personalized Lip-To-Speech Synthesis with Face Image Based Voice Control
  711. Zero-Shot Sound Event Classification Using a Sound Attribute Vector with Global and Local Feature Learning
  712. Zero-Shot Speech Emotion Recognition Using Generative Learning with Reconstructed Prototypes
  713. Zone Plate Virtual Lenses for Memory-Constrained NLOS Imaging
  714. ifUNet++: Iterative Feedback UNet++ for Infrared Small Target Detection
  715. jaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus
  716. mmSense: Detecting Concealed Weapons with a Miniature Radar Sensor
  717. mmWave Wi-Fi Trajectory Estimation with Continuous-Time Neural Dynamic Learning
  718. ψ-Net: Point Structural Information Network for No-Reference Point Cloud Quality Assessment

Looking for submission deadlines instead? See the conference deadline calendar.