← All conferences

ICASSP 2022 Accepted Papers

The full list of 1,864 papers accepted at ICASSP 2022 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. 3D Texture Super Resolution via the Rendering Loss
  2. 3d Cross-Scale Feature Transformer Network for Brain Mr Image Super-Resolution
  3. 4D Convolutional Neural Networks for Multi-Spectral and Multi-Temporal Remote Sensing Data Classification
  4. A Bayesian Permutation Training Deep Representation Learning Method for Speech Enhancement with Variational Autoencoder
  5. A Benchmark of State-of-the-Art Sound Event Detection Systems Evaluated on Synthetic Soundscapes
  6. A Bert Based Joint Learning Model with Feature Gated Mechanism for Spoken Language Understanding
  7. A Bridge between Features and Evidence for Binary Attribute-Driven Perfect Privacy
  8. A Byzantine-Resilient Dual Subgradient Method for Vertical Federated Learning
  9. A CRLB Analysis of AoA Estimation Using Bluetooth 5
  10. A Channel Attention Based MLP-Mixer Network for Motor Imagery Decoding With EEG
  11. A Character-Level Span-Based Model for Mandarin Prosodic Structure Prediction
  12. A Closer Look at Autoencoders for Unsupervised Anomaly Detection
  13. A Clustering-based ML Scheme for Capacity Approaching Soft Level Sensing in 3D TLC NAND
  14. A Commonsense Knowledge Enhanced Network with Retrospective Loss for Emotion Recognition in Spoken Dialog
  15. A Communication Efficient Quasi-Newton Method for Large-Scale Distributed Multi-Agent Optimization
  16. A Comparison of Discrete and Soft Speech Units for Improved Voice Conversion
  17. A Complex Spectral Mapping with Inplace Convolution Recurrent Neural Networks For Acoustic Echo Cancellation
  18. A Configurable Multilingual Model is All You Need to Recognize All Languages
  19. A Convex Formulation for the Robust Estimation of Multivariate Exponential Power Models
  20. A DNN Based Post-Filter to Enhance the Quality of Coded Speech in MDCT Domain
  21. A Data-Driven Approach for Acoustic Parameter Similarity Estimation of Speech Recording
  22. A Data-Driven Cognitive Salience Model for Objective Perceptual Audio Quality Assessment
  23. A Data-Driven Quantization Design for Distributed Testing Against Independence with Communication Constraints
  24. A Deep Hierarchical Fusion Network for Fullband Acoustic Echo Cancellation
  25. A Differentiable Optimisation Framework for The Design of Individualised DNN-based Hearing-Aid Strategies
  26. A Dilated Residual Vision Transformer for Atrial Fibrillation Detection from Stacked Time-Frequency ECG Representations
  27. A Domain Transfer Based Data Augmentation Method for Automated Respiratory Classification
  28. A Dynamic Reweighting Strategy For Fair Federated Learning
  29. A Fast and Efficient Network for Single Image Shadow Detection
  30. A Few-Sample Strategy for Guitar Tablature Transcription Based on Inharmonicity Analysis and Playability Constraints
  31. A Frame Loss of Multiple Instance Learning for Weakly Supervised Sound Event Detection
  32. A Framework for Private Communication with Secret Block Structure
  33. A Gaussian Mixture Model for Dialogue Generation with Dynamic Parameter Sharing Strategy
  34. A General Framework For Incomplete Cross-Modal Retrieval With Missing Labels And Missing Modalities
  35. A Generalized Hierarchical Nonnegative Tensor Decomposition
  36. A Generalized Kernel Risk Sensitive Loss for Robust Two-Dimensional Singular Value Decomposition
  37. A Generic Method to Estimate Camera Extrinsic Parameters
  38. A Glance-and-Gaze Network for Respiratory Sound Classification
  39. A Global to Local Guiding Network for Missing Data Imputation
  40. A Graph Attention Interactive Refine Framework with Contextual Regularization for Jointing Intent Detection and Slot Filling
  41. A Hybrid Approach to Combine Wireless and Earcup Microphones for ANC Headphones with Error Separation Module
  42. A Hybrid Learning Framework for Deep Spiking Neural Networks with One-Spike Temporal Coding
  43. A Knowledge/Data Enhanced Method for Joint Event and Temporal Relation Extraction
  44. A Light Weight Model for Video Shot Occlusion Detection
  45. A Lightweight Instrument-Agnostic Model for Polyphonic Note Transcription and Multipitch Estimation
  46. A Lightweight Self-Supervised Training Framework for Monocular Depth Estimation
  47. A Likelihood Ratio Based Domain Adaptation Method for E2E Models
  48. A Low-Parametric Model for Bit-Rate Estimation of VVC Residual Coding
  49. A Maximal Correlation Approach to Imposing Fairness in Machine Learning
  50. A Melody-Unsupervision Model for Singing Voice Synthesis
  51. A Method For Estimating The Grouping Of Participants In Classroom Group Work Using Only Audio Information
  52. A Method for Detecting Coronary Artery Disease using Noisy Ultrashort Electrocardiogram Recordings
  53. A Method to Reveal Speaker Identity in Distributed ASR Training, and How to Counter IT
  54. A Minimally Supervised Approach for Medical Image Quality Assessment in Domain Shift Settings
  55. A Model for Assessor Bias in Automatic Pronunciation Assessment
  56. A Multi Domain Knowledge Enhanced Matching Network for Response Selection in Retrieval-Based Dialogue Systems
  57. A Multi-Resolution Low-Rank Tensor Decomposition
  58. A Multi-Task Learning Framework for Chinese Medical Procedure Entity Normalization
  59. A Multi-Task Learning Method for Weakly Supervised Sound Event Detection
  60. A Multiscale Gradient-Backpropagation Optimization Framework for Deformable Convolution Based Compressed Video Enhancement
  61. A Multitask Learning Framework for Speaker Change Detection with Content Information from Unsupervised Speech Decomposition
  62. A Mutual Learning Framework for Few-Shot Sound Event Detection
  63. A Neural Network-based Howling Detection Method for Real-Time Communication Applications
  64. A Neural Prosody Encoder for End-to-End Dialogue Act Classification
  65. A New Coprime-Array-based Configuration with Augmented Degrees of Freedom and Reduced Mutual Coupling
  66. A New Data Augmentation Method for Intent Classification Enhancement and its Application on Spoken Conversation Datasets
  67. A New Deep Learning Method for Multispectral Image Time Series Completion Using Hyperspectral Data
  68. A New Framework for Multiple Deep Correlation Filters Based Object Tracking
  69. A Noise-Robust Self-Supervised Pre-Training Model Based Speech Representation Learning for Automatic Speech Recognition
  70. A Non-Convex Proximal Approach for Centroid-Based Classification
  71. A Non-Hierarchical Attention Network with Modality Dropout for Textual Response Generation in Multimodal Dialogue Systems
  72. A Nonlinear Steerable Complex Wavelet Decomposition of Images
  73. A Note on Totally Symmetric Equi-Isoclinic Tight Fusion Frames
  74. A Novel 1D State Space for Efficient Music Rhythmic Analysis
  75. A Novel Angular Estimation Method in the Presence of Nonuniform Noise
  76. A Novel Convolutional Neural Network Based on Adaptive Multi-Scale Aggregation and Boundary-Aware for Lateral Ventricle Segmentation on MR images
  77. A Novel Lightweight Network for Fast Monocular Depth Estimation
  78. A Novel Micro-Expression Recognition Approach Using Attention-Based Magnification-Adaptive Networks
  79. A Novel Negative ℓ1 Penalty Approach for Multiuser One-Bit Massive MIMO Downlink with PSK Signaling
  80. A Novel Part Feature Integration and Fusion Method for Fine-Grained Vehicle Recognition
  81. A Novel Sequential Monte Carlo Framework for Predicting Ambiguous Emotion States
  82. A Novel Unsupervised Autoencoder-Based HFOs Detector in Intracranial EEG Signals
  83. A Performance Analysis for Multi-Ris-Assisted Full Duplex Wireless Communication System
  84. A Pre-Trained Audio-Visual Transformer for Emotion Recognition
  85. A Priori SNR Estimation for Speech Enhancement Based on PESQ-Induced Reinforcement Learning
  86. A Question-Oriented Propagation Network for News Reading Comprehension
  87. A Remedy For Distributional Shifts Through Expected Domain Translation
  88. A Robust Contrastive Alignment Method for Multi-Domain Text Classification
  89. A Robust Deep Audio Splicing Detection Method via Singularity Detection Feature
  90. A Robust Object Segmentation Network for UnderWater Scenes
  91. A Self-Supervised Pre-Training Framework for Vision-Based Seizure Classification
  92. A Semi-Handcrafted Keypoint Detector with Discriminative Feature Encoding
  93. A Set-Theoretic Approach to Mimo Detection
  94. A Simple Formula for the Moments of Unitarily Invariant Matrix Distributions
  95. A Simple Graph Neural Network via Layer Sniffer
  96. A Simple Hybrid Filter Pruning for Efficient Edge Inference
  97. A Slide-Save Based Framework for Multi-Source DOA Extraction with Closely Spaced Sources
  98. A Stimuli-Relevant Directed Dependency Index for Time Series
  99. A Study of Designing Compact Audio-Visual Wake Word Spotting System Based on Iterative Fine-Tuning in Neural Network Pruning
  100. A Study of The Robustness of Raw Waveform Based Speaker Embeddings Under Mismatched Conditions
  101. A Study on the Efficacy of Model Pre-Training In Developing Neural Text-to-Speech System
  102. A Style Transfer Mapping and Fine-Tuning Subject Transfer Framework Using Convolutional Neural Networks for Surface Electromyogram Pattern Recognition
  103. A Test for Conditional Correlation Between Random Vectors Based on Weighted U-Statistics
  104. A Time Domain Progressive Learning Approach with SNR Constriction for Single-Channel Speech Enhancement and Recognition
  105. A Time Encoding Approach to Training Spiking Neural Networks
  106. A Track-Wise Ensemble Event Independent Network for Polyphonic Sound Event Localization and Detection
  107. A Trainable Bounded Denoiser Using Double Tight Frame Network for Snapshot Compressive Imaging
  108. A Training Framework for Stereo-Aware Speech Enhancement Using Deep Neural Networks
  109. A Transfer Learning Approach for Pronunciation Scoring
  110. A Two-Stage Contrastive Learning Framework For Imbalanced Aerial Scene Recognition
  111. A Two-Stage U-Net for High-Fidelity Denoising of Historical Recordings
  112. A Two-Step Approach to Leverage Contextual Data: Speech Recognition in Air-Traffic Communications
  113. A Two-Step Backward Compatible Fullband Speech Enhancement System
  114. A Two-Stream Information Fusion Approach to Abnormal Event Detection in Video
  115. A Unified Two-Stage Model for Separating Superimposed Images
  116. A Universal Ordinal Regression for Assessing Phoneme-Level Pronunciation
  117. A Variational Bayesian Approach to Learning Latent Variables for Acoustic Knowledge Transfer
  118. A Wavelet-Based Dual-Stream Network for Underwater Image Enhancement
  119. A free lunch from ViT: adaptive attention multi-scale fusion Transformer for fine-grained visual recognition
  120. A-PixelHop: A Green, Robust and Explainable Fake-Image Detector
  121. AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks
  122. ACP: Adaptive Channel Pruning for Efficient Neural Networks
  123. ADA-VAD: Unpaired Adversarial Domain Adaptation for Noise-Robust Voice Activity Detection
  124. ADD 2022: the first Audio Deep Synthesis Detection Challenge
  125. ADIMA: Abuse Detection In Multilingual Audio
  126. ADMM-DAD Net: A Deep Unfolding Network for Analysis Compressed Sensing
  127. ADT: Anti-Deepfake Transformer
  128. AECMOS: A Speech Quality Assessment Metric for Echo Impairment
  129. AIMNet: Adaptive Image-Tag Merging Network For Automatic Medical Report Generation
  130. AISHELL-NER: Named Entity Recognition from Chinese Speech
  131. ALSNet: A Dilated 1-D CNN for Identifying ALS from Raw EMG Signal
  132. APPLADE: Adjustable Plug-and-Play Audio Declipper Combining DNN with Sparse Optimization
  133. ARM 4-BIT PQ: SIMD-Based Acceleration for Approximate Nearest Neighbor Search on ARM
  134. ASR Error Correction with Dual-Channel Self-Supervised Learning
  135. ASR-Aware End-to-End Neural Diarization
  136. ASSEM-VC: Realistic Voice Conversion by Assembling Modern Speech Synthesis Techniques
  137. Accelerated Intravascular Ultrasound Imaging using Deep Reinforcement Learning
  138. Accelerating ILL-Conditioned Robust Low-Rank Tensor Regression
  139. Access Control for Privacy-Preserving Gaussian Process Regression
  140. Accurate Inference of Unseen Combinations of Multiple Rootcauses with Classifier Ensemble
  141. Accurate Instance Segmentation Via Collaborative Learning
  142. Accurate Multiscale Selective Fusion of CT and Video Images for Real-Time Endoscopic Camera 3D Tracking in Robotic Surgery
  143. Accurate and Resource-Efficient Lipreading with Efficientnetv2 and Transformers
  144. Acoustic Application of Phase Reconstruction Algorithms in Optics
  145. Acoustic Comparison of Physical Vocal Tract Models with Hard and Soft Walls
  146. Acoustic Imaging Aboard The International Space Station (ISS): Challenges and Preliminary Results
  147. Acoustic-to-Articulatory Inversion Based on Speech Decomposition and Auxiliary Feature
  148. Ada-JSR: Sample Efficient Adaptive Joint Support Recovery From Extremely Compressed Measurement Vectors
  149. Ada-STNet: A Dynamic AdaBoost Spatio-Temporal Network for Traffic Flow Prediction
  150. AdaPID: An Adaptive PID Optimizer for Training Deep Neural Networks
  151. Adapting Speech Separation to Real-World Meetings using Mixture Invariant Training
  152. Adaptive Actor-Critic Bilateral Filter
  153. Adaptive Attention Graph Capsule Network
  154. Adaptive Diffusion with Compressed Communication
  155. Adaptive Discounting of Implicit Language Models in RNN-Transducers
  156. Adaptive Group Testing with Mismatched Models
  157. Adaptive Identification of Underwater Acoustic Channel with a Mix of Static and Time-Varying Parameters
  158. Adaptive Intra-Group Aggregation for Co-Saliency Detection
  159. Adaptive Matching Strategy for Multi-Target Multi-Camera Tracking
  160. Adaptive Node Participation for Straggler-Resilient Federated Learning
  161. Adaptive Pseudo Labeling for Source-Free Domain Adaptation in Medical Image Segmentation
  162. Adaptive Variational Nonlinear Chirp Mode Decomposition
  163. Adaptive Weighted Network With Edge Enhancement Module For Monocular Self-Supervised Depth Estimation
  164. Adaptive Wireless Power Allocation with Graph Neural Networks
  165. AdderIC: Towards Low Computation Cost Image Compression
  166. Adjacency Pairs-Aware Hierarchical Attention Networks for Dialogue Intent Classification
  167. Advancing Momentum Pseudo-Labeling with Conformer and Initialization Strategy
  168. AdverFacial: Privacy-Preserving Universal Adversarial Perturbation Against Facial Micro-Expression Leakages
  169. AdverSparse: An Adversarial Attack Framework for Deep Spatial-Temporal Graph Neural Networks
  170. Adversarial Audio Synthesis Using a Harmonic-Percussive Discriminator
  171. Adversarial Examples Detection Based on Error Level Analysis and Space Mapping
  172. Adversarial Examples for Image Cropping in Social Media
  173. Adversarial Input Ablation for Audio-Visual Learning
  174. Adversarial Learning Enhancement for 3D Human Pose and Shape Estimation
  175. Adversarial Learning in Transformer Based Neural Network in Radio Signal Classification
  176. Adversarial Linear Quadratic Regulator under Falsified Actions
  177. Adversarial Mask Transformer for Sequential Learning
  178. Adversarial Robustness by Design Through Analog Computing And Synthetic Gradients
  179. Adversarial Sample Detection for Speaker Verification by Neural Vocoders
  180. Adversary Distillation for One-Shot Attacks on 3D Target Tracking
  181. Advin: Automatically Discovering Novel Domains and Intents from User Text Utterances
  182. Aerial Base Station Placement Leveraging Radio Tomographic Maps
  183. Against Backdoor Attacks In Federated Learning With Differential Privacy
  184. Agcyclegan: Attention-Guided Cyclegan for Single Underwater Image Restoration
  185. Airborne Mimo Radar Transmit-Receive Design Under Spectral Constraint in Signal-Dependent Clutter
  186. Alarm Sound Detection Using Topological Signal Processing
  187. Alignment-Learning Based Single-Step Decoding for Accurate and Fast Non-Autoregressive Speech Recognition
  188. All-Neural Beamformer for Continuous Speech Separation
  189. Alleviating the Loss-Metric Mismatch in Supervised Single-Channel Speech Enhancement
  190. Ambiguity Modelling with Label Distribution Learning for Music Classification
  191. Amicable Examples for Informed Source Separation
  192. Amicable Examples for Informed Source Separation
  193. An Accelerated Rank-(L, L, 1, 1) Block Term Decomposition Of Multi-Subject Fmri Data Under Spatial Orthonormality Constraint
  194. An Adapter Based Pre-Training for Efficient and Scalable Self-Supervised Speech Representation Learning
  195. An Adaptive Orientational Beamforming Technique for Narrowband Interference Rejection
  196. An Anomaly Detection Method Based on Self-Supervised Learning with Soft Label Assignment for Defect Visual Inspection
  197. An Approach to Mispronunciation Detection and Diagnosis with Acoustic, Phonetic and Linguistic (APL) Embeddings
  198. An Asymptotically Optimal Approximation of the Conditional Mean Channel Estimator Based on Gaussian Mixture Models
  199. An Audio-Saliency Masking Transformer for Audio Emotion Classification in Movies
  200. An Effective Steganalysis for Robust Steganography with Repetitive JPEG Compression
  201. An Efficient DP-SGD Mechanism for Large Scale NLU Models
  202. An Efficient Framework for Detection and Recognition of Numerical Traffic Signs
  203. An Efficient Method For Generic Dsp Implementation Of Dilated Convolution
  204. An Efficient Method for Model Pruning Using Knowledge Distillation with Few Samples
  205. An Embarrassingly Simple Model for Dialogue Relation Extraction
  206. An End-to-End Chinese Text Normalization Model Based on Rule-Guided Flat-Lattice Transformer
  207. An End-to-End Deep Learning Framework For Multiple Audio Source Separation And Localization
  208. An End-to-End Deep Learning Speech Coding and Denoising Strategy for Cochlear Implants
  209. An Enhanced Deep Learning Approach for Tectonic Fault and Fracture Extraction in Very High Resolution Optical Images
  210. An Error Correction Scheme for Improved Air-Tissue Boundary in Real-Time MRI Video for Speech Production
  211. An Experimental Study on Transferring Data-Driven Image Compressive Sensing to Bioelectric Signals
  212. An Exploration of Hubert with Large Number of Cluster Units and Model Assessment Using Bayesian Information Criterion
  213. An Implicit Gradient-Type Method for Linearly Constrained Bilevel Problems
  214. An Information Maximization Based Blind Source Separation Approach for Dependent and Independent Sources
  215. An Investigation of Streaming Non-Autoregressive sequence-to-sequence Voice Conversion
  216. An Investigation of the Effectiveness of Phase for Audio Classification
  217. An Online Throughput Maximization Algorithm for Green Coordinated Multi-Point Systems
  218. An Overview of the FIRST ICASSP Special Session on Computer Audition for Healthcare
  219. Analyzing The Robustness of Unsupervised Speech Recognition
  220. Annihilation Filter Approach for Estimating Graph Dynamics from Diffusion Processes
  221. Anno-MI: A Dataset of Expert-Annotated Counselling Dialogues
  222. Anomalous Sound Detection Using Spectral-Temporal Information Fusion
  223. Applying Deep Learning to Known-Plaintext Attack on Chaotic Image Encryption Schemes
  224. Applying Differential Privacy to Tensor Completion
  225. Approaches Toward Physical and General Video Anomaly Detection
  226. Approximating The Likelihood Ratio in Linear-Gaussian State-Space Models for Change Detection
  227. Architecture for Variable Bitrate Neural Speech Codec with Configurable Computation Complexity
  228. Are GAN-based morphs threatening face recognition?
  229. Asd-Transformer: Efficient Active Speaker Detection Using Self And Multimodal Transformers
  230. Atomic Norm Based Localization and Orientation Estimation for Millimeter-Wave MIMO OFDM Systems
  231. Attachment Recognition in School-Age Children: A Multimodal Approach Based on Language and Paralanguage Analysis
  232. Attention Back-End for Automatic Speaker Verification with Multiple Enrollment Utterances
  233. Attention Guided Invariance Selection for Local Feature Descriptors
  234. Attention Probe: Vision Transformer Distillation in the Wild
  235. Attention-Based Dual-Stream Vision Transformer for Radar Gait Recognition
  236. Attention-Based Fusion for Bone-Conducted and Air-Conducted Speech Enhancement in the Complex Domain
  237. Attention-based Adversarial Partial Domain Adaptation
  238. Attentional Gated Res2net for Multivariate Time Series Classification
  239. Attentionpit: Soft Permutation Invariant Training for Audio Source Separation with Attention Mechanism
  240. Attentive Max Feature Map and Joint Training for Acoustic Scene Classification
  241. Attenuation Of Acoustic Early Reflections In Television Studios Using Pretrained Speech Synthesis Neural Network
  242. Attributable Watermarking of Speech Generative Models
  243. Attribute-Conditioned Face Swapping Network for Low-Resolution Images
  244. Audio Deepfake Detection System with Neural Stitching for ADD 2022
  245. Audio Peak Reduction Using a Synced allpass Filter
  246. Audio Signal Processing for Telepresence Based on Wearable Array in Noisy and Dynamic Scenes
  247. Audio-Text Retrieval in Context
  248. Audio-To-Symbolic Arrangement Via Cross-Modal Music Representation Learning
  249. Audio-Visual Multi-Channel Speech Separation, Dereverberation and Recognition
  250. Audio-Visual Object Classification for Human-Robot Collaboration
  251. Audio-Visual Scene-Aware Dialog and Reasoning Using Audio-Visual Transformers with Joint Student-Teacher Learning
  252. Audio-Visual Tracking of Multiple Speakers Via a PMBM Filter
  253. Audio-Visual Wake Word Spotting System for MISP Challenge 2021
  254. Audioclip: Extending Clip to Image, Text and Audio
  255. Auditory-Based Data Augmentation for end-to-end Automatic Speech Recognition
  256. Augmentation Strategy Optimization for Language Understanding
  257. Augmenting Molecular Deep Generative Models with Topological Data Analysis Representations
  258. Automated Audio Captioning Using Transfer Learning and Reconstruction Latent Space Similarity Regularization
  259. Automated Prosody Classification for Oral Reading Fluency with Quadratic Kappa Loss and Attentive X-Vectors
  260. Automatic Assessment of the Degree of Clinical Depression from Speech Using X-Vectors
  261. Automatic DJ Transitions with Differentiable Audio Effects and Generative Adversarial Networks
  262. Automatic Depression Detection: an Emotional Audio-Textual Corpus and A Gru/Bilstm-Based Model
  263. Automatic Depression Level Assessment from Speech By Long-Term Global Information Embedding
  264. Automatic Respiratory Sound Classification Via Multi-Branch Temporal Convolutional Network
  265. Autoregressive Variational Autoencoder with a Hidden Semi-Markov Model-Based Structured Attention for Speech Synthesis
  266. AuxFormer: Robust Approach to Audiovisual Emotion Recognition
  267. Auxiliary Loss of Transformer with Residual Connection for End-to-End Speaker Diarization
  268. Avqvc: One-Shot Voice Conversion By Vector Quantization With Applying Contrastive Learning
  269. Axonal Delay as a Short-Term Memory for Feed Forward Deep Spiking Neural Networks
  270. BNU: A Balance-Normalization-Uncertainty Model for Incremental Event Detection
  271. BSOLO: Boundary-Aware One-Stage Instance Segmentation SOLO
  272. Balanced Ranking and Sorting For Class Incremental Object Detection
  273. Balanced Stripe-Wise Pruning In The Filter
  274. Bayesian Continual Imputation and Prediction For Irregularly Sampled Time Series Data
  275. Being Greedy Does Not Hurt: Sampling Strategies for End-To-End Speech Recognition
  276. Best of Both Worlds: Multi-Task Audio-Visual Automatic Speech Recognition and Active Speaker Detection
  277. Bi-Directional Modality Fusion Network For Audio-Visual Event Localization
  278. Bi-Directional Normalization and Color Attention-Guided Generative Adversarial Network for Image Enhancement
  279. BiP-Net: Bidirectional Perspective Strategy Based Arbitrary-Shaped Text Detection Network
  280. Bilevel Learning of ℓ1 Regularizers with Closed-Form Gradients (BLORC)
  281. Bilingual End-to-End ASR with Byte-Level Subwords
  282. Binary Dense Predictors for Human Pose Estimation Based on Dynamic Thresholds and Filtering
  283. Blind Equalization of Moving Average Channels Over Galois Fields
  284. Blind Extraction of Equitable Partitions from Graph Signals
  285. Blind Modulo Analog-to-Digital Conversion of Vector Processes
  286. Blind Reverberation Time Estimation in Dynamic Acoustic Conditions
  287. Blind Separation of Linear-Quadratic Mixtures of Mutually Independent and Autocorrelated Sources
  288. Blind Source Separation via a Weak Exclusion Principle
  289. Blind Unmixing Using A Double Deep Image Prior
  290. Block-Activated Algorithms For Multicomponent Fully Nonsmooth Minimization
  291. Block-Coordinate Frank-Wolfe Algorithm And Convergence Analysis For Semi-Relaxed Optimal Transport Problem
  292. Block-Sparse Adversarial Attack to Fool Transformer-Based Text Classifiers
  293. Bloom-Net: Blockwise Optimization for Masking Networks Toward Scalable and Efficient Speech Enhancement
  294. Bona Fide Riesz Projections for Density Estimation
  295. Boost Ensemble Learning for Classification of CTG SIGNALS
  296. Boundary-Aware Bias Loss for Transformer-Based Aerial Image Segmentation Model
  297. Bounded Simplex-Structured Matrix Factorization
  298. Bounding Box Distribution Learning and Center Point Calibration for Robust Visual Tracking
  299. Building Robust Spoken Language Understanding by Cross Attention Between Phoneme Sequence and ASR Hypothesis
  300. Bundle ICP with Virtual Depth for Hand-Held 3d Scanner
  301. Bytecover2: Towards Dimensionality Reduction of Latent Embedding for Efficient Cover Song Identification
  302. Byzantine-Resilient Decentralized Collaborative Learning
  303. Byzantine-Resilient Decentralized Resource Allocation
  304. Byzantine-Robust Aggregation with Gradient Difference Compression and Stochastic Variance Reduction for Federated Learning
  305. Byzantine-Robust Federated Deep Deterministic Policy Gradient
  306. Byzantine-Robust and Communication-Efficient Distributed Non-Convex Learning Over Non-IID Data
  307. CDMA: Cross-Domain Distance Metric Adaptation for Speaker Verification
  308. CDX-NET: Cross-Domain Multi-Feature Fusion Modeling Via Deep Neural Networks for Multivariate Time Series Forecasting in AIOps
  309. CF-Net: Complementary Fusion Network for Rotation Invariant Point Cloud Completion
  310. CLIPCAM: A Simple Baseline For Zero-Shot Text-Guided Object And Action Localization
  311. CLseg: Contrastive Learning of Story Ending Generation
  312. CNN-Aided Factor Graphs with Estimated Mutual Information Features for Seizure Detection
  313. CNN-Transformer with Self-Attention Network for Sound Event Detection
  314. CPD Computation via Recursive Eigenspace Decompositions
  315. CPT: Cross-Modal Prefix-Tuning for Speech-To-Text Translation
  316. CRPN: Distinguish Novel Categories Via Class-Relevant Region Proposal Network for Few-Shot Object Detection
  317. CS-GResNet: A Simple and Highly Efficient Network for Facial Expression Recognition
  318. CS-REP: Making Speaker Verification Networks Embracing Re-Parameterization
  319. CSI Clustering with Variational Autoencoding
  320. Cache: Modeling Contribution-Aware Context Hierarchically for Long-Range Dialogue State Tracking
  321. Caching Networks: Capitalizing on Common Speech for ASR
  322. Call-Sign Recognition and Understanding for Noisy Air-Traffic Transcripts Using Surveillance Information
  323. Camera Calibration Through Camera Projection Loss
  324. Can Audio Captions Be Evaluated With Image Caption Metrics?
  325. Capitalization Normalization for Language Modeling with an Accurate and Efficient Hierarchical RNN Model
  326. Carina - A Corpus of Aligned German Read Speech Including Annotations
  327. Cascade Multi-Channel Noise Reduction and Acoustic Feedback Cancellation
  328. Cascading Bandit Under Differential Privacy
  329. Category-Adapted Sound Event Enhancement with Weakly Labeled Data
  330. Category-Adaptive Domain Adaptation for Semantic Segmentation
  331. Causal Alignment Based Fault Root Causes Localization for Wireless Network
  332. Causal Linear Topological Filters Over A 2-Simplex
  333. Cell-Free Massive Mimo: Exploiting The Wax Decomposition
  334. Channel Redundancy and Overlap in Convolutional Neural Networks with Channel-Wise NNK Graphs
  335. Channel-Wise AV-Fusion Attention for Multi-Channel Audio-Visual Speech Recognition
  336. Characterizing the Adversarial Vulnerability of Speech self-Supervised Learning
  337. Chinese Spelling Text Generation of Mathematical Formulas
  338. Chunkfusion: A Learning-Based RGB-D 3D Reconstruction Framework Via Chunk-Wise Integration
  339. Classical-To-Quantum Transfer Learning for Spoken Command Recognition Based on Quantum Neural Networks
  340. Climate and Weather: Inspecting Depression Detection via Emotion Recognition
  341. Cloning One's Voice Using Very Limited Data in the Wild
  342. Closed-Form Single Source Direction-of-Arrival Estimator Using First-Order Relative Harmonic Coefficients
  343. Closing the Sim-to-Real Gap in Guided Wave Damage Detection with Adversarial Training of Variational Auto-Encoders
  344. Clustering Complex Subspaces in Large Dimensions
  345. Clustering and Separating Similarities for Deep Unsupervised Hashing
  346. Cmri2spec: Cine MRI Sequence to Spectrogram Synthesis via A Pairwise Heterogeneous Translator
  347. Co-Attention-Guided Bilinear Model for Echo-Based Depth Estimation
  348. Coarray Manifold Separation In The Spherical Harmonics Domain For Enhanced Source Localization
  349. Coarse-To-Fine Unsupervised Change Detection for Remote Sensing Images Via Object-Based MRF and Inception UNET
  350. Cognitive Coding Of Speech
  351. Collaborative Object Detectors Adaptive to Bandwidth and Computation
  352. Combating False Sense of Security: Breaking the Defense of Adversarial Training Via Non-Gradient Adversarial Attack
  353. Combining Multiple Style Transfer Networks and Transfer Learning For LGE-CMR Segmentation
  354. Combining Unsupervised and Text Augmented Semi-Supervised Learning For Low Resourced Autoregressive Speech Recognition
  355. Communication-Efficient Distributed MAX-VAR Generalized CCA via Error Feedback-Assisted Quantization
  356. Communication-Efficient Online Federated Learning Framework for Nonlinear Regression
  357. Comparison of Boundary Artifact Removal Methods in Coding of Generalized Cubemap Projection Using VVC
  358. Competitive Multi-Agent Reinforcement Learning with Self-Supervised Representation
  359. Complex IRM-Aware Training for Voice Activity Detection Using Attention Model
  360. Complex-Valued Spatial Autoencoders for Multichannel Speech Enhancement
  361. Composing Graphical Models with Generative Adversarial Networks for EEG Signal Modeling
  362. Compressed Data Sharing Based On Information Bottleneck Model
  363. Compressing Transformer-Based ASR Model by Task-Driven Loss and Attention-Based Multi-Level Feature Distillation
  364. Compression-Aware Projection with Greedy Dimension Reduction for Convolutional Neural Network Activations
  365. Compressive Phase Retrieval Based On Sparse Latent Generative Priors
  366. Compressive Scanning Transmission Electron Microscopy
  367. Computationally Efficient Fixed-Filter ANC for Speech Based on Long-Term Prediction for Headphone Applications
  368. Conditional Diffusion Probabilistic Model for Speech Enhancement
  369. Conditionally Factorized Variational Bayes with Importance Sampling
  370. Coneface: Approximate Pairwise Loss for Face Recognition
  371. Confidence Estimation for Speech Emotion Recognition Based on the Relationship Between Emotion Categories and Primitives
  372. Confidence-Aware Multi-Teacher Knowledge Distillation
  373. Conformer-Based Hybrid ASR System For Switchboard Dataset
  374. Conformer-Based Self-Supervised Learning For Non-Speech Audio Tasks
  375. Conformer-Based Speech Recognition with Linear Nyström Attention and Rotary Position Embedding
  376. Conjugate Augmented Spatial-Temporal Near-Field Sources Localization with Cross Array
  377. Connecting Targets via Latent Topics And Contrastive Learning: A Unified Framework For Robust Zero-Shot and Few-Shot Stance Detection
  378. Considering User Agreement in Learning to Predict the Aesthetic Quality
  379. Consistent Training and Decoding for End-to-End Speech Recognition Using Lattice-Free MMI
  380. Constant Q Cepstral coefficients for classification of normal vs. Pathological infant cry
  381. Container Localisation and Mass Estimation with an RGB-D Camera
  382. Content Preserving Scale Space Network for Fast Image Restoration from Noisy-Blurry Pairs
  383. Context Modeling with Evidence Filter for Multiple Choice Question Answering
  384. Context-Adaptive Document-Level Neural Machine Translation
  385. Context-Aware Graph-Based Self-Supervised Learning of Whole Slide Images
  386. Context-Aware Mask Prediction Network for End-to-End Text-Based Speech Editing
  387. Contextual Adapters for Personalized Speech Recognition in Neural Transducers
  388. Continual Learning Using Lattice-Free MMI for Speech Recognition
  389. Continual Self-Training With Bootstrapped Remixing For Speech Enhancement
  390. Continuous Speech Separation with Recurrent Selective Attention Network
  391. Continuous Streaming Multi-Talker ASR with Dual-Path Transducers
  392. Contrastive Heartbeats: Contrastive Learning for Self-Supervised ECG Representation and Phenotyping
  393. Contrastive Knowledge Graph Attention Network for Request-Based Recipe Recommendation
  394. Contrastive Prediction Strategies for Unsupervised Segmentation and Categorization of Phonemes and Words
  395. Contrastive Predictive Coding for Anomaly Detection of Fetal Health from the Cardiotocogram
  396. Contrastive Sensor Transformer for Predictive Maintenance of Industrial Assets
  397. Contrastive Siamese Network for Semi-Supervised Speech Recognition
  398. Contrastive Translation Learning For Medical Image Segmentation
  399. Contrastive-mixup Learning for Improved Speaker Verification
  400. Controllable Speech Representation Learning Via Voice Conversion and AIC Loss
  401. Controlled Sensing and Anomaly Detection Via Soft Actor-Critic Reinforcement Learning
  402. Controlling Smart Propagation Environments: Long-Term Versus Short-Term Phase Shift Optimization
  403. Controlling The Fréchet Variance Improves Batch Normalization on the Symmetric Positive Definite Manifold
  404. Conversational Speech Recognition by Learning Conversation-Level Characteristics
  405. Convex Clustering for Autocorrelated Time Series
  406. Convmixer: Feature Interactive Convolution with Curriculum Learning for Small Footprint and Noisy Far-Field Keyword Spotting
  407. Convoluational Transformer With Adaptive Position Embedding For Covid-19 Detection From Cough Sounds
  408. Convolutional Beamspace Using IIR Filters
  409. Convolutional Filtering in Simplicial Complexes
  410. Convolutional ISTA Network with Temporal Consistency Constraints for Video Reconstruction from Event Cameras
  411. Convolutional Weighted Minimum Mean Square Error Filter for Joint Source Separation and Dereverberation
  412. Coughtrigger: Earbuds IMU Based Cough Detection Activator Using An Energy-Efficient Sensitivity-Prioritized Time Series Classifier
  413. Counting the Number of Different Scaling Exponents in Multivariate Scale-Free Dynamics: Clustering by Bootstrap in the Wavelet Domain
  414. Coupled Feature Learning Via Structured Convolutional Sparse Coding for Multimodal Image Fusion
  415. Cramer-Rao Bound Analysis of Distributed DOA Estimation Exploiting Mixed-Precision Covariance Matrix
  416. Cramer-Rao Bound for the Time-Varying Poisson
  417. Cramér-Rao Bound and Antenna Selection Optimization for Dual Radar-Communication Design
  418. Cross-Channel Attention-Based Target Speaker Voice Activity Detection: Experimental Results for the M2met Challenge
  419. Cross-Domain Few-Shot Learning for Rare-Disease Skin Lesion Segmentation
  420. Cross-Domain Speech Enhancement with a Neural Cascade Architecture
  421. Cross-Layer Aggregation with Transformers for Multi-Label Image Classification
  422. Cross-Modal Knowledge Distillation For Vision-To-Sensor Action Recognition
  423. Cross-Modal Knowledge Distillation in Multi-Modal Fake News Detection
  424. Cross-Speaker Style Transfer for Text-to-Speech Using Data Augmentation
  425. Cross-Target Stance Detection Via Refined Meta-Learning
  426. Csenet: Complex Squeeze-and-Excitation Network for Speech Depression Level Prediction
  427. Curriculum Optimization for Low-Resource Speech Recognition
  428. Custom Attribution Loss for Improving Generalization and Interpretability of Deepfake Detection
  429. Customer Satisfaction Estimation Using Unsupervised Representation Learning with Multi-Format Prediction Loss
  430. Customizable End-To-End Optimization Of Online Neural Network-Supported Dereverberation For Hearing Devices
  431. Cut And Continuous Paste Towards Real-Time Deep Fall Detection
  432. Cyber-Threat Propagation over Network-Slicing Architectures
  433. DAM-GAN : Image Inpainting Using Dynamic Attention Map Based on Fake Texture Detection
  434. DCNGAN: A Deformable Convolution-Based GAN with QP Adaptation for Perceptual Quality Enhancement of Compressed Video
  435. DCSN: Deformable Convolutional Semantic Segmentation Neural Network for Non-Rigid Scenes
  436. DGC-Vector: A New Speaker Embedding for Zero-Shot Voice Conversion
  437. DHWP: Learning High-Quality Short Hash Codes Via Weight Pruning
  438. DMANET: Deep Learning-Based Differential Microphone Arrays for Multi-Channel Speech Separation
  439. DNN Based Multiframe Single-Channel Noise Reduction Filters
  440. DOA M-Estimation Using Sparse Bayesian Learning
  441. DOMAINDESC: Learning Local Descriptors With Domain Adaptation
  442. DP-DWA: Dual-Path Dynamic Weight Attention Network With Streaming Dfsmn-San For Automatic Speech Recognition
  443. DPCCN: Densely-Connected Pyramid Complex Convolutional Network for Robust Speech Separation and Extraction
  444. DPT-FSNet: Dual-Path Transformer Based Full-Band and Sub-Band Fusion Network for Speech Enhancement
  445. DRC-NET: Densely Connected Recurrent Convolutional Neural Network for Speech Dereverberation
  446. DRVC: A Framework of Any-to-Any Voice Conversion with Self-Supervised Learning
  447. Data Agnostic Filter Gating For Efficient Deep Networks
  448. Data Augmentation for Long-Tailed and Imbalanced Polyphone Disambiguation in Mandarin
  449. Data Efficient Support Vector Machine Training Using the Minimum Description Length Principle
  450. Data Incubation - Synthesizing Missing Data for Handwriting Recognition
  451. Data Shapley Value for Handling Noisy Labels: An Application in Screening Covid-19 Pneumonia from Chest CT Scans
  452. Data-Driven Algorithms for Gaussian Measurement Matrix Design in Compressive Sensing
  453. Data-Driven Approach for the Floquet Propagator Inverse Problem Solution
  454. Data-Driven Optimization for Zero-Delay Lossy Source Coding with Side Information
  455. Data-Driven Spatially Dependent PDE Identification
  456. Decentralized Bilevel Optimization for Personalized Client Learning
  457. Decentralized Learning in the Presence of Low-Rank Noise
  458. Deep Actor-Critic for Continuous 3D Motion Control in Mobile Relay Beamforming Networks
  459. Deep Adaptation Control for Acoustic Echo Cancellation
  460. Deep Adaptive Aec: Hybrid of Deep Learning and Adaptive Acoustic Echo Cancellation
  461. Deep Augmented Music Algorithm for Data-Driven Doa Estimation
  462. Deep Deterministic Independent Component Analysis for Hyperspectral Unmixing
  463. Deep Hashing with Hash Center Update for Efficient Image Retrieval
  464. Deep Impulse Responses: Estimating and Parameterizing Filters with Deep Networks
  465. Deep Initialization for Guaranteed Unimodular Quadratic Programming
  466. Deep Iterative Phase Retrieval for Ptychography
  467. Deep Joint Source-Channel Coding for Wireless Image Transmission with Adaptive Rate Control
  468. Deep Kernel Learning Networks with Multiple Learning Paths
  469. Deep Learning Based Off-Angle Iris Recognition
  470. Deep Learning Based Passive Beamforming for IRS-Assisted Monostatic Backscatter Systems
  471. Deep Learning for Location Based Beamforming with Nlos Channels
  472. Deep Learning for Prominence Detection In Children's Read Speech
  473. Deep Learning on the Sphere for Multi-model Ensembling of Significant Wave Height
  474. Deep Markov Clustering for Panoptic Segmentation
  475. Deep Neural Network (DNN) Audio Coder Using A Perceptually Improved Training Method
  476. Deep Object Detection with Example Attribute Based Prediction Modulation
  477. Deep Performer: Score-to-Audio Music Performance Synthesis
  478. Deep Piecewise Hashing for Efficient Hamming Space Retrieval
  479. Deep Proximal Unfolding For Image Recovery from Under-Sampled Channel Data in Intravascular Ultrasound
  480. Deep Rank Cross-Modal Hashing with Semantic Consistent for Image-Text Retrieval
  481. Deep Residual Echo Suppression and Noise Reduction: A Multi-Input FCRN Approach in a Hybrid Speech Enhancement System
  482. Deep Scale-Aware Image Smoothing
  483. Deep Sequential Beamformer Learning for Multipath Channels in Mmwave Communication Systems
  484. Deep Spatio-Temporal Wind Power Forecasting
  485. Deep Temporal Interpolation of Radar-Based Precipitation
  486. Deep Video Inpainting Guided by Audio-Visual Self-Supervision
  487. Deep Video Inpainting Localization Using Spatial and Temporal Traces
  488. Deep-Learning-Assisted Configuration of Reconfigurable Intelligent Surfaces in Dynamic Rich-Scattering Environments
  489. Deep-MLE: Fusion between a Neural Network and MLE for A Single Snapshot DOA Estimation
  490. DeepGBASS: Deep Guided Boundary-Aware Semantic Segmentation
  491. DeepHull: Fast Convex Hull Approximation in High Dimensions
  492. Deepchorus: A Hybrid Model of Multi-Scale Convolution And Self-Attention for Chorus Detection
  493. Deepfake Speech Detection Through Emotion Recognition: A Semantic Approach
  494. Deepfilternet: A Low Complexity Speech Enhancement Framework for Full-Band Audio Based On Deep Filtering
  495. Defending Against Universal Attack Via Curvature-Aware Category Adversarial Training
  496. Deformable Convolution Dense Network for Compressed Video Quality Enhancement
  497. Deformable VisTR: Spatio Temporal Deformable Attention for Video Instance Segmentation
  498. Delay-Oriented Distributed Scheduling Using Graph Neural Networks
  499. Deliberation of Streaming RNN-Transducer by Non-Autoregressive Decoding
  500. Delta Distancing: A Lifting Approach to Localizing Items from User Comparisons
  501. Dementia Detection by Fusing Speech and Eye-Tracking Representation
  502. Demon: Improved Neural Network Training With Momentum Decay
  503. Denoising-Guided Deep Reinforcement Learning For Social Recommendation
  504. Denoising-Oriented Deep Hierarchical Reinforcement Learning for Next-Basket Recommendation⋆
  505. Depth Pruning with Auxiliary Networks for Tinyml
  506. Depth Removal Distillation for RGB-D Semantic Segmentation
  507. Depth-Based Ensemble Learning Network For Face Anti-Spoofing
  508. Deriving Explainable Discriminative Attributes Using Confusion About Counterfactual Class
  509. Design of Real-Time System Based on Machine Learning for Snoring and OSA Detection
  510. Designing a QAM Signal Detector for Massive Mimo Systems via PS-ADMM Approach
  511. Detail Generation and Fusion Networks for Image Inpainting
  512. Detecting Anomaly in Chemical Sensors via Regularized Contrastive Learning
  513. Detecting Backdoor Attacks against Point Cloud Classifiers
  514. Detection of COPD Exacerbation from Speech: Comparison of Acoustic Features and Deep Learning Based Speech Breathing Models
  515. Detection of Covid-19 from Joint Time and Frequency Analysis of Speech, Breathing and Cough Audio
  516. Determining Joint Periodicities in Multi-Time Data with Sampling Uncertainties
  517. Determining the best Acoustic Features for Smoker Identification
  518. Deterministic Transform Based Weight Matrices for Neural Networks
  519. Dictionary Learning with Uniform Sparse Representations for Anomaly Detection
  520. Differentiable Digital Signal Processing Mixture Model for Synthesis Parameter Extraction from Mixture of Harmonic Sounds
  521. Differentiable Programming A La Moreau
  522. Differentiable Wavetable Synthesis
  523. Differentiate-and-Fire Time-Encoding of Finite-Rate-of-Innovation Signals
  524. Difficulty-Aware Neural Band-to-Piano Score Arrangement based on Note- and Statistic-Level Criteria
  525. Dilated Convolutional Neural Network-Based Deep Reference Picture Generation for Video Compression
  526. Direct Design of Biquad Filter Cascades with Deep Learning by Sampling Random Polynomials
  527. Direct Localization: An Ising Model Approach
  528. Direct Noisy Speech Modeling for Noisy-To-Noisy Voice Conversion
  529. Discourse-Level Prosody Modeling with a Variational Autoencoder for Non-Autoregressive Expressive Speech Synthesis
  530. Discrete Multi-Kernel K-Means with Diverse and Optimal Kernel Learning
  531. Disentangled Feature-Guided Multi-Exposure High Dynamic Range Imaging
  532. Disentangled Speaker Embedding for Robust Speaker Verification
  533. Disentangling Content and Fine-Grained Prosody Information Via Hybrid ASR Bottleneck Features for Voice Conversion
  534. Dispeech: A Synthetic Toy Dataset for Speech Disentangling
  535. Distilhubert: Speech Representation Learning by Layer-Wise Distillation of Hidden-Unit Bert
  536. Distributed Audio-Visual Parsing Based On Multimodal Transformer and Deep Joint Source Channel Coding
  537. Distributed Graph Learning With Smooth Data Priors
  538. Distributed Hybrid Beamforming for Mmwave Cell-Free Massive MIMO
  539. Distributed Image Transmission Using Deep Joint Source-Channel Coding
  540. Distributed Label Dequantized Gaussian Process Latent Variable Model for Multi-View Data Integration
  541. Distributed Link Sparsification for Scalable Scheduling Using Graph Neural Networks
  542. Distributed Particle Filters for State Tracking on the Stiefel Manifold Using Tangent Space Statistics
  543. Distribution Augmentation for Low-Resource Expressive Text-To-Speech
  544. Distribution Learning for Age Estimation from Speech
  545. Divergence-Guided Feature Alignment for Cross-Domain Object Detection
  546. Diverse Audio Captioning Via Adversarial Training
  547. Diversity-Controllable and Accurate Audio Captioning Based on Neural Condition
  548. Dnsmos P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors
  549. Do You Live a Healthy Life? Analyzing Lifestyle by Visual Life Logging
  550. Doa Estimation Via Coarray Tensor Completion with Missing Slices
  551. Document-Level Event Extraction via Human-Like Reading Process
  552. Domain Adaptation for Speaker Recognition in Singing and Spoken Voice
  553. Domain Adaptation via Mutual Information Maximization for Handwriting Recognition
  554. Domain Decomposition Algorithms for Real-Time Homogeneous Diffusion Inpainting in 4K
  555. Domain Generalized Few-Shot Image Classification via Meta Regularization Network
  556. Domain Robust Deep Embedding Learning for Speaker Recognition
  557. Domain-Agnostic Meta-Learning for Cross-Domain Few-Shot Classification
  558. Domain-Invariant Feature Learning for Cross Corpus Speech Emotion Recognition
  559. Domain-Invariant Representation Learning from EEG with Private Encoders
  560. Don't Separate, Learn To Remix: End-To-End Neural Remixing With Joint Optimization
  561. Don't Speak Too Fast: The Impact of Data Bias on Self-Supervised Speech Models
  562. Double Closed-Loop Network for Image Deblurring
  563. Double Noise Mean Teacher Self-Ensembling Model for Semi-Supervised Tumor Segmentation
  564. Double-RIS Versus Single-RIS Aided Systems: Tensor-Based Mimo Channel Estimation and Design Perspectives
  565. Downstream Augmentation Generation For Contrastive Learning
  566. Dual Active Noise Control with Common Sensors
  567. Dual Attention Pooling Network for Recording Device Classification Using Neutral and Whispered Speech
  568. Dual Graph Cross-Domain Few-Shot Learning for Hyperspectral Image Classification
  569. Dual Path Graph Convolutional Networks
  570. Dual-Attention Network for Few-Shot Segmentation
  571. Dual-Branch Attention-In-Attention Transformer for Single-Channel Speech Enhancement
  572. Dual-Domain Low-Rank Fusion Deep Metric Learning for Off-the-Person ECG Biometrics
  573. Duration Modeling of Neural TTS for Automatic Dubbing
  574. DynSNN: A Dynamic Approach to Reduce Redundancy in Spiking Neural Networks
  575. Dynamic Binary Neural Network by Learning Channel-Wise Thresholds
  576. Dynamic Multi-Scale Loss Balance for Object Detection
  577. Dynamic Point Cloud Interpolation
  578. Dynamic Portfolio Cuts: A Spectral Approach to Graph-Theoretic Diversification
  579. Dynamic Resource Optimization for Adaptive Federated Learning Empowered by Reconfigurable Intelligent Surfaces
  580. Dynamic Sliding Window for Realtime Denoising Networks
  581. Dynamic Texture Recognition Using PDV Hashing and Dictionary Learning on Multi-Scale Volume Local Binary Pattern
  582. Dynamically Pruning Segformer for Efficient Semantic Segmentation
  583. Dynimp: Dynamic Imputation for Wearable Sensing Data through Sensory and Temporal Relatedness
  584. Dysfluency Classification in Stuttered Speech Using Deep Learning for Real-Time Applications
  585. EAD-Conformer: a Conformer-Based Encoder-Attention-Decoder-Network for Multi-Task Audio Source Separation
  586. EMGSE: Acoustic/EMG Fusion for Multimodal Speech Enhancement
  587. EMOQ-TTS: Emotion Intensity Quantization for Fine-Grained Controllable Emotional Text-to-Speech
  588. ER-PIQA: A Task-Guided Pedestrian Image Quality Assessment Via Embedding Reconstruction
  589. ESPnet-SLU: Advancing Spoken Language Understanding Through ESPnet
  590. Echo-Aware Adaptation of Sound Event Localization and Detection in Unknown Environments
  591. Eco-Fedsplit: Federated Learning with Error-Compensated Compression
  592. Economics of Semantic Communication System in Wireless Powered Internet of Things
  593. Edge Sampling of Graphs Based on Edge Smoothness
  594. Effect of Noise Suppression Losses on Speech Distortion and ASR Performance
  595. Effective and Inconspicuous Over-the-Air Adversarial Examples with Adaptive Filtering
  596. Efficient Adapter Transfer of Self-Supervised Speech Models for Automatic Speech Recognition
  597. Efficient Identity-Based Chameleon Hash for Mobile Devices
  598. Efficient Monaural Speech Separation with Multiscale Time-Delay Sampling
  599. Efficient Sequence Training of Attention Models Using Approximative Recombination
  600. Efficient Two-Stage Beam Training and Channel Estimation for Ris-Aided Mmwave Systems Via Fast Alternating Least Squares
  601. Efficient Universal Shuffle Attack for Visual Object Tracking
  602. Efficient and Stable Information Directed Exploration for Continuous Reinforcement Learning
  603. Efficiently and Globally Solving Joint Beamforming and Compression Problem in the Cooperative Cellular Network Via Lagrangian Duality
  604. Embedding Signals on Graphs with Unbalanced Diffusion Earth Mover's Distance
  605. Embedding and Beamforming: All-Neural Causal Beamformer for Multichannel Speech Enhancement
  606. Emotionflow: Capture the Dialogue Level Emotion Transitions
  607. Enabling On-Device Training of Speech Recognition Models With Federated Dropout
  608. Encrypted Image Visual Security Index via Non-Local Recognizable Degree Evaluation
  609. Encryption Resistant Deep Neural Network Watermarking
  610. End-To-End Alexa Device Arbitration
  611. End-To-End Deep Learning-Based Adaptation Control for Frequency-Domain Adaptive System Identification
  612. End-To-End Multi-Modal Speech Recognition with Air and Bone Conducted Speech
  613. End-To-End Music Remastering System Using Self-Supervised And Adversarial Training
  614. End-To-End Neural Coreference Resolution Revisited: A Simple Yet Effective Baseline
  615. End-To-End Speech Recognition with Joint Dereverberation of Sub-Band Autoregressive Envelopes
  616. End-to-End ASR-Enhanced Neural Network for Alzheimer's Disease Diagnosis
  617. End-to-End Complex-Valued Multidilated Convolutional Neural Network for Joint Acoustic Echo Cancellation and Noise Suppression
  618. End-to-End Keyword Spotting Using Neural Architecture Search and Quantization
  619. End-to-End Low Resource Keyword Spotting Through Character Recognition and Beam-Search Re-Scoring
  620. End-to-End Network Based on Transformer for Automatic Detection of Covid-19
  621. End-to-End Neural Speech Coding for Real-Time Communications
  622. End-to-End Speech Recognition from Federated Acoustic Models
  623. End-to-End Speech Summarization Using Restricted Self-Attention
  624. Endpoint Detection for Streaming End-to-End Multi-Talker ASR
  625. Energy Alignment for Bias Rectification in Class Incremental Learning
  626. Enhance Rnnlms with Hierarchical Multi-Task Learning for ASR
  627. Enhancing Affective Representations Of Music-Induced Eeg Through Multimodal Supervision And Latent Domain Adaptation
  628. Enhancing Class Understanding Via Prompt-Tuning For Zero-Shot Text Classification
  629. Enhancing Contextual Encoding With Stage-Confusion and Stage-Transition Estimation for EEG-Based Sleep Staging
  630. Enhancing Contrastive Learning with Temporal Cognizance for Audio-Visual Representation Generation
  631. Enhancing Privacy Through Domain Adaptive Noise Injection For Speech Emotion Recognition
  632. Enhancing Prototypical Few-Shot Learning By Leveraging The Local-Level Strategy
  633. Enhancing Speaking Styles in Conversational Text-to-Speech Synthesis with Graph-Based Multi-Modal Context Modeling
  634. Enhancing Utility In The Watchdog Privacy Mechanism
  635. Enhancing and Dissecting Crowd Counting by Synthetic Data
  636. Enrich Features for Few-Shot Point Cloud Classification
  637. Entrainment Analysis for Assessment of Autistic Speech Prosody Using Bottleneck Features of Deep Neural Network
  638. Environmental Sound Extraction Using Onomatopoeic Words
  639. Epileptic Spike Detection by Recurrent Neural Networks with Self-Attention Mechanism
  640. Equal Loss: A Simple Loss Function for Noise Robust Learning
  641. Estimating the Confidence of Speech Spoofing Countermeasure
  642. Estimation Of Channels In Systems With Intelligent Reflecting Surfaces
  643. Estimation of the Admittance Matrix in Power Systems Under Laplacian and Physical Constraints
  644. Evaluation of Orthogonal Chirp Division Multiplexing for Automotive Integrated Sensing and Communications
  645. Evaluation of Video Coding for Machines without Ground Truth
  646. Event-Based Multimodal Spiking Neural Network with Attention Mechanism
  647. Evolutionary Neural Architecture Design of Liquid State Machine for Image Classification
  648. Exact Partitioning of High-Order Planted Models with A Tensor Nuclear Norm Constraint
  649. Exact Sparse Super-Resolution Via Model Aggregation
  650. Expectation Consistent Plug-and-Play for MRI
  651. Experimental Investigation on STFT Phase Representations for Deep Learning-Based Dysarthric Speech Detection
  652. Experts Versus All-Rounders: Target Language Extraction for Multiple Target Languages
  653. Explainable Artificial Intelligence for Authorship Attribution on Social Media
  654. Explainable Fact-Checking Through Question Answering
  655. Explaining Deep Learning Models for Spoofing and Deepfake Detection with Shapley Additive Explanations
  656. Explicitly Modeling Importance and Coherence for Timeline Summarization
  657. Exploiting Annotators' Typed Description of Emotion Perception to Maximize Utilization of Ratings for Speech Emotion Recognition
  658. Exploiting Caption Diversity for Unsupervised Video Summarization
  659. Exploiting Cross Domain Acoustic-to-Articulatory Inverted Features for Disordered Speech Recognition
  660. Exploiting Hybrid Models of Tensor-Train Networks For Spoken Command Recognition
  661. Exploiting Language Model For Efficient Linguistic Steganalysis
  662. Explore Relative and Context Information with Transformer for Joint Acoustic Echo Cancellation and Speech Enhancement
  663. Exploring Auditory Acoustic Features for The Diagnosis of Covid-19
  664. Exploring Category Consistency for Weakly Supervised Semantic Segmentation
  665. Exploring Complementarity of Global and Local Spatiotemporal Information for Fake Face Video Detection
  666. Exploring Deeper Graph Convolutions for Semi-Supervised Node Classification
  667. Exploring Dementia Detection from Speech: Cross Corpus Analysis
  668. Exploring Dual Stream Global Information For Image Captioning
  669. Exploring Effective Data Utilization for Low-Resource Speech Recognition
  670. Exploring Heterogeneous Characteristics of Layers in ASR Models for More Efficient Training
  671. Exploring Machine Speech Chain For Domain Adaptation
  672. Exploring Non-Autoregressive End-to-End Neural Modeling for English Mispronunciation Detection and Diagnosis
  673. Exploring Transferability Measures and Domain Selection in Cross-Domain Slot Filling
  674. Exploring Transformer's Potential on Automatic Piano Transcription
  675. Exploring the Effect of ℓ0/ℓ2 Regularization in Neural Network Pruning using the LC Toolkit
  676. Extended Graph Temporal Classification for Multi-Speaker End-to-End ASR
  677. Extending the Use of MDL for High-Dimensional Problems: Variable Selection, Robust Fitting, and Additive Modeling
  678. Extracting and Distilling Direction-Adaptive Knowledge for Lightweight Object Detection in Remote Sensing Images
  679. Extreme-Point Pursuit for Unit-Modulus Optimization
  680. Eyes Tell All: Irregular Pupil Shapes Reveal GAN-Generated Faces
  681. FAZ-BV: A Diabetic Macular Ischemia Grading Framework Combining Faz Attention Network and Blood Vessel Enhancement Filters
  682. FB-MSTCN: A Full-Band Single-Channel Speech Enhancement Method Based on Multi-Scale Temporal Convolutional Network
  683. FDSNeT: An Accurate Real-Time Surface Defect Segmentation Network
  684. FINT: Field-Aware Interaction Neural Network for Click-Through Rate Prediction
  685. FOV-Based Coding Optimization for 360-Degree Virtual Reality Videos
  686. FRCRN: Boosting Feature Representation Using Frequency Recurrence for Monaural Speech Enhancement
  687. FRE-GAN 2: Fast and Efficient Frequency-Consistent Audio Synthesis
  688. FSM: Feature Sampling Module for Object Detection
  689. FSOINET: Feature-Space Optimization-Inspired Network For Image Compressive Sensing
  690. Factorized Neural Transducer for Efficient Language Model Adaptation
  691. Fairness-Aware Selective Sampling on Attributed Graphs
  692. Fake Audio Detection Based On Unsupervised Pretraining Models
  693. Fast Contextual Adaptation with Neural Associative Memory for On-Device Personalized Speech Recognition
  694. Fast Fault Diagnosis Method Of Rolling Bearings In Multi-Sensor Measurement Enviroment
  695. Fast Graph Sampling for Short Video Summarization Using Gershgorin Disc Alignment
  696. Fast Learning of Fast Transforms, with Guarantees
  697. Fast Low Rank Column-Wise Compressive Sensing For Accelerated Dynamic MRI
  698. Fast Multiscale Diffusion On Graphs
  699. Fast Task-Specific Adaptation in Spoken Language Assessment with Meta-Learning
  700. Fast Video Object Segmentation via Dynamic YOLACT
  701. Fast and Stable Convergence of Online SGD for CV@R-Based Risk-Aware Learning
  702. Fast-Rir: Fast Neural Diffuse Room Impulse Response Generator
  703. Fast-Slow Transformer for Visually Grounding Speech
  704. FastAudio: A Learnable Audio Front-End For Spoof Speech Detection
  705. Feature Augmentation Learning for Few-Shot Palmprint Image Recognition With Unconstrained Acquisition
  706. Feature Imitating Networks
  707. Feature Space Message Passing Network for Medical Image Semantic Segmentation
  708. Feature-Based Sensing Matrix Design for Analog to Information Converters
  709. FedClean: A Defense Mechanism against Parameter Poisoning Attacks in Federated Learning
  710. Federated Learning Challenges and Opportunities: An Outlook
  711. Federated Multi-Armed Bandit Via Uncoordinated Exploration
  712. Federated Over-Air Robust Subspace Tracking from Missing Data
  713. Federated Self-Supervised Learning for Acoustic Event Classification
  714. Federated Self-Training for Data-Efficient Audio Recognition
  715. Federated Stochastic Gradient Descent Begets Self-Induced Momentum
  716. Few-Shot Gaze Estimation with Model Offset Predictors
  717. Few-Shot Generation By Modeling Stereoscopic Priors
  718. Few-Shot Learning with Improved Local Representations via Bias Rectify Module
  719. Few-Shot Musical Source Separation
  720. Few-Shot Object Detection with Local Correspondence RPN and Attentive Head
  721. Few-Shot One-Class Domain Adaptation Based On Frequency For Iris Presentation Attack Detection
  722. Filteraugment: An Acoustic Environmental Data Augmentation Method
  723. Find The Way Back: Invertible Kernel Estimator For Blind Image Super-Resolution
  724. Fine-Grained Dynamic Loss for Accurate Single-Image Super-Resolution
  725. Fine-Grained Style Control In Transformer-Based Text-To-Speech Synthesis
  726. Fine-Tuning Wav2Vec2 for Speaker Recognition
  727. Fldp: Flexible Strategy For Local Differential Privacy
  728. Floor Plan Reconstruction with High-Precision Rf-Based Tracking
  729. Flow-Based Fast Multichannel Nonnegative Matrix Factorization for Blind Source Separation
  730. Flow-Based Point Cloud Completion Network with Adversarial Refinement
  731. FlowDT: A Flow-Aware Digital Twin for Computer Networks
  732. Forensic Analysis and Localization of Multiply Compressed MP3 Audio Using Transformers
  733. Fostering The Robustness Of White-Box Deep Neural Network Watermarks By Neuron Alignment
  734. Fracture Detection and Localization in Chest X-Rays Using Semi-Supervised Learning with Dynamic Sharpening
  735. Fraug: A Frame Rate Based Data Augmentation Method for Depression Detection from Speech Signals
  736. Free Lunch for Cross-Domain Occluded Face Recognition without Source Data
  737. Frequency-Specific Non-Linear Granger Causality in a Network of Brain Signals
  738. From Bottom-Up To Top-Down: Characterization Of Training Process In Gaze Modeling
  739. From Shallow to Deep: Compositional Reasoning over Graphs for Visual Question Answering
  740. Frontend Attributes Disentanglement for Speech Emotion Recognition
  741. FullSubNet+: Channel Attention Fullsubnet with Complex Spectrograms for Speech Enhancement
  742. Fusing ASR Outputs in Joint Training for Speech Emotion Recognition
  743. Fusion and Orthogonal Projection for Improved Face-Voice Association
  744. Fusion of Modulation Spectral and Spectral Features with Symptom Metadata for Improved Speech-Based Covid-19 Detection
  745. Fusion-Id: A Photoplethysmography and Motion Sensor Fusion Biometric Authenticator With Few-Shot on-Boarding
  746. GAZEATTENTIONNET: Gaze Estimation with Attentions
  747. GOS: A Large-Scale Annotated Outdoor Scene Synthetic Dataset
  748. GPU-Accelerated Forward-Backward Algorithm with Application to Lattice-Free MMI
  749. Gan-Based Joint Activity Detection and Channel Estimation for Grant-Free Random Access
  750. Ganet: Unary Attention Reaches Pairwise Attention Via Implicit Group Clustering in Light-Weight CNNs
  751. Gated Multimodal Fusion with Contrastive Learning for Turn-Taking Prediction in Human-Robot Dialogue
  752. Generalization Ability of MOS Prediction Networks
  753. Generalized Autocorrelation Analysis for Multi-Target Detection
  754. Generalized Face Anti-Spoofing via Cross-Adversarial Disentanglement with Mixing Augmentation
  755. Generalized Matching Pursuits for the Sparse Optimization of Separable Objectives
  756. Generalized Sliced Probability Metrics
  757. Generalized Time Domain Velocity Vector
  758. Generalized Zero-Shot Learning Using Conditional Wasserstein Autoencoder
  759. Generating Disentangled Arguments with Prompts: A Simple Event Extraction Framework That Works
  760. Generation for Unsupervised Domain Adaptation: A Gan-Based Approach for Object Classification with 3D Point Cloud Data
  761. Generation of Personal Sound Fields in Reverberant Environments Using Interframe Correlation
  762. Generative Adversarial Network Including Referring Image Segmentation For Text-Guided Image Manipulation
  763. Genre-Conditioned Acoustic Models for Automatic Lyrics Transcription of Polyphonic Music
  764. Genre-Conditioned Long-Term 3D Dance Generation Driven by Music
  765. Geometric Low-Rank Tensor Approximation for Remotely Sensed Hyperspectral And Multispectral Imagery Fusion
  766. Glassoformer: A Query-Sparse Transformer for Post-Fault Power Grid Voltage Prediction
  767. Global Evolution Neural Network for Segmentation of Remote Sensing Images
  768. Global Optimization Solution for Dynamic Adaptive 360-Degree Streaming
  769. Global-Local Feature Enhancement Network for Robust Object Detection using mmWave Radar and Camera
  770. Goal-Oriented Communication for Edge Learning Based On the Information Bottleneck
  771. Gradient Staleness in Asynchronous Optimization Under Random Communication Delays
  772. Gradient Variance Loss for Structure-Enhanced Image Super-Resolution
  773. Gradient-Weighted Class Activation Mapping for Spatio Temporal Graph Convolutional Network
  774. Gradual Surrogate Gradient Learning in Deep Spiking Neural Networks
  775. Graph Attentive Feature Aggregation for Text-Independent Speaker Verification
  776. Graph Convolution for Re-Ranking in Person Re-Identification
  777. Graph Convolutional Network Based Semi-Supervised Learning on Multi-Speaker Meeting Data
  778. Graph Convolutional Networks With Autoencoder-Based Compression And Multi-Layer Graph Learning
  779. Graph Fine-Grained Contrastive Representation Learning
  780. Graph Learning Based Autoencoder for Hyperspectral Band Selection
  781. Graph Learning From Multivariate Dependent Time Series Via A Multi-Attribute Formulation
  782. Graph Learning Information Criterion
  783. Graph-Based Point Cloud Denoising Using Shape-Aware Consistency For Free-Viewpoint Video
  784. Graph-Structured Sparse Regularization Via Convex Optimization
  785. Graphon-Aided Joint Estimation of Multiple Graphs
  786. Grassmannian Dimensionality Reduction Using Triplet Margin Loss for Ume Classification of 3d Point Clouds
  787. Gridless DOA Estimation Under the Multi-Frequency Model
  788. Group-Wise Feature Selection for Supervised Learning
  789. HBP: An Efficient Block Permutation Solver Using Hungarian Algorithm and Spectrogram Inpainting for Multichannel Audio Source Separation
  790. HGCN: Harmonic Gated Compensation Network for Speech Enhancement
  791. HIRL: Hybrid Image Restoration Based on Hierarchical Deep Reinforcement Learning via Two-Step Analysis
  792. HOQRI: Higher-Order QR Iteration for Scalable Tucker Decomposition
  793. HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection
  794. Half Inverted Nested Arrays with Large Hole-Free Fourth-Order Difference Co-Arrays
  795. Hand Gesture Recognition Using Temporal Convolutions and Attention Mechanism
  796. Harmonic Gated Compensation Network Plus for ICASSP 2022 DNS Challenge
  797. Harmonic and Percussive Sound Separation Based on Mixed Partial Derivative of Phase Spectrogram
  798. Harmonicity Plays a Critical Role in DNN Based Versus in Biologically-Inspired Monaural Speech Segregation Systems
  799. Harvesting Partially-Disjoint Time-Frequency Information for Improving Degenerate Unmixing Estimation Technique
  800. Have Best of Both Worlds: Two-Pass Hybrid and E2E Cascading Framework for Speech Recognition
  801. Heart Rate and Oxygen Saturation Estimation from Facial Video with Multimodal Physiological Data Generation
  802. Heterogeneous Graph Node Classification With Multi-Hops Relation Features
  803. Heuristic Dropout: An Efficient Regularization Method for Medical Image Segmentation Models
  804. HiFi-SVC: Fast High Fidelity Cross-Domain Singing Voice Conversion
  805. HiFiDenoise: High-Fidelity Denoising Text to Speech with Adversarial Networks
  806. Hierarchical Classification of Singing Activity, Gender, and Type in Complex Music Recordings
  807. Hierarchical Conditional End-to-End ASR with CTC and Multi-Granular Subword Units
  808. Hierarchical Deep Learning Model with Inertial and Physiological Sensors Fusion for Wearable-Based Human Activity Recognition
  809. Hierarchical Feature Aggregation Network for Deep Image Compression
  810. Hierarchical Graph-Based Neural Network for Singing Melody Extraction
  811. Hierarchical Prosody Modeling and Control in Non-Autoregressive Parallel Neural TTS
  812. Hierarchical Signal Fusion Network for Pulsar Detection with Phase-Correlation and Signal Attentions
  813. Hierarchical and Multi-View Dependency Modelling Network for Conversational Emotion Recognition
  814. High-Dimensional Sparse Bayesian Learning without Covariance Matrices
  815. High-Fidelity Portrait Editing Via Exploring Differentiable Guided Sketches from the Latent Space
  816. High-Quality Self-Supervised Snapshot Hyperspectral Imaging
  817. Histogram-Guided Semantic-Aware Colorization
  818. Histokt: Cross Knowledge Transfer in Computational Pathology
  819. Hodgelets: Localized Spectral Representations of Flows On Simplicial Complexes
  820. Holistic Semi-Supervised Approaches for EEG Representation Learning
  821. How Can a Cognitive Radar Mask its Cognition?
  822. How Neural Processes Improve Graph Link Prediction
  823. How Secure Are The Adversarial Examples Themselves?
  824. Human Decision Making with Bounded Rationality
  825. Human Emotion Recognition Using Multi-Modal Biological Signals Based On Time Lag-Considered Correlation Maximization
  826. Hybrid Attention-Based Prototypical Networks for Few-Shot Sound Classification
  827. Hybrid RNN-T/Attention-Based Streaming ASR with Triggered Chunkwise Attention and Dual Internal Language Model Integration
  828. Hybrid Weighting Loss for Precipitation Nowcasting from Radar Images
  829. Hybrid sub-word segmentation for handling long tail in morphologically rich low resource languages
  830. Hypergraph-Based Reinforcement Learning for Stock Portfolio Selection
  831. Hypergraphs with Edge-Dependent Vertex Weights: Spectral Clustering Based on the 1-Laplacian
  832. Hyperspectral Image Classification Based on Co-Learning Through Dual-Architecture Ensemble
  833. Hyperspectral Image Super-Resolution with Deep Priors and Degradation Model Inversion
  834. ICASSP 2022 Acoustic Echo Cancellation Challenge
  835. ICASSP 2022 L3DAS22 Challenge: Ensemble of Resnet-Conformers with Ambisonics Data Augmentation for Sound Event Localization and Detection
  836. ICASSP-SPGC 2022: Root Cause Analysis for Wireless Network Fault Localization
  837. IMPQ: Reduced Complexity Neural Networks Via Granular Precision Assignment
  838. ISDA: Position-Aware Instance Segmentation with Deformable Attention
  839. ISOMETRIC MT: Neural Machine Translation for Automatic Dubbing
  840. ISTFTNET: Fast and Lightweight Mel-Spectrogram Vocoder Incorporating Inverse Short-Time Fourier Transform
  841. Icassp 2022 Deep Noise Suppression Challenge
  842. Identification of Pulse Streams Of Unknown Shape From Time Encoding Machine Samples
  843. Image Denoising with Deep Unfolding And Normalizing Flows
  844. Image Steganalysis with Convolutional Vision Transformer
  845. Image-Text Alignment and Retrieval Using Light-Weight Transformer
  846. Image-to-Graph Transformers for Chemical Structure Recognition
  847. Image-to-Video Re-Identification via Mutual Discriminative Knowledge Transfer
  848. Importance Sampling Cams For Weakly-Supervised Segmentation
  849. Importance of Switch Optimization Criterion in Switching WPE Dereverberation
  850. Importantaug: A Data Augmentation Agent for Speech
  851. Improve Few-Shot Voice Cloning Using Multi-Modal Learning
  852. Improve Image Captioning Via Relation Modeling
  853. Improved Beamforming Encoding for Joint Radar and Communication
  854. Improved Language Identification Through Cross-Lingual Self-Supervised Learning
  855. Improved Meta Learning for Low Resource Speech Recognition
  856. Improved Representation Learning For Acoustic Event Classification Using Tree-Structured Ontology
  857. Improved Simulation of Realistically-Spatialised Simultaneous Speech Using Multi-Camera Analysis in The Chime-5 Dataset
  858. Improved Singing Voice Separation with Chromagram-Based Pitch-Aware Remixing
  859. Improving Actor-Critic Reinforcement Learning Via Hamiltonian Monte Carlo Method
  860. Improving Adversarial Waveform Generation Based Singing Voice Conversion with Harmonic Signals
  861. Improving Anomaly Detection with a Self-Supervised Task Based on Generative Adversarial Network
  862. Improving BCI-based Color Vision Assessment Using Gaussian Process Regression
  863. Improving Biomedical Named Entity Recognition with a Unified Multi-Task MRC Framework
  864. Improving Bird Classification with Unsupervised Sound Separation
  865. Improving Brain Decoding Methods and Evaluation
  866. Improving CTC-Based Speech Recognition Via Knowledge Transferring from Pre-Trained Language Models
  867. Improving Character Error Rate is Not Equal to Having Clean Speech: Speech Enhancement for ASR Systems with Black-Box Acoustic Models
  868. Improving Class Activation Map for Weakly Supervised Object Localization
  869. Improving Confidence Estimation on Out-of-Domain Data for End-to-End Speech Recognition
  870. Improving Contextual Coherence in Variational Personalized and Empathetic Dialogue Agents
  871. Improving Cross-Lingual Speech Synthesis with Triplet Training Scheme
  872. Improving Cross-Modal Understanding in Visual Dialog Via Contrastive Learning
  873. Improving Dialogue Generation via Proactively Querying Grounded Knowledge
  874. Improving Dual-Microphone Speech Enhancement by Learning Cross-Channel Features with Multi-Head Attention
  875. Improving Dynamic Graph Convolutional Network with Fine-Grained Attention Mechanism
  876. Improving Emotional Speech Synthesis by Using SUS-Constrained VAE and Text Encoder Aggregation
  877. Improving End-To-End Speech Translation Model with Bert-Based Contextual Information
  878. Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge Selection
  879. Improving End-to-end Models for Set Prediction in Spoken Language Understanding
  880. Improving Factored Hybrid HMM Acoustic Modeling without State Tying
  881. Improving Fairness in Speaker Verification via Group-Adapted Fusion Network
  882. Improving Fastspeech TTS with Efficient Self-Attention and Compact Feed-Forward Network
  883. Improving Feature Generalizability with Multitask Learning in Class Incremental Learning
  884. Improving Generalization of Deep Networks for Estimating Physical Properties of Containers and Fillings
  885. Improving Inference for Spatial Signals by Contextual False Discovery Rates
  886. Improving Joint Sparse Hyperspectral Unmixing by Simultaneously Clustering Pixels According To Their Mixtures
  887. Improving Lyrics Alignment Through Joint Pitch Detection
  888. Improving Maximum Likelihood Difference Scaling Method To Measure Inter Content Scale
  889. Improving Noise Robustness of Contrastive Speech Representation Learning with Speech Reconstruction
  890. Improving Non-Autoregressive End-to-End Speech Recognition with Pre-Trained Acoustic and Language Models
  891. Improving Phase-Rectified Signal Averaging for Fetal Heart Rate Analysis
  892. Improving Phonetic Realizations in its by Using Phoneme-Aligned Graphemes
  893. Improving Pseudo-Label Training For End-To-End Speech Recognition Using Gradient Mask
  894. Improving Recognition-Synthesis Based any-to-one Voice Conversion with Cyclic Training
  895. Improving Reference-Based Image Colorization For Line Arts Via Feature Aggregation And Contrastive Learning
  896. Improving Self-Supervised Learning for Speech Recognition with Intermediate Layer Supervision
  897. Improving Separation-Based Speaker Diarization Via Iterative Model Refinement And Speaker Embedding Based Post-Processing
  898. Improving Source Separation by Explicitly Modeling Dependencies between Sources
  899. Improving Spoken Language Understanding by Enhancing Text Representation
  900. Improving The Latency And Quality Of Cascaded Encoders
  901. Improving Ultrasound Image Classification with Local Texture Quantisation
  902. Improving the Classification of Phonetic Segments from Raw Ultrasound Using Self-Supervised Learning and Hard Example Mining
  903. Improving the Fusion of Acoustic and Text Representations in RNN-T
  904. In Pursuit of Preserving the Fidelity of Adversarial Images
  905. Incipient Fault Severity Estimation Using Local Mahalanobis Distance
  906. Incoherent Synthesis of Sparse Broadband Arrays based on a Parameter-Free Subspace Clustering
  907. Incorporating End-to-End Framework Into Target-Speaker Voice Activity Detection
  908. Incorporating Gaze Behavior Using Joint Embedding With Scene Context for Driver Takeover Detection
  909. Increasing Loudness in Audio Signals: A Perceptually Motivated Approach to Preserve Audio Quality
  910. Incremental Context Aware Attentive Knowledge Tracing
  911. Incremental User Embedding Modeling for Personalized Text Classification
  912. Independent Vector Analysis Based Subgroup Identification from Multisubject fMRI Data
  913. Individualized Hear-Through For Acoustic Transparency Using PCA-Based Sound Pressure Estimation At The Eardrum
  914. Infant Crying Detection In Real-World Environments
  915. Infergrad: Improving Diffusion Models for Vocoder by Considering Inference in Training
  916. Inferring Camera Intrinsics Based on Surfaces of Revolution: A Single Image Geometric Network Approach for Camera Calibration
  917. Information Theoretic Limits For Standard and One-Bit Compressed Sensing with Graph-Structured Sparsity
  918. Informative Attention Supervision for Grounded Video Description
  919. Initialization-Free Implicit-Focusing (IF2) for Wideband Direction-of-Arrival Estimation
  920. Injecting Text and Cross-Lingual Supervision in Few-Shot Learning from Self-Supervised Models
  921. Instantaneous Linear Dimensionality Reduction of Multichannel Time-Series Signal for Array Signal Processing
  922. Integer-Only Zero-Shot Quantization for Efficient Speech Recognition
  923. Integrated Sensing and Communications Via 5G NR Waveform: Performance Analysis
  924. Integrating Dependency Tree into Self-Attention for Sentence Representation
  925. Integrating Multiple ASR Systems into NLP Backend with Attention Fusion
  926. Integrating Pretrained Language Model for Dialogue Policy Evaluation
  927. Integrating Statistical Uncertainty into Neural Network-Based Speech Enhancement
  928. Integrating Text Inputs for Training and Adapting RNN Transducer ASR Models
  929. Integration of Anomaly Machine Sound Detection into Active Noise Control to Shape the Residual Sound
  930. Integration of Pre-Trained Networks with Continuous Token Interface for End-to-End Spoken Language Understanding
  931. Intelligent Wi-Fi Based Child Presence Detection System
  932. Interactive Feature Fusion for End-to-End Noise-Robust Speech Recognition
  933. Interactive Multi-Level Prosody Control for Expressive Speech Synthesis
  934. Intermix: An Interference-Based Data Augmentation and Regularization Technique for Automatic Deep Sound Classification
  935. Internet Streaming Audio Based Speech Reception Threshold Measurement in Cochlear Implant Users
  936. Interpretable Image Classification Using Sparse Oblique Decision Trees
  937. Interpreting Intermediate Convolutional Layers In Unsupervised Acoustic Word Classification
  938. Inverse Imaging with Generative Priors Via Langevin Dynamics
  939. Investigating Robustness of Biological vs. Backprop Based Learning
  940. Investigating Self-Supervised Learning for Speech Enhancement and Separation
  941. Investigating Sequence-Level Normalisation For CTC-Like End-to-End ASR
  942. Investigating the Potential of Auxiliary-Classifier Gans for Image Classification in Low Data Regimes
  943. Investigation And Comparison of Optimization Methods for Variational Autoencoder-Based Underdetermined Multichannel Source Separation
  944. Investigation of Robustness of Hubert Features from Different Layers to Domain, Accent and Language Variations
  945. Invisible and Efficient Backdoor Attacks for Compressed Deep Neural Networks
  946. Is Cross-Attention Preferable to Self-Attention for Multi-Modal Emotion Recognition?
  947. Iterative Channel Estimation and Data Detection Algorithm For OTFS Modulation
  948. Iterative Learning for Distorted Image Restoration
  949. Iterative Re-weighted Least Squares Algorithms for Non-negative Sparse and Group-sparse Recovery
  950. Iterative Self Knowledge Distillation - from Pothole Classification to Fine-Grained and Covid Recognition
  951. ItôWave: Itô Stochastic Differential Equation is all You Need for Wave Generation
  952. JE2NET: Joint Exploitation and Exploration in Reinforcement Learning Based Image Restoration
  953. Jmpnet: Joint Motion Prediction for Learning-Based Video Compression
  954. Joint Beam Selection and Precoding Based on Differential Evolution for Millimeter-Wave Massive MIMO Systems
  955. Joint Calibration and Mapping of Satellite Altimetry Data Using Trainable Variational Models
  956. Joint Centrality Estimation and Graph Identification from Mixture of Low Pass Graph Signals
  957. Joint Dual-Domain Matrix Factorization for ECG Biometric Recognition
  958. Joint Ego-Noise Suppression and Keyword Spotting on Sweeping Robots
  959. Joint Far- and Near-End Speech Intelligibility Enhancement Based on the Approximated Speech Intelligibility Index
  960. Joint Global-Local Alignment for Domain Adaptive Semantic Segmentation
  961. Joint Hypoglycemia Prediction and Glucose Forecasting via Deep Multi-Task Learning
  962. Joint Inference of Multiple Graphs with Hidden Variables from Stationary Graph Signals
  963. Joint Learning for Addressee Selection and Response Generation in Multi-Party Conversation
  964. Joint Learning of Feature Extraction and Cost Aggregation for Semantic Correspondence
  965. Joint Magnitude Estimation and Phase Recovery Using Cycle-In-Cycle GAN for Non-Parallel Speech Enhancement
  966. Joint Model Order Estimation for Multiple Tensors with A Coupled Mode and Applications to the Joint Decomposition of EEG, MEG Magnetometer, and Gradiometer Tensors
  967. Joint Modeling of Code-Switched and Monolingual ASR via Conditional Factorization
  968. Joint Multiple Intent Detection and Slot Filling Via Self-Distillation
  969. Joint Normality Test Via Two-Dimensional Projection
  970. Joint Radar-Communications Processing from A Dual-Blind Deconvolution Perspective
  971. Joint Source Localization and Association Through Overcomplete Representation Under Multipath Propagation Environment
  972. Joint Speech Recognition and Audio Captioning
  973. Joint Temporal Convolutional Networks and Adversarial Discriminative Domain Adaptation for EEG-Based Cross-Subject Emotion Recognition
  974. Joint Unsupervised and Supervised Training for Multilingual ASR
  975. Joint and Adversarial Training with ASR for Expressive Speech Synthesis
  976. K-Converter: An Unsupervised Singing Voice Conversion System
  977. KaraSinger: Score-Free Singing Voice Synthesis with VQ-VAE Using Mel-Spectrograms
  978. Kernel Estimation Network for Blind Super-Resolution
  979. Key-Sparse Transformer for Multimodal Speech Emotion Recognition
  980. Knowledge Augmented Bert Mutual Network in Multi-Turn Spoken Dialogues
  981. Knowledge Distillation for Neural Transducers from Large Self-Supervised Pre-Trained Models
  982. Knowledge Distillation from Language Model to Acoustic Model: A Hierarchical Multi-Task Learning Approach
  983. Knowledge Transfer from Large-Scale Pretrained Language Models to End-To-End Speech Recognizers
  984. L-SpEx: Localized Target Speaker Extraction
  985. L3DAS22 Challenge: Learning 3D Audio Sources in a Real Office Environment
  986. LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech
  987. LERPS: Lighting Estimation and Relighting for Photometric Stereo
  988. LETR: A Lightweight and Efficient Transformer for Keyword Spotting
  989. LIGHT-SERNET: A Lightweight Fully Convolutional Neural Network for Speech Emotion Recognition
  990. LMS and NLMS Algorithms for the Identification of Impulse Responses with Intrinsic Symmetric or Antisymmetric Properties
  991. LPC Augment: an LPC-based ASR Data Augmentation Algorithm for Low and Zero-Resource Children's Dialects
  992. LRPD: Large Replay Parallel Dataset
  993. Label Propagation Across Graphs: Node Classification Using Graph Neural Tangent Kernels
  994. Label-Aware Ranked Loss for Robust People Counting Using Automotive In-Cabin Radar
  995. Label-Occurrence-Balanced Mixup for Long-Tailed Recognition
  996. Language Adaptive Cross-Lingual Speech Representation Learning with Sparse Sharing Sub-Networks
  997. Large-Scale ASR Domain Adaptation Using Self- and Semi-Supervised Learning
  998. Large-Scale Independent Component Analysis By Speeding Up Lie Group Techniques
  999. Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification
  1000. Latent Space Slicing for Enhanced Entropy Modeling In Learning-Based Point Cloud Geometry Compression

Looking for submission deadlines instead? See the conference deadline calendar.