← All conferences

ICASSP 2022 Accepted Papers

The full list of 1,864 papers accepted at ICASSP 2022 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. Lattention: Lattice-Attention in ASR Rescoring
  2. Lattice Rescoring Based on Large Ensemble of Complementary Neural Language Models
  3. LatticeBART: Lattice-to-Lattice Pre-Training for Speech Recognition
  4. Learnable Hypergraph Laplacian for Hypergraph Learning
  5. Learnable Nonlinear Compression for Robust Speaker Verification
  6. Learnable Wavelet Packet Transform for Data-Adapted Spectrograms
  7. Learning Acoustic Frame Labeling for Phoneme Segmentation with Regularized Attention Mechanism
  8. Learning Adjustable Image Rescaling with Joint Optimization of Perception and Distortion
  9. Learning Approach For Fast Approximate Matrix Factorizations
  10. Learning Common Dependency Structure for Unsupervised Cross-Domain Ner
  11. Learning Continuous Representation of Audio for Arbitrary Scale Super Resolution
  12. Learning Correlation for Online Multiple Object Tracking
  13. Learning Decoupling Features Through Orthogonality Regularization
  14. Learning Deep Pathological Features for WSI-Level Cervical Cancer Grading
  15. Learning Domain-Invariant Transformation for Speaker Verification
  16. Learning Expanding Graphs for Signal Interpolation
  17. Learning Filterbanks for End-to-End Acoustic Beamforming
  18. Learning Gaussian Graphical Models with Differing Pairwise Sample Sizes
  19. Learning Monocular 3D Human Pose Estimation With Skeletal Interpolation
  20. Learning Monocular Mesh Recovery of Multiple Body Parts Via Synthesis
  21. Learning Multiple Explainable and Generalizable Cues for Face Anti-Spoofing
  22. Learning Music Audio Representations Via Weak Language Supervision
  23. Learning Music Sequence Representation From Text Supervision
  24. Learning Semantic-Aligned Feature Representation for Text-Based Person Search
  25. Learning Sound Localization Better from Semantically Similar Samples
  26. Learning Sparse Graphs with a Core-Periphery Structure
  27. Learning Structured Sparsity For Time-Frequency Reconstruction
  28. Learning Subject-Invariant Representations from Speech-Evoked EEG Using Variational Autoencoders
  29. Learning Task-Specific Representation for Video Anomaly Detection with Spatial-Temporal Attention
  30. Learning To Integrate Vision Data Into Road Network Data
  31. Learning to Enhance or Not: Neural Network-Based Switching of Enhanced and Observed Signals for Overlapping Speech Recognition
  32. Learning to Fuse Heterogeneous Features for Low-Light Image Enhancement
  33. Learning to Predict Speech in Silent Videos Via Audiovisual Analogy
  34. Learning to Sample for Sparse Signals
  35. Learning-Aided Initialization for Variational Bayesian DOA Estimation
  36. Learning-Based Personal Speech Enhancement for Teleconferencing by Exploiting Spatial-Spectral Features
  37. Learning-Based Resource Allocation with Dynamic Data Rate Constraints
  38. Learnings from Federated Learning in The Real World
  39. Leveraging Bilinear Attention to Improve Spoken Language Understanding
  40. Leveraging Local Temporal Information for Multimodal Scene Classification
  41. Leveraging Sparse Coding for EEG Based Emotion Recognition in Shooting
  42. LightPose: A Lightweight and Efficient Model with Transformer for Human Pose Estimation
  43. Linear-Time Sampling on Signed Graphs Via Gershgorin Disc Perfect Alignment
  44. Lipreading Model Based On Whole-Part Collaborative Learning
  45. Listen, Know and Spell: Knowledge-Infused Subword Modeling for Improving ASR Performance of OOV Named Entities
  46. LiteHAR: Lightweight Human Activity Recognition from WIFI Signals with Random Convolution Kernels
  47. LocUNet: Fast Urban Positioning Using Radio Maps and Deep Learning
  48. Local Context Interaction-Aware Glyph-Vectors for Chinese Sequence Tagging
  49. Local Information Modeling with Self-Attention for Speaker Verification
  50. Local and Global Alignments for Generalizable Sensor-Based Human Activity Recognition
  51. Local-Global Feature Aggregation for Light Field Image Super-Resolution
  52. Localization based Sequential Grouping for Continuous Speech Separation
  53. Localizing More Sources than Sensors in Presence of Coherent Sources
  54. Locate This, Not that: Class-Conditioned Sound Event DOA Estimation
  55. Location-Based Training for Multi-Channel Talker-Independent Speaker Separation
  56. Look, Listen and Pay More Attention: Fusing Multi-Modal Information for Video Violence Detection
  57. Low Complex Accurate Multi-Source RTF Estimation
  58. Low Complexity Equalization for Afdm In Doubly Dispersive Channels
  59. Low Precision Local Learning for Hardware-Friendly Neuromorphic Visual Recognition
  60. Low Resources Online Single-Microphone Speech Enhancement with Harmonic Emphasis
  61. Low-Complexity Attention Modelling via Graph Tensor Networks
  62. Low-Complexity Multi-Model CNN in-Loop Filter for AVS3
  63. Low-Latency Human-Computer Auditory Interface Based on Real-Time Vision Analysis
  64. Low-Light Image Enhancement via Feature Restoration
  65. Low-Rank Phase Retrieval with Structured Tensor Models
  66. M2Met: The Icassp 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
  67. MA-NET: Multi-Scale Attention-Aware Network for Optical Flow Estimation
  68. MAG+: An Extended Multimodal Adaptation Gate for Multimodal Sentiment Analysis
  69. MAKD: MULTIPLE Auxiliary Knowledge Distillation
  70. MANNER: Multi-View Attention Network For Noise Erasure
  71. MBA-RainGAN: A Multi-Branch Attention Generative Adversarial Network for Mixture of Rain Removal
  72. MBNet: A Multi-Resolution Branch Network for Semantic Segmentation Of Ultra-High Resolution Images
  73. MEJIGCLU: More Effective Jigsaw Clustering For Unsupervised Visual Representation Learning
  74. MFA: TDNN with Multi-Scale Frequency-Channel Attention for Text-Independent Speaker Verification with Short Utterances
  75. MLP-SVNET: A Multi-Layer Perceptrons Based Network for Speaker Verification
  76. MM-DFN: Multimodal Dynamic Fusion Network for Emotion Recognition in Conversations
  77. MOS Predictor for Synthetic Speech with I-Vector Inputs
  78. MRI Recovery with a Self-Calibrated Denoiser
  79. MS-ROCANet: Multi-Scale Residual Orthogonal-Channel Attention Network for Scene Text Detection
  80. MSDTRON: A High-Capability Multi-Speaker Speech Synthesis System for Diverse Data Using Characteristic Information
  81. MTAF: Shopping Guide Micro-Videos Popularity Prediction Using Multimodal and Temporal Attention Fusion Approach
  82. Magic Dust for Cross-Lingual Adaptation of Monolingual Wav2vec-2.0
  83. Making The Unknown More Certain: A Stacked Ensemble Classifier for Open Gesture Recognition with a Social Robot
  84. Manifold Learning-Supported Estimation of Relative Transfer Functions For Spatial Filtering
  85. Mannet: A Large-Scale Manipulated Image Detection Dataset And Baseline Evaluations
  86. Map: Multispectral Adversarial Patch to Attack Person Detection
  87. Mask-Based Attention Parallel Network for in-the-Wild Facial Expression Recognition
  88. Masked Acoustic Unit for Mispronunciation Detection and Correction
  89. Massive Unsourced Random Access Based on Bilinear Vector Approximate Message Passing
  90. Massively Multilingual ASR: A Lifelong Learning Solution
  91. Matching Point Sets with Quantum Circuit Learning
  92. Material-Guided Siamese Fusion Network for Hyperspectral Object Tracking
  93. Matrix Decomposition on Graphs: A Simplified Functional View
  94. Maximizing Audio Event Detection Model Performance on Small Datasets Through Knowledge Transfer, Data Augmentation, and Pretraining: an Ablation Study
  95. Maximum Batch Frobenius Norm for Multi-Domain Text Classification
  96. Melons: Generating Melody With Long-Term Structure Using Transformers And Structure Graph
  97. Memobert: Pre-Training Model with Prompt-Based Learning for Multimodal Emotion Recognition
  98. Memory in Echo State Networks and the Controllability Matrix Rank
  99. Memory-Based Message Passing: Decoupling the Message for Propagation from Discrimination
  100. Message Passing-Based Cooperative Localization with Embedded Particle Flow
  101. Meta Talk: Learning To Data-Efficiently Generate Audio-Driven Lip-Synchronized Talking Face With High Definition
  102. MetricGAN-U: Unsupervised Speech Enhancement/ Dereverberation Based Only on Noisy/ Reverberated Speech
  103. Metricbert: Text Representation Learning Via Self-Supervised Triplet Training
  104. Mimo Detection by Variational Posterior Inference
  105. Minimizing Residuals for Native-Nonnative Voice Conversion in a Sparse, Anchor-Based Representation of Speech
  106. Minimum Word Error Training For Non-Autoregressive Transformer-Based Code-Switching ASR
  107. Mining Hard Samples Locally And Globally For Improved Speech Separation
  108. Mismatched Supervised Learning
  109. Mitigating Closed-Model Adversarial Examples with Bayesian Neural Modeling for Enhanced End-to-End Speech Recognition
  110. Mixed In Time And Modality: Curse Or Blessingƒ Cross-Instance Data Augmentation for Weakly Supervised Multimodal Temporal Fusion
  111. Mixed Knowledge Relation Transformer for Image Captioning
  112. Mixed Precision DNN Quantization for Overlapped Speech Separation and Recognition
  113. Mixed Transformer U-Net for Medical Image Segmentation
  114. Mixer-TTS: Non-Autoregressive, Fast and Compact Text-to-Speech Model Conditioned on Language Model Embeddings
  115. Mixture Model Auto-Encoders: Deep Clustering Through Dictionary Learning
  116. Mmlatch: Bottom-Up Top-Down Fusion For Multimodal Sentiment Analysis
  117. Model Selection via Misspecified Cramér-Rao Bound Minimization
  118. Model-Based Approach for Measuring the Fairness in ASR
  119. Model-Based Online Learning for Resource Sharing in Joint Radar-Communication Systems
  120. Model-Based Reconstruction for Collimated Beam Ultrasound Systems
  121. Modeling Beats and Downbeats with a Time-Frequency Transformer
  122. Modeling Human Memory in Multi-Object Tracking with Transformers
  123. Modeling Intention, Emotion and External World in Dialogue Systems
  124. Modeling The Detection Capability Of High-Speed Spiking Cameras
  125. Modeling of Pre-Trained Neural Network Embeddings Learned From Raw Waveform for COVID-19 Infection Detection
  126. Modernn: Towards Fine-Grained Motion Details for Spatiotemporal Predictive Learning
  127. Modulo Event-Driven Sampling: System Identification and Hardware Experiments
  128. Monocular Vehicle 3D Bounding Box Estimation Using Homograhy and Geometry in Traffic Scene
  129. Monotonic Generalized Nash Games with Application to the Management of Energy-Aware Aloha Networks
  130. Motif-Topology and Reward-Learning Improved Spiking Neural Network for Efficient Multi-Sensory Integration
  131. Multi-ACCDOA: Localizing And Detecting Overlapping Sounds From The Same Class With Auxiliary Duplicating Permutation Invariant Training
  132. Multi-Channel Attentive Graph Convolutional Network with Sentiment Fusion for Multimodal Sentiment Analysis
  133. Multi-Channel End-To-End Neural Diarization with Distributed Microphones
  134. Multi-Channel Multi-Speaker ASR Using 3D Spatial Feature
  135. Multi-Channel Narrow-Band Deep Speech Separation with Full-Band Permutation Invariant Training
  136. Multi-Channel Speaker Diarization Using Spatial Features for Meetings
  137. Multi-Channel Speaker Verification with Conv-Tasnet Based Beamformer
  138. Multi-Channel Speech Denoising for Machine Ears
  139. Multi-Domain Unpaired Ultrasound Image Artifact Removal Using a Single Convolutional Neural Network
  140. Multi-Domain Unsupervised Image-to-Image Translation with Appearance Adaptive Convolution
  141. Multi-Feature Integration for Speaker Embedding Extraction
  142. Multi-Focus Guided Semantic Aggregation for Video Object Detection
  143. Multi-Frame Full-Rank Spatial Covariance Analysis for Underdetermined BSS in Reverberant Environments
  144. Multi-Frame Super-Resolution With Raw Images Via Modified Deformable Convolution
  145. Multi-Head Relu Implicit Neural Representation Networks
  146. Multi-Hierarchy Proxy Structure for Deep Metric Learning
  147. Multi-Level Contrastive Learning for Cross-Lingual Alignment
  148. Multi-Level Relation Aware Network for Person Re-Identification
  149. Multi-Level Spatial-Temporal Adaptation Network for Motor Imagery Classification
  150. Multi-Lingual Multi-Task Speech Emotion Recognition Using wav2vec 2.0
  151. Multi-Modal Acoustic-Articulatory Feature Fusion For Dysarthric Speech Recognition
  152. Multi-Modal Emotion Recognition with Self-Guided Modality Calibration
  153. Multi-Modal Learning with Text Merging for TEXTVQA
  154. Multi-Modal Pre-Training for Automated Speech Recognition
  155. Multi-Modal Recurrent Fusion for Indoor Localization
  156. Multi-Pose Virtual Try-On Via Self-Adaptive Feature Filtering
  157. Multi-Query Multi-Head Attention Pooling and Inter-Topk Penalty for Speaker Verification
  158. Multi-Relation Message Passing for Multi-Label Text Classification
  159. Multi-Role Event Argument Extraction as Machine Reading Comprehension with Argument Match Optimization
  160. Multi-Sample Subband Wavernn Via Multivariate Gaussian
  161. Multi-Scale Refinement Network Based Acoustic Echo Cancellation
  162. Multi-Scale Reinforcement Learning Strategy for Object Detection
  163. Multi-Scale Speaker Embedding-Based Graph Attention Networks For Speaker Diarisation
  164. Multi-Scale Temporal Frequency Convolutional Network With Axial Attention for Speech Enhancement
  165. Multi-Scale Temporal Frequency Convolutional Network with Axial Attention for Multi-Channel Speech Enhancement
  166. Multi-Speaker Pitch Tracking via Embodied Self-Supervised Learning
  167. Multi-Stage Graph Representation Learning for Dialogue-Level Speech Emotion Recognition
  168. Multi-Stage and Multi-Loss Training for Fullband Non-Personalized and Personalized Speech Enhancement
  169. Multi-Task Deep Residual Echo Suppression with Echo-Aware Loss
  170. Multi-Task Gaussian Process Regression for the Detection of Sleep Cycles in Premature Infants
  171. Multi-Task Learning Improves Synthetic Speech Detection
  172. Multi-Task Learning Improves the Brain Stoke Lesion Segmentation
  173. Multi-Task RNN-T with Semantic Decoder for Streamable Spoken Language Understanding
  174. Multi-Task Voice Activated Framework Using Self-Supervised Learning
  175. Multi-Task fMRI Data Fusion Using IVA and PARAFAC2
  176. Multi-Turn Incomplete Utterance Restoration As Object Detection
  177. Multi-Turn RNN-T for Streaming Recognition of Multi-Party Speech
  178. Multi-View And Multi-Modal Event Detection Utilizing Transformer-Based Multi-Sensor Fusion
  179. Multi-View Data Representation Via Deep Autoencoder-Like Nonnegative Matrix Factorization
  180. Multi-View Information Bottleneck Without Variational Approximation
  181. Multi-View Learning Based on Non-Redundant Fusion for Icu Patient Mortality Prediction
  182. Multi-View Self-Attention Based Transformer for Speaker Recognition
  183. Multiband Image Fusion with Controllable Error Guarantees
  184. Multichannel Noise Reduction Using Dilated Multichannel U-Net and Pre-Trained Single-Channel Network
  185. Multichannel Speech Enhancement Without Beamforming
  186. Multilingual Second-Pass Rescoring for Automatic Speech Recognition Systems
  187. Multilingual Text-To-Speech Training Using Cross Language Voice Conversion And Self-Supervised Learning Of Speech Representations
  188. Multimodal Depression Classification using Articulatory Coordination Features and Hierarchical Attention Based text Embeddings
  189. Multimodal Emotion Recognition with Surgical and Fabric Masks
  190. Multimodal Evaluation Method for Sound Event Detection
  191. Multimodal Graph Signal Denoising Via Twofold Graph Smoothness Regularization with Deep Algorithm Unrolling
  192. Multimodal Sentiment Analysis on Unaligned Sequences Via Holographic Embedding
  193. Multimodal Transformer with Learnable Frontend and Self Attention for Emotion Recognition
  194. Multiple Instance Learning with Task-Specific Multi-Level Features for Weakly Annotated Histopathological Image Classification
  195. Multiple Kernel K-Means Clustering with Simultaneous Spectral Rotation
  196. Multiple Offsets Multilateration: A New Paradigm for Sensor Network Calibration with Unsynchronized Reference Nodes
  197. Multiple Patch-Aware Network for Faster Real-World Image Dehazing
  198. Multiple Temporal Context Embedding Networks for Unsupervised time Series Anomaly Detection
  199. Multiplication-Avoiding Variant of Power Iteration with Applications
  200. Multiscale Attention Aggregation Network for 2D Vessel Segmentation
  201. Multiscale Crowd Counting and Localization By Multitask Point Supervision
  202. Multistream Neural Architectures for Cued Speech Recognition Using a Pre-Trained Visual Feature Extractor and Constrained CTC Decoding
  203. Multisv: Dataset for Far-Field Multi-Channel Speaker Verification
  204. Multitask Gaussian Process With Hierarchical Latent Interactions
  205. Multitask Sparse Neural Network for Hyperspectral Image Denoising
  206. Multivariate Multiscale Cosine Similarity Entropy
  207. Multiview Long-Short Spatial Contrastive Learning For 3D Medical Image Analysis
  208. Music Enhancement via Image Translation and Vocoding
  209. Music Identification Using Brain Responses to Initial Snippets
  210. Music Phrase Inpainting Using Long-Term Representation and Contrastive Loss
  211. Music Source Separation With Deep Equilibrium Models
  212. Musicyolo: A Sight-Singing Onset/Offset Detection Framework Based on Object Detection Instead of Spectrum Frames
  213. NEX+: Novel View Synthesis with Neural Regularisation Over Multi-Plane Images
  214. NFT-K: Non-Fungible Tangent Kernels
  215. NN3A: Neural Network Supported Acoustic Echo Cancellation, Noise Suppression and Automatic Gain Control for Real-Time Communications
  216. NVC-Net: End-To-End Adversarial Voice Conversion
  217. Natural-Looking Adversarial Examples from Freehand Sketches
  218. Navigating Audio-Visual Event Detection Across Mismatched Modalities
  219. Nearest Subspace Search in The Signed Cumulative Distribution Transform Space For 1d Signal Classification
  220. Neartracker: Acoustic 2-D Target Tracking with Nearby Reflector in Siso System
  221. Neighbor-Augmented Transformer-Based Embedding for Retrieval
  222. Netrca: An Effective Network Fault Cause Localization Algorithm
  223. Neufa: Neural Network Based End-to-End Forced Alignment with Bidirectional Attention Mechanism
  224. Neural Architecture Search for Speech Emotion Recognition
  225. Neural Audio-To-Score Music Transcription For Unconstrained Polyphony Using Compact Output Representations
  226. Neural Cascade Architecture for Joint Acoustic Echo and Noise Suppression
  227. Neural Collapse in Deep Homogeneous Classifiers and The Role of Weight Decay
  228. Neural Grapheme-To-Phoneme Conversion with Pre-Trained Grapheme Models
  229. Neural HMMS Are All You Need (For High-Quality Attention-Free TTS)
  230. Neural Network-Based Compression Framework for DOA Estimation Exploiting Distributed Array
  231. Neural Speech Synthesis on a Shoestring: Improving the Efficiency of Lpcnet
  232. Neural-FST Class Language Model for End-to-End Speech Recognition
  233. New Improved Criterion for Model Selection in Sparse High-Dimensional Linear Regression Models
  234. News Recommendation Via Multi-Interest News Sequence Modelling
  235. No More Than 6ft Apart: Robust K-Means via Radius Upper Bounds
  236. No-Reference Quality Assessment of Variable Frame-Rate Videos Using Temporal Bandpass Statistics
  237. Node Slicing Broad Learning System for Text Classification
  238. Node-Screening Tests For The L0-Penalized Least-Squares Problem
  239. Noise Suppression for Improved Few-Shot Learning
  240. Noise-Robust Speech Recognition With 10 Minutes Unparalleled In-Domain Data
  241. Non-Autoregressive ASR with Self-Conditioned Folded Encoders
  242. Non-Autoregressive End-To-End Automatic Speech Recognition Incorporating Downstream Natural Language Processing
  243. Non-Autoregressive Transformer with Unified Bidirectional Decoder for Automatic Speech Recognition
  244. Non-Invasive Blood Pressure Monitoring with Multi-Modal In-Ear Sensing
  245. Non-Rigid Transformation Based Adversarial Attack Against 3d Object Tracking
  246. Nonlinear Signal Decomposition Based on Block Sparse Approximation
  247. Nonverbal Sound Detection for Disordered Speech
  248. Not All Features are Equal: Selection of Robust Features for Speech Emotion Recognition in Noisy Environments
  249. Novel Class Discovery: A Dependency Approach
  250. Novel Instance Mining with Pseudo-Margin Evaluation for Few-Shot Object Detection
  251. OPTE: Online Per-Title Encoding for Live Video Streaming
  252. ORCA-PARTY: An Automatic Killer Whale Sound Type Separation Toolkit Using Deep Learning
  253. OT Cleaner: Label Correction as Optimal Transport
  254. Object Detection and Tracking in Ultrasound Scans Using an Optical Flow and Semantic Segmentation Framework Based on Convolutional Neural Networks
  255. Object-Oriented Backdoor Attack Against Image Captioning
  256. Occluded Person Re-Identification Via Relational Adaptive Feature Correction Learning
  257. Off-The-Grid Covariance-Based Super-Resolution Fluctuation Microscopy
  258. Off-the-Shelf Deep Integration For Residual-Echo Suppression
  259. Omni-Sparsity DNN: Fast Sparsity Optimization for On-Device Streaming E2E ASR Via Supernet
  260. On Adversarial Robustness Of Large-Scale Audio Visual Learning
  261. On Continuous-Domain Inverse Problems with Sparse Superpositions of Decaying Sinusoids as Solutions
  262. On Federated Learning with Energy Harvesting Clients
  263. On Identifiable Polytope Characterization for Polytopic Matrix Factorization
  264. On Language Model Integration for RNN Transducer Based Speech Recognition
  265. On Loss Functions and Evaluation Metrics for Music Source Separation
  266. On Mini-Batch Training with Varying Length Time Series
  267. On Spectral and Temporal Sparsification of Speech Signals for the Improvement of Speech Perception in CI Listeners
  268. On Submodular Set Cover Problems for Near-Optimal Online Kernel Basis Selection
  269. On Synchronization of Wireless Acoustic Sensor Networks in the Presence of Time-Varying Sampling Rate Offsets and Speaker Changes
  270. On The Convergence of ADAM-Type Algorithms for Solving Structured Single Node and Decentralized Min-Max Saddle Point Games
  271. On The Effectiveness of Active Learning by Uncertainty Sampling in Classification of High-Dimensional Gaussian Mixture Data
  272. On The Impact of Normalization Strategies in Unsupervised Adversarial Domain Adaptation for Acoustic Scene Classification
  273. On The Observability in Visual Slam Networks
  274. On The Relaxation of Orthogonal Tensor Rank and Its Nonconvex Riemannian Optimization for Tensor Completion
  275. On the Acquisition of Stationary Signals Using Uniform ADCS
  276. On the False Alarm Probability of the Normalized Matched Filter for Off-Grid Target Detection
  277. On the Importance of Different Frequency Bins for Speaker Verification
  278. On the Interplay between Sparsity, Naturalness, Intelligibility, and Prosody in Speech Synthesis
  279. On the Potential of Spatially-Spread Orthogonal Time Frequency Space Modulation for ISAC Transmissions
  280. On the Prediction of the Frequency Response of a Wooden Plate from Its Mechanical Parameters
  281. On the Stability of Low Pass Graph Filter with a Large Number of Edge Rewires
  282. On the Use of Component Structural Characteristics for Voxel Segmentation in Semicon 3D Images
  283. On the Use of Geodesic Triangles between Gaussian Distributions for Classification Problems
  284. One Model to Enhance Them All: Array Geometry Agnostic Multi-Channel Personalized Speech Enhancement
  285. One TTS Alignment to Rule Them All
  286. One-Shot Voice Conversion For Style Transfer Based On Speaker Adaptation
  287. Online Continual Learning Using Enhanced Random Vector Functional Link Networks
  288. Online Detection of Scalp-Invisible Mesial-Temporal Brain Interictal Epileptiform Discharges from EEG
  289. Online Ecg Biometrics Via Hadamard Code
  290. Online Learning for Latent Yule-Simon Processes
  291. Online Learning with Probabilistic Feedback
  292. OpenFEAT: Improving Speaker Identification by Open-Set Few-Shot Embedding Adaptation with Transformer
  293. Operator Formulation for Linear Transformations and Signal Estimation in the Joint Spatial-Slepian Domain
  294. Optimal Combination Policies for Adaptive Social Learning
  295. Optimal Qos-Aware Network Slicing for Service-Oriented Networks with Flexible Routing
  296. Optimal Resource Allocation and Beamforming for Two-User Miso WPCNS for a Non-Linear Circuit-Based EH Model : (Invited Paper)
  297. Optimization Guarantees for ISTA and ADMM Based Unfolded Networks
  298. Optimization of Compressive Light Field Display in Dual-Guided Learning
  299. Optimization of a Fixed Virtual Sensing Feedback ANC Controller For In-Ear Headphones with Multiple Loudspeakers
  300. Optimize Wav2vec2s Architecture for Small Training Set Through Analyzing its Pre-Trained Models Attention Pattern
  301. Optimizing Alignment of Speech and Language Latent Spaces for End-To-End Speech Recognition and Understanding
  302. Optimizing Latent Space Directions for Gan-Based Local Image Editing
  303. Optimizing The Consumption Of Spiking Neural Networks With Activity Regularization
  304. Optm3sec: Optimizing Multicast Irs-Aided Multiantenna Dfrc Secrecy Channel With Multiple Eavesdroppers
  305. Orthogonal Nonnegative Matrix Tri-Factorization for Community Detection in Multiplex Networks
  306. Out-Of-Distribution As A Target Class in Semi-Supervised Learning
  307. Over-Parameterized Network Solves Phase Retrieval Effectively
  308. Over-the-Air Personalized Federated Learning
  309. PAMA-TTS: Progression-Aware Monotonic Attention for Stable SEQ2SEQ TTS with Accurate Phoneme Duration Control
  310. PDD-Net: A Precise Defect Detection Network Based on Point Set Representation
  311. PEAR: Photographic Embedding for Aesthetic Rating
  312. PGTRNET: Two-Phase Weakly Supervised Object Detection with Pseudo Ground Truth Refinement
  313. PMP-NET: Rethinking Visual Context for Scene Graph Generation
  314. POPO: Pessimistic Offline Policy Optimization
  315. PU-Refiner: A Geometry Refiner with Adversarial Learning for Point Cloud Upsampling
  316. PVAE-TTS: Adaptive Text-to-Speech via Progressive Style Adaptation
  317. PYXIS: An Open-Source Performance Dataset Of Sparse Accelerators
  318. Pair-Level Supervised Contrastive Learning for Natural Language Inference
  319. Panchromatic Imagery Copy-Paste Localization Through Data-Driven Sensor Attribution
  320. Parallel Composition of Weighted Finite-State Transducers
  321. Parameter Estimation in Sparse Inverse Problems Using Bernoulli-Gaussian Prior
  322. Parameter-Free Style Projection for Arbitrary Image Style Transfer
  323. Parametric Modeling of Human Wrist for Bioimpedance-Based Physiological Sensing
  324. Parametric Models for Doa Trajectory Localization
  325. Part-of-Speech Models Compression Methods for on-Device Grapheme-to-Phoneme Conversion
  326. Partial Arithmetic Consensus based Distributed Intensity Particle Flow SMC-PHD Filter for Multi-Target Tracking
  327. Partial Variable Training for Efficient on-Device Federated Learning
  328. Partially Fake Audio Detection by Self-Attention-Based Fake Span Discovery
  329. Partially Relaxed Orthogonal Least Squares Weighted Subspace Fitting Direction-of-Arrival Estimation
  330. Pas-Mef: Multi-Exposure Image Fusion Based On Principal Component Analysis, Adaptive Well-Exposedness And Saliency Map
  331. Passtrans: An Improved Password Reuse Model Based on Transformer
  332. Patch Steganalysis: A Sampling Based Defense Against Adversarial Steganography
  333. Path Signatures for Non-Intrusive Load Monitoring
  334. Peer Collaborative Learning for Polyphonic Sound Event Detection
  335. Perfect Reconstruction of Classes of Non-Bandlimited Signals from Projections with Unknown Angles
  336. Performance Optimization for Wireless Semantic Communications over Energy Harvesting Networks
  337. Performance-Efficiency Trade-Offs in Unsupervised Pre-Training for Speech Recognition
  338. Personalized Automatic Speech Recognition Trained on Small Disordered Speech Datasets
  339. Personalized Pagerank Graph Attention Networks
  340. Personalized speech enhancement: new models and Comprehensive evaluation
  341. Phase Continuity: Learning Derivatives of Phase Spectrum for Speech Enhancement
  342. Phase Control of Parametric Array Loudspeaker by Optimizing Sideband Weights
  343. Phase Shifted Bedrosian Filterbank: An Interpretable Audio Front-End for Time-Domain Audio Source Separation
  344. Phase-Only Reconfigurable Sparse Array Beamforming Using Deep Learning
  345. Phone-Informed Refinement of Synthesized Mel Spectrogram for Data Augmentation in Speech Recognition
  346. Phone-to-Audio Alignment without Text: A Semi-Supervised Approach
  347. Phoneme Mispronunciation Detection By Jointly Learning To Align
  348. Phonology Recognition in American Sign Language
  349. Phonotactic Language Recognition Using A Universal Phoneme Recognizer and A Transformer Architecture
  350. Photon-Limited Deblurring Using Algorithm Unrolling
  351. Physical Layer Anonymous Communications: An Anonymity Entropy Oriented Precoding Design (Invited Paper)
  352. Picknet: Real-Time Channel Selection for Ad Hoc Microphone Arrays
  353. Pixel-Level and Affinity-Level Knowledge Distillation for Unsupervised Segmentation of Covid-19 Lesions
  354. Pixinwav: Residual Steganography for Hiding Pixels in Audio
  355. Plug-and-Play and Relay Regularizations on Noisy Low Rank Tensor Completion for Snapshot Multispectral Image Restoration
  356. Point Cloud Attribute Compression Via Chroma Subsampling
  357. Point Cloud Denoising Using Normal Vector-Based Graph Wavelet Shrinkage
  358. Point-Mass Filter with Decomposition of Transient Density
  359. Polyphone Disambiguation and Accent Prediction Using Pre-Trained Language Models in Japanese TTS Front-End
  360. Polyphonic Audio Event Detection: Multi-Label or Multi-Class Multi-Task Classification Problem?
  361. Position-Invariant Adversarial Attacks on Neural Modulation Recognition
  362. PostGAN: A GAN-Based Post-Processor to Enhance the Quality of Coded Speech
  363. Power Allocation for Wireless Federated Learning Using Graph Neural Networks
  364. Power-Efficient Hybrid MIMO Receiver with Task-Specific Beamforming using Low-Resolution ADCs
  365. Predicting Flat-Fading Channels via Meta-Learned Closed-Form Linear Filters and Equilibrium Propagation
  366. Predicting Human Motion Using Key Subsequences
  367. Predicting the Generalization Gap in Deep Models using Anchoring
  368. Preliminary Results on the Generation of Artificial Handwriting Data Using a Decomposition-Recombination Strategy
  369. Preserving Trajectory Privacy in Driving Data Release
  370. Prime Knowledge with Local Pattern Consistency for Knowledge Distillation
  371. Prior-Bert and Multi-Task Learning for Target-Aspect-Sentiment Joint Detection
  372. Privacy Attacks for Automatic Speech Recognition Acoustic Models in A Federated Learning Framework
  373. Privacy Protection In Learning Fair Representations
  374. Privacy Sensitive Speech Analysis Using Federated Learning to Assess Depression
  375. Privacy-Aware Communication over a Wiretap Channel with Generative Networks
  376. Privacy-Enhancing Appliance Filtering For Smart Meters
  377. Privacy-Preserving Action Recognition
  378. Privacy-Preserving Distributed Expectation Maximization for Gaussian Mixture Model Using Subspace Perturbation
  379. Privacy-Preserving Federated Multi-Task Linear Regression: A One-Shot Linear Mixing Approach Inspired By Graph Regularization
  380. Private Learning Via Knowledge Transfer with High-Dimensional Targets
  381. Probabilistic Fine-Grained Urban Flow Inference with Normalizing Flows
  382. Probably Pleasant? A Neural-Probabilistic Approach to Automatic Masker Selection for Urban Soundscape Augmentation
  383. Progressive Continual Learning for Spoken Keyword Spotting
  384. Progressive Image Super-Resolution via Neural Differential Equation
  385. Progressive Multi-Stage Neural Audio Coding with Guided References
  386. Progressive Teacher-Student Training Framework for Music Tagging
  387. Progressive-Granularity Retrieval Via Hierarchical Feature Alignment for Person Re-Identification
  388. Prosodyspeech: Towards Advanced Prosody Model for Neural Text-to-Speech
  389. Prosospeech: Enhancing Prosody with Quantized Vector Pre-Training in Text-To-Speech
  390. Prototype Learning for Interpretable Respiratory Sound Analysis
  391. Prototype-Based Inter-Camera Learning for Person Re-Identification
  392. Provable Sample Complexity Guarantees For Learning Of Continuous-Action Graphical Games With Nonparametric Utilities
  393. Provable Second-Order Riemannian Gauss-Newton Method for Low-Rank Tensor Estimation ‖
  394. Proximal-Based Adaptive Simulated Annealing for Global Optimization
  395. Pseudo Strong Labels for Large Scale Weakly Supervised Audio Tagging
  396. Pseudo-Interacting Guided Network for Few-Shot Segmentation
  397. Pseudo-Label Transfer from Frame-Level to Note-Level in a Teacher-Student Framework for Singing Transcription from Polyphonic Music
  398. Pseudo-Labeling for Massively Multilingual Speech Recognition
  399. Punctuation Prediction for Streaming On-Device Speech Recognition
  400. Pyramid Fusion Attention Network For Single Image Super-Resolution
  401. QA4QG: Using Question Answering to Constrain Multi-Hop Question Generation
  402. Qrelation: an Agent Relation-Based Approach for Multi-Agent Reinforcement Learning Value Function Factorization
  403. Quantifying Discriminability between NMF Bases
  404. Quantization-Aware Precoding For Mu-Mimo With Limited-Capacity Fronthaul
  405. Quantized Winograd Acceleration for CONV1D Equipped ASR Models on Mobile Devices
  406. Quantum Federated Learning with Quantum Data
  407. Quantum Long Short-Term Memory
  408. Quickest Detection of Composite and Non-Stationary Changes with Application to Pandemic Monitoring
  409. RCANet: Row-Column Attention Network for Semantic Segmentation
  410. RIS-Aided Monostatic Mimo Radar with Co-Located Antennas
  411. RTSNet: Deep Learning Aided Kalman Smoothing
  412. Randomized Smoothing Under Attack: How Good is it in Practice?
  413. Rangeinet: Fast Lidar Point Cloud Temporal Interpolation
  414. Rank-Based Loss For Learning Hierarchical Representations
  415. Rate Coding Or Direct Coding: Which One Is Better For Accurate, Robust, And Energy-Efficient Spiking Neural Networks?
  416. Rate Control for Learned Video Compression
  417. Rational Arrays for DOA Estimation
  418. Raw Plenoptic Video Coding Under Hexagonal Lattice Resolution of Motion Vectors
  419. Raw Source and Filter Modelling for Dysarthric Speech Recognition
  420. RawNeXt: Speaker Verification System For Variable-Duration Utterances With Deep Layer Aggregation And Extended Dynamic Scaling Policies
  421. Rawboost: A Raw Data Boosting and Augmentation Method Applied to Automatic Speaker Verification Anti-Spoofing
  422. Real Additive Margin Softmax for Speaker Verification
  423. Real-M: Towards Speech Separation on Real Mixtures
  424. Real-Time Fall Detection Using Mmwave Radar
  425. Real-World Adversarial Examples Via Makeup
  426. Real-World On-Board Uav Audio Data Set For Propeller Anomalies
  427. Realistic Monocular-To-3d Virtual Try-On Via Multi-Scale Characteristics Capture
  428. Recognition Of Silently Spoken Word From Eeg Signals Using Dense Attention Network (DAN)
  429. Recovery of Graph Signals From Sign Measurements
  430. Recovery of Noisy Pooled Tests via Learned Factor Graphs with Application to COVID-19 Testing
  431. Recurrent Design of Probing Waveform for Sparse Bayesian Learning Based DOA Estimation
  432. Referee: Towards Reference-Free Cross-Speaker Style Transfer with Low-Quality Data for Expressive Speech Synthesis
  433. Reference Microphone Selection and Low-Rank Approximation Based Multichannel Wiener Filter with Application to Speech Recognition
  434. Reformulating Speaker Diarization As Community Detection With Emphasis On Topological Structure
  435. Region-to-Region Kernel Interpolation of Acoustic Transfer Function with Directional Weighting
  436. Regression Assisted Matrix Completion for Reconstructing a Propagation Field with Application to Source Localization
  437. Regularization Using Denoising: Exact and Robust Signal Recovery
  438. Regularized Latent Space Exploration for Discriminative Face Super-Resolution
  439. Relation Discovery in Nonlinearly Related Large-Scale Settings
  440. Relative Viewpoint Estimation Based on Structured 3d Representation Alignment
  441. Remix-Cycle-Consistent Learning on Adversarially Learned Separator for Accurate and Stable Unsupervised Speech Separation
  442. Repeat after Me: Self-Supervised Learning of Acoustic-to-Articulatory Mapping by Vocal Imitation
  443. Repetition Assessment for Speech and Language Disorders: A Study of the Logopenic Variant of Primary Progressive Aphasia
  444. Representation Learning Through Cross-Modal Conditional Teacher-Student Training For Speech Emotion Recognition
  445. RescoreBERT: Discriminative Speech Recognition Rescoring With Bert
  446. Residual Recovery Algorithm for Modulo Sampling
  447. Residual-Guided Personalized Speech Synthesis based on Face Image
  448. Restless Multi-Armed Bandits under Exogenous Global Markov Process
  449. Rethinking Computer-Aided Pelvis Segmentation
  450. Rethinking Two-B-Real Net for Real-Time Salient Object Detection
  451. Retrieval Bias Aware Ensemble Model for Conditional Sentence Generation
  452. Retrieval Enhanced Segment Generation Neural Network for Task-Oriented Dialogue Systems
  453. Retrieving Speaker Information from Personalized Acoustic Models for Speech Recognition
  454. Robust Adaptive Beamforming Based on Power Method Processing and Spatial Spectrum Matching
  455. Robust Adaptive Beamforming Maximizing the Worst-Case SINR Over Distributional Uncertainty Sets for Random INC Matrix And Signal Steering Vector
  456. Robust Adaptive Noise Canceller Algorithm with Snr-Based Stepsize Control and Noise-Path Gain Compensation
  457. Robust Bayesian Reconstruction of Multispectral Single-Photon 3D Lidar Data with Non-Uniform Background
  458. Robust Classification with Flexible Discriminant Analysis in Heterogeneous Data
  459. Robust Collaborative Learning for Sequence Modelling
  460. Robust Disentangled Variational Speech Representation Learning for Zero-Shot Voice Conversion
  461. Robust High-Order Tensor Recovery Via Nonconvex Low-Rank Approximation
  462. Robust Nonparametric Distribution Forecast with Backtest-Based Bootstrap and Adaptive Residual Selection
  463. Robust Parameter Estimation Based on the K-Divergence
  464. Robust Pressure Matching with ATF Perturbation Constraints for Sound Field Control
  465. Robust Self-Supervised Speaker Representation Learning Via Instance Mix Regularization
  466. Robust Signal Processing Over Simplicial Complexes
  467. Robust Speaker Verification Using Population-Based Data Augmentation
  468. Robust Speaker Verification with Joint Self-Supervised and Supervised Learning
  469. Robust Thermal Infrared Pedestrian Detection By Associating Visible Pedestrian Knowledge
  470. Robust Unstructured Knowledge Access in Conversational Dialogue with ASR Errors
  471. Robust Video Hashing Based on Local Fluctuation Preserving for Tracking Deep Fake Videos
  472. Robust and Efficient Uncertainty Aware Biosignal Classification via Early Exit Ensembles
  473. Run-and-Back Stitch Search: Novel Block Synchronous Decoding For Streaming Encoder-Decoder ASR
  474. S-DCCRN: Super Wide Band DCCRN with Learnable Complex Feature for Speech Enhancement
  475. S2 Reducer: High-Performance Sparse Communication to Accelerate Distributed Deep Learning
  476. S3PRL-VC: Open-Source Voice Conversion Framework with Self-Supervised Speech Representations
  477. S3T: Self-Supervised Pre-Training with Swin Transformer For Music Classification
  478. SA-SDR: A Novel Loss Function for Separation of Meeting Style Data
  479. SADN: Learned Light Field Image Compression with Spatial-Angular Decorrelation
  480. SAGA: Self-Augmentation with Guided Attention for Representation Learning
  481. SALSA-Lite: A Fast and Effective Feature for Polyphonic Sound Event Localization and Detection with Microphone Arrays
  482. SDETR: Attention-Guided Salient Object Detection with Transformer
  483. SDNET: Lightweight Facial Expression Recognition For Sample Disequilibrium
  484. SDR - Medium Rare with Fast Computations
  485. SERAB: A Multi-Lingual Benchmark for Speech Emotion Recognition
  486. SIG-VC: A Speaker Information Guided Zero-Shot Voice Conversion System for Both Human Beings and Machines
  487. SLUE: New Benchmark Tasks For Spoken Language Understanding Evaluation on Natural Speech
  488. SODA: Self-Organizing Data Augmentation in Deep Neural Networks Application to Biomedical Image Segmentation Tasks
  489. SP Attack: Single-Perspective Attack for Generating Adversarial Omnidirectional Images
  490. SQAPP: No-Reference Speech Quality Assessment Via Pairwise Preference
  491. SRP-DNN: Learning Direct-Path Phase Difference for Multiple Moving Sound Source Localization
  492. SRU++: Pioneering Fast Recurrence with Attention for Speech Recognition
  493. SYNT++: Utilizing Imperfect Synthetic Data to Improve Speech Recognition
  494. Safari from Visual Signals: Recovering Volumetric 3d Shapes
  495. Safeguarding UAV Networks through Integrated Sensing, Jamming, and Communications
  496. Sain: Similarity-Aware Video Frame Interpolation
  497. Sampling Set Selection for Graph Signals under Arbitrary Signal Priors
  498. Sar-Shipnet: Sar-Ship Detection Neural Network via Bidirectional Coordinate Attention and Multi-Resolution Feature Fusion
  499. Scalable Data Association and Multi-Target Tracking Under a Poisson Mixture Measurement Process
  500. Scalable Neural Architectures for End-to-End Environmental Sound Classification
  501. Scalable Ridge Leverage Score Sampling for the Nyström Method
  502. Scattering Statistics of Generalized Spatial Poisson Point Processes
  503. Score Difficulty Analysis for Piano Performance Education based on Fingering
  504. Screen & Relax: Accelerating The Resolution Of Elastic-Net By Safe Identification of The Solution Support
  505. SecMPNN: 3-Party Privacy-Preserving Molecular Structure Properties Inference
  506. Seed: Sound Event Early Detection Via Evidential Uncertainty
  507. SegNet-Based Deep Representation Learning for Dysphagia Classification
  508. Seismic Fault Identification Using Graph High-Frequency Components as Input to Graph Convolutional Network
  509. Selective Multi-Task Learning For Speech Emotion Recognition Using Corpora Of Different Styles
  510. Selective Mutual Learning: An Efficient Approach for Single Channel Speech Separation
  511. Selective Scale Cascade Attention Network for Breast Cancer Histopathology Image Classification
  512. Self Supervised Representation Learning with Deep Clustering for Acoustic Unit Discovery from Raw Speech
  513. Self-Attention for Incomplete Utterance Rewriting
  514. Self-Critical Sequence Training for Automatic Speech Recognition
  515. Self-Ensemble Variance Regularization for Domain Adaptation
  516. Self-Knowledge Distillation based Self-Supervised Learning for Covid-19 Detection from Chest X-Ray Images
  517. Self-Knowledge Distillation via Feature Enhancement for Speaker Verification
  518. Self-Learned Video Super-Resolution with Augmented Spatial and Temporal Context
  519. Self-Supervised Acoustic Anomaly Detection Via Contrastive Learning
  520. Self-Supervised Contrastive Learning for Cross-Domain Hyperspectral Image Representation
  521. Self-Supervised Learning Method Using Multiple Sampling Strategies for General-Purpose Audio Representation
  522. Self-Supervised Learning for Sentiment Analysis via Image-Text Matching
  523. Self-Supervised Learning on A Lightweight Low-Light Image Enhancement Model with Curve Refinement
  524. Self-Supervised Representation Learning for Unsupervised Anomalous Sound Detection Under Domain Shift
  525. Self-Supervised Speaker Recognition Training using Human-Machine Dialogues
  526. Self-Supervised Speaker Recognition with Loss-Gated Learning
  527. Self-Supervised Speaker Verification with Simple Siamese Network and Self-Supervised Regularization
  528. Semantic Association Network for Video Corpus Moment Retrieval
  529. Semantically Proportional Patchmix for Few-Shot Learning
  530. Semi-Supervised 360° Depth Estimation from Multiple Fisheye Cameras with Pixel-Level Selective Loss
  531. Semi-Supervised Gaussian Mixture Variational Autoencoder for Pulse Shape Discrimination
  532. Semi-Supervised Source Localization With Residual Physical Learning
  533. Semi-Supervised Standardized Detection of Periodic Signals with Application to Exoplanet Detection
  534. Semidefinite Relaxation Method for Moving Object Localization Using a Stationary Transmitter at Unknown Position
  535. Sensing-Assisted Beam Tracking in V2I Networks: Extended Target Case
  536. Sensors to Sign Language: A Natural Approach to Equitable Communication
  537. Sentiment-Aware Automatic Speech Recognition Pre-Training for Enhanced Speech Emotion Recognition
  538. Sentiment-Aware Distillation for Bitcoin Trend Forecasting Under Partial Observability
  539. Sequence Transduction with Graph-Based Supervision
  540. Sequential MCMC Methods for Audio Signal Enhancement
  541. Shared Transformer Encoder with Mask-Based 3d Model Estimation for Container Mass Estimation
  542. Short-and-Sparse Deconvolution Via Rank-One Constrained Optimization (Roco)
  543. Signal Compression via Neural Implicit Representations
  544. Signal Processing On Cell Complexes
  545. Signal Recovery from Inconsistent Nonlinear Observations
  546. Simple Attention Module Based Speaker Verification with Iterative Noisy Label Detection
  547. Simpler is Better: Spectral Regularization and Up-Sampling Techniques for Variational Autoencoders
  548. Simplicial Convolutional Neural Networks
  549. Simulation-and-Mining: Towards Accurate Source-Free Unsupervised Domain Adaptive Object Detection
  550. Simultaneous Nonlocal Low-Rank And Deep Priors For Poisson Denoising
  551. Single Image De-Raining with High-Low Frequency Guidance
  552. Single-Shot Balanced Detector for Geospatial Object Detection
  553. Sketch Storytelling
  554. Sketched RT3D: How to Reconstruct Billions of Photons Per Second
  555. Skim: Skipping Memory Lstm for Low-Latency Real-Time Continuous Speech Separation
  556. SleepGAN: Towards Personalized Sleep Therapy Music
  557. Slim: Explicit Slot-Intent Mapping with Bert for Joint Multi-Intent Detection and Slot Filling
  558. Social Welfare Maximization in Cross-Silo Federated Learning
  559. Solving The Long-Tailed Problem Via Intra- And Inter-Category Balance
  560. Sound Event Detection Guided by Semantic Contexts of Scenes
  561. Source Separation By Steering Pretrained Music Models
  562. Spain-Net: Spatially-Informed Stereophonic Music Source Separation
  563. Sparse Adversarial Attack For Video Via Gradient-Based Keyframe Selection
  564. Sparse Array Source Enumeration Via Coarray Subspace Optimization
  565. Sparse Modeling of The Early Part of Noisy Room Impulse Responses with Sparse Bayesian Learning
  566. Sparse Multi-Reference Alignment: Sample Complexity and Computational Hardness
  567. Sparse Recovery of Acoustic Waves
  568. Sparse Self-Attention for Semi-Supervised Sound Event Detection
  569. Sparse Subspace Tracking in High Dimensions
  570. Sparse-Group Log-Sum Penalized Graphical Model Learning For Time Series
  571. SparseBFA: Attacking Sparse Deep Neural Networks with the Worst-Case Bit Flips on Coordinates
  572. Sparsity Improves Unsupervised Attribute Discovery in Stylegan
  573. Sparsity-Based Sound Field Separation in the Spherical Harmonics Domain
  574. Spatial Active Noise Control Based on Individual Kernel Interpolation of Primary and Secondary Sound Fields
  575. Spatial Active Noise Control with the Remote Microphone Technique: an Approach with a Moving Higher Order Microphone
  576. Spatial Data Augmentation with Simulated Room Impulse Responses for Sound Event Localization and Detection
  577. Spatial Mixup: Directional Loudness Modification as Data Augmentation for Sound Event Localization and Detection
  578. Spatial Processing Front-End for Distant ASR Exploiting Self-Attention Channel Combinator
  579. Spatial-Context-Aware Deep Neural Network for Multi-Class Image Classification
  580. Spatial-Temporal Graph Convolution Network for Multichannel Speech Enhancement
  581. Spatio-Temporal Attention Graph Convolution Network for Functional Connectome Classification
  582. Spatio-Temporal Graph Complementary Scattering Networks
  583. Spatio-Temporal Graph Convolutional Networks for Continuous Sign Language Recognition
  584. Spatio-Temporal Motion Aggregation Network for Video Action Detection
  585. Spatio-Temporal PRRS Epidemic Forecasting via Factorized Deep Generative Modeling
  586. Speaker Embedding Conversion for Backward and Cross-Channel Compatibility
  587. Speaker Generation
  588. Speaker Identity Preservation in Dysarthric Speech Reconstruction by Adversarial Speaker Adaptation
  589. Speaker Normalization for Self-Supervised Speech Emotion Recognition
  590. Speaker Reinforcement Using Target Source Extraction for Robust Automatic Speech Recognition
  591. Speaker-Targeted Audio-Visual Speech Recognition Using a Hybrid CTC/Attention Model with Interference Loss
  592. Specialised Video Quality Model For Enhanced User Generated Content (UGC) With Special Effects
  593. Spectral Permutation Test on Persistence Diagrams
  594. Spectral-Spatial Symmetrical Aggregation Cross-Linking Multi-Modal Data Fusion Network
  595. Speech Denoising in the Waveform Domain With Self-Attention
  596. Speech Emotion Recognition Using Self-Supervised Features
  597. Speech Emotion Recognition with Co-Attention Based Multi-Level Acoustic Information
  598. Speech Emotion Recognition with Global-Aware Fusion on Multi-Scale Feature Representation
  599. Speech Enhancement for Low Bit Rate Speech Codec
  600. Speech Enhancement with Neural Homomorphic Synthesis
  601. Speech Pattern Based Black-Box Model Watermarking for Automatic Speech Recognition
  602. Speech Recognition Using Biologically-Inspired Neural Networks
  603. Speech Recovery For Real-World Self-Powered Intermittent Devices
  604. Speech Tasks Relevant to Sleepiness Determined With Deep Transfer Learning
  605. SpeechSplit2.0: Unsupervised Speech Disentanglement for Voice Conversion without Tuning Autoencoder Bottlenecks
  606. Speechmoe2: Mixture-of-Experts Model with Improved Routing
  607. Spell My Name: Keyword Boosted Speech Recognition
  608. Spherical Convolutional Recurrent Neural Network for Real-Time Sound Source Tracking
  609. Spoken Language Recognition with Cluster-Based Modeling
  610. Stability Analysis of Unfolded WMMSE for Power Allocation
  611. Stability of Neural Networks on Manifolds to Relative Perturbations
  612. Stable and Transferable Wireless Resource Allocation Policies Via Manifold Neural Networks
  613. Stacked Multi-Scale Attention Network for Image Colorization
  614. Statistical Pyramid Dense Time Delay Neural Network for Speaker Verification
  615. Statistical, Spectral and Graph Representations for Video-Based Facial Expression Recognition in Children
  616. Stealthy Backdoor Attack with Adversarial Training
  617. Stgat-Mad : Spatial-Temporal Graph Attention Network For Multivariate Time Series Anomaly Detection
  618. Stpointgcn: Spatial Temporal Graph Convolutional Network for Multiple People Recognition Using Millimeter-Wave Radar
  619. Streaming Transformer Transducer based Speech Recognition Using Non-Causal Convolution
  620. Streaming on-Device Detection of Device Directed Speech from Voice and Touch-Based Invocation
  621. Structural Prior Models for 3-D Deep Vessel Segmentation
  622. Study of Positional Encoding Approaches for Audio Spectrogram Transformers
  623. Study of the Null Directions on The Performance of Differential Beamformers
  624. Study on Time-of-Flight Estimation in Ultrasonic Well Logging Tool: Model-Driven Transfer Learning
  625. Studying Three Families of Divergences to Compare Wide-Sense Stationary Gaussian Arma Processes
  626. Stylegan-Induced Data-Driven Regularization for Inverse Problems
  627. Subgraph Representation Learning with Hard Negative Samples for Inductive Link Prediction
  628. Subjective And Objective Quality Assessment Of Mobile Gaming Video
  629. Subspace Clustering Using Unsupervised Data Augmentation
  630. Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
  631. Super-Resolution of Satellite Images by two-Dimensional RRDB and Edge-Enhancement Generative Adversarial Network
  632. Superresolution and Segmentation of OCT Scans Using Multi-Stage Adversarial Guided Attention Training
  633. Supervised Attention in Sequence-to-Sequence Models for Speech Recognition
  634. Supervised Learning Based Sparse Channel Estimation For RIS Aided Communications
  635. Supervised Training of Siamese Spiking Neural Networks with Earth Mover's Distance
  636. Supervised and Self-Supervised Pretraining Based Covid-19 Detection Using Acoustic Breathing/Cough/Speech Signals
  637. Symbol-Level Online Channel Tracking for Deep Receivers
  638. Synergistic Network Learning and Label Correction for Noise-Robust Image Classification
  639. Synpose: A Large-Scale and Densely Annotated Synthetic Dataset for Human Pose Estimation in Classroom
  640. Syntax-Based Graph Matching for Knowledge Base Question Answering
  641. Synthesis of Adversarial Samples in Two-Stage Classifiers
  642. Synthesizing Dysarthric Speech Using Multi-Speaker Tts For Dysarthric Speech Recognition
  643. T-NGA: Temporal Network Grafting Algorithm for Learning to Process Spiking Audio Sensor Events
  644. T-SVD Based Broadband Non-Synchronous Measurements
  645. TCRNet: Make Transformer, CNN and RNN Complement Each Other
  646. TEA-PSE: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System for ICASSP 2022 DNS Challenge
  647. TED Talk Teaser Generation with Pre-Trained Models
  648. TFPSNet: Time-Frequency Domain Path Scanning Network for Speech Separation
  649. TH-Net: A Method Of Single 3d Object Tracking Based On Transformers And Hausdorff Distance
  650. TINYS2I: A Small-Footprint Utterance Classification Model with Contextual Support for On-Device SLU
  651. TNTC: Two-Stream Network with Transformer-Based Complementarity for Gait-Based Emotion Recognition
  652. TP-VIT: A Two-Pathway Vision Transformer for Video Action Recognition
  653. TPARN: Triple-Path Attentive Recurrent Network for Time-Domain Multichannel Speech Enhancement
  654. Tackling Data Scarcity in Speech Translation Using Zero-Shot Multilingual Machine Translation Techniques
  655. Tackling the Score Shift in Cross-Lingual Speaker Verification by Exploiting Language Information
  656. TalkingFlow: Talking Facial Landmark Generation with Multi-Scale Normalizing Flow Network
  657. Target-Aware Auto-Augmentation for Unsupervised Domain Adaptive Object Detection
  658. TargetDrop: A Targeted Regularization Method for Convolutional Neural Networks
  659. Teaching CNNs to Mimic Human Visual Cognitive Process & Regularise Texture-Shape Bias
  660. Tempo: Improving Training Performance in Cross-Silo Federated Learning
  661. Temporal Contrastive-Loss for Audio Event Detection
  662. Temporal Cross-Graph Network for Brain Functional Activity Prediction
  663. Temporal Dynamic Convolutional Neural Network for Text-Independent Speaker Verification and Phonemic Analysis
  664. Temporal Early Exiting for Streaming Speech Commands Recognition
  665. Temporal Knowledge Distillation for on-device Audio Classification
  666. Tensor-Based Orthogonal Matching Pursuit with Phase Rotation for Channel Estimation In Hybrid Beamforming Mimo-Ofdm Systems
  667. Terahertz Image Restoration Benchmarking Dataset
  668. Test-Time Detection of Backdoor Triggers for Poisoned Deep Neural Networks
  669. Text Adaptive Detection for Customizable Keyword Spotting
  670. Text-Free Non-Parallel Many-To-Many Voice Conversion Using Normalising Flow
  671. Text-Image De-Contextualization Detection Using Vision-Language Models
  672. Text2Poster: Laying Out Stylized Texts on Retrieved Images
  673. Text2video: Text-Driven Talking-Head Video Synthesis with Personalized Phoneme - Pose Dictionary
  674. Texture Information Boosts Video Quality Assessment
  675. The CUHK-Tencent Speaker Diarization System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
  676. The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks
  677. The Coral++ Algorithm for Unsupervised Domain Adaptation of Speaker Recognition
  678. The DKU Audio-Visual Wake Word Spotting System for the 2021 MISP Challenge
  679. The Data/Identity Tradeoff with Censored Sensors
  680. The Dawn of Quantum Natural Language Processing
  681. The First Multimodal Information Based Speech Processing (Misp) Challenge: Data, Tasks, Baselines And Results
  682. The Impact of JPEG Compression on Prior Image Noise
  683. The Impact of Removing Head Movements on Audio-Visual Speech Enhancement
  684. The Mirrornet : Learning Audio Synthesizer Controls Inspired by Sensorimotor Interaction
  685. The PCG-AIID System for L3DAS22 Challenge: MIMO and MISO Convolutional Recurrent Network for Multi Channel Speech Enhancement and Speech Recognition
  686. The Prototype Co-Prime Array with a Robust Difference Co-Array
  687. The Representation Jensen-Rényi Divergence
  688. The Royalflush System of Speech Recognition for M2met Challenge
  689. The Second Dicova Challenge: Dataset and Performance Analysis for Diagnosis of Covid-19 Using Acoustics
  690. The Sjtu System For Multimodal Information Based Speech Processing Challenge 2021
  691. The USTC-Ximalaya System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription (M2met) Challenge
  692. The Vicomtech Audio Deepfake Detection System Based on Wav2vec2 for the 2022 ADD Challenge
  693. The Volcspeech System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
  694. The impact of cross language on acoustic-to-articulatory inversion and its influence on articulatory speech synthesis
  695. Thin Slices of Depression: Improving Depression Detection Performance Through Data Segmentation
  696. Threshold Independent Evaluation of Sound Event Detection Scores
  697. Tie Your Embeddings Down: Cross-Modal Latent Spaces for End-to-end Spoken Language Understanding
  698. Tight Integration Of Neural- And Clustering-Based Diarization Through Deep Unfolding Of Infinite Gaussian Mixture Model
  699. Time Domain Adversarial Voice Conversion for ADD 2022
  700. Time Domain Radial Filter Design for Spherical Waves
  701. Time-Balanced Focal Loss for Audio Event Detection
  702. Time-Domain Acoustic Contrast Control with A Spatial Uniformity Constraint for Personal Audio Systems
  703. Time-Domain Audio-Visual Speech Separation on Low Quality Videos
  704. Time-Frequency Attention for Monaural Speech Enhancement
  705. Time-Frequency and Geometric Analysis of Task-Dependent Learning in Raw Waveform Based Acoustic Models
  706. TitaNet: Neural Model for Speaker Representation with 1D Depth-Wise Separable Convolutions and Global Context
  707. To Catch A Chorus, Verse, Intro, or Anything Else: Analyzing a Song with Structural Functions
  708. Tonet: Tone-Octave Network for Singing Melody Extraction from Polyphonic Music
  709. Topological Correlation of Brain Signals
  710. Torchaudio: Building Blocks for Audio and Speech Processing
  711. Toward Degradation-Robust Voice Conversion
  712. Toward mmWave-Based Sound Enhancement and Separation
  713. Towards A Common Speech Analysis Engine
  714. Towards Accurate Cross-Domain in-Bed Human Pose Estimation
  715. Towards Automatic Transcription of Polyphonic Electric Guitar Music: A New Dataset and a Multi-Loss Transformer Model
  716. Towards Better Meta-Initialization with Task Augmentation for Kindergarten-Aged Speech Recognition
  717. Towards Closed-Loop Speech Synthesis from Stereotactic EEG: A Unit Selection Approach
  718. Towards Controllable and Physical Interpretable Underwater Scene Simulation
  719. Towards End-to-End Integration of Dialog History for Improved Spoken Language Understanding
  720. Towards Expressive Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis
  721. Towards Fast And Convenient End-To-End HRTF Personalization
  722. Towards Faster Continuous Multi-Channel HRTF Measurements Based On Learning System Models
  723. Towards Identity Preserving Normal to Dysarthric Voice Conversion
  724. Towards Interpretability of Speech Pause in Dementia Detection Using Adversarial Learning
  725. Towards Interpreting Deep Learning Models to Understand Loss of Speech Intelligibility in Speech Disorders Step 2: Contribution of the Emergence of Phonetic Traits
  726. Towards Joint Frame-Level and MOS Quality Predictions with Low-Complexity Objective Models
  727. Towards Learning Universal Audio Representations
  728. Towards Lifelong Learning of Multilingual Text-to-Speech Synthesis
  729. Towards Lightweight Applications: Asymmetric Enroll-Verify Structure for Speaker Verification
  730. Towards Low-Distortion Multi-Channel Speech Enhancement: The ESPNET-Se Submission to the L3DAS22 Challenge
  731. Towards Measuring Fairness in Speech Recognition: Casual Conversations Dataset Transcriptions
  732. Towards Practical and Efficient Long Video Summary
  733. Towards Reducing the Need for Speech Training Data to Build Spoken Language Understanding Systems
  734. Towards Robust Speech-to-Text Adversarial Attack
  735. Towards Robust Visual Transformer Networks via K-Sparse Attention
  736. Towards Speaker Age Estimation With Label Distribution Learning
  737. Towards Transferable Speech Emotion Representation: On Loss Functions for Cross-Lingual Latent Representations
  738. Towards Using Clothes Style Transfer for Scenario-Aware Person Video Generation
  739. Towards end-to-end Speaker Diarization with Generalized Neural Speaker Clustering
  740. Tracking the Dimensions of Latent Spaces of Gaussian Process Latent Variable Models
  741. Training Privacy-Preserving Video Analytics Pipelines by Suppressing Features That Reveal Information About Private Attributes
  742. Training Robust Zero-Shot Voice Conversion Models with Self-Supervised Features
  743. Training Stable Graph Neural Networks Through Constrained Learning
  744. Training Strategies for Automatic Song Writing: A Unified Framework Perspective
  745. Training Strategies for Improved Lip-Reading
  746. Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers Using End-to-End Speaker-Attributed ASR
  747. Transducer-Based Streaming Deliberation for Cascaded Encoders
  748. Transductive Clip with Class-Conditional Contrastive Learning
  749. Transformer-Based Domain Adaptation for Event Data Classification
  750. Transformer-Based Estimation of Spoken Sentences Using Electrocorticography
  751. Transformer-Based Multi-Aspect Multi-Granularity Non-Native English Speaker Pronunciation Assessment
  752. Transformer-Based Person Search Model with Symmetric Online Instance Matching
  753. Transformer-Based Streaming ASR with Cumulative Attention
  754. Transformer-S2A: Robust and Efficient Speech-to-Animation
  755. Transient Analysis of Clustered Multitask Diffusion RLS Algorithm
  756. Transient Detection with Unknown Statistics Via Source Coding
  757. Transmit Beamforming with Fixed Covariance for Integrated MIMO Radar and Multiuser Communications
  758. Transtl: Spatial-Temporal Localization Transformer for Multi-Label Video Classification
  759. TriBYOL: Triplet BYOL for Self-Supervised Representation Learning
  760. Tts4pretrain 2.0: Advancing the use of Text and Speech in ASR Pretraining with Consistency and Contrastive Losses
  761. Tunet: A Block-Online Bandwidth Extension Model Based On Transformers And Self-Supervised Pretraining
  762. Turn-to-Diarize: Online Speaker Diarization Constrained by Transformer Transducer Speaker Turn Detection
  763. Two Strategies Toward Lightweight Image Super-Resolution
  764. Two-Path GMM-ResNet and GMM-SENet for ASV Spoofing Detection
  765. Two-Snapshot DOA Estimation Via Hankel-Structured Matrix Completion
  766. Type-Aware Medical Visual Question Answering
  767. U-GAT-VC: Unsupervised Generative Attentional Networks for Non-Parallel Voice Conversion
  768. UNET-TTS: Improving Unseen Speaker and Style Transfer in One-Shot Voice Cloning
  769. Ubilung: Multi-Modal Passive-Based Lung Health Assessment
  770. Ubiquitous Physiological Prediction of SUD Patients' Wellness State Using Memory-Based Convolutional Models
  771. Uformer: A Unet Based Dilated Complex & Real Dual-Path Conformer Network for Simultaneous Speech Enhancement and Dereverberation
  772. Uncertainty Estimation with a VAE-Classifier Hybrid Model
  773. Uncertainty in Data-Driven Kalman Filtering for Partially Known State-Space Models
  774. Underdetermined Two-Dimensional Localization for Wideband Sources Based on Distributed Sensor Array Networks
  775. Underwater Image Enhancement Via Learning Water Type Desensitized Representations
  776. Underwater Small Target Detection Based on Deformable Convolutional Pyramid
  777. Underwater Stereo Matching Via Unsupervised Appearance And Feature Adaptation Networks
  778. Unfolding Model-Based Beamforming for High Quality Ultrasound Imaging
  779. Unified Matrix Coding for NN Originated MIP in H.266/VVC
  780. Unified Multimodal Punctuation Restoration Framework for Mixed-Modality Corpus
  781. Unified Speculation, Detection, and Verification Keyword Spotting
  782. Unimodular Waveform Design with Low Correlation Levels: A Fast Algorithm Development to Support Large-Scale Code Lengths
  783. Unispeech-Sat: Universal Speech Representation Learning With Speaker Aware Pre-Training
  784. Universal Efficient Variable-Rate Neural Image Compression
  785. Universal Paralinguistic Speech Representations Using self-Supervised Conformers
  786. Unlimited Sampling with Local Averages
  787. Unlimited Sampling with Sparse Outliers: Experiments with Impulsive and Jump or Reset Noise
  788. Unrolling Particles: Unsupervised Learning of Sampling Distributions
  789. Unsupervised Anomaly Detection for Container Cloud Via BILSTM-Based Variational Auto-Encoder
  790. Unsupervised Audio-Caption Aligning Learns Correspondences Between Individual Sound Events and Textual Phrases
  791. Unsupervised Clustering and Analysis of Contraction-Dependent Fetal Heart Rate Segments
  792. Unsupervised Contrastive Hashing for Cross-Modal Retrieval in Remote Sensing
  793. Unsupervised Data Selection for Speech Recognition with Contrastive Loss Ratios
  794. Unsupervised Deep Learning Network for Deformable Fundus Image Registration
  795. Unsupervised Hierarchical Translation-Based Model for Multi-Modal Medical Image Registration
  796. Unsupervised Model Adaptation for End-to-End ASR
  797. Unsupervised Speech Enhancement with Speech Recognition Embedding and Disentanglement Losses
  798. Unsupervised Word-Level Prosody Tagging for Controllable Speech Synthesis
  799. Unsupervised and Untrained Underwater Image Restoration Based on Physical Image Formation Model
  800. Upmixing Via Style Transfer: A Variational Autoencoder for Disentangling Spatial Images And Musical Content
  801. Urban Sound & Sight: Dataset And Benchmark For Audio-Visual Urban Scene Understanding
  802. User Scheduling Using Graph Neural Networks for Reconfigurable Intelligent Surface Assisted Multiuser Downlink Communications
  803. Using Acoustic Deep Neural Network Embeddings to Detect Multiple Sclerosis From Speech
  804. Using Multiple Reference Audios and Style Embedding Constraints for Speech Synthesis
  805. Using Spectral Sequence-to-Sequence Autoencoders to Assess Mild Cognitive Impairment
  806. Using a Single Input to Forecast Human Action Keystates in Everyday Pick and Place Actions
  807. Usted: Improving ASR with a Unified Speech and Text Encoder-Decoder
  808. VADOI: Voice-Activity-Detection Overlapping Inference for End-To-End Long-Form Speech Recognition
  809. VCD: View-Constraint Disentanglement for Action Recognition
  810. VCVTS: Multi-Speaker Video-to-Speech Synthesis Via Cross-Modal Knowledge Transfer from Voice Conversion
  811. VISinger: Variational Inference with Adversarial Learning for End-to-End Singing Voice Synthesis
  812. VQA-BC: Robust Visual Question Answering Via Bidirectional Chaining
  813. VR-FAM: Variance-Reduced Encoder with Nonlinear Transformation for Facial Attribute Manipulation
  814. VSEGAN: Visual Speech Enhancement Generative Adversarial Network
  815. VU-BERT: A Unified Framework for Visual Dialog
  816. VarArray: Array-Geometry-Agnostic Continuous Speech Separation
  817. Variable Span Trade-Off Filter for Sound Zone Control with Kernel Interpolation Weighting
  818. Variance Reduction-Boosted Byzantine Robustness in Decentralized Stochastic Optimization
  819. Varianceflow: High-Quality and Controllable Text-to-Speech using Variance Information via Normalizing Flow
  820. Variational Bayesian Framework for Advanced Image Generation with Domain-Related Variables
  821. Variational Bayesian Graph Convolutional Network for Robust Collaborative Filtering
  822. Variational Bayesian Tensor Networks with Structured Posteriors
  823. Video Anomaly Detection via Prediction Network with Enhanced Spatio-Temporal Memory Exchange
  824. Video Frame Interpolation via Local Lightweight Bidirectional Encoding with Channel Attention Cascade
  825. Violinist Identification Using Note-Level Timbre Feature Distributions
  826. Vision Transformer Equipped With Neural Resizer On Facial Expression Recognition Task
  827. Vision Transformer-Based Retina Vessel Segmentation with Deep Adaptive Gamma Correction
  828. Visual Representation Learning with Self-Supervised Attention for Low-Label High-Data Regime
  829. Visualtts: TTS with Accurate Lip-Speech Synchronization for Automatic Voice Over
  830. Vocalsound: A Dataset for Improving Human Vocal Sounds Recognition
  831. Vocbench: A Neural Vocoder Benchmark for Speech Synthesis
  832. Voice Filter: Few-Shot Text-to-Speech Speaker Adaptation Using Voice Conversion as a Post-Processing Module
  833. W-ART: Action Relation Transformer for Weakly-Supervised Temporal Action Localization
  834. WENETSPEECH: A 10000+ Hours Multi-Domain Mandarin Corpus for Speech Recognition
  835. WLS Design of Arma Graph Filters Using Iterative Second-Order Cone Programming
  836. Wasserstein Cross-Lingual Alignment For Named Entity Recognition
  837. Wassertrain: An Adversarial Training Framework Against Wasserstein Adversarial Attacks
  838. Watermarking Images in Self-Supervised Latent Spaces
  839. Wav2CLIP: Learning Robust Audio Representations from Clip
  840. Wav2vec-Switch: Contrastive Learning from Original-Noisy Speech Pairs for Robust Speech Recognition
  841. Wave-Domain Approach for Cancelling Noise Entering Open Windows
  842. Wavebender GAN: An Architecture for Phonetically Meaningful Speech Manipulation
  843. Waveform Optimization for Wireless Power Transfer with Power Amplifier and Energy Harvester Non-linearities
  844. Wavelet-Based Unsupervised Label-to-Image Translation
  845. Weak Target Detection in Massive MIMO Radar via an Improved Reinforcement Learning Approach
  846. Weakly Supervised Point Cloud Upsampling VIA Optimal Transport
  847. Wearable Seld Dataset: Dataset For Sound Event Localization And Detection Using Wearable Devices Around Head
  848. Weighted Graph Embedded Low-Rank Projection Learning for Feature Extraction
  849. Weighted Wavelet-Based Spectral-Spatial Transforms For CFA-Sampled Raw Camera Image Compression Considering Image Features
  850. What Is The Patient Looking At? Robust Gaze-Scene Intersection Under Free-Viewing Conditions
  851. When BERT Meets Quantum Temporal Convolution Learning for Text Classification in Heterogeneous Computing
  852. When Does Backdoor Attack Succeed in Image Reconstruction? A Study of Heuristics vs. Bi-Level Solution
  853. Wide-Sense Stationarity and Spectral Estimation for Generalized Graph Signal
  854. Wikitag: Wikipedia-Based Knowledge Embeddings Towards Improved Acoustic Event Classification
  855. Win The Lottery Ticket Via Fourier Analysis: Frequencies Guided Network Pruning
  856. Wishart Localization Prior On Spatial Covariance Matrix In Ambisonic Source Separation Using Non-Negative Tensor Factorization
  857. Wlinker: Modeling Relational Triplet Extraction As Word Linking
  858. Word Order does not Matter for Speech Recognition
  859. WordMarkov: A New Password Probability Model of Semantics
  860. Zero-Shot Cross-Lingual Transfer Using Multi-Stream Encoder and Efficient Speaker Representation
  861. Zeroth-Order Randomized Subspace Newton Methods
  862. nnSpeech: Speaker-Guided Conditional Variational Autoencoder for Zero-Shot Multi-speaker text-to-speech
  863. r-G2P: Evaluating and Enhancing Robustness of Grapheme to Phoneme Conversion by Controlled Noise Introducing and Contextual Information Incorporation
  864. r-Local Unlabeled Sensing: Improved Algorithm and Applications

Looking for submission deadlines instead? See the conference deadline calendar.