ICASSP 2023 Accepted Papers
The full list of 2,718 papers accepted at ICASSP 2023 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- Flowgrad: Using Motion for Visual Sound Source Localization
- Flowpose: Conditional Normalizing Flows for 3D Human Pose and Shape Estimation from Monocular Videos
- Flowreg: Latent Space Regularization Using Normalizing Flow For Limited Samples Learning
- Focusing on Targets for Improving Weakly Supervised Visual Grounding
- Forecasting of Breathing Events from Speech for Respiratory Support
- Forensics for Adversarial Machine Learning Through Attack Mapping Identification
- Frame-Level Multi-Label Playing Technique Detection Using Multi-Scale Network and Self-Attention Mechanism
- Frame-Wise and Overlap-Robust Speaker Embeddings for Meeting Diarization
- Framewise Multiple Sound Source Localization and Counting Using Binaural Spatial Audio Signals
- Framewise Wavegan: High Speed Adversarial Vocoder In Time Domain With Very Low Computational Complexity
- Free-View Expressive Talking Head Video Editing
- Freevc: Towards High-Quality Text-Free One-Shot Voice Conversion
- Frequency Bin-Wise Single Channel Speech Presence Probability Estimation Using Multiple DNNS
- Frequency Reciprocal Action and Fusion for Single Image Super-Resolution
- Frequency and Scale Perspectives of Feature Extraction
- Frequency-Aware Attentional Feature Fusion for Deepfake Detection
- Frequency-Selective Hybrid Beamforming For Mmwave Full-Duplex
- Fretnet: Continuous-Valued Pitch Contour Streaming For Polyphonic Guitar Tablature Transcription
- From Easy to Hard: Two-Stage Selector and Reader for Multi-Hop Question Answering
- From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech Recognition
- Front-End Adapter: Adapting Front-End Input of Speech Based Self-Supervised Learning for Speech Recognition
- Full-Band General Audio Synthesis with Score-Based Diffusion
- Fully Complex-Valued Deep Learning Model for Visual Perception
- Fully Distributed Federated Learning with Efficient Local Cooperations
- Fully Unsupervised Topic Clustering of Unlabelled Spoken Audio Using Self-Supervised Representation Learning and Topic Model
- G2CNN: Geometric Prior Based GCNN for Single-View 3D Reconstruction with Loop Subdivision
- G2PL: Lexicon Enhanced Chinese Polyphone Disambiguation Using Bert Adapter with a New Dataset
- GANStrument: Adversarial Instrument Sound Synthesis with Pitch-Invariant Instance Conditioning
- GCC-Speaker: Target Speaker Localization with Optimal Speaker-Dependent Weighting in Multi-Speaker Scenarios
- GOP-Based Latent Refinement for Learned Video Coding
- GSWIN: Gated MLP Vision Model with Hierarchical Structure of Shifted Window
- GTN-Bailando: Genre Consistent long-Term 3D Dance Generation Based on Pre-Trained Genre Token Network
- GaPP: Multi-Target Tracking with Gaussian Processes
- Gaitcotr: Improved Spatial-Temporal Representation for Gait Recognition with a Hybrid Convolution-Transformer Framework
- Gaitmixer: Skeleton-Based Gait Representation Learning Via Wide-Spectrum Multi-Axial Mixer
- Gated Contextual Adapters For Selective Contextual Biasing In Neural Transducers
- Gated Enhanced RPN and Hybrid-View for Few-Shot Object Detection
- Gator: Graph-Aware Transformer with Motion-Disentangled Regression for Human Mesh Recovery from a 2D Pose
- Gaussian Prior Reinforcement Learning for Nested Named Entity Recognition
- Gaussian Process Dynamical Modeling for Adaptive Inference Over Graphs
- Gaze Pre-Train For Improving Disparity Estimation Networks
- Gct: Gated Contextual Transformer for Sequential Audio Tagging
- Gender-Cartoon: Image Cartoonization Method Based on Gender Classification
- General Category Network: Handwritten Mathematical Expression Recognition with Coarse-Grained Recognition Task
- General or Specific? Investigating Effective Privacy Protection in Federated Learning for Speech Emotion Recognition
- Generalized Invariant Matching Property Via Lasso
- Generalized Relative Harmonic Coefficients
- Generalized Two-Stage Particle Filter for High Dimensions
- Generative Model based Highly Efficient Semantic Communication Approach for Image Transmission
- Generative Modeling Based Manifold Learning for Adaptive Filtering Guidance
- Generic Dependency Modeling for Multi-Party Conversation
- Geogcn: Geometric Dual-Domain Graph Convolution Network For Point Cloud Denoising
- Geometric Matrix Completion with Collaborative Routing Between Capsules
- Geometry-Aware DOA Estimation Using a Deep Neural Network with Mixed-Data Input Features
- Gesper: A Unified Framework for General Speech Restoration
- Glacier: Glass-Box Transformer for Interpretable Dynamic Neuroimaging
- Global HRTF Interpolation Via Learned Affine Transformation of Hyper-Conditioned Features
- Global Localisation in Continuous Magnetic Vector Fields Using Gaussian Processes
- Global Matching-Optimization Network for Stereo Depth Estimation
- Global and Nodal Mutual Information Maximization in Heterogeneous Graphs
- Global-Context Aware Generative Protein Design
- Gluformer: Transformer-based Personalized glucose Forecasting with uncertainty quantification
- Going in Style: Audio Backdoors Through Stylistic Transformations
- Good Neighbors are All You Need for Chinese Grapheme-To-Phoneme Conversion
- Grad-CAM-Inspired Interpretation of Nearfield Acoustic Holography using Physics-Informed Explainable Neural Network
- Grad-StyleSpeech: Any-Speaker Adaptive Text-to-Speech Synthesis with Diffusion Models
- Gradient Remedy for Multi-Task Learning in End-to-End Noise-Robust Speech Recognition
- Graph Based Semantic Ensemble of Riemannian Neural Structured Learning for BCI-EEG Signal Classification
- Graph Contrastive Learning with Learnable Graph Augmentation
- Graph Learning from Gaussian and Stationary Graph Signals
- Graph Neural Networks for Object Type Classification Based on Automotive Radar Point Clouds and Spectra
- Graph Neural Networks for Sound Source Localization on Distributed Microphone Networks
- Graph Representation Learning For Stroke Recurrence Prediction
- Graph Signal Processing For Neurogimaging to Reveal Dynamics of Brain Structure-Function Coupling
- Graph Signal Processing for Narrowband Direction of Arrival Estimation
- Graph Wavelet-Based Point Cloud Geometric Denoising with Surface-Consistent Non-Negative Kernel Regression
- Graph-Based Point Cloud Color Denoising with 3-Dimensional Patch-Based Similarity
- Graph-Based Spectro-Temporal Dependency Modeling for Anti-Spoofing
- Graph-Graph Context Dependency Attention for Graph Edit Distance
- Graphit: Iterative Reweighted ℓ1 Algorithm for Sparse Graph Inference in State-Space Models
- Graphmad: Graph Mixup for Data Augmentation Using Data-Driven Convex Clustering
- Gridless Target Localization for FDA-Mimo Radar with Sparse Arrays
- Group Personalized Federated Learning
- Group-Wise Co-Salient Object Detection with Siamese Transformers Via Brownian Distance Covariance Matching
- Guide and Select: A Transformer-Based Multimodal Fusion Method for Points of Interest Description Generation
- Guided Speech Enhancement Network
- HAG: Hierarchical Attention with Graph Network for Dialogue Act Classification in Conversation
- HARQ Delay Minimization of 5G Wireless Network with Imperfect Feedback
- HDNet: Hierarchical Dynamic Network for Gait Recognition using Millimeter-wave radar
- HEiMDaL: Highly Efficient Method for Detection and Localization of Wake-Words
- HIFI++: A Unified Framework for Bandwidth Extension and Speech Enhancement
- HIPI: A Hierarchical Performer Identification Model Based on Symbolic Representation of Music
- HITSZ TMG at ICASSP 2023 SPGC Shared Task: Leveraging Pre-Training and Distillation Method for Title Generation with Limited Resource
- HPFTN: Hierarchical Progressive Fusion Transformer Network for Video Denoising
- HQP-MVS:High-Quality Plane Priors Assisted Multi-View Stereo for Low-Textured Areas
- HRTF Field: Unifying Measured HRTF Magnitude Representation with Neural Fields
- HTNet: Human Topology aware network for 3d Human pose estimation
- HYDRA-HGR: A Hybrid Transformer-Based Architecture for Fusion of Macroscopic and Microscopic Neural Drive Information
- Hadamard Layer to Improve Semantic Segmentation
- Half-Temporal and Half-Frequency Attention U2Net for Speech Signal Improvement
- Halluaudio: Hallucinate Frequency as Concepts For Few-Shot Audio Classification
- Hankel Structured Low Rank and Sparse Representation Via L0-Norm Optimization for Compressed Ultrasound Plane Wave Signal Reconstruction
- HappyQuokka System for ICASSP 2023 Auditory EEG Challenge
- Hardware Friendly Spline Sketched Lidar
- Hardware-Limited Non-Uniform Task-Based Quantizers
- He-Gan: Differentially Private Gan Using Hamiltonian Monte Carlo Based Exponential Mechanism
- HeMPPCAT: Mixtures of Probabilistic Principal Component analysers for data with heteroscedastic noise
- Healthcall Corpus and Transformer Embeddings from Healthcare Customer-Agent Conversations
- Hearing and Seeing Abnormality: Self-Supervised Audio-Visual Mutual Learning for Deepfake Detection
- Heart Rate Estimation and Performance Analysis using MIMO Radar with Dispersed Antennas
- Heart Rate Extraction from Abdominal Audio Signals
- Hearttoheart: The Arts of Infant Versus Adult-Directed Speech Classification
- Heterogeneous Graph Learning for Acoustic Event Classification
- Heuristic Masking for Text Representation Pretraining
- HiSSNet: Sound Event Detection and Speaker Identification via Hierarchical Prototypical Networks for Low-Resource Headphones
- Hiding Speaker's Sex in Speech Using Zero-Evidence Speaker Representation in an Analysis/Synthesis Pipeline
- Hierarchical Diffusion Models for Singing Voice Neural Vocoder
- Hierarchical Filtering With Online Learned Priors for ECG Denoising
- Hierarchical Graph Learning for Stock Market Prediction Via a Domain-Aware Graph Pooling Operator
- Hierarchical Hypergraph Recurrent Attention Network for Temporal Knowledge Graph Reasoning
- Hierarchical Interactive Reconstruction Network for Video Compressive Sensing
- Hierarchical Multi-Agent Reinforcement Learning with Intrinsic Reward Rectification
- Hierarchical Multi-Task Learning for Fabric Component Analysis Based on NIR Spectral Signals
- Hierarchical Network with Decoupled Knowledge Distillation for Speech Emotion Recognition
- Hierarchical Pronunciation Assessment with Multi-Aspect Attention
- Hierarchical Softmax for End-To-End Low-Resource Multilingual Speech Recognition
- Hierarchical Spatial-Temporal Transformer with Motion Trajectory for Individual Action and Group Activity Recognition
- Hierarchical Spatiotemporal Feature Fusion Network For Video Saliency Prediction
- Hierarchical Transformer for Multi-Label Trailer Genre Classification
- High Quality Audio Coding with Mdctnet
- High-Acoustic Fidelity Text To Speech Synthesis With Fine-Grained Control Of Speech Attributes
- High-Dimensional Confidence Regions in Sparse MRI
- High-Dynamic Range ADC for Finite-Rate-of-Innovation Signals
- High-Frequency Transformer Network Based on Window Cross-Attention for Pansharpening
- High-Level Feature Fusion Network for Session-Based Social Recommendation
- High-Resolution Embedding Extractor for Speaker Diarisation
- High-Resolution Neural Network Processing of LFM Radar Pulses
- High-Speed Drone Detection Based On Yolo-V8
- Higher-Order Link Prediction Via Learnable Maximum Mean Discrepancy
- Higher-Order Sparse Convolutions in Graph Neural Networks
- Higher-Order Spatio-Temporal Neural Networks for Covid-19 Forecasting
- Hindi as a Second Language: Improving Visually Grounded Speech with Semantically Similar Samples
- Hint-Dynamic Knowledge Distillation
- History, Present and Future: Enhancing Dialogue Generation with Few-Shot History-Future Prompt
- How to Push the Fastest Model 50x Faster: Streaming Non-Autoregressive Speech Synthesis on Resouce-Limited Devices
- HuBERT-AGG: Aggregated Representation Distillation of Hidden-Unit Bert for Robust Speech Recognition
- Human Pose Estimation from Ambiguous Pressure Recordings with Spatio-Temporal Masked Transformers
- Hybrid Neural Network with Cross- and Self-Module Attention Pooling for Text-Independent Speaker Verification
- Hybrid Ris-Assisted Interference Mitigation for Spectrum Sharing
- Hybrid Transformers for Music Source Separation
- Hybridformer: Improving Squeezeformer with Hybrid Attention and NSR Mechanism
- Hyneter: Hybrid Network Transformer for Object Detection
- HyperSteg: Hyperbolic Learning for Deep Steganography
- Hyperbolic Audio Source Separation
- Hypernetwork-Based Adaptive Image Restoration
- Hyperspectral Image Denoising Via Nonlocal Rank Residual Modeling
- Hypothesis Test for Leakage Detection in Water Pipelines with High-Dimensional Sensor Signals
- I Hear Your True Colors: Image Guided Audio Generation
- I See What You Hear: A Vision-Inspired Method to Localize Words
- I-Tuning: Tuning Frozen Language Models with Image for Lightweight Image Captioning
- I3D: Transformer Architectures with Input-Dependent Dynamic Depth for Speech Recognition
- IAST: Instance Association Relying on Spatio-Temporal Features for Video Instance Segmentation
- ICASSP 2023 Auditory EEG Decoding Challenge
- ICASSP 2023 Spoken Language Understanding Grand Challenge
- ICCRN: Inplace Cepstral Convolutional Recurrent Neural Network for Monaural Speech Enhancement
- ICEL: Learning with Inconsistent Explanations
- ICStega: Image Captioning-based Semantically Controllable Linguistic Steganography
- IQGAN: Robust Quantum Generative Adversarial Network for Image Synthesis On NISQ Devices
- IR-ECG: Invertible Reconstruction of ECG
- ISmallNet: Densely Nested Network with Label Decoupling for Infrared Small Target Detection
- ITER-SIS: Robust Unlimited Sampling Via Iterative Signal Sieving
- Ideal: Improved Dense Local Contrastive Learning For Semi-Supervised Medical Image Segmentation
- Identifiable Bounded Component Analysis Via Minimum Volume Enclosing Parallelotope
- Identifying Coordination in a Cognitive Radar Network - A Multi-Objective Inverse Reinforcement Learning Approach
- Identifying Entrainment in Task-Oriented Conversations
- Identifying Opinion Influencers over Social Networks
- Identifying Source Speakers for Voice Conversion Based Spoofing Attacks on Speaker Verification Systems
- Image Adversarial Steganography Based on Joint Distortion
- Image Completion Via Dual-Path Cooperative Filtering
- Image Fusion Via Slice-Based Convolutional Sparse Representation
- Image Generation is May All You Need for VQA
- Image Inpainting with Semantic-Aware Transformer
- Image Reconstruction without Explicit Priors
- Image Segmentation for Improved Lossless Screen Content Compression
- Image Sharing Chain Detection VIA Sequence-To-Sequence Model
- Image Source Method Based on the Directional Impulse Responses
- Imaginary Voice: Face-Styled Diffusion Model for Text-to-Speech
- ImagineNet: Target Speaker Extraction with Intermittent Visual Cue Through Embedding Inpainting
- Immersive Enhancement and Removal of Loudspeaker Sound Using Wireless Assistive Listening Systems and Binaural Hearing Devices
- Implementing Continuous HRTF Measurement in Near-Field
- Implicit Bayes Adaptation: A Collaborative Transport Approach
- Implicit Vehicle Positioning with Cooperative Lidar Sensing
- Implicitly Rotation Equivariant Neural Networks
- Importance of Different Temporal Modulations of Speech: a Tale of two Perspectives
- Improved Acoustic-to-Articulatory Inversion Using Representations from Pretrained Self-Supervised Learning Models
- Improved Appliance Transient Feature Extraction Via Template Matching
- Improved Belief Propagation Decoding of Turbo Codes
- Improved Deep Speaker Localization and Tracking: Revised Training Paradigm and Controlled Latency
- Improved Indoor Localization With NLOS Signal Propagations
- Improved Mask-Based Neural Beamforming for Multichannel Speech Enhancement by Snapshot Matching Masking
- Improved Projection Learning for Lower Dimensional Feature Maps
- Improved Small Sample Hypothesis Testing Using the Uncertain Likelihood Ratio
- Improved Training Of Mixture-Of-Experts Language GANs
- Improved Wifi-Based Respiration Tracking via Contrast Enhancement
- Improved Wordpcfg for Passwords with Maximum Probability Segmentation
- Improvements to Embedding-Matching Acoustic-to-Word ASR Using Multiple-Hypothesis Pronunciation-Based Embeddings
- Improving Accented Speech Recognition with Multi-Domain Training
- Improving Acoustic Echo Cancellation by Mixing Speech Local and Global Features with Transformer
- Improving Adversarial Robustness with Hypersphere Embedding and Angular-Based Regularizations
- Improving Audio Captioning Using Semantic Similarity Metrics
- Improving Automatic Sleep Staging Via Temporal Smoothness Regularization
- Improving Bert Fine-Tuning via Stabilizing Cross-Layer Mutual Information
- Improving CTC-Based ASR Models With Gated Interlayer Collaboration
- Improving Contextual Biasing with Text Injection
- Improving Contextual Spelling Correction by External Acoustics Attention and Semantic Aware Data Augmentation
- Improving Disfluency Detection with Multi-Scale Self Attention and Contrastive Learning
- Improving Dropout in Graph Convolutional Networks for Recommendation via Contrastive Loss
- Improving EEG-based Emotion Recognition by Fusing Time-Frequency and Spatial Representations
- Improving Electric Load Demand Forecasting with Anchor-Based Forecasting Method
- Improving Fairness and Robustness in End-to-End Speech Recognition Through Unsupervised Clustering
- Improving Few-Shot Learning for Talking Face System with TTS Data Augmentation
- Improving Heart Rate and Heart Rate Variability Estimation from Video Through a HR-RR-Tuned Filter
- Improving Image Captioning with Control Signal of Sentence Quality
- Improving Knowledge Distillation for Non-Intrusive Load Monitoring Through Explainability Guided Learning
- Improving Learning Objectives for Speaker Verification from the Perspective of Score Comparison
- Improving Massively Multilingual ASR with Auxiliary CTC Objectives
- Improving Music Genre Classification from multi-modal Properties of Music and Genre Correlations Perspective
- Improving Noisy Student Training on Non-Target Domain Data for Automatic Speech Recognition
- Improving Non-Autoregressive Speech Recognition with Autoregressive Pretraining
- Improving Occluded Human Pose Estimation Via Linked Joints
- Improving Performance of Real-Time Full-Band Blind Packet-Loss Concealment with Predictive Network
- Improving Phase-Vocoder-Based Time Stretching by Time-Directional Spectrogram Squeezing
- Improving Prosody for Cross-Speaker Style Transfer by Semi-Supervised Style Extractor and Hierarchical Modeling in Speech Synthesis
- Improving Retrieval-Based Dialogue System Via Syntax-Informed Attention
- Improving Scheduled Sampling for Neural Transducer-Based ASR
- Improving Self-Supervised Learning for Audio Representations by Feature Diversity and Decorrelation
- Improving Sentence Similarity Estimation for Unsupervised Extractive Summarization
- Improving Speech Enhancement via Event-Based Query
- Improving Speech Prosody of Audiobook Text-To-Speech Synthesis with Acoustic and Textual Contexts
- Improving Speech-to-Speech Translation Through Unlabeled Text
- Improving Spoken Language Identification with Map-Mix
- Improving Text-Audio Retrieval by Text-Aware Attention Pooling and Prior Matrix Revised Loss
- Improving Transformer-Based End-to-End Speaker Diarization by Assigning Auxiliary Losses to Attention Heads
- Improving Transformer-Based Networks with Locality for Automatic Speaker Verification
- Improving Weakly Supervised Sound Event Detection with Causal Intervention
- Improving fast-slow Encoder based Transducer with Streaming Deliberation
- Improving the Modality Representation with multi-view Contrastive Learning for Multimodal Sentiment Analysis
- Improving the Stochastic Gradient Descent's Test Accuracy by Manipulating the ℓ∞ Norm of its Gradient Approximation
- Improving the out-of-Distribution Generalization Capability of Language Models: Counterfactually-Augmented Data is not Enough
- In Search of Strong Embedding Extractors for Speaker Diarisation
- In-Sensor & Neuromorphic Computing Are all You Need for Energy Efficient Computer Vision
- Incorporating Lip Features into Audio-Visual Multi-Speaker DOA Estimation by Gated Fusion
- Incorporating Reliability in Graph Information Propagation by Fluid Dynamics Diffusion: A case of Multimodal Semisupervised Deep Learning
- Incorporating Uncertainty from Speaker Embedding Estimation to Speaker Verification
- Incorporating Visual Information Reconstruction into Progressive Learning for Optimizing audio-visual Speech Enhancement
- Independent Vector Analysis with Multivariate Gaussian Model: a Scalable Method by Multilinear Regression
- Individual Sub-Band Estimation Approach to Bandwidth Extension and Enhancement of Coded Speech
- Inductive Relation Prediction from Relational Paths and Context with Hierarchical Transformers
- InfoShape: Task-Based Neural Data Shaping via Mutual Information
- Information Extraction from Pill Bottle Images via Text Stitching
- Information and Sensing Beamforming Optimization for Multi-User Multi-Target MIMO ISAC Systems
- Infrared and Visible Image Fusion by Using Multi-Scale Transformation and Fractional-Order Gradient Information
- Inplace Cepstral Speech Enhancement System for the ICASSP 2023 Clarity Challenge
- Input-Dependent Dynamical Channel Association For Knowledge Distillation
- Instance-Aware Hierarchical Structured Policy for Prompt Learning in Vision-Language Models
- Int-GNN: A User Intention Aware Graph Neural Network for Session-Based Recommendation
- Integrated Sensing and Full-Duplex Communication: Joint Transceiver Beamforming and Power Allocation
- Integrating Syntactic and Semantic Knowledge in AMR Parsing with Heterogeneous Graph Attention Network
- Integrating the Sensing and Radio Communications Channel Modelling From Radar Mutual Interference
- Intent Does Matter! Propagating High-Order Relations for Exploring Interest Preferences
- Inter-Pulse Estimation for Sperm Whale Click Detection
- Inter-Scale Sure-Let Denoise with Structured Deep Image Prior: Interpretable Self-Supervised Learning
- Inter-Subnet: Speech Enhancement with Subband Interaction
- Interaction-Assisted Multi-Modal Representation Learning for Recommendation
- Interference Leakage Minimization in RIS-Assisted MIMO Interference Channels
- Intermediate Fine-Tuning Using Imperfect Synthetic Speech for Improving Electrolaryngeal Speech Recognition
- Intermpl: Momentum Pseudo-Labeling With Intermediate CTC Loss
- Internal Language Model Estimation Based Adaptive Language Model Fusion for Domain Adaptation
- Interpolation Filter Model For Ramanujan Subspace Signals
- Interpolation of Spatial Room Impulse Responses Using Partial Optimal Transport
- Interpretability in the Context of Sequential Cost-Sensitive Feature Acquisition
- Interpretable Multi-Scale Neural Network for Granger Causality Discovery
- Interpretable Nonnegative Incoherent Deep Dictionary Learning for FMRI Data Analysis
- Interpretable, Unrolled Deep Radar Beampattern Design
- Interpretation of Neural Networks is Susceptible to Universal Adversarial Perturbations
- Interweaved Graph and Attention Network for 3D Human Pose Estimation
- Introducing Topography in Convolutional Neural Networks
- Inv-Senet: Invariant Self Expression Network for Clustering Under Biased Data
- Invariant Adversarial Imitation Learning From Visual Inputs
- Inverse Quadratic Transform for Minimizing A Sum of Ratios
- Inverse Reinforcement Learning with Graph Neural Networks for IoT Resource Allocation
- Investigating Content-Aware Neural Text-to-Speech MOS Prediction Using Prosodic and Linguistic Features
- Investigating SINDy as a Tool for Causal Discovery in Time Series Signals
- Investigation into Phone-Based Subword Units for Multilingual End-to-End Speech Recognition
- IoU-Aware Multi-Expert Cascade Network Via Dynamic Ensemble for Long-Tailed Object Detection
- Is Multi-Task Learning an Upper Bound for Continual Learning?
- Is Quality Enoughƒ Integrating Energy Consumption in a Large-Scale Evaluation of Neural Audio Synthesis Models
- Iterative Shallow Fusion of Backward Language Model for End-To-End Speech Recognition
- Iterative Water-Filling Power and Subcarrier Allocation for Multicarrier NOMA Downlink
- JEIT: Joint End-to-End Model and Internal Language Model Training for Speech Recognition
- JNDMix: Jnd-Based Data Augmentation for No-Reference Image Quality Assessment
- JPEG Pleno Call for Proposals Responses Quality Assessment
- JSV-VC: Jointly Trained Speaker Verification and Voice Conversion Models
- Jamming Source Localization Using Augmented Physics-Based Model
- Jazznet: A Dataset of Fundamental Piano Patterns for Music Audio Machine Learning Research
- Jeffreys Divergence-Based Regularization of Neural Network Output Distribution Applied to Speaker Recognition
- Joint Angle and Respiration Estimation for Passive and Device-Free Respiration Monitoring
- Joint Ann-SNN Co-training for Object Localization and Image Segmentation
- Joint Antenna Selection and Beamforming in Integrated Automotive Radar Sensing-Communications with Quantized Double Phase Shifters
- Joint Channel and Direction Estimation for Ground-to-UAV Communications Enabled by a Simultaneous Reflecting and Sensing RIS
- Joint Compression and Demosaicking For Satellite Images
- Joint Cryo-ET Alignment and Reconstruction with Neural Deformation Fields
- Joint Data Association, NLOS Mitigation, and Clutter Suppression for Networked Device-Free Sensing in 6G Cellular Network
- Joint Discriminator and Transfer Based Fast Domain Adaptation For End-To-End Speech Recognition
- Joint Estimation of Clustered user Activity and Correlated Channels with Unknown Covariance in mMTC
- Joint Estimation of DOA and Distance in Noisy Reverberant Conditions
- Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection
- Joint Human Orientation-Activity Recognition Using WIFI Signals for Human-Machine Interaction
- Joint Microstrip Selection and Beamforming Design for MmWave Systems with Dynamic Metasurface Antennas
- Joint Millimeter-Wave AoD and AoA Estimation Using one OFDM Symbol and Frequency-Dependent Beams
- Joint Modeling for ASR Correction and Dialog State Tracking
- Joint Modelling of Spoken Language Understanding Tasks with Integrated Dialog History
- Joint Multi-Level Feature Network for Lightweight Person Re-Identification
- Joint Neural Representation for Multiple Light Fields
- Joint Noise Reduction and Listening Enhancement for Full-End Speech Enhancement
- Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation
- Joint Robust Representation And Generalization Enhancement For Cross-Modality Person Re-Identification
- Joint Symbol-Level Precoding and Sub-Block-Level RIS Design for Dual-Function Radar-Communications
- Joint Training and Decoding for Multilingual End-to-End Simultaneous Speech Translation
- Joint Training of Hierarchical GANs and Semantic Segmentation for Expression Translation
- Joint Unmixing And Demosaicing Methods For Snapshot Spectral Images
- Joint Unsupervised and Supervised Learning for Context-Aware Language Identification
- Joint Waveform and Passive Beamformer Design in Multi-IRS-Aided Radar
- Jointly Visual- and Semantic-Aware Graph Memory Networks for Temporal Sentence Localization in Videos
- K2NN: Self-Supervised Learning with Hierarchical Nearest Neighbors for Remote Sensing
- KEPS-NET: Robust Parking slot Detection based Keypoint estimation for High Localization Accuracy
- KG-ECO: Knowledge Graph Enhanced Entity Correction For Query Rewriting
- Kalmanbot: Kalmannet-Aided Bollinger Bands for Pairs Trading
- Kernel Estimation and Deconvolution for Blind Image Super-Resolution
- Kernel Interpolation of Acoustic Transfer Functions with Adaptive Kernel for Directed and Residual Reverberations
- Kernel Ridge Regression for Generalized Graph Signal Processing
- Keyword-Specific Acoustic Model Pruning for Open-Vocabulary Keyword Spotting
- Knowledge Distillation with Active Exploration and Self-Attention Based Inter-Class Variation Transfer for Image Segmentation
- Knowledge Transfer for on-Device Speech Emotion Recognition With Neural Structured Learning
- Knowledge-Augmented Frame Semantic Parsing with Hybrid Prompt-Tuning
- Knowledge-Aware Bayesian Co-Attention for Multimodal Emotion Recognition
- Knowledge-Aware Few Shot Learning for Event Detection from Short Texts
- Knowledge-Aware Graph Convolutional Network with Utterance-Specific Window Search for Emotion Recognition In Conversations
- Knowledge-Graph Augmented Music Representation for Genre Classification
- LA-VOCE: LOW-SNR Audio-Visual Speech Enhancement Using Neural Vocoders
- LABANet: Lead-Assisting Backbone Attention Network for Oral Multi-Pathology Segmentation
- LDTSF: A Label-Decoupling Teacher-Student Framework for Semi-Supervised Echocardiography Segmentation
- LE-DTA: Local Extrema Convolution for Drug Target Affinity Prediction
- LEAPT: Learning Adaptive Prefix-to-Prefix Translation For Simultaneous Machine Translation
- LED: Label Correlation Enhanced Decoder for Multi-Label Text Classification
- LGVIT: Local-Global Vision Transformer for Breast Cancer Histopathological Image Classification
- LIMI-VC: A Light Weight Voice Conversion Model with Mutual Information Disentanglement
- LINK: Linguistic Steganalysis Framework with External Knowledge
- LMBAO: A Landmark Map for Bundle Adjustment Odometry in LiDAR SLAM
- LMCodec: A Low Bitrate Speech Codec with Causal Transformer Models
- LP-IOANet: Efficient High Resolution Document Shadow Removal
- LQGNET: Hybrid Model-Based and Data-Driven Linear Quadratic Stochastic Control
- LSSED: A Robust Segmentation Network for Inflamed Appendix from CT Images
- LSTM-Based Video Quality Prediction Accounting for Temporal Distortions in Videoconferencing Calls
- Label-Efficient and Robust Learning from Multiple Experts
- Label-Guided Contrastive Learning for Out-of-Domain Detection
- Large Covariance Matrix Estimation with Oracle Statistical Rate
- Large Dimensional Analysis of LS-SVM Transfer Learning: Application to Polsar Classification
- Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
- Large-Scale Language Model Rescoring on Long-Form Data
- Large-Scale Nonverbal Vocalization Detection Using Transformers
- Laryngeal Leukoplakia Classification Via Dense Multiscale Feature Extraction in White Light Endoscopy Images
- Lasso-Based Fast Residual Recovery For Modulo Sampling
- Last: Scalable Lattice-Based Speech Modelling in Jax
- Latent Iterative Refinement for Modular Source Separation
- Lattice-Free Sequence Discriminative Training for Phoneme-Based Neural Transducers
- LeanSpeech: The Microsoft Lightweight Speech Synthesis System for Limmits Challenge 2023
- Learn Topological Representation with Flexible Manifold Layer
- Learnable Flow Model Conditioned on Graph Representation Memory for Anomaly Detection
- Learnable Frontends That Do Not Learn: Quantifying Sensitivity To Filterbank Initialisation
- Learned Generative Misspecified Lower Bound
- Learned Kalman Filtering in Latent Space with High-Dimensional Data
- Learned Video Coding with Motion Compensation Mixture Model
- Learning 3D Human Pose and Shape Estimation Using Uncertainty-Aware Body Part Segmentation
- Learning ASR Pathways: A Sparse Multilingual ASR Model
- Learning Audio-Visual Dereverberation
- Learning Causal Representations for Generalizable Face Anti Spoofing
- Learning Cross-Lingual Visual Speech Representations
- Learning Cross-Modal Audiovisual Representations with Ladder Networks for Emotion Recognition
- Learning Dependencies of Discrete Speech Representations with Neural Hidden Markov Models
- Learning Dynamic Graphs under Partial Observability
- Learning Environmental Structure Using Acoustic Probes with a Deep Neural Network
- Learning Expressive And Generalizable Motion Features For Face Forgery Detection
- Learning From Label Proportion with Online Pseudo-Label Decision by Regret Minimization
- Learning From Positive and Unlabeled Data Using Observer-GAN
- Learning From Single-Expert Annotated Labels for Automatic Sleep Staging
- Learning From Yourself: A Self-Distillation Method For Fake Speech Detection
- Learning Generalizable Light Field Networks from Few Images
- Learning Gradients of Convex Functions with Monotone Gradient Networks
- Learning Graph Laplacian from Intrinsic Patterns via Gaussian Process
- Learning How to Learn Domain-Invariant Parameters for Domain Generalization
- Learning Hybrid Representations of Semantics and Distortion for Blind Image Quality Assessment
- Learning Hypergraphs From Signals With Dual Smoothness Prior
- Learning Interpretable Filters In Wav-UNet For Speech Enhancement
- Learning Properties of Holomorphic Neural Networks of Dual Variables
- Learning Quantum Entanglement Distillation With Noisy Classical Communications
- Learning Robust Self-Attention Features for Speech Emotion Recognition with Label-Adaptive Mixup
- Learning Scene Flow from 3d Point Clouds with Cross-Transformer and Global Motion Cues
- Learning Silhouettes with Group Sparse Autoencoders
- Learning Sparse Alignments via Optimal Transport for Cross-Domain Fake News Detection
- Learning Sparse auto-Encoders for Green AI image coding
- Learning Speech Representations with Flexible Hidden Feature Dimensions
- Learning Supervised Covariation Projection Through General Covariance
- Learning Task-Aligned Mask Query for Instance Segmentation
- Learning To Generate 3d Representations of Building Roofs Using Single-View Aerial Imagery
- Learning To Locate Visual Answer In Video Corpus Using Question
- Learning To Regularized Resource Allocation with Budget Constraints
- Learning Unbiased Rewards with Mutual Information in Adversarial Imitation Learning
- Learning a Weight Map for Weakly-Supervised Localization
- Learning from the Raw Domain: Cross Modality Distillation for Compressed Video Action Recognition
- Learning on Entropy Coded Images with CNN
- Learning on Graphs under Label Noise
- Learning to Auto-Correct for High-Quality Spectrograms
- Learning to Balance the Global Coherence and Informativeness in Knowledge-Grounded Dialogue Generation
- Learning to Build Reasoning Chains by Reliable Path Retrieval
- Learning to Detect Novel and Fine-Grained Acoustic Sequences Using Pretrained Audio Representations
- Learning to Explain: a Gradient-based Attribution Method for Interpreting Super-Resolution Networks
- Learning to Locate the Text Forgery in Smartphone Screenshots
- Learning to Personalize Equalization for High-Fidelity Spatial Audio Reproduction
- Learning to Reconnect Interrupted Trajectories for Weakly Supervised Multi-Object Tracking
- Learning with Multigraph Convolutional Filters
- Learnt Mutual Feature Compression for Machine Vision
- Lego-Features: Exporting Modular Encoder Features for Streaming and Deliberation ASR
- Less Is More: A Unified Architecture for Device-Directed Speech Detection with Multiple Invocation Types
- Level-Line Guided Edge Drawing for Robust Line Segment Detection
- Leveraging Heteroscedastic Uncertainty in Learning Complex Spectral Mapping for Single-Channel Speech Enhancement
- Leveraging Label Correlations in a Multi-Label Setting: a Case Study in Emotion
- Leveraging Language Embeddings for Cross-Lingual Self-Supervised Speech Representation Learning
- Leveraging Large Text Corpora For End-To-End Speech Summarization
- Leveraging Multiple Sources in Automatic African American English Dialect Detection for Adults and Children
- Leveraging Neural Koopman Operators to Learn Continuous Representations of Dynamical Systems from Scarce Data
- Leveraging Phone-Level Linguistic-Acoustic Similarity For Utterance-Level Pronunciation Scoring
- Leveraging Positional-Related Local-Global Dependency for Synthetic Speech Detection
- Leveraging Pretrained Representations With Task-Related Keywords for Alzheimer's Disease Detection
- Leveraging Sparsity with Spiking Recurrent Neural Networks for Energy-Efficient Keyword Spotting
- Lexicon-injected Semantic Parsing for Task-Oriented Dialog
- LiNuIQA: Lightweight No-Reference Image Quality Assessment Based on Non-Uniform Weighting
- LiQuiD-MIMO Radar: Distributed MIMO Radar with Low-Bit Quantization
- Light Field Compression Via Compact Neural Scene Representation
- Light Projection-Based Physical-World Vanishing Attack Against Car Detection
- Light-Weight CNN-Attention Based Architecture for Hand Gesture Recognition Via Electromyography
- Light-Weight Sequential SBL Algorithm: An Alternative to OMP
- LightGrad: Lightweight Diffusion Probabilistic Model for Text-to-Speech
- Lightvessel: Exploring Lightweight Coronary Artery Vessel Segmentation Via Similarity Knowledge Distillation
- Lightweight Annotation and Class Weight Training for Automatic Estimation of Alarm Audibility in Noise
- Lightweight Feature Encoder for Wake-Up Word Detection Based on Self-Supervised Speech Representation
- Lightweight Fisher Vector Transfer Learning for Video Deduplication
- Lightweight Machine Learning for Seizure Detection on Wearable Devices
- Lightweight Portrait Segmentation Via Edge-Optimized Attention
- Lightweight Prosody-TTS for Multi-Lingual Multi-Speaker Scenario
- Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform
- Lightweight, Multi-Speaker, Multi-Lingual Indic Text-to-Speech
- Line Segment Matching Based on Intersection-Enhanced Point Correspondences
- Linear Microphone Array Parallel to the Driving Direction for in-Car Speech Enhancement
- Lip-to-Speech Synthesis in the Wild with Multi-Task Learning
- Lit the Darkness: Three-Stage Zero-Shot Learning for Low-Light Enhancement with Multi-Neighbor Enhancement Factors
- LiteG2P: A Fast, Light and High Accuracy Model for Grapheme-to-Phoneme Conversion
- Liveness Score-Based Regression Neural Networks for Face Anti-Spoofing
- Local Feature Enhanced Adversarial Network for the Blind Image Quality Assessment
- Local Graph-Homomorphic Processing for Privatized Distributed Systems
- Local to global prior Learning for blind Unsupervised Image super Resolution
- Local-Global Progressive U-Transformers for Accurate Hepatic and Portal Veins Segmentation in Abdominal MR Images
- Local-Global Siamese Network with Efficient Inter-Scale Feature Learning for Change Detection in VHR Remote Sensing Images
- Locale Encoding for Scalable Multilingual Keyword Spotting Models
- Locality Preserving Multiview Graph Hashing For Large Scale Remote Sensing Image Search
- Log-Can: Local-Global Class-Aware Network For Semantic Segmentation of Remote Sensing Images
- Logo-Former: Local-Global Spatio-Temporal Transformer for Dynamic Facial Expression Recognition
- Logovit: Local-Global Vision Transformer for Object Re-Identification
- Long Range Imaging Using Multispectral Fusion of RGB and NIR Images
- Long-Memory Message-Passing for Spatially Coupled Systems
- Long-Short Attention Network For The Spectral Super-Resolution Of Multispectral Images
- Long-Tailed Image Recognition with Dynamic Re-Weighting
- Long-Tailed Recognition with Causal Invariant Transformation
- Long-Term Synchronization of Wireless Acoustic Sensor Networks with Nonpersistent Acoustic Activity Using Coherence State
- LongFNT: Long-Form Speech Recognition with Factorized Neural Transducer
- Longshortnet: Exploring Temporal and Semantic Features Fusion In Streaming Perception
- Look and Think: Intrinsic Unification of Self-Attention and Convolution for Spatial-Channel Specificity
- Loss Function Design for DNN-Based Sound Event Localization and Detection on Low-Resource Realistic Data
- Lost In Translation: Generating Adversarial Examples Robust to Round-Trip Translation
- Low Precision Representations for High Dimensional Models
- Low in Resolution, High in Precision: UAV Detection with Super-Resolution and Motion Information Extraction
- Low-Bitrate Redundancy Coding of Speech Using A Rate-Distortion-Optimized Variational Autoencoder
- Low-Complexity Acoustic Echo Cancellation with Neural Kalman Filtering
- Low-Dose CT Reconstruction Via Optimization-Inspired GAN
- Low-Latency Electrolaryngeal Speech Enhancement Based on Fastspeech2-Based Voice Conversion and Self-Supervised Speech Representation
- Low-Rank Constrained Memory Autoencoder for Hyperspectral Anomaly Detection
- Low-Rank Plus Sparse Trajectory Decomposition for Direct Exoplanet Imaging
- Low-Rank Tensor Decompositions for Quaternion Multiway Arrays
- Low-Resource Music Genre Classification with Cross-Modal Neural Model Reprogramming
- Lyapunov-Driven Deep Reinforcement Learning for Edge Inference Empowered by Reconfigurable Intelligent Surfaces
- M-CTRL: A Continual Representation Learning Framework with Slowly Improving Past Pre-Trained Model
- M-SpeechCLIP: Leveraging Large-Scale, Pre-Trained Models for Multilingual Speech to Image Retrieval
- M2-CTTS: End-to-End Multi-Scale Multi-Modal Conversational Text-to-Speech Synthesis
- M22: Rate-Distortion Inspired Gradient Compression
- M2TSR: Multi-Range and Mix-Grained Transformer for Single Image Super-Resolution
- M3ST: Mix at Three Levels for Speech Translation
- MADI: Inter-Domain Matching and Intra-Domain Discrimination for Cross-Domain Speech Recognition
- MAID: A Conditional Diffusion Model for Long Music Audio Inpainting
- MASKED-AP: Attention Pyramid Convolutional Neural Network with Mask for Cervical Cell Classification
- MAST: Multiscale Audio Spectrogram Transformers
- MCKD: Mutually Collaborative Knowledge Distillation For Federated Domain Adaptation And Generalization
- MCNET: Fuse Multiple Cues for Multichannel Speech Enhancement
- MCNeT: Measurement-Consistent Networks Via A Deep Implicit Layer For Solving Inverse Problems
- MDR-MFI:Multi-Branch Decoupled Regression and Multi-Scale Feature Interaction for Partial-to-Partial Cloud Registration
- MEET: A Monte Carlo Exploration-Exploitation Trade-Off for Buffer Sampling
- MFAT: A Multi-Level Feature Aggregated Transformer for Person Re-Identification
- MFCCGAN: A Novel MFCC-Based Speech Synthesizer Using Adversarial Learning
- MGAT: Multi-Granularity Attention Based Transformers for Multi-Modal Emotion Recognition
- MHLAT: Multi-Hop Label-Wise Attention Model for Automatic ICD Coding
- MHSCNET: A Multimodal Hierarchical Shot-Aware Convolutional Network for Video Summarization
- MID-Attribute Speaker Generation Using Optimal-Transport-Based Interpolation of Gaussian Mixture Models
- MLCGAN: Multi-Lead ECG Synthesis with Multi Label Conditional Generative Adversarial Network
- MLP-GAN for Brain Vessel Image Segmentation
- MMATR: A Lightweight Approach for Multimodal Sentiment Analysis Based on Tensor Methods
- MMCosine: Multi-Modal Cosine Loss Towards Balanced Audio-Visual Fine-Grained Learning
- MODEFORMER: Modality-Preserving Embedding For Audio-Video Synchronization Using Transformers
- MPE4G : Multimodal Pretrained Encoder for Co-Speech Gesture Generation
- MPS-AMS: Masked Patches Selection and Adaptive Masking Strategy Based Self-Supervised Medical Image Segmentation
- MRML: Multimodal Rumor Detection by Deep Metric Learning
- MRNET: Multi-Refinement Network for Dual-Pixel Images Defocus Deblurring
- MSFORMER: Multi-Scale Transformer with Neighborhood Consensus for Feature Matching
- MSN-net: Multi-Scale Normality Network for Video Anomaly Detection
- MSNet: A Deep Architecture Using Multi-Sentiment Semantics for Sentiment-Aware Image Style Transfer
- MSP-Former: Multi-Scale Projection Transformer for Single Image Desnowing
- MTDL-NET: Morphological and Temporal Discriminative Learning for Heartbeat Classification
- MTFD: Multi-Teacher Fusion Distillation for Compressed Video Action Recognition
- MUG: A General Meeting Understanding and Generation Benchmark
- Mabnet: Master Assistant Buddy Network With Hybrid Learning for Image Retrieval
- Machine Learning Based Early Debris Detection Using Automotive Low Level Radar Data
- Machine Learning-Aided Piece-Wise Modeling Technique of Power Amplifier for Digital Predistortion
- Make More of Your Data: Minimal Effort Data Augmentation for Automatic Speech Recognition and Translation
- Make Your Enemy Your Friend: Improving Image Rotation Angle Estimation with Harmonics
- Making Synchrosqueezing Locally Adaptive in The Time-Frequency Plane
- Managing Information Updating with Edge Computing: A Distributed and Learning Approach
- Margin-Mixup: A Method for Robust Speaker Verification In Multi-Speaker Audio
- MarginNCE: Robust Sound Localization with a Negative Margin
- Mask Guided Selective Context Decoding for Handwritten Chinese Text Recognition
- Mask the Bias: Improving Domain-Adaptive Generalization of CTC-Based ASR with Internal Language Model Estimation
- Maskdul: Data Uncertainty Learning in Masked Face Recognition
- Masked Autoencoders are Articulatory Learners
- Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input
- Masked Spectrogram Prediction for Self-Supervised Audio Pre-Training
- Masked Token Similarity Transfer for Compressing Transformer-Based ASR Models
- Masking Speech Contents by Random Splicing: is Emotional Expression Preserved?
- Massively Multilingual ASR on 70 Languages: Tokenization, Architecture, and Generalization Capabilities
- Massively Multilingual Shallow Fusion with Large Language Models
- Matching-Based Term Semantics Pre-Training for Spoken Patient Query Understanding
- Matrix Low-Rank Approximation for Policy Gradient Methods
- Matrix Recovery using Deep Generative Priors with Low-Rank Deviations
- Matrix Resolvent Eigenembeddings for Dynamic Graphs
- Maximum Likelihood Distillation for Robust Modulation Classification
- Mcrood: Multi-Class Radar Out-Of-Distribution Detection
- Measure and Countermeasure of the Capsulation Attack Against Backdoor-Based Deep Neural Network Watermarks
- Measuring Deviation from Stochasticity in Time-Series Using Autoencoder Based Time-Invariant Representation: Application to Black Hole Data
- Measuring the Transferability of ℓ∞ Attacks by the ℓ2 Norm
- Medleyvox: An Evaluation Dataset for Multiple Singing Voices Separation
- Meeting Action Item Detection with Regularized Context Modeling
- Memory-Augmented Contrastive Learning for Talking Head Generation
- Memory-Augmented U-Transformer For Multivariate Time Series Anomaly Detection
- Mendam: Multi-Expert Network with Distribution-Aware Momentum for Long-Tailed Recognition
- Meta Learning for Domain Agnostic Soft Prompt
- Meta Learning with Adaptive Loss Weight for Low-Resource Speech Recognition
- Meta++ Network for Few-Shot Aerospace Crack Segmentation
- Meta-Dag: Meta Causal Discovery Via Bilevel Optimization
- Meta-Learning for Image-Guided Millimeter-Wave Beam Selection in Unseen Environments
- Metric Learning for User-Defined Keyword Spotting
- Metric-Oriented Speech Enhancement Using Diffusion Probabilistic Model
- Mimo Radar Transmit Beampattern Matching Via Manifold Optimization
- Mingling or Misalignment? Temporal Shift for Speech Emotion Recognition with Pre-Trained Representations
- Minimising Distortion for GAN-Based Facial Attribute Manipulation
- Misspecified Cramér-Rao Bound of RIS-Aided Localization Under Geometry Mismatch
- Mitigating Domain Dependency for Improved Speech Enhancement Via SNR Loss Boosting
- Mitigating Unintended Memorization in Language Models Via Alternating Teaching
- Mixed Far-field and Near-field Source Localization Based on Low-Rank Matrix Reconstruction
- Mixed Sample Augmentation for Online Distillation
- Mixer: DNN Watermarking using Image Mixup
- MoLE : Mixture Of Language Experts For Multi-Lingual Automatic Speech Recognition
- Modaldrop: Modality-Aware Regularization for Temporal-Spectral Fusion in Human Activity Recognition
- Model Fingerprinting with Benign Inputs
- Model-Based Spectral Reconstruction Of Interferometric Acquisitions
- Model-Free Learning of Optimal Beamformers for Passive IRS-Assisted Sumrate Maximization
- Model-Free Online Learning for Waveform Optimization In Integrated Sensing And Communications
- Model-Matching Principle Applied to the Design of an Array-Based All-Neural Binaural Rendering System for Audio Telepresence
- Model-based vs. Data-driven Approaches for Predicting Rain-induced Attenuation in Commercial Microwave Links: A Comparative Empirical Study
- Modeling Global Latent Semantic in Multi-Turn Conversations with Random Context Reconstruction
- Modeling Turn-Taking in Human-To-Human Spoken Dialogue Datasets Using Self-Supervised Features
- Modeling the Wave Equation Using Physics-Informed Neural Networks Enhanced With Attention to Loss Weights
- Modelling Black-Box Audio Effects with Time-Varying Feature Modulation
- Modelling Low-Resource Accents Without Accent-Specific TTS Frontend
- Modify: Model-Driven Face Stylization Without Style Images
- Modular Conformer Training for Flexible End-to-End ASR
- Modulation-Based Center Alignment and Motion Mining for Spatial Temporal Action Detection
- Modulo EEG Signal Recovery Using Transformer
- Monocular 3D Human Pose Estimation Based on Global Temporal-Attentive and Joints-Attention In Video
- More Speaking or More Speakers?
- MossFormer: Pushing the Performance Limit of Monaural Speech Separation Using Gated Single-Head Transformer with Convolution-Augmented Joint Self-Attentions
- Motion Matters: A Novel Motion Modeling for Cross-View Gait Feature Learning
- Motion-Aware Video Paragraph Captioning via Exploring Object-Centered Internal Knowledge
- Motor Activity Recognition Using Eeg Data and Ensemble of Stacked BLSTM-LSTM Network and Transformer Model
- Mouth Breathing Detection Using Audio Captured Through Earbuds
- Movienet-PS: A Large-Scale Person Search Dataset in the Wild
- Moving Towards Non-Binary Gender Identification Via Analysis of System Errors in Binary Gender Classification
- Multi-Agent Adversarial Training Using Diffusion Learning
- Multi-Agent Reinforcement Learning for Covert Semantic Communications over Wireless Networks
- Multi-Aspect Interest Neighbor-Augmented Network for Next-Basket Recommendation
- Multi-Blank Transducers for Speech Recognition
- Multi-Carrier Wideband OCDM-Based THZ Automotive Radar
- Multi-Channel Audio Signal Generation
- Multi-Channel Speaker Extraction with Adversarial Training: The Wavlab Submission to The Clarity ICASSP 2023 Grand Challenge
- Multi-Dimensional Frequency Dynamic Convolution with Confident Mean Teacher for Sound Event Detection
- Multi-Dimensional Signal Recovery Using Low-Rank Deconvolution
- Multi-Dimensional and Multi-Scale Modeling for Speech Separation Optimized by Discriminative Learning
- Multi-Functional Reconfigurable Intelligent Surface
- Multi-Head Attention and GRU for Improved Match-Mismatch Classification of Speech Stimulus and EEG Response
- Multi-Head Feature Pyramid Networks for Breast Mass Detection
- Multi-Head Uncertainty Inference for Adversarial Attack Detection
- Multi-Label Temporal Evidential Neural Networks for Early Event Detection
- Multi-Layer Feature Division Transferable Adversarial Attack
- Multi-Layer Seasonal Perception Network for Time Series Forecasting
- Multi-Level Fusion for Burst Super-Resolution with Deep Permutation-Invariant Conditioning
- Multi-Lingual Pronunciation Assessment with Unified Phoneme Set and Language-Specific Embeddings
- Multi-Local Attention for Speech-Based Depression Detection
- Multi-Microphone Speaker Separation by Spatial Regions
- Multi-Modal Domain Generalization for Cross-Scene Hyperspectral Image Classification
- Multi-Modal Food Classification in a Diet Tracking System with Spoken and Visual Inputs
- Multi-Object Localization and Irrelevant-Semantic Separation for Nuclei Segmentation in Histopathology Images
- Multi-Observation Hidden Semi-Markov Model for Photoplethysmogram Signal Semantic Segmentation
- Multi-Output RNN-T Joint Networks for Multi-Task Learning of ASR and Auxiliary Tasks
- Multi-Rate Adaptive Transform Coding for Video Compression
- Multi-Resolution Convolutional Dictionary Learning for Riverbed Dynamics Modeling
- Multi-Resolution Location-Based Training for Multi-Channel Continuous Speech Separation
- Multi-Resolution Sequence Aggregation and Model-Agnostic Framework for Time-Series Forecasting
- Multi-Scale Compositional Constraints for Representation Learning on Videos
- Multi-Scale Receptive Field Graph Model for Emotion Recognition in Conversations
- Multi-Source Templates Learning for Real-Time Aerial Tracking
- Multi-Speaker Data Augmentation for Improved end-to-end Automatic Speech Recognition
- Multi-Speaker End-to-End Multi-Modal Speaker Diarization System for the MISP 2022 Challenge
- Multi-Speaker Expressive Speech Synthesis via Multiple Factors Decoupling
- Multi-Speaker Multi-Lingual VQTTS System for LIMMITS 2023 Challenge
- Multi-Speaker Speech Synthesis from Electromyographic Signals by Soft Speech Unit Prediction
- Multi-Speaker and Wide-Band Simulated Conversations as Training Data for End-to-End Neural Diarization
- Multi-Stage Aggregation Transformer for Medical Image Segmentation
- Multi-Stream Facial Adaptive Network for Expression Recognition from a Single Image
- Multi-Task Bias-Variance Trade-Off Through Functional Constraints
- Multi-Task Sub-Band Network For Deep Residual Echo Suppression
- Multi-Task Transformer with Relation-Attention and Type-Attention for Named Entity Recognition
- Multi-Temporal Lip-Audio Memory for Visual Speech Recognition
- Multi-User Data Detection in Massive MIMO with 1-Bit ADCS
- Multi-User Methods for Vibrational Radar Backscatter Communications
- Multi-View Graph Regularized Deep Autoencoder-Like NMF Framework
- Multi-View K-Means with Laplacian Embedding
- Multi-View Learning for Speech Emotion Recognition with Categorical Emotion, Categorical Sentiment, and Dimensional Scores
- Multi-View Millimeter-Wave Imaging Over Wireless Cellular Network
- Multi-modal ASR error correction with joint ASR error detection
- Multicast Beamformer Design for Mimo Coded Caching Systems
- Multichannel Time-Encoding of Finite-Rate-of-Innovation Signals
- Multilayer Subspace Learning With Self-Sparse Robustness for Two-Dimensional Feature Extraction
- Multilevel FISTA for Image Restoration
- Multilevel Transformer for Multimodal Emotion Recognition
- Multilingual Alzheimer's Dementia Recognition through Spontaneous Speech: A Signal Processing Grand Challenge
- Multilingual End-To-End Spoken Language Understanding For Ultra-Low Footprint Applications
- Multilingual Query-by-Example Keyword Spotting with Metric Learning and Phoneme-to-Embedding Mapping
- Multilingual Word Error Rate Estimation: E-Wer3
- Multimodal Dyadic Impression Recognition via Listener Adaptive Cross-Domain Fusion
- Multimodal Emotion Recognition Based on Deep Temporal Features Using Cross-Modal Transformer and Self-Attention
- Multimodal Facial Action unit Detection with Physiological Signals
- Multimodal Knowledge Distillation for Arbitrary-Oriented Object Detection in Aerial Images
- Multimodal Microscopy Image Alignment Using Spatial and Shape Information and a Branch-and-Bound Algorithm
- Multimodal Propaganda Detection Via Anti-Persuasion Prompt enhanced contrastive learning
- Multiple Access Computation Offloading for the K-User Case
- Multiple Acoustic Features Speech Emotion Recognition Using Cross-Attention Transformer
- Multiple Contrastive Learning for Multimodal Sentiment Analysis
- Multiple Domain-Adversarial Ensemble Learning for Domain Generalization
- Multiple Signed Graph Learning for Gene Regulatory Network Inference
- Multiple Target Measurements: Bayesian Framework for Moving Object Detection in Mimo Radar
- Multiresolution Signal Processing of Financial Market Objects
- Multiscale Audio Spectrogram Transformer for Efficient Audio Classification
- Multispectral Image Fusion based on Super Pixel Segmentation
- Multistage Spatial Context Models for Learned Image Compression
- Multitask Detection of Speaker Changes, Overlapping Speech and Voice Activity Using Wav2vec 2.0
- Multitrack Music Transcription with a Time-Frequency Perceiver
- Multitrack Music Transformer
- Music Mixing Style Transfer: A Contrastive Learning Approach to Disentangle Audio Effects
- Music Rearrangement Using Hierarchical Segmentation
- Mutual Information Based Reweighting for Precipitation Nowcasting
- Mutually Guided Few-Shot Learning For Relational Triple Extraction
- MvCo-DoT: Multi-View Contrastive Domain Transfer Network for Medical Report Generation
- Möbius Total Variation for Directed Acyclic Graphs
- N2MVSNet: Non-Local Neighbors Aware Multi-View Stereo Network
- NAS-DYMC: NAS-Based Dynamic Multi-Scale Convolutional Neural Network for Sound Event Detection
- NBA-OMP: Near-Field Beam-Split-Aware Orthogonal Matching Pursuit for Wideband THz Channel Estimation
- NC-WAMKD: Neighborhood Correction Weight-Adaptive Multi-Teacher Knowledge Distillation for Graph-Based Semi-Supervised Node Classification
- NCL: Textual Backdoor Defense Using Noise-Augmented Contrastive Learning
- NF-PCAC: Normalizing Flow Based Point Cloud Attribute Compression
- NL-DSE: Non-Local Neural Network with Decoder-Squeeze-and-Excitation for Monocular Depth Estimation
- NNSVS: A Neural Network-Based Singing Voice Synthesis Toolkit
- NRTSI: Non-Recurrent Time Series Imputation
- NSV-TTS: Non-Speech Vocalization Modeling And Transfer In Emotional Text-To-Speech
- NVOC-22: A Low Cost Mel Spectrogram Vocoder for Mobile Devices
- Named Entity Detection and Injection for Direct Speech Translation
- Narrow Down Before Selection: A Dynamic Exclusion Model for Multiple-Choice QA
- Nasty-SFDA: Source Free Domain Adaptation from a Nasty Model
- Native Multi-Band Audio Coding Within Hyper-Autoencoded Reconstruction Propagation Networks
- Naturalistic Head Motion Generation from Speech
- Navigating and Reaching Therapeutic Goals with Dynamical Systems in Conversation-Based Interventions
- Near-field Localization with Dynamic Metasurface Antennas
- Neighborhood Information-Based Label Refinement for Person Re-Identification with Label Noise
- Nested Attention Network with Graph Filtering for Visual Question and Answering
- Networked Policy Gradient Play in Markov Potential Games
- Neural Architecture Search with Multimodal Fusion Methods for Diagnosing Dementia
- Neural Architecture of Speech
- Neural Band-to-Piano Score Arrangement with Stepless Difficulty Control
- Neural Diarization with Non-Autoregressive Intermediate Attractors
- Neural Feature Predictor and Discriminative Residual Coding for Low-Bitrate Speech Coding
- Neural Fourier Shift for Binaural Speech Rendering
- Neural Maximum-a-Posteriori Beamforming for Ultrasound Imaging
- Neural Mode Estimation
- Neural Network Models with Integrated Training and Adaptation For Nonlinear Acoustic System Identification
- Neural Networks with Quantization Constraints
- Neural Optimization Of Geometry And Fixed Beamformer For Linear Microphone Arrays
- Neural Source Coding For Bandwidth-Efficient Brain-Computer Interfacing With Wireless Neuro-Sensor Networks
- Neural Speech Phase Prediction Based on Parallel Estimation Architecture and Anti-Wrapping Losses
- Neural Transducer Training: Reduced Memory Consumption with Sample-Wise Computation
- Neural-AFC: Learning-Based Step-Size Control for Adaptive Feedback Cancellation with Closed-Loop Model Training
- Neurally Augmented State Space Model for Simultaneous Communication and Tracking with Low Complexity Receivers
- New Interpretable Patterns and Discriminative Features from Brain Functional Network Connectivity using Dictionary Learning
- Newton-Based Trainable Learning Rate
- Next-Speaker Prediction Based on Non-Verbal Information in Multi-Party Video Conversation
- No Reference Quality Assessment for Screen Content Images Based on Entire and High-Influence Regions
- Node-Wise Domain Adaptation Based on Transferable Attention for Recognizing Road Rage via EEG
- Noise PSD Insensitive RTF Estimation in a Reverberant and Noisy Environment
- Noise-Aware Target Extension with Self-Distillation for Robust Speech Recognition
- Noise-Disentanglement Metric Learning for Robust Speaker Verification
- Non-Convex Approaches for Low-Rank Tensor Completion under Tubal Sampling
- Noncoherent Multiuser Grassmannian Constellations for the Mimo Multiple Access Channel
- Nonnegative Block-Term Decomposition with the β-Divergence: Joint Data Fusion and Blind Spectral Unmixing
- Nonparallel Emotional Voice Conversion for Unseen Speaker-Emotion Pairs Using Dual Domain Adversarial Network & Virtual Domain Pairing
- Nonparallel High-Quality Audio Super Resolution with Domain Adaptation and Resampling CycleGANs
- Nord: Non-Matching Reference Based Relative Depth Estimation from Binaural Speech
- Not All Classes are Equal: Adaptively Focus-Aware Confidence for Semi-Supervised Object Detection
- Note and Playing Technique Transcription of Electric Guitar Solos in Real-World Music Performance
- Nowcasting of Extreme Precipitation Using Deep Generative Models
- Numerical Semantic Modeling for Implicit Discourse Relation Recognition
- OAFormer: Learning Occlusion Distinguishable Feature for Amodal Instance Segmentation
- OPT: One-shot Pose-Controllable Talking Head Generation
- OTW: Optimal Transport Warping for Time Series
- Oct Image Blind Despeckling Based on Gradient Guided Filter with Speckle Statistical Prior
- On Adversarial Robustness of Audio Classifiers
- On Batching Variable Size Inputs for Training End-to-End Speech Enhancement Systems
- On Bidirectional Preestimates and Their Application to Identification of fast Time-Varying Systems
- On Cross-Layer Alignment for Model Fusion of Heterogeneous Neural Networks
- On Crowdsourcing-Design with Comparison Category Rating for Evaluating Speech Enhancement Algorithms
- On Designing A 3d Imaging Summer Project For Ontario's High School Students During Covid-19 Pandemic
- On Designing Light-Weight Object Trackers Through Network Pruning: Use CNNS or Transformers?
- On Minimal Variations for Unsupervised Representation Learning
- On Multiple-Input/Binaural-Output Antiphasic Speaker Signal Extraction
- On Negative Sampling for Contrastive Audio-Text Retrieval
- On Neural Architectures for Deep Learning-Based Source Separation of Co-Channel OFDM Signals
- On Out-of-Distribution Detection for Audio with Deep Nearest Neighbors
- On Parametric Misspecified Bayesian Cramér-Rao Bound: An Application to Linear/Gaussian Systems
- On Super-Resolution with Separation Prior
- On The Design and Training Strategies for Rnn-Based Online Neural Speech Separation Systems
- On The Detection of Synthetic Images Generated by Diffusion Models
- On The Fairness of Multitask Representation Learning
- On The Primal and Dual Formulations Of The Discrete Mumford-Shah Functional
- On Tracking a Stochastically Time-Varying Subspace
- On Unsupervised Uncertainty-Driven Speech Pseudo-Label Filtering and Model Calibration
- On Using the UA-Speech and Torgo Databases to Validate Automatic Dysarthric Speech Classification Approaches
- On Weighted Cross-Entropy for Label-Imbalanced Separable Data: An Algorithmic-Stability Study
- On Word Error Rate Definitions and Their Efficient Computation for Multi-Speaker Speech Recognition Systems
- On the Effectiveness of Monoaural Target Source Extraction for Distant end-to-end Automatic Speech Recognition
- On the Importance of Different Cough Phases for COVID-19 Detection
- On the Joint Estimation of Phase Noise and time-Varying Channels for OFDM under High-Mobility Conditions
- On the Minimum Perimeter Criterion for Bounded Component Analysis
- On the Quantization of Recurrent Neural Networks for Smiles Generation
- On the Reduction of Large-Scale Room Acoustic Models
- On the Relevance of the Differences Between HRTF Measurement Setups for Machine Learning
- On the Robustness of Non-Intrusive Speech Quality Model by Adversarial Examples
- On the Role of LIP Articulation in Visual Speech Perception
- On the Role of Visual Context in Enriching Music Representations
- On the Value of Stochastic Side Information in Online Learning
- On-the-Fly Text Retrieval for end-to-end ASR Adaptation
- Once-for-All Sequence Compression for Self-Supervised Speech Models
- One-Shot Action Detection via Attention Zooming In
- One-Shot Medical Action Recognition With A Cross-Attention Mechanism And Dynamic Time Warping
- One-Shot Neural Band Selection for Spectral Recovery
- Online Binaural Speech Separation Of Moving Speakers With A Wavesplit Network
- Online Caching with Fetching cost for Arbitrary Demand Pattern: a Drift-Plus-Penalty Approach
- Online Edge Flow Prediction Over Expanding Simplicial Complexes
- Online Learning-Based Waveform Selection for Improved Vehicle Recognition in Automotive Radar
- Online Model Compression for Federated Learning with Large Models
- Online Residual-Based Key Frame Sampling with Self-Coach Mechanism and Adaptive Multi-Level Feature Fusion
- Online Vector Autoregressive Models Over Expanding Graphs
- Ontology-Aware Network for Zero-Shot Sketch-Based Image Retrieval
- Open-Set Automatic Target Recognition
- Optimal Carrier Frequency Design for Frequency Diverse Array Mimo Radar
- Optimal Compression for Minimizing Classification Error Probability: An Information-Theoretic Approach
- Optimal Condition Training for Target Source Separation
- Optimal Kernel for Real-Time Arbitrary-Shaped Text Detection
- Optimal Mixed-ADC Arrangement for DOA Estimation Via CRB Using ULA
- Optimal Transport in Diffusion Modeling for Conversion Tasks in Audio Domain
- Optimal Transport with a Diversified Memory Bank for Cross-Domain Speaker Verification
- Optimising Different Feature Types for Inpainting-Based Image Representations
- Optimization for Robustness Evaluation Beyond ℓp Metrics
- Optimization of Sensor Configurations for Fault Identification in Smart Buildings
- Optimization of the Deep Neural Networks for Seizure Detection
- Optimized Dithering for Quantization Index Modulation
- Optimized Quality Feature Learning for Video Quality Assessment
- Optimizing Distributed Multi-Sensor Multi-Target Tracking Algorithm Based On Labeled Multi-Bernoulli Filter
- Optimizing Quantum Federated Learning Based on Federated Quantum Natural Gradient Descent
- Optimizing Vision Transformers for Medical Image Segmentation
- Order Reduction of Multi-Channel FIR Filters by Balanced Truncation
- Outlier-Insensitive Kalman Filtering Using NUV Priors
- Output-Dependent Gaussian Process State-Space Model
- Outside Knowledge Visual Question Answering Version 2.0
- Overcoming Posterior Collapse in Variational Autoencoders Via EM-Type Training
- Overcoming the Seesaw in Monocular 3D Object Detection Via Language Knowledge Transferring
- Overlay Cognitive Radio Using Symbol Level Precoding With Quantized CSI
- Overview of the 2023 ICASSP SP Clarity Challenge: Speech Enhancement for Hearing Aids
- Overview of the ICASSP 2023 General Meeting Understanding and Generation Challenge (MUG)
- Overview of the L3DAS23 Challenge on Audio-Visual Extended Reality
- PAGE: A Position-Aware Graph-Based Model for Emotion Cause Entailment in Conversation
- PCF: ECAPA-TDNN with Progressive Channel Fusion for Speaker Verification
- PCQA-Graphpoint: Efficient Deep-Based Graph Metric for Point Cloud Quality Assessment
- PCSalmix: Gradient Saliency-Based Mix Augmentation for Point Cloud Classification
- PFT-SSR: Parallax Fusion Transformer for Stereo Image Super-Resolution
- PI-Trans: Parallel-Convmlp and Implicit-Transformation Based Gan for Cross-View Image Translation
- PMMSD: Development of the Matrix Sentence Intelligibility Dataset for Mandarin with Lombard Effect
- PMNet: Large-Scale Channel Prediction System for ICASSP 2023 First Pathloss Radio Map Prediction Challenge
- POINTACL: Adversarial Contrastive Learning for Robust Point Clouds Representation Under Adversarial Attack
- PQLM - Multilingual Decentralized Portable Quantum Language Model
- PRIME: 3D Human Pose and Body Shape Recovery with Perspective Projection
- PRRD: Pixel-Region Relation Distillation For Efficient Semantic Segmentation
- PU-Edgeformer: Edge Transformer for Dense Prediction in Point Cloud Upsampling
- PUFFIN: Pitch-Synchronous Neural Waveform Generation for Fullband Speech on Modest Devices
- Paaploss: A Phonetic-Aligned Acoustic Parameter Loss for Speech Enhancement
- Pair DETR: Toward Faster Convergent DETR
- Papez: Resource-Efficient Speech Separation with Auditory Working Memory
- Parafac2-Based Coupled Matrix and Tensor Factorizations
- Parallel 2D Seismic Ray Tracing Using Cuda on a Jetson Nano
- Parallel Sentence-Level Explanation Generation for Real-World Low-Resource Scenarios
- Parameter Efficient Transfer Learning for Various Speech Processing Tasks
- Parameter-Efficient Transfer Learning of Pre-Trained Transformer Models for Speaker Verification Using Adapters
- Parasympathetic-Sympathetic Causal Interactions and Perceived Workload for Varying Difficulty Affective Computing Tasks
- Partially Adaptive Multichannel Joint Reduction of Ego-Noise and Environmental Noise
- Particle Flow Gaussian Sum Particle Filter
- Passive Acoustic Tracking of Whales in 3-D
- Passive Detection of Rank-One Gaussian Signals for Known Channel Subspaces and Arbitrary Noise
- Peak-First CTC: Reducing the Peak Latency of CTC Models by Applying Peak-First Regularization
- Perceive and Predict: Self-Supervised Speech Representation Based Loss Functions for Speech Enhancement
- Perceptual Analysis of Speaker Embeddings for Voice Discrimination between Machine And Human Listening
- Perceptual Quality Assessment for Digital Human Heads
- Perceptual-Neural-Physical Sound Matching
- Performance Above All? Energy Consumption vs. Performance, a Study on Sound Event Detection with Heterogeneous Data
- Performance Comparison of TTS Models for Brazilian Portuguese to Establish a Baseline
- Performance of Social Machine Learning Under Limited Data
- Performing Neural Architecture Search Without Gradients
- Period VITS: Variational Inference with Explicit Pitch Modeling for End-To-End Emotional Speech Synthesis
- Permutation Invariant Training for Paraphrase Identification
- Person Identification with Wearable Sensing Using Missing Feature Encoding and Multi-Stage Modality Fusion
- Personalized Federated Learning on Long-Tailed Data via Adversarial Feature Augmentation
- Personalized Lightweight Text-to-Speech: Voice Cloning with Adaptive Structured Pruning
- Personalized Speech Enhancement Combining Band-Split RNN and Speaker Attentive Module
- Personalized Task Load Prediction in Speech Communication
- Personalizing Federated Learning with Over-The-Air Computations
- Perspective Projection-Based 3d CT Reconstruction from Biplanar X-Rays
- Phase Retrieval for Rydberg Quantum Arrays
- Phase Unwrapping in Correlated Noise for FMCW Lidar Depth Estimation
- Phase-Aware Spoof Speech Detection Based On Res2net with Phase Network
- PhaseAug: A Differentiable Augmentation for Speech Synthesis to Simulate One-to-Many Mapping
- Phonation Mode Detection in Singing: A Singer Adapted Model
- Phoneix: Acoustic Feature Processing Strategy for Enhanced Singing Pronunciation With Phoneme Distribution Predictor
- Phoneme-Level Bert for Enhanced Prosody of Text-To-Speech with Grapheme Predictions
- Phonetic Anchor-Based Transfer Learning to Facilitate Unsupervised Cross-Lingual Speech Emotion Recognition
- Phonetic RNN-Transducer for Mispronunciation Diagnosis
- Physics-Informed Transfer Learning for Voltage Stability Margin Prediction
- Picking the Underused Heads: A Network Pruning Perspective of Attention Head Selection for Fusing Dialogue Coreference Information
- Piecewise Position Encoding in Convolutional Neural Network for Cough-Based Covid-19 Detection
- Pitch Mark Detection from Noisy Speech Waveform Using Wave-U-Net
- Play It Back: Iterative Attention For Audio Recognition
- Polarized Signal Singular Spectrum Analysis with Complex SSA
- Police: Provably Optimal Linear Constraint Enforcement For Deep Neural Networks
- Pondering About Task Spatial Misalignment: Classification-Localization Equilibrated Object Detection
- Pooling Strategies for Simplicial Convolutional Networks
- Pop2Piano : Pop Audio-Based Piano Cover Generation
- Position-Aware Graph-Based Learning of Whole Slide Images
- Positive-Pair Redundancy Reduction Regularisation for Speech-Based Asthma Diagnosis Prediction
- Possibilistic Bernoulli Filter for Extended Target Tracking
- Post-Trained Language Model Adaptive to Extractive Summarization of Long Spoken Documents
- Powerful and Extensible WFST Framework for Rnn-Transducer Losses
- Practice of the Conformer Enhanced Audio-Visual Hubert on Mandarin and English
- Pre-Trained Model Representations and Their Robustness Against Noise for Speech Emotion Analysis
- Pre-Training Strategies Using Contrastive Learning and Playlist Information for Music Classification and Similarity
- Precognition in Contextual Spoken Language Understanding via Knowledge Distillation
- Predicting Brain Age Using Transferable Covariance Neural Networks
- Predicting Multi-Codebook Vector Quantization Indexes for Knowledge Distillation
- Predictive Skim: Contrastive Predictive Coding for Low-Latency Online Speech Separation
- Prefallkd: Pre-Impact Fall Detection Via CNN-ViT Knowledge Distillation
- Prefix Tuning for Automated Audio Captioning
- Prefix-Level Detection and Autocorrection of Keyboard Input Errors
- Preformer: Predictive Transformer with Multi-Scale Segment-Wise Correlations for Long-Term Time Series Forecasting
- Preserving Background Sound in Noise-Robust Voice Conversion Via Multi-Task Learning
- Pretrained Transformers for Seizure Detection
- Pretraining Conformer with ASR for Speaker Verification
- Prior-Enhanced Temporal Action Localization Using Subject-Aware Spatial Attention
- Priv-Aug-Shap-ECGResNet: Privacy Preserving Shapley-Value Attributed Augmented Resnet for Practical Single-Lead Electrocardiogram Classification
- Privacy Preserving Face Recognition with Lensless Camera
- Privacy-Enhanced Federated Learning Against Attribute Inference Attack for Speech Emotion Recognition
- Privacy-Preserving Automatic Speaker Diarization
- Privacy-Preserving Occupancy Estimation
- Probabilistic Back-ends for Online Speaker Recognition and Clustering
- Procontext: Exploring Progressive Context Transformer for Tracking
- Procter: Pronunciation-Aware Contextual Adapter For Personalized Speech Recognition In Neural Transducers
- Product Graph Learning From Multi-Attribute Graph Signals with Inter-Layer Coupling
- Progressive Diversifying Policy for Multi-Agent Reinforcement Learning
- Progressive Meta-Pooling Learning for Lightweight Image Classification Model
- Progressive Multi-Stage Neural Audio Codec with Psychoacoustic Loss and Discriminator
- Progressive Perception Learning for Distribution Modulation in Siamese Tracking
- Progressive Refinement Learning Based on Feature Cross Perception for Residential Areas Semantic Segmentation
- Projected Hierarchical ALS for Generalized Boolean Matrix Factorization
- Promoting Cooperation in Multi-Agent Reinforcement Learning via Mutual Help
- Prompt Makes mask Language Models Better Adversarial Attackers
- Prompt-Distiller: Few-Shot Knowledge Distillation for Prompt-Based Language Learners with Dual Contrastive Learning
- Prompttts: Controllable Text-To-Speech With Text Descriptions
- Prosody Is Not Identity: A Speaker Anonymization Approach Using Prosody Cloning
- Prosody-Aware Speecht5 for Expressive Neural TTS
- Prosody-Controllable Spontaneous TTS with Neural HMMS
- Prototype Knowledge Distillation for Medical Segmentation with Missing Modality
- Prototype-Based Layered Federated Cross-Modal Hashing
- Provable Computational and Statistical Guarantees for Efficient Learning of Continuous-Action Graphical Games
- Provably Convergent Plug & Play Linearized ADMM, Applied to Deblurring Spatially Varying Kernels
- Prune Then Distill: Dataset Distillation with Importance Sampling
- Pseudo Multi-Source Domain Extension and Selective Pseudo-Labeling for Unsupervised Domain Adaptive Medical Image Segmentation
- Pseudo-Inverted Bottleneck Convolution for Darts Search Space
- Pseudo-Query Generation For Semi-Supervised Visual Grounding With Knowledge Distillation
- Pushing the Limits of Self-Supervised Speaker Verification using Regularized Distillation Framework
- Pyramid Dynamic Inference: Encouraging Faster Inference Via Early Exit Boosting
- Pyramid Spatial Feature Transform and Shared-Offsets Deformable Alignment Based Convolutional Network for HDR Imaging
- QI-TTS: Questioning Intonation Control for Emotional Speech Synthesis
- QTROJAN: A Circuit Backdoor Against Quantum Neural Networks
- Quantifying Catastrophic Forgetting in Continual Federated Learning
- Quantile Online Learning for Semiconductor Failure Analysis
- Quantitative Evidence on Overlooked Aspects of Enrollment Speaker Embeddings for Target Speaker Separation
- Quantized Precoding and RIS-Assisted Modulation for Integrated Sensing and Communications Systems
- Quantpipe: Applying Adaptive Post-Training Quantization For Distributed Transformer Pipelines In Dynamic Edge Environments
- Quantum Deep Recurrent Reinforcement Learning
- Quantum Graph Transformers
- Quantum Transfer Learning Using the Large-Scale Unsupervised Pre-Trained Model Wavlm-Large for Synthetic Speech Detection
- Quantum Variational Bayes on Manifolds
- Quaternion Orthogonal Transformer for Facial Expression Recognition in the Wild
- Query-Utterance Attention With Joint Modeuing For Query-Focused Meeting Summarization
- Question Answering System with Sparse and Noisy Feedback
- Quickest Change Detection with Leave-one-out Density Estimation
- RAT: Radial Attention Transformer for Singing Technique Recognition
- RCDPT: Radar-Camera Fusion Dense Prediction Transformer
- RD-NAS: Enhancing One-Shot Supernet Ranking Ability Via Ranking Distillation From Zero-Cost Proxies
- RDO Candidate Selection for Maximizing Coding Efficiency in a Practical HEVC Encoder
- RGB-D Based Pose-Invariant Face Recognition Via Attention Decomposition Module
- RIS Reflection and Placement Optimisation for Underlay D2D Communications in Cognitive Cellular Networks
- RIS-Aided Wideband DFRC with Reconfigurable Holographic Surface
- RL-IFF: Indoor Localization via Reinforcement Learning-Based Information Fusion
- RNN-Based Step-Size Estimation for the RLS Algorithm with Application to Acoustic Echo Cancellation
- ROI-Based Deep Image Compression with Swin Transformers
- Radar Clutter Covariance Estimation: A Nonlinear Spectral Shrinkage Approach
- Radio Map Based UAV Target Localization
- Radio Sensing with Large Intelligent Surface for 6G
- Radio-Astronomy Imaging and Interference Excision Using Tensor Decomposition and Canonical Correlation Analysis
- Rain2Avoid: Self-Supervised Single Image Deraining
- Raising The Limit of Image Rescaling Using Auxiliary Encoding
- Randmasking Augment: A Simple and Randomized Data Augmentation For Acoustic Scene Classification
- Random Projector: Efficient Deep Image Prior
- Range-ISL Minimization and Spectral Shaping in MIMO Radar Systems via Waveform Design
- Rapid Audiometric Evaluation for Personalized Headphone Listening
- Rate Region Characterization for Semantics and Bits based Multiuser Communications
- Rate Splitting and Precoding Strategies for Multi-User MIMO Broadcast Channels with Common and Private Streams
- Rate-Distortion Optimization with Alternative References for UGC Video Compression
- Rate-Distortion Optimized Variable-Node-size Trisoup for Point Cloud Coding
- Raw Ultrasound-Based Phonetic Segments Classification Via Mask Modeling
- Real-Time Audio-Visual End-To-End Speech Enhancement
- Real-Time Human Reconstruction Based on Human Pose Prior and Epipolar Refinement
- Real-Time MRI Video Synthesis from Time Aligned Phonemes with Sequence-to-Sequence Networks
- Real-Time Modelling of Observation Filter in the Remote Microphone Technique for an Active Noise Control Application
- Real-Time Multichannel Speech Separation and Enhancement Using a Beamspace-Domain-Based Lightweight CNN
- Real-Time Speech Enhancement with Dynamic Attention Span
- Real-Time Speech Interruption Analysis: from Cloud to Client Deployment
- Real-Time Target Sound Extraction
- Real-Time Wireless ECG-Derived Respiration Rate Estimation using an Autoencoder with a DCT Layer
- Received Power Maximization with Practical Phase-Dependent Amplitude Response in RIS-Aided OFDM Wireless Communications
- Receptive Field Reliant Zero-Cost Proxies for Neural Architecture Search
- Recouple Event Field via Probabilistic Bias for Event Extraction
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.