ICASSP 2024 Accepted Papers
The full list of 2,679 papers accepted at ICASSP 2024 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- "It os Okay to be Uncommon": Quantizing Sound Event Detection Networks on Hardware Accelerators with Uncommon Sub-Byte Support
- 1-D Spatial Attention in Binarized Convolutional Neural Networks
- 2D Human Pose Estimation Calibration and Keypoint Visibility Classification
- 3-D Near-Field Localization by Jointly Exploiting Spatial and Temporal Information Based on a Nonuniform Cross Array
- 3D Automated Quantitative Calculations Based on CT Images of the Hip Joint
- 3D Hand Joint and Grasping Estimation for Teleoperation System
- 3D Parallelism for Transformers via Integer Programming
- 3D Point Cloud Semantic Segmentation Based on Diffusion Model
- 3D Pose Estimation from Monocular Video with Camera-Bone Angle Regularization on the Image Feature
- 3DSAM: Segment Anything in NeRF
- 3M-Transformer: A Multi-Stage Multi-Stream Multimodal Transformer for Embodied Turn-Taking Prediction
- 3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications
- 6DoF SELD: Sound Event Localization and Detection Using Microphones and Motion Tracking Sensors on Self-Motioning Human
- A 3D Virtual Try-On Method with Global-Local Alignment and Diffusion Model
- A Bayesian Approach to High-Order Link Prediction
- A Bi-Pyramid Multimodal Fusion Method for the Diagnosis Of Bipolar Disorders
- A Binary BP Decoding Using Posterior Adjustment for Quantum LDPC Codes
- A Birgat Model for Multi-Intent Spoken Language Understanding with Hierarchical Semantic Frames
- A CCM-Based Joint DOA-Frequency Estimation and Signal Recovery with Efficient Sub-Nyquist Sampling
- A Chat about Boring Problems: Studying GPT-Based Text Normalization
- A Closer Look at Wav2vec2 Embeddings for On-Device Single-Channel Speech Enhancement
- A Codec-Based Approach for Video Life-Cycle Characterization in Social Networks
- A Comparative Analysis of Poetry Reading Audio: Singing, Narrating, or Somewhere in Between?
- A Comparative Study on Annotation Quality of Crowdsourcing and LLm Via Label Aggregation
- A Comparison of Parameter-Efficient ASR Domain Adaptation Methods for Universal Speech and Language Models
- A Complete Method for the 3D Reconstruction of Axonal Pathways from 2 Orthogonal 3D OCT Images of the Lamina Cribrosa
- A Comprehensive Framework for Occluded Human Pose Estimation
- A Computationally Efficient Semi-Blind Source Separation Approach for Nonlinear Echo Cancellation Based on an Element-Wise Iterative Source Steering
- A Concept for a Slam Back End Hardware Accelerator
- A Contrario Paradigm for Yolo-Based Infrared Small Target Detection
- A Convergent Primal-Dual Deep Plug-and-Play Algorithm for Constrained Image Restoration
- A Counterfactual Inspired Framework For Quantifying Edge Effects On Gnns Fairness
- A Cross Search Method for Data Augmentation in Neural Machine Translation
- A Crowdsourcing Approach to Video Quality Assessment
- A Deep Representation Learning-Based Speech Enhancement Method Using Complex Convolution Recurrent Variational Autoencoder
- A DenseNet-Based Method for Decoding Auditory Spatial Attention with EEG
- A Density-Guided Temporal Attention Transformer for Indiscernible Object Counting in Underwater Videos
- A Detailed Audio-Text Data Simulation Pipeline Using Single-Event Sounds
- A Distributed Joint Integrated Probabilistic Data Association (JIPDA) Filter with Soft Object Association
- A Dual-Path Framework with Frequency-and-Time Excited Network for Anomalous Sound Detection
- A Facial Expression Transfer Method Based on 3DMM and Diffusion Models
- A Fast Blind Deblurring Algorithm Using Local Gradient Product Prior
- A Fast, Performant, Secure Distributed Training Framework For LLM
- A Federated Graph to Embedding Approach for Knowledge Graph Completion
- A Fine-Grained Attribute Pre-Labeling Method Based on Label Dependency and Feature Similarity Dynamics
- A Fine-Grained Tri-Modal Interaction Model for Multimodal Sentiment Analysis
- A Flexible Online Framework for Projection-Based Stft Phase Retrieval
- A Foundation Model for Music Informatics
- A Framework for Portrait Stylization with Skin-Tone Awareness and Nudity Identification
- A Fully Differentiable Model for Unsupervised Singing Voice Separation
- A General Framework for Rotation Invariant Point Cloud Analysis
- A Generative Adversarial Framework for Dialogue Generation with Neural Architecture Search
- A Gibbs Sampler for Bayesian Nonparametric State-Space Models
- A Graph Neural Network Based Approach for Fault Delineation in Seismic Data using Graph Total Variation and Multigraph
- A Graph Neural Network Based Fusion of MRI-Derived Brain Network and Clinical Data for Glioblastoma Survival Prediction
- A Graph-Prediction-Based Approach for Debiasing Underreported Data
- A Green Learning Approach to Spoofed Speech Detection
- A Guided Upsampling Network for Short wave Infrared Images Using Graph Regularization
- A Hierarchical Multi-Proxy Loss with Dynamic Main-Proxy for Deep Metric Learning
- A Hybrid CNN-Transformer for Focal Liver Lesion Classification
- A Hybrid Deep-Online Learning Based Method for Active Noise Control in Wave Domain
- A Hybrid Slow-Time Coding Framework for Automotive MIMO Radar
- A Joint Data Compression and Time-Delay Estimation Distributed Systems via Extremum Encoding
- A Joint Look on Lunar Satellite and Cooperative Surface PNT
- A Keyless Extraction Framework Targeting at Deep Learning Based Image-Within-Image Models
- A Learning Resource Recommendation Algorithm Based on Online Learning Behavior
- A Learning-Based Multi-Node Fusion Positioning Method Using Wearable Inertial Sensors
- A Learning-Based System for Automatic Intentional Non-Adherence Detection from Dosing Videos
- A Light-Weight State Detection Model for Kalman-Filter-Based Acoustic Feedback Cancellation with Rapid Recovery from Abrupt Path Changes
- A Lightweight Change Detection Method Based on Feature Interaction and Transformer for High Resolution Remote Sensing Images
- A Lightweight Hybrid Multi-Channel Speech Extraction System with Directional Voice Activity Detection
- A Low-Latency Fft-Ifft Cascade Architecture
- A Machine-Learning Model for Detecting Depression, Anxiety, and Stress from Speech
- A Meta-Preconditioning Approach for Deep Q-Learning
- A Method for Bilevel Optimization with Convex Lower-Level Problem
- A Method for X-Ray Image Landmarks Localization using Cyclic Coordinate-Guided Strategy
- A Modified Cramér-Rao Bound for Discrete-Time Markovian Dynamic Systems
- A Multi-Carrier Information Hiding Algorithm Based on Layered Compression of 3d Point Cloud Model
- A Multi-Scale Bimodal Fusion Network for Robust and Accurate Online Handwriting Recognition
- A Multimodal Approach to Device-Directed Speech Detection with Large Language Models
- A Multiscale Objective Function for Camera Color Correction
- A Near-Field Source Localization Method for Uniform/Sparse Centrally Symmetric Rectangular Arrays
- A Neural Syntax Parser for Coronary Artery Anatomical Labeling in Coronary CT Angiography
- A Neurophysiological-Auditory "Listen Receipt" for Communication Enhancement
- A New Fourth-Order Sparse Array Generator Based on Sum-Difference Co-Array Analysis
- A New Perspective on Understanding Resolution Limit Via an Asymptotic Study of Christoffel-Darboux Kernel Based Spectrum Estimator
- A New Pre-Training Paradigm for Offline Multi-Agent Reinforcement Learning with Suboptimal Data
- A New Similarity-Based Relational Knowledge Distillation Method
- A Novel 3-D Focusing Scheme for Distributed SAR Tomography
- A Novel Cascade Instruction Tuning Method for Biomedical NER
- A Novel Contrastive Diffusion Graph Convolutional Network for Few-Shot Skeleton-Based Action Recognition
- A Novel Cross-Sensor Self-Supervised Learning Method for Rotating Machinery Fault Diagnosis
- A Novel Demodulation and Selection Pilot Power Trade-Off for Codebook-Based IRS with Imperfect Channel Estimates
- A Novel Discrete Fractional Complex Hadamard Transform for Medical Image Encryption
- A Novel Iterative Thresholding Algorithm for Arctangent Regularization Problem
- A Novel Local-Global Feature Fusion Framework for Body-Weight Exercise Recognition with Pressure Mapping Sensors
- A Novel Medical Image Fusion Framework Integrating Multi-scale Encoder-Decoder with Discrete Wavelet Decomposition
- A Novel Multi-Atlas Fusion Model Based On Contrastive Learning For Functional Connectivity Graph Diagnosis
- A Novel Multimodal Sentiment Analysis Model Based on Gated Fusion and Multi-Task Learning
- A Novel Residual-Guided Learning Method for Image Steganography
- A One-Class Approach to Detect Super-Resolution Satellite Imagery with Spectral Features
- A PLS-Integrated Lasso Method With Application in Index Tracking
- A Parameterized Generative Adversarial Network Using Cyclic Projection for Explainable Medical Image Classifications
- A Practical Online Multichannel Dereverberation Approach with Data-Reuse Technique
- A Prior Driven Semi-Supervised ViTGAN for Image Recolorization
- A Probability Gradient Based Approach for Sampling Boundaries of In-Domain Data
- A Prompt-Based Method with Multi-View Optimization for Open Relation Extraction
- A Property-Guided Diffusion Model For Generating Molecular Graphs
- A Real-Time Active Speaker Detection System Integrating an Audio-Visual Signal with a Spatial Querying Mechanism
- A Real-Time Lyrics Alignment System Using Chroma and Phonetic Features for Classical Vocal Performance
- A Real-Time Video Quality Metric for HTTP Adaptive Streaming
- A Reconstruction-Based Feature Adaptation for Anomaly Detection with Self-Supervised Multi-Scale Aggregation
- A Reduced-Reference Quality Assessment Metric for Textured Mesh Digital Humans
- A Relation-Aware Heterogeneous Graph Transformer on Dynamic Fusion for Multimodal Classification Tasks
- A Riemannian-Based Joint Design Framework of Mimo Radar Transmit Waveform And Receive Filter Via Information Theory
- A Robust Audio Deepfake Detection System via Multi-View Feature
- A Robust GLRT Detector Against Missing Data in Cooperative Sensing
- A Robust Pitch-Fusion Model for Speech Emotion Recognition in Tonal Languages
- A Robust Quantile Huber Loss with Interpretable Parameter Adjustment in Distributional Reinforcement Learning
- A Robust and Scalable Method with an Analytic Solution for Multi-Subject FMRI Data Analysis
- A Saliency Enhanced Feature Fusion Based Multiscale RGB-D Salient Object Detection Network
- A Scalable Sparse Transformer Model for Singing Melody Extraction
- A Self-Supervised Pressure Map Human Keypoint Detection Approch: Optimizing Generalization and Computational Efficiency Across Datasets
- A Separation Priority Pipeline for Single-Channel Speech Separation in Noisy Environments
- A Sequential Averaging Plug-and-Play Method for Image Restoration Via Fixed-Point Projection
- A Simple and Effective Method for Anomaly Detection on Attributed Graphs via Feature Consistency
- A Smoothed Bregman Proximal Gradient Algorithm for Decentralized Nonconvex Optimization
- A Soft Contrastive Learning-Based Prompt Model for Few-Shot Sentiment Analysis
- A Sound Approach: Using Large Language Models to Generate Audio Descriptions for Egocentric Text-Audio Retrieval
- A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition
- A Speaker Recognition Method Based on Stable Learning
- A Spectral Analysis of Graph Neural Networks on Dense and Sparse Graphs
- A Statistical Characterization Of Communication Performance In RIS-Aided Networks
- A Steered Response Power Approach with Bilinear Prediction-Based Trade-Off Prewhitening for Speaker Localization
- A Stochastic Gradient Approach for Communication Efficient Confederated Learning
- A Stochastic Proximal WMMSE for Ergodic Sum Rate Maximization
- A Study of Mispronunciation Detection and Diagnosis Based on Meta-Learning
- A Study of Multichannel Spatiotemporal Features and Knowledge Distillation on Robust Target Speaker Extraction
- A Study on Combining Non-Parallel and Parallel Methodologies for Mandarin-English Cross-Lingual Voice Conversion
- A Study on Graph Embedding for Speaker Recognition
- A Study on the Adverse Impact of Synthetic Speech on Speech Recognition
- A Supervised Information Enhanced Multi-Granularity Contrastive Learning Framework for EEG Based Emotion Recognition
- A Targeted Adversarial Attack Method for Multi-Classification Malicious Traffic Detection
- A Transformer Approach for Polyphonic Audio-to-Score Transcription
- A Tri-Dynamic Preprocessing Framework for UGC Video Compression
- A Two-Stage Dehazing Framework Based on Inverted Image Curve-Enhancement
- A Two-Stage Framework in Cross-Spectrum Domain for Real-Time Speech Enhancement
- A Unified DNN-Based System for Industrial Pipeline Segmentation
- A Unified Framework for Multi-Intent Spoken Language Understanding with Prompting
- A Unified Front-End Framework for English Text-to-Speech Synthesis
- A Unified Loss Function to Tackle Inter-Class and Intra-Class Data Imbalance in Sound Event Detection
- A Variable Smoothing for Nonconvexly Constrained Nonsmooth Optimization with Application to Sparse Spectral Clustering
- A Wasserstein Graph Distance Based on Distributions of Probabilistic Node Embeddings
- A Weighted-Variance Variational Autoencoder Model for Speech Enhancement
- AAT: Adapting Audio Transformer for Various Acoustics Recognition Tasks
- ADHD Diagnosis and Biomarker Detection Based on Multimodal Graph Convolutional Neural Network
- ADIFT: Zero-Shot Generative Model Adaption Via Adaptive Domain-Invariant Feature Transfer
- ADVSV: An Over-the-Air Adversarial Attack Dataset for Speaker Verification
- AEAM3D: Adverse Environment-Adaptive Monocular 3D Object Detection via Feature Extraction Regularization
- AEGIS-Net: Attention-Guided Multi-Level Feature Aggregation for Indoor Place Recognition
- AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition
- AHRNET: Attention and Heatmap-Based Regressor for Hand Pose Estimation and Mesh Recovery
- ANM-Based Source Localization Under Mixed Field
- AQF: Assessing the Quality of Hyperspectral Reconstruction with a Learnable Metric
- ARFA: An Asymmetric Receptive Field Autoencoder Model for Spatiotemporal Prediction
- AS-pVAD: A Frame-Wise Personalized Voice Activity Detection Network with Attentive Score Loss
- ASPED: An Audio Dataset for Detecting Pedestrians
- AUTOSGM: A Unified Lowpass Regularization Framework for Accelerated Learning
- AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models
- AV2WAV: Diffusion-Based Re-Synthesis from Continuous Self-Supervised Features for Audio-Visual Speech Enhancement
- Accelerated Recovery of Spectrally Sparse Signals Viamodified Proximal Gradient in Hankel Space
- Accelerating Gradient Descent for Over-Parameterized Asymmetric Low-Rank Matrix Sensing via Preconditioning
- Accent-Specific Vector Quantization for Joint Unsupervised and Supervised Training in Accent Robust Speech Recognition
- Accurate Gigapixel Crowd Counting by Iterative Zooming and Refinement
- Accurate Interpolation of Scattered Data Via Learning Relation Graph
- Accurate and Robust Scene Text Recognition via Adversarial Training
- Acoustic BPE for Speech Generation with Discrete Tokens
- Activation Compression of Graph Neural Networks Using Block-Wise Quantization with Improved Variance Minimization
- Active Explainable Recommendation with Limited Labeling Budgets
- Active Learning for Sound Event Classification Using Bayesian Neural Networks with Gaussian Variational Posterior
- Active Learning with Core-Set Sampling and Scale-Sensitive Loss for 3D Object Detection
- Active Noise Control Over 3D Space with A Dynamic Noise Source
- Active Noise Control Over A Large Region with Multiple Spherical Microphone Arrays In Wave Domain
- Activity Recognition Method Based on Kernel Supervised Laplacian Eigenmaps
- AdaFL: Adaptive Client Selection and Dynamic Contribution Evaluation for Efficient Federated Learning
- AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition
- AdaPlus: Integrating Nesterov Momentum and Precise Stepsize Adjustment on Adamw Basis
- Adapter-Based Incremental Learning for Face Forgery Detection
- Adapting Frechet Audio Distance for Generative Music Evaluation
- Adapting Large Language Model with Speech for Fully Formatted End-to-End Speech Recognition
- Adapting Pitch-Based Self Supervised Learning Models for Tempo Estimation
- Adaptive Chroma Block Vector Derivation from Luma for Screen Content Coding
- Adaptive Confidence Multi-View Hashing for Multimedia Retrieval
- Adaptive Data Augmentation for Aspect Sentiment Quad Prediction
- Adaptive Fourier Decomposition Based Signal Extraction on Weak Electromagnetic Field
- Adaptive Gaussian Regularization Constrained Sparse Subspace Clustering for Image Segmentation
- Adaptive Grid 2-D Direction of Arrival Estimation Method Using an Integrated Dictionary
- Adaptive Head Pose Estimation with Real-Time Structured Light
- Adaptive Image-Enhanced Knowledge Graph Completion
- Adaptive Joint Channel Estimation/Data Detection in Flexible Multicarrier Mimo Systems - A Tensor-Based Approach
- Adaptive Kalmannet: Data-Driven Kalman Filter with Fast Adaptation
- Adaptive Multi-Armed Bandit Learning for Task Offloading in Mobile Edge Computing
- Adaptive Multi-Exposure Fusion for Enhanced Neural Radiance Fields
- Adaptive Multi-View Joint Contrastive Learning on Graphs
- Adaptive Multiview Community-Preserved Graph Convolutional Network for Multiatlas-Based Functional Connectivity Analysis
- Adaptive Order Aggregator and Extractor Graph Neural Network
- Adaptive Parameter Sharing for Multi-Agent Reinforcement Learning
- Adaptive Pedestrian Trajectory Prediction via Target-Directed Angle Augmentation
- Adaptive Prompt Construction Method for Relation Extraction
- Adaptive Quantization with Mixed-Precision Based on Low-Cost Proxy
- Adaptive Reweighted Sparse Belief Propagation Decoding for Polar Codes
- Adaptive Secondary Transform Sets for Video Coding Beyond AV1
- Adaptive Sensor Selection with Deterministic Priors for DoA Tracking
- Adaptive Spatial-Temporal Hypergraph Fusion Learning for Next POI Recommendation
- Adaptive Speech Emotion Representation Learning Based On Dynamic Graph
- Adaptive Super Resolution for One-Shot Talking-Head Generation
- Adaptive Video Watermarking with Perceptual Guarantee and Efficiency Optimization
- Adaptive-Avg-Pooling Based Attention Vision Transformer for Face Anti-Spoofing
- Addressing Confounds in Functional Connectivity Analyses of Calcium Imaging
- Addressing Data Scarcity in Voice Disorder Detection with Self-Supervised Models
- AdvShadow: Evading DeepFake Detection via Adversarial Shadow Attack
- AdvTTS: Adversarial Text-to-Speech Synthesis Attack on Speaker Identification Systems
- Advancing Acoustic Howling Suppression Through Recursive Training of Neural Networks
- Adversarial Domain Adaptation for Classification with Nested Dichotomies
- Adversarial Jamming for Autoencoder Distribution Matching
- Adversarial Learning on Compressed Posterior Space for Non-Iterative Score-based End-to-End Text-to-Speech
- Adversarial Robustness of Convolutional Models Learned in the Frequency Domain
- Adversarial Speech for Voice Privacy Protection from Personalized Speech Generation
- Aerial-IRS-Assisted Load Balancing In Downlink Networks
- Ainur: Harmonizing Speed and Quality in Deep Music Generation Through Lyrics-Audio Embeddings
- Align, Adapt and Inject: Audio-Guided Image Generation, Editing and Stylization
- All Neural Kronecker Product Beamforming for Speech Extraction with Large-Scale Microphone Arrays
- Alleviating Hallucinations Via Supportive Window Indexing in Abstractive Summarization
- Alpharotate: A Rotation Detection Benchmark Using Tensorflow
- Ambisonics Networks - The Effect of Radial Functions Regularization
- An Accurate and Efficient Neural Network for OCTA Vessel Segmentation and a New Dataset
- An Active Noise Control System Based On Soundfield Interpolation Using A Physics-Informed Neural Network
- An Adapter-Based Unified Model for Multiple Spoken Language Processing Tasks
- An Adaptive Algorithm for Tracking Third-Order Coupled Canonical Polyadic Decomposition
- An Anchor Learning Approach for Citation Field Learning
- An Asymptotically Achievable Rate Bound for Establishing High-Fidelity Entanglements in Quantum Networks
- An Attention-Enhanced Retentive Broad Learning System for Subject-Generic Emotion Recognition from EEG Signals
- An Audio-Textual Diffusion Model for Converting Speech Signals into Ultrasound Tongue Imaging Data
- An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement
- An Efficient Algorithm For Clustered Multi-Task Compressive Sensing
- An Efficient Algorithm for Multiuser Sum-Rate Maximization of Large-Scale Active RIS-Aided MIMO System
- An Efficient Alternating Riemannian/Projected Gradient Descent Ascent Algorithm for Fair Principal Component Analysis
- An Efficient Hierarchical Block Coordinate Descent Method for Time-Varying Graphical Lasso
- An Efficient Temporary Deepfake Location Approach Based Embeddings for Partially Spoofed Audio Detection
- An Efficient Transformer For Demosaicing Via Compressed Multi-Branch Attention Mechanism
- An Efficient and Interpre Table Speech Enhancement Network Via Deep Dictionary Learning
- An Empirical Investigation of Domain Adaptation Ability for Chinese Spelling Check Models
- An Empirical Study on the Impact of Positional Encoding in Transformer-Based Monaural Speech Enhancement
- An End-to-End EEG Channel Selection Method with Residual Gumbel Softmax for Brain-Assisted Speech Enhancement
- An Error Self-Corrected DOA Estimation Model for Sparse Array Based on ANM
- An Experimental Comparison of Multi-View Self-Supervised Methods for Music Tagging
- An Experimental Comparison of Noise-Robust Text-To-Speech Synthesis Systems Based On Self-Supervised Representation
- An Explainable Proxy Model for Multilabel Audio Segmentation
- An Explicit Multi-Modal Fusion Method for Sign Language Translation
- An Initial Investigation of Neural Replay Simulator for Over-The-Air Adversarial Perturbations to Automatic Speaker Verification
- An Interpretable and Generalizable Speech Detector Based on a CNN-LSTM Framework
- An Investigation of Distribution Alignment in Multi-Genre Speaker Recognition
- An MVDR-Embedded U-Net Beamformer for Effective and Robust Multichannel Speech Enhancement
- An Optimized Interleaved OFDM Chirp Orthogonal Waveform Design for Dechirped Miniature MMW MIMO Radar
- An Unsupervised Segmentation of Vocal Breath Sounds
- Analysis of High-Order Brain Networks Resolved in Time and Frequency Using CP Decomposition
- Analysis of an Elliptic Localization Algorithm Using Fixed Point Iteration
- Analysis of the Memorization and Generalization Capabilities of AI Agents: are Continual Learners Robust?
- Analysis of the SINR in LEO-PNT Systems with 5G PRS Multiplexing: Integration of PRS and NTN
- Analyzing Adversarial Vulnerabilities of Graph Lottery Tickets
- Anchor-Guided GAN with Contrastive Loss for Low-Resource Out-of-Domain Detection
- Anim-400K: A Large-Scale Dataset for Automated End to End Dubbing of Video
- Anomalous Sound Detection by Feature-Level Anomaly Simulation
- Anomaly Detection from a Frequency Perspective: M-Band Wavelet Packet Anomaly Detection Network
- Anomaly-Aware Semantic Self-Alignment Framework for Video-Based Person Re-Identification
- Anonymizing Speaker Voices: Easy to Imitate, Difficult to Recognize?
- Anti-Deception Jamming Power Optimization Strategy for Multi-Target Tracking Tasks in Multi-Radar Systems
- Apollo's Unheard Voices: Graph Attention Networks for Speaker Diarization and Clustering for Fearless Steps Apollo Collection
- Application of SNNS Model Based On Multi-Dimensional Attention In Drone Radio Frequency Signal Classification
- Applying Hybrid Quantum LSTM for Indoor Localization Based on RSSI
- Arbitrary Style Transfer Based on Content Integrity and Style Consistency Enhancement
- Arbitrary Style Transfer with Prototype-Based Channel Alignment
- Architecture-Agnostic Iterative Black-Box Certified Defense Against Adversarial Patches
- Are Deep Neural Networks Robust to Named Entities? An Adversarial Attack and Defense Perspective
- Are SNNs Truly Energy-efficient? - A Hardware Perspective
- Are Soft Prompts Good Zero-Shot Learners for Speech Recognition?
- Array Geometry Optimization for Region-of-Interest Near-Field Beamforming
- Asformer: Learning From Adjacent Scale
- Assessing GNSS Carrier-to-Noise-Density Ratio Estimation in The Presence of Meaconer Interference
- Assessing Vibroacoustic Sound Massage Through The Biosignal of Human Speech: Evidence of Improved Wellbeing
- Asymmetric Clean Segments-Guided Self-Supervised Learning for Robust Speaker Verification
- Asymptotic Behavior of Super-Resolution Sparse Bayesian Learning
- Asymptotically Tight Misspecified Bayesian Cramér-Rao Bound
- Asynchronous Diffusion Learning with Agent Subsampling and Local Updates
- AttA-NET: Attention Aggregation Network for Audio-Visual Emotion Recognition
- AttHear: Explaining Audio Transformers Using Attention-Aware NMF
- Attention Decoupling for Query-Based Object Detection
- Attention Is All You Need For Blind Room Volume Estimation
- Attention-Based Spatial-Frequency Information Network for Underwater Single Image Super-Resolution
- Attention-Driven Multichannel Speech Enhancement in Moving Sound Source Scenarios
- Attention-Guided Adaptation for Code-Switching Speech Recognition
- AttentionLUT: Attention Fusion-Based Canonical Polyadic LUT for Real-Time Image Enhancement
- Attr-Int: A Simple and Effective Entity Alignment Framework for Heterogeneous Knowledge Graphs
- Attribute-Aware Amplification of Facial Feature Sequences for Facial Emotion Recognition
- Attribute-Aware Head Swapping Guided by 3d Modeling
- Attribution-Based Scanline Perturbation Attack on 3d Detectors of Lidar Point Clouds
- Audio Deepfake Detection With Self-Supervised Wavlm And Multi-Fusion Attentive Classifier
- Audio Difference Learning for Audio Captioning
- Audio Match Cutting: Finding and Creating Matching Audio Transitions in Movies and Videos
- Audio Prompt Tuning for Universal Sound Separation
- Audio Transformer for Synthetic Speech Detection via Formant Magnitude and Phase Analysis
- Audio-Aided Learning Framework for Image Classification with Limited Training Images
- Audio-Free Prompt Tuning for Language-Audio Models
- Audio-Journey: Open Domain Latent Diffusion Based Text-To-Audio Generation
- Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-Talker Speech
- Audio-Visual Child-Adult Speaker Classification in Dyadic Interactions
- Audio-Visual Speech Recognition In-The-Wild: Multi-Angle Vehicle Cabin Corpus and Attention-Based Method
- Audiosr: Versatile Audio Super-Resolution at Scale
- Audiovisual Speaker Separation with Full- and Sub-Band Modeling in the Time-Frequency Domain
- Auditory Cortex-Inspired Spectral Attention Modulation for Binaural Sound Localization in HRTF Mismatch
- AugSumm: Towards Generalizable Speech Summarization Using Synthetic Labels from Large Language Models
- Augment on Manifold: Mixup Regularization with UMAP
- Augmenting Conformers With Structured State-Space Sequence Models For Online Speech Recognition
- Augmenting Transformer Autoencoders with Phenotype Classification for Robust Detection of Psychotic Relapses
- AutoCali: Enhancing AoA-based Indoor Localization through Automatic Phase Calibration
- AutoFGNN: A Framework for Extracting All Frequency Information from Large-Scale Graphs
- AutoPrep: An Automatic Preprocessing Framework for In-The-Wild Speech Data
- AutoSen: Improving Automatic WiFi Human Sensing through Cross-Modal Autoencoder
- Automated Labeling of Automotive Radar Azimuth Multipath
- Automatic Channel Selection and Spatial Feature Integration for Multi-Channel Speech Recognition Across Various Array Topologies
- Automatic Design of Adapter Architectures for Enhanced Parameter-Efficient Fine-Tuning
- Automatic Detection Of Sleepiness-Related Syndromes and Symptoms Using Voice and Speech Biomarkers
- Automatic Recognition of Gesture Identity and Onset of Cued-Speech
- Automatic Speech Recognition Tuned for Child Speech in the Classroom
- Automatic Temporal Alignment for Pitch Estimation Evaluation
- Automotive Radar Interference Characterization: FMCW or PMCW?
- Automotive Radar Interference Mitigation Via SINR Maximization
- Automotive Radar Point Cloud Parametric Density Estimation using Camera Images
- Autonomous Generative Feature Replay for Non-Exemplar Class-Incremental Learning
- Autoregressive 3D Shape Completion via Sphere-Guided Disentangled Representation
- Autost: Training-Free Neural Architecture Search For Spiking Transformers
- Axis Order Invariance Learned from Point Clouds
- BAE-Net: a Low Complexity and High Fidelity Bandwidth-Adaptive Neural Network for Speech Super-Resolution
- BCC: Bidirectional Consistency Constraint Method for Hierarchical Text Classification
- BEVLOC: End-to-End 6-DoF Localization Via Cross-Modality Correlation Under Bird's Eye View
- BEVoxSeg: BEV-Voxel Representation for Fast and Accurate Camera-Based 3D Segmentation
- BFRFormer: Transformer-Based Generator for Real-World Blind Face Restoration
- BIGVSAN: Enhancing Gan-Based Neural Vocoders with Slicing Adversarial Network
- BNMTrans: A Brain Network Sequence-Driven Manifold-Based Transformer for Cognitive Impairment Detection Using EEG
- BPDO: Boundary Points Dynamic Optimization for Arbitrary Shape Scene Text Detection
- BRAVEn: Improving Self-supervised pre-training for Visual and Auditory Speech Recognition
- BWSNET: Automatic Perceptual Assessment of Audio Signals
- Balanced And Discriminative Contrastive Learning For Class-Imbalanced Medical Images
- Balanced Learning for Multi-Domain Long-Tailed Speaker Recognition
- Balancing Easy and Hard Distortions: A Multi-Rate Knowledge Distillation Strategy for Blind Image Quality Assessment
- Balancing Representation Abstractions and Local Details Preservation for 3d Point Cloud Quality Assessment
- Balancing Speaker-Rater Fairness for Gender-Neutral Speech Emotion Recognition
- Ballistocardiogram-Based Heart Rate Variability Estimation for Stress Monitoring using Consumer Earbuds
- Bandwidth-Efficient Inference for Nerual Image Compression
- Bass Accompaniment Generation Via Latent Diffusion
- Batch Substitution Calibration of a Mems Microphone Array : Impact of Sensor Performance Dispersion on Directivity Estimation
- Bayesian Activity Detection for Massive Connectivity in Cell-Free IoT Networks
- Bayesian Learning-Based Kalman Smoothing For Linear Dynamical Systems With Unknown Sparse Inputs
- Bayesian Optimization with Gaussian Processes for Robust Localization
- Bayesian Topology Inference on Partially Known Networks from Input-Output Pairs
- Bayesian-Boosted MetaLoc: Efficient Training and Guaranteed Generalization for Indoor Localization
- Beamforming Design and Performance Evaluation for RIS-Aided Localization Using LEO Satellite Signals
- Beamforming Through Online Convex Combination of Differential Beamformers
- Beast: Online Joint Beat and Downbeat Tracking Based on Streaming Transformer
- Benchmarking Adversarial Robustness of Image Shadow Removal with Shadow-Adaptive Attacks
- Beta Quantile Regression for Robust Estimation of Uncertainty in the Presence of Outliers
- Beyond Empirical Windowing: An Attention-Based Approach for Trust Prediction In Autonomous Vehicles
- Beyond Simple Text Style Transfer: Unveiling Compound Text Style Transfer with Prompt-Based Pre-Trained Language Models
- Beyond the Limit of Weight-Sharing: Pioneering Space-Evolving NAS with Large Language Models
- Beyond the Snowfall: Enhancing Snowy Day Object Detection Through Progressive Restoration and Multi-Feature Fusion
- Bi-Directional Motion Attention with Contrastive Learning for few-shot Action Recognition
- Binary Signal Alignment: Optimal Solution is Polynomial-Time and Linear-Time Solution is Quasi-Optimal
- Binaural Angular Separation Network
- Binaural Rendering of Heterogeneous Sound Sources with Extent
- Binaural Room Transfer Function Interpolation Via System Inversion
- Binaural Sound Source Localization Using a Hybrid Time and Frequency Domain Model
- Binaural Speech Enhancement Using Deep Complex Convolutional Transformer Networks
- Binauralmusic: A Diverse Dataset for Improving Cross-Modal Binaural Audio Generation
- Biomimetic Mappings for Active Sonar Object Recognition in Clutter
- Blenda: Domain Adaptive Object Detection Through Diffusion-Based Blending
- Blind Beamforming for Intelligent Reflecting Surface: A Reinforcement Learning Approach
- Blind Deconvolution of Sparse Graph Signals in the Presence of Perturbations
- Blind Estimation of Audio Effects Using an Auto-Encoder Approach and Differentiable Digital Signal Processing
- Blind Inpainting with Object-Aware Discrimination for Artificial Marker Removal
- Blind Separation of Noisy Mixtures Over Galois Fields
- Block Adaptive Subspace Pursuit Method for Wall Clutter Mitigation
- Boosting Adversarial Robustness Distillation Via Hybrid Decomposed Knowledge
- Boosting End-to-End Multilingual Phoneme Recognition Through Exploiting Universal Speech Attributes Constraints
- Boosting Image Quality Assessment Performance: Unsupervised Score Fusion by Deep Maximum a Posteriori Estimation
- Boosting LLMS with Ontology-Aware Prompt for Ner Data Augmentation
- Boosting Pruned Networks with Linear Over-Parameterization
- Boosting Speech Enhancement with Clean Self-Supervised Features Via Conditional Variational Autoencoders
- Boosting Unknown-Number Speaker Separation with Transformer Decoder-Based Attractor
- Boosting Zero-Shot Human-Object Interaction Detection with Vision-Language Transfer
- Boosting Zero-Shot Node Classification via Dependency Capture and Discriminative Feature Learning
- Boosting of Implicit Neural Representation-Based Image Denoiser
- Bootstrap Predictive Coding: Investigating a Non-Contrastive Self-Supervised Learning Approach
- Boundary-Driven Active Learning for Anomaly Detection in Time Series Data Streams
- Bounding Box-Guided Pseudo Point Clouds Early-Fusion and Density Optimize for 3D Object Detection
- Brain Structure-Function Interaction Network for Fluid Cognition Prediction
- BrainFC-CGAN: A Conditional Generative Adversarial Network for Brain Functional Connectivity Augmentation and Aging Synthesis
- Branchformer-Based TDNN for Automatic Speaker Verification
- Breaking Speaker Recognition with Paddingback
- Breaking the Barrier: Selective Uncertainty-Based Active Learning for Medical Image Segmentation
- Breast Ultrasound Computer-Aided Diagnosis Using Structure-Aware Triplet Path Networks
- Bregman Graph Neural Network
- Bridging The Domain Gap Arising from Text Description Differences for Stable Text-To-Image Generation
- Bridging the Gap: A Self-Learning Model Using Implicit Knowledge for Chinese Spelling Correction
- Bridging the Gap: Sketch to Color Diffusion Model with Semantic Prompt Learning
- Bridging the Gaps of Both Modality and Language: Synchronous Bilingual CTC for Speech Translation and Speech Recognition
- Bringing the Discussion of Minima Sharpness to the Audio Domain: A Filter-Normalised Evaluation for Acoustic Scene Classification
- Broadband Personal Sound Zone Control in the Presence of Nonlinearities
- Buffered Gaussian Modeling for Vectorized HD Map Construction
- Build a 50+ Hours Chinese Mandarin Corpus for Children's Speech Recognition
- Building Lane-Level Maps from Aerial Images
- ByteHum: Fast and Accurate Query-by-Humming in the Wild
- C-CLAPA: Improving Text-Audio Cross Domain Retrieval with Captioning and Augmentations
- CAG-FPN: Channel Self-Attention Guided Feature Pyramid Network for Object Detection
- CAGEN: Controllable Anomaly Generator using Diffusion Model
- CALSeg: Improving Calibration of Medical Image Segmentation Via Variational Label Smoothing
- CC-DA: Cross-Domain Consistency Data Augmentation for 3D Tumor Segmentation
- CDA-MBPO: Corrected Data Aggregation for Model-Based Policy Optimization
- CDCNet: A Fast and Lightweight Dehazing Network with Color Distortion Correction
- CDUMA: An Adaptive Approach for Mitigating Confounder for MCQA
- CED: Consistent Ensemble Distillation for Audio Tagging
- CEDNet: A Continuous Emotion Detection Network for Naturalistic Stimuli Using MEG Signals
- CEMOAE: A Dynamic Autoencoder with Masked Channel Modeling for Robust EEG-Based Emotion Recognition
- CENet: Content-Aware Enhanced Network for Practical Scene Parsing
- CGN: A Simple Yet Effective Multi-Channel Gated Network for Long-Term Time Series Forecasting
- CIF-RNNT: Streaming ASR Via Acoustic Word Embeddings with Continuous Integrate-and-Fire and RNN-Transducers
- CIF-T: A Novel CIF-Based Transducer Architecture for Automatic Speech Recognition
- CKT-RCM: Clip-Based Knowledge Transfer and Relational Context Mining for Unbiased Panoptic Scene Graph Generation
- CLAF: Contrastive Learning with Augmented Features for Imbalanced Semi-Supervised Learning
- CLAP4Emo: ChatGPT-Assisted Speech Emotion Retrieval with Natural Language Supervision
- CLIP-Font: Sementic Self-Supervised Few-Shot Font Generation with Clip
- CLIP-MSA: Incorporating Inter-Modal Dynamics and Common Knowledge to Multimodal Sentiment Analysis With Clip
- CLPSD: Detecting Ethereum Phishing Scams based on Curriculum Learning
- CLT: Cooperative Lottery Ticket Hypothesis in Live Streaming Sales Prediction
- CM-PIE: Cross-Modal Perception for Interactive-Enhanced Audio-Visual Video Parsing
- CNFA: Conditional Normalizing Flow for Query-Limited Attack
- COLLD: Contrastive Layer-to-Layer Distillation for Compressing Multilingual Pre-Trained Speech Encoders
- COLORFLOW: A Conditional Normalizing Flow for Image Colorization
- COPHTC: Contrastive Learning with Prompt Tuning for Hierarchical Text Classification
- CORAAL QA: A Dataset and Framework for Open Domain Spontaneous Speech Question Answering from Long Audio Files
- CPAUG: Refining Copy-Paste Augmentation for Speech Anti-Spoofing
- CPMSVD: Cross-Project Multiclass Software Vulnerability Detection Via Fused Deep Feature and Domain Adaptation
- CRC-Aided Learned Ensembles of Belief-Propagation Polar Decoders
- CROCFUN: Cross-Modal Conditional Fusion Network for Pansharpening
- CROSSWORD: A Semantic Approach To Text Compression Via Masking
- CReStyler: Text-Guided Single Image Style Transfer Method Based on CNN and Restormer
- CSCNet: Class-Specified Cascaded Network for Compositional Zero-Shot Learning
- CSI-Free Over-The-Air Decentralized Learning Over Frequency Selective Channels
- CSNet: Contrastive Siamese Network for Robust SLU
- CST-Former: Transformer with Channel-Spectro-Temporal Attention for Sound Event Localization and Detection
- CT and MRI Fusion with Anisotropic Guided Filtering
- CUTDEM: Depth-Aware Enhanced Multi-View Image Mixing for Light Field Super-Resolution
- Camera Calibration using a Single View of a Symmetric Object
- Camera-Radar Association for Data Annotation
- Can ChatGPT Serve as a Multi-Criteria Decision Maker? A Novel Approach to Supplier Evaluation
- Can LLM Find the Green Circle? Investigation and Human-Guided Tool Manipulation for Compositional Generalization
- Can Large-Scale Vocoded Spoofed Data Improve Speech Spoofing Countermeasure with a Self-Supervised Front End?
- Can Synthetic Data Boost the Training of Deep Acoustic Vehicle Counting Networks?
- Can We Trust Explainable AI Methods on ASR? An Evaluation on Phoneme Recognition
- Can Whisper Perform Speech-Based In-Context Learning?
- Caption Unification for Multi-View Lifelogging Images Based on In-Context Learning with Heterogeneous Semantic Contents
- Capturing Detail Variations for Lightweight Neural Radiance Fields
- Cardinality-Constrained Binary Quadratic Optimization via Extreme Point Pursuit, with Application to the Densest K-Subgraph Problem
- CartoonDiff: Training-free Cartoon Image Generation with Diffusion Transformer Models
- Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning for Visual Story Synthesis
- CausalME: Balancing bi-modalities in Visual Question Answering
- Causality-Inspired Single-Source Domain Generalization for Face Anti-Spoofing
- Causally Uncovering Bias in Video Micro-Expression Recognition
- Center of Pressure Estimation by Analyzing Walking Videos
- Changenet: Multi-Temporal Asymmetric Change Detection Dataset
- Channel Estimation and Prediction in Wireless Communications Assisted by Semi-Passive RIS
- Channel Estimation in Underdetermined Systems Utilizing Variational Autoencoders
- Channel-Spatial Transformer for Efficient Image Super-Resolution
- Character Attribute Extraction from Movie Scripts Using LLMs
- Chat: Cascade Hole-Aware Transformers with Geometric Spatial Consistency for Accurate Monocular Endoscopic Depth Estimation
- Child FER: Domain-Agnostic Facial Expression Recognition in Children Using a Secondary Image Diffusion Model
- Chunked Attention-Based Encoder-Decoder Model for Streaming Speech Recognition
- Circular Decomposition and Cross-Modal Recombination for Multimodal Sentiment Analysis
- Class-Incremental Learning for Multi-Label Audio Classification
- Class-Wise Buffer Management for Incremental Object Detection: An Effective Buffer Training Strategy
- Class: Continual Learning Approach for Speech Super-Resolution
- Classification-Oriented Semantic Wireless Communications
- Client-Free Federated Unlearning via Training Reconstruction with Anchor Subspace Calibration
- Clinical Scores Prediction and Medication Adjustment for Course of Parkinson's Disease
- Clip-Based Synergistic Knowledge Transfer for text-based Person Retrieval
- Cliprerank: An Extremely Simple Method For Improving Ad-Hoc Video Search
- Close-Range Direction of Arrival Estimation in the Presence of Clock Jitter
- Co-Occurrence Graph-Enhanced Hierarchical Prediction of ICD Codes
- Co-Salient Object Detection via Discriminative Prototypes Contrast
- CoQ: AN Empirical Framework for Multi-hop Question Answering Empowered by Large Language Models
- CoSLR: Contrastive Chinese Sign Language Recognition with prior knowledge And Multi-Tasks Joint Learning
- Coding for the Unsourced B-Channel with Erasures: Enhancing the Linked Loop Code
- Cognitive Virtual Sensing Technique for Feedforward Active Noise Control
- Collaborative Watermarking for Adversarial Speech Synthesis
- Color Agnostic Cross-Spectral Disparity Estimation
- Combining Conformer and Dual-Path-Transformer Networks for Single Channel Noisy Reverberant Speech Separation
- CommIN: Semantic Image Communications as an Inverse Problem with INN-Guided Diffusion Models
- Communication Efficient Private Federated Learning Using Dithering
- Communication-Efficient Decentralized Dynamic Kernel Learning
- Communication-Efficient Federated Learning Through Adaptive Weight Clustering And Server-Side Distillation
- Communication-Efficient Federated Optimization over Semi-Decentralized Networks
- Communication-Efficient Laplace Mechanism for Differential Privacy via Random Quantization
- Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks
- Communication-Oriented Automatic Assessment System for Accented Spoken Chinese in Read-Aloud Tasks
- Compact and De-Biased Negative Instance Embedding for Multi-Instance Learning on Whole-Slide Image Classification
- Comparable Demonstrations Are Important In In-Context Learning: A Novel Perspective On Demonstration Selection
- Comparative Study of Tokenization Algorithms for End-to-End Open Vocabulary Keyword Detection
- Comparing and Combining Audio Processing and Deep Learning Features for Classification of Heartbeat Sounds
- Comparing data-Driven and Handcrafted Features for Dimensional Emotion Recognition
- Comparison Of Frequency-Fusion Mechanisms For Binaural Direction-Of-Arrival Estimation For Multiple Speakers
- Comparison of Conditions for Omnidirectional Video with Spatial Audio in Terms of Subjective Quality and Impacts on Objective Metrics Resolving Power
- Complementary Fusion Network Based on Frequency Hybrid Attention for Pansharpening
- Complex Bounded Component Analysis: Identifiability and Algorithm
- Complexity Reduction of Template Matching-Based Reference Picture Padding in Video Coding
- Complexity Scaling for Speech Denoising
- Composite Federated Learning with Heterogeneous Data
- Computational Complexity of Asynchronous Policy Iteration for Two-Player Zero-Sum Markov Games
- Computing an Entire Solution Path of a Nonconvexly Regularized Convex Sparse Model
- Concealing Medical Condition by Node Toggling in ASR for Dementia Patients
- Concentrated Reasoning and Unified Reconstruction for Multi-Modal Media Manipulation
- Concss: Contrastive-based Context Comprehension for Dialogue-Appropriate Prosody in Conversational Speech Synthesis
- Confidence-Aware Spatial-Temporal Attention Graph Convolutional Network for Skeleton-Based Expert-Novice Level Classification
- Conformalized Multimodal Uncertainty Regression and Reasoning
- Conformer is All You Need for Visual Speech Recognition
- Congestion-Aware Distributed Task Offloading in Wireless Multi-Hop Networks Using Graph Neural Networks
- Conjugate Gradient Based Adaptive Algorithm for Nonlinear AEC
- Connecting Speech Encoder and Large Language Model for ASR
- Considering Temporal Connection between Turns for Conversational Speech Synthesis
- Consistent and Relevant: Rethink the Query Embedding in General Sound Separation
- Contactless Radar Heart Rate Variability Monitoring Via Deep Spatio-Temporal Modeling
- Content-Based Objective Evaluation of Artificially Generated Sign Language Videos
- Context-Aware Dual Attention Network for Multimodal Sarcasm Detection
- Context-Aware Preference Learning System Based on Dirichlet Process Gaussian Mixture Model
- Context-Aware Transformer for Single Image Rain Streaks Removal
- Context-Aware and Contrastiveness-Driven Feature Learning for Cross-Domain Few-Shot Hyperspectral Image Classification
- Context-Guided and Syntactic Augmented Dual Graph Convolutional Network for Aspect-Based Sentiment Analysis
- Contextual Biasing Methods for Improving Rare Word Detection in Automatic Speech Recognition
- Contextual Biasing of Named-Entities with Large Language Models
- Contextual Human Object Interaction Understanding from Pre-Trained Large Language Model
- Contextualized Automatic Speech Recognition With Attention-Based Bias Phrase Boosted Beam Search
- Continual Learning with Class-Level Minimally Interfered Update
- Continuous Review and Timely Correction: Enhancing the Resistance to Noisy Labels via Self-Not-True Distillation
- Contrastive Deep Nonnegative Matrix Factorization For Community Detection
- Contrastive Learning for Regression on Hyperspectral Data
- Contrastive Learning with Audio Discrimination for Customizable Keyword Spotting in Continuous Speech
- Contrastive Learning with Bidirectional Transformers for Knowledge Tracing
- Contrastive Learning with High-Quality and Low-Quality Augmented Data for Query-Focused Summarization
- Contrastive Loss Based Frame-Wise Feature Disentanglement for Polyphonic Sound Event Detection
- Contrastive Speaker Embedding With Sequential Disentanglement
- Contrmix: Progressive Mixed Contrastive Learning for Semi-Supervised Medical Image Segmentation
- ControlCap: Controllable Captioning via No-Fuss Lexicon
- Controllable Prosody Generation with Partial Inputs
- Controllable Semantic Linguistic Steganography via Summarization Generation
- Controllable Speaking Styles Using A Large Language Model
- Convergent Plug-And-Play Using Contractive Denoisers
- Conversation Clique-Based Model for Emotion Recognition In Conversation
- Conversational Co-Speech Gesture Generation via Modeling Dialog Intention, Emotion, and Context with Diffusion Models
- Convnext-TTS And Convnext-VC: Convnext-Based Fast End-To-End Sequence-To-Sequence Text-To-Speech And Voice Conversion
- Cooking-Clip: Context-Aware Language-Image Pretraining for Zero-Shot Recipe Generation
- Cooperative Sensing Via Matrix Factorization of the Partially Received Sample Covariance Matrix
- Coordinate-Based Neural Network for Fourier Phase Retrieval
- Core Body Temperature and its Role in Detecting Acute Stress: A Feasibility Study
- Corn: Co-Trained Full- and No-Reference Speech Quality Assessment
- Corner Detection Based on a Rotation-Invariant and Noise-Insensitive Curvature Measurement
- Corpus Synthesis for Zero-Shot ASR Domain Adaptation Using Large Language Models
- Correcting Faulty Road Maps by Image Inpainting
- Correction Focused Language Model Training For Speech Recognition
- Correlation-Based Machine Learning Techniques for Channel Estimation with Fluid Antennas
- Cost Aware Untargeted Poisoning Attack Against Graph Neural Networks
- Counting Network for Learning from Majority Label
- Coupled Block-Term Tensor Decomposition for Near-Field Localization in multi-static MIMO Radar Systems
- Coupling Self-Supervised and Supervised Contrastive Learning for Multiple Classification of Cervical Cytological Whole Slide Images
- Coverage Analysis For mmWAVE UAV Networks with Static and Dynamic Blockages
- Cramer-Rao Bound for Admittance Matrix Estimation under Laplacian Constraints
- Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction Loss
- Credible Teacher for Semi-Supervised Object Detection in Open Scene
- Cross Branch Feature Fusion Decoder for Consistency Regularization-Based Semi-Supervised Change Detection
- Cross Modal Training for ASR Error Correction with Contrastive Learning
- Cross Pseudo-Labeling for Semi-Supervised Audio-Visual Source Localization
- Cross-Age Contrastive Learning for Age-Invariant Face Recognition
- Cross-Attention watermarking of Large Language Models
- Cross-Camera Human Motion Transfer by Time Series Analysis
- Cross-Domain Cross-Task Transfer Mobile Touch-Stroke Authentication
- Cross-Image Distillation for Semi-Supervised Semantic Segmentation
- Cross-Lingual Learning in Multilingual Scene Text Recognition
- Cross-Modal Alignment for End-to-End Spoken Language Understanding Based on Momentum Contrastive Learning
- Cross-Modal Multi-Tasking for Speech-to-Text Translation via Hard Parameter Sharing
- Cross-Modal Multiscale Difference-Aware Network for Joint Moment Retrieval and Highlight Detection
- Cross-Modal Parallel Training for Improving end-to-end Accented Speech Recognition
- Cross-Modal Synthesis of Structural MRI and Functional Connectivity Networks via Conditional ViT-GANs
- Cross-Modality and Within-Modality Regularization for Audio-Visual Deepfake Detection
- Cross-Speaker Encoding Network for Multi-Talker Speech Recognition
- Cross-Subject EEG Emotion Recognition Based on Interconnected Dynamic Domain Adaptation
- Cross-Target Stance Detection by Exploiting Target Analytical Perspectives
- Cross-Triggering Issue in Audio Event Detection and Mitigation
- Crowd Modeling and Control Via Cooperative Adaptive Filtering
- Crowdsourced Multilingual Speech Intelligibility Testing
- Crowdsourced and Automatic Speech Prominence Estimation
- CryCeleb: A Speaker Verification Dataset Based on Infant Cry Sounds
- Crypto-Mine: Cryptanalysis Via Mutual Information Neural Estimation
- Cubic Knowledge Distillation for Speech Emotion Recognition
- Cuffless Blood Pressure Estimation Using Magnetic Flux In A Ring Form Factor
- Curricular Contrastive Regularization for Speech Enhancement with Self-Supervised Representations
- Customising General Large Language Models for Specialised Emotion Recognition Tasks
- Customized Treatment Per Pixel for Blind Image Super-Resolution
- Cutransnet: Transformers to Make Strong Encoders for Multi-Task Vision Perception of Autonomous Driving
- Cyclic Misspecified Cramer-Rao Bound for Periodic Parameter Estimation
- D3: Dual-Domain Defenses for Byzantine-Resilient Decentralized Resource Allocation
- DACR: Distribution-Augmented Contrastive Reconstruction for Time-Series Anomaly Detection
- DAMP: Distribution-Aware Magnitude Pruning for Budget-Sensitive Graph Convolutional Networks
- DAP: Domain-Aware Prompt Learning for Vision-and-Language Navigation
- DBS: Differentiable Budget-Aware Searching For Channel Pruning
- DCL-Net: Dual Contrastive Learning Network for Semi-Supervised Multi-Organ Segmentation
- DCS: Debiased Contrastive Learning with Weak Supervision for Time Series Classification
- DCTTS: Discrete Diffusion Model with Contrastive Learning for Text-to-Speech Generation
- DDD: A Perceptually Superior Low-Response-Time DNN-Based Declipper
- DDI-CoCo: A Dataset for Understanding the Effect of Color Contrast in Machine-Assisted Skin Disease Detection
- DDN-Net: Deep Residual Shrinkage Denoising Networks with Channel-Wise Adaptively Soft Thresholds for Automated Major Depressive Disorder Identification
- DEEPOREDNET: Contrastive Learning-Based Attention-Weighted Dual Channel Residual Network for Ocular Redness Assessment
- DEGAN: Discrimination Enhanced GAN for Perceptual-Oriented Super-Resolution
- DETS: End-to-End Single-Stage Text-to-Speech Via Hierarchical Diffusion Gan Models
- DF-VTON: Dense Flow Guided Virtual Try-On Network
- DG-RainDiff: Depth-Guided Dynamic Message Passing Diffusion Model for Mixture of Rain Removal
- DGLP: Incorporating Orientation Information for Enhanced Link Prediction in Directed Graphs
- DI-MVS: Learning Efficient Multi-View Stereo With Depth-Aware Iterations
- DIB-X: Formulating Explainability Principles for a Self-Explainable Model Through Information Theoretic Learning
- DIFFSC: Semantic Communication Framework With Enhanced Denoising Through Diffusion Probabilistic Models
- DITW: A High-Performance Deep-Independent Template-Based Watermarking
- DJCM: A Deep Joint Cascade Model for Singing Voice Separation and Vocal Pitch Estimation
- DMEL: The Differentiable Log-Mel Spectrogram as a Trainable Layer in Neural Networks
- DMKD: Improving Feature-Based Knowledge Distillation for Object Detection Via Dual Masking Augmentation
- DMT: Comprehensive Distillation with Multiple Self-Supervised Teachers
- DOA Estimation for Switch-Element Arrays Based on Sparse Representation
- DONE: Dynamic Neural Representation Via Hyperplane Neural ODE
- DPM-TSE: A Diffusion Probabilistic Model for Target Sound Extraction
- DROPFL: Client Dropout Attacks Against Federated Learning Under Communication Constraints
- DRSM: Efficient Neural 4D Decomposition for Dynamic Reconstruction in Stationary Monocular Cameras
- DSIS: A Novel (K, N) Threshold Deniable Secret Image Sharing Scheme with Lossless Recovery
- DT-NeRF: Decomposed Triplane-Hash Neural Radiance Fields For High-Fidelity Talking Portrait Synthesis
- DURRNET: Deep Unfolded Single Image Reflection Removal Network with Joint Prior
- Darkshot: Lighting Dark Images with Low-Compute and High-Quality
- Data Augmentation via Subgroup Mixup for Improving Fairness
- Data Driven Grapheme-to-Phoneme Representations for a Lexicon-Free Text-to-Speech
- Data-Aided Channel Estimation Utilizing Gaussian Mixture Models
- Data-Driven Convex Regularizers for Inverse Problems
- Data-Driven Lattices for Vector Quantization
- Data-Free Watermark for Deep Neural Networks by Truncated Adversarial Distillation
- Data-Scarce Condition Modeling Requires Model-Based Prior Regularization
- Dataset Distillation with Channel Efficient Process
- De Novo Molecule Generation with Graph Latent Diffusion Model
- Debiasing Recommenders Through Personalized Popularity-Aware Margins
- Debris Sensing Based on Leo Constellation: An Intersatellite Channel Parameter Estimation Approach
- Decentralized Generalized Approximate Message-Passing for Tree-Structured Networks
- Decentralized Low Rank Matrix Recovery from Column-Wise Projections by Alternating GD and Minimization
- Decentralizing Coherent Joint Transmission Precoding Via Deterministic Equivalents
- Decoupled Self-Adaptive Distribution Regularization for Few-Shot Image Classification
- Decoupled Spatial and Temporal Processing for Resource Efficient Multichannel Speech Enhancement
- Decoupling and Refilling: A Simple Data Augmentation Method for Aspect Term Extraction
- Deep Convolution Network Based Super Resolution DOA Estimation with Toeplitz and Sparse Prior
- Deep Fusion of Shifted MLP and CNN for Medical Image Segmentation
- Deep INCM Reconstruction for Adaptive Beamforming
- Deep Learning AMR Model Inference Acceleration with CFU for Edge Systems
- Deep Learning Based Single-Shot Profilometry by Three-Channel Binary-Defocused Projection
- Deep Learning Inversion of Ocean Wave Spectrum from SAR Satellite Observations
- Deep Manifold Transformation for Protein Representation Learning
- Deep Neighbor Layer Aggregation for Lightweight Self-Supervised Monocular Depth Estimation
- Deep Neural Network Models Trained with a Fixed Random Classifier Transfer Better Across Domains
- Deep Optimization of Relay Networks-Using Relays as Neurons
- Deep Plug-and-Play Algorithm for Unsaturated Imaging
- Deep Regression for Biological Age Estimation in Multiple Organs: Investigations on 40, 000 Subjects of the UK Biobank
- Deep Reinforcement Learning for Energy Minimization in Multi-RIS-Aided Cell-Free MEC Networks
- Deep Residual W-Unit Learning with Semantic Embedding for Automatic Pulmonary CT Artery-Vein Separation
- Deep Unfolded Annealed Stein Particle Filter for Vehicle Tracking
- Deep Unrolling Network for SAR Image Despeckling
- Deep Variational Privacy Funnel: General Modeling with Applications in Face Recognition
- Deep Versatile Hyperspectral Reconstruction Model from A Snapshot Measurement with Arbitrary Masks
- DeepGRE: Global Robustness Evaluation of Deep Neural Networks
- Defending against Clean-Image Backdoor Attack in Multi-Label Classification
- DefocusSR: An Efficient Framework for Defocus Image Super-Resolution Guided by Depth Information
- DeformMLP: Dynamic Large-Scale Receptive Field MLP Networks for Human Motion Prediction
- Deformation And Penetration Hybrid Detection-Net For Parcels Inspection In Industrial Supply Chain
- Delay Embedding for Matrix Graphical Model Learning from Dependent Data
- Delineation of Prostate Cancer Via Enhanced AI-Based Algorithm In Ultrasound Images
- Delving Deeper Into Vulnerable Samples in Adversarial Training
- Dementia Assessment Using Mandarin Speech with an Attention-Based Speech Recognition Encoder
- Denoising Diffusion Probabilistic Models for Action-Conditioned 3D Motion Generation
- Depth-Guided Dominant Plane Perception for Unsupervised Homography Estimation
- Design of Spatial-Slow-Time Constant-Modulus Waveform Transmission and Receive Adaptive Filter for Dual-Function Radar Communications with Reconfigurable Intelligent Surface
- Detecting Check-Worthy Claims in Political Debates, Speeches, and Interviews Using Audio Data
- Detecting Continuous Gravitational Waves Using Generated Training Data
- Detection and Attribution of Models Trained on Generated Data
- Detection in Complex Scenes Using Rgb and Depth Multimodal Feature Fusion
- Detection of Epileptic Seizures in Long Eeg Recordings Using an Anomaly Detector with Artifact Rejection
- Detector Design for Distributed Multichannel Radar Sensors in Colored Interference Environments
- Determined BSS by Combination of IVA and DNN via Proximal Average
- Diacorrect: Error Correction Back-End for Speaker Diarization
- Diagnosis of Autism Spectrum Disorder Based on Contrastive Functional Connectivity Graph Learning Network
- Diagonalize Integral Graph by DCT
- DialCLIP: Empowering Clip As Multi-Modal Dialog Retriever
- Dialog Modeling in Audiobook Synthesis
- Diarist: Streaming Speech Translation with Speaker Diarization
- Dicetrack: Lightweight Dice Classification on Resource-Constrained Platforms with Optimized Deep Learning Models
- Diff-HOD: Diffusion Model for Object Detection in Hazy Weather Conditions
- Diff-SV: A Unified Hierarchical Framework for Noise-Robust Speaker Verification Using Score-Based Diffusion Probabilistic Models
- DiffDub: Person-Generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-Encoder
- DiffRENT: A Diffusion Model for Recording Environment Transfer of Speech
- Differentiable Quantum Architecture Search For Job Shop Scheduling Problem
- Differentiable Resolution Compression and Alignment for Efficient Video Classification and Retrieval
- Differential Beamforming with Null Constraints for Spherical Microphone Arrays
- Differentially Private Federated Frank-Wolfe
- Diffevent: Event Residual Diffusion for Image Deblurring
- Diffradar: High-Quality Mmwave Radar Perception With Diffusion Probabilistic Model
- Diffstock: Probabilistic Relational Stock Market Predictions Using Diffusion Models
- Diffusion Models for Audio Semantic Communication
- Diffusion Optimistic Learning for Min-Max Optimization
- Diffusion-Based Adversarial Purification for Robust Deep Mri Reconstruction
- Diffusion-Based Pose Refinement and Multi-Hypothesis Generation for 3D Human Pose Estimation
- Diffusion-Based Speech Enhancement in Matched and Mismatched Conditions Using a Heun-Based Sampler
- Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
- Diffusion-Based Speech Enhancement with a Weighted Generative-Supervised Learning Loss
- Diffusioninst: Diffusion Model for Instance Segmentation
- Digital Pathology Image Deblurring Via Local Focus Quality Assessment
- Digital Task-Oriented Communication with Hardware-Limited Task-Based Quantization
- Direct Position Determination by Covariance-Fitting on the Riemannian Manifold of Hermitian Positive Definite Matrices
- Directed Scattering for Knowledge Graph-Based Cellular Signaling Analysis
- Directional Gain Based Noise Covariance Matrix Estimation for MVDR Beamforming
- Discovering Malicious Signatures in Software from Structural Interactions
- Discrete Audio Representation as an Alternative to Mel-Spectrograms for Speaker and Speech Recognition
- Discriminative Semi-Supervised Feature Selection Via a Class-Credible Pseudo-Label Learning Framework
- Discriminative Training of VBx Diarization
- Disentangle Estimation of Causal Effects from Cross-Silo Data
- Disentangled Graph Representation with Contrastive Learning for Rumor Detection
- Disentanglement Network: Disentangle the Emotional Features from Acoustic Features for Speech Emotion Recognition
- Disentangling the Spectral Properties of the Hodge Laplacian: not all small Eigenvalues are Equal
- Distill Vision Transformers to CNNs via Teacher Collaboration
- Distilling Distributional Uncertainty from a Gaussian Process
- Distilling Hubert with LSTMs via Decoupled Knowledge Distillation
- Distributed Decision-Making for Community Structured Networks
- Distributed Stochastic Contextual Bandits for Protein Drug Interaction
- Distributed Vector Approximate Message Passing
- Distribution-Aware Contrastive Learning for Robust Medical Image Segmentation
- Diversifying Cross-Domain Few-Shot Learning via Multimodal Image Editing
- Diversity-Aware Buffer for Coping with Temporally Correlated Data Streams in Online Test-Time Adaptation
- Diversity-Based Core-Set Selection for Text-to-Speech with Linguistic and Acoustic Features
- Do Learned Speech Symbols Follow Zipf's Law?
- Do Self-Supervised Speech and Language Models Extract Similar Representations as Human Brain?
- Does Audio Deepfake Detection Rely on Artifacts?
- Does Video Summarization Require Videos? Quantifying the Effectiveness of Language in Video Summarization
- Domain Adaptive Graph Classification
- Domain Generalization with fourier Transform and soft thresholding
- Domain-Adaptive Semantic Segmentation Emerges From Vision-Language Supervised Domain-Debiased Self-Training
- Domain-Adaptive and Subgroup-Specific Cascaded Temperature Regression for Out-of-Distribution Calibration
- Domain-Slot Aware Contrastive Learning for Improved Dialogue State Tracking
- Domain-Wise Invariant Learning for Panoptic Scene Graph Generation
- Domaindiff: Boost out-of-Distribution Generalization with Synthetic Data
- Double Reverse Regularization Network Based on Self-Knowledge Distillation for SAR Object Classification
- Driver Scanpath Prediction Based On Inverse Reinforcement Learning
- Drop Sparse Convolution for 3D Object Detection
- Dropout Multi-Head Attention for Single Image Super-Resolution
- DuNet: A Robust End-to-End Deep Neural Network Framework for Imbalanced Classification
- Dual Contrastive Learning Guided Pathological Image Re-Staining
- Dual Directional Complementary Gradient Fusion and Deep Refinement for Hyperspectral Image Super Resolution
- Dual Level Intent-Slot Interaction for Improved Multi-Intent Spoken Language Understanding
- Dual Parameter-Efficient Fine-Tuning for Speaker Representation Via Speaker Prompt Tuning and Adapters
- Dual Rank-1 Tensor Attention Module for Convolutional Neural Networks
- Dual-Channel Unlimited Sampling for Bandpass Signals
- Dual-Color Granularity Alignment for Text-Based Person Search
- Dual-Mix for Cross-Modal Retrieval with Noisy Labels
- Dual-Path Minimum-Phase and All-Pass Decomposition Network for Single Channel Speech Dereverberation
- Dual-Stream Contrastive Predictive Network with Joint Handcrafted Feature View for SAR Ship Classification
- DualGCN-MIL: Whole Slide Image Classification Based on Double Relationship Graph Learning
- Dualvc 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice Conversion
- DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis
- Dust: Dual-Grained Syntax-Aware Transformer Network for Chinese Named Entity Recognition
- Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of a Multilingual ASR Model
- Dynamic Bandwidth Variational Mode Decomposition
- Dynamic Clustering and Cluster Contrastive Learning for Unsupervised Person Re-Id With Feature Distribution Alignment
- Dynamic Data Sampler for Cross-Language Transfer Learning in Large Language Models
- Dynamic Frequency Domain Graph Convolutional Network for Traffic Forecasting
- Dynamic Label Smoothing Strategy for Biosignal Classification
- Dynamic Model Structure Adjustment to Realize Quantum Continual Learning Based on Quantum Data
- Dynamic Multi-Scale Context Aggregation for Conversational Aspect-Based Sentiment Quadruple Analysis
- Dynamic Mutual-Activated Transformer for Human Motion Prediction
- Dynamic Privacy Allocation for Locally Differentially Private Federated Learning with Composite Objectives
- Dynamic Random Feature Gaussian Processes for Bayesian Optimization of Time-Varying Functions
- Dynamic Replay Training for Class-Incremental Learning
- Dynamic Speech Emotion Recognition Using A Conditional Neural Process
- Dynamic Video Frame Interpolation with Integrated Difficulty Pre-Assessment
- Dynamic-Superb: Towards a Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark For Speech
- EC-NAS: Energy Consumption Aware Tabular Benchmarks for Neural Architecture Search
- ECIL-MU: Embedding Based Class Incremental Learning and Machine Unlearning
- ECM-OPCC: Efficient Context Model for Octree-Based Point Cloud Compression
- ECPNet: An Enhanced Curve Perception Network for Lane Detection
- ED-TTS: Multi-Scale Emotion Modeling Using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
- EDM: Synthetic Data from Exemplar Diffusion Model Improves Non-Communicable Diseases Detection
- EEG Emotion Recognition Based on Dynamical Graph Attention Network
- EEG-Based Fast Auditory Attention Detection in Real-Life Scenarios Using Time-Frequency Attention Mechanism
- EK-Net: Real-Time Scene Text Detection with Expand Kernel Distance
- EMALG: An Enhanced Mandarin Lombard Grid Corpus with Meaningful Sentences
- EMOCONV-Diff: Diffusion-Based Speech Emotion Conversion for Non-Parallel and in-the-Wild Data
- EOFD-Net: Edge Optimization and Feature Denoising for Weakly Supervised Deep Nuclei Segmentation with Point Annotations
- EPA: Neural Collapse Inspired Robust Out-of-distribution Detector
- ESA: Expert-and-Samples-Aware Incremental Learning Under Longtail Distribution
- ESTGN: Enhanced Self-Mined Text Guided Super-Resolution Network for Superior Image Super Resolution
- ESVC: Combining Adaptive Style Fusion and Multi-Level Feature Disentanglement for Expressive Singing Voice Conversion
- ETP: Learning Transferable ECG Representations via ECG-Text Pre-Training
- Early Diagnosing Parkinson's Disease Via a Deep Learning Model Based on Augmented Facial Expression Data
- Echocardiography Video Synthesis from End Diastolic Semantic Map Via Diffusion Model
- Edge Attention Learning for Efficient Camouflaged Object Detection
- Edge Deployable Distributed Evolutionary Optimization based Calibration method for Neural Quantization
- Effect of Beampattern on Matrix Completion with Sparse Arrays
- Effect of Target Signals and Delays on Spatially Selective Active Noise Control for Open-Fitting Hearables
- Effective Connectivity-Based Multi-View Feature Learning Method for Dementia Diagnosis with FNIRS Signal
- Effective Image Tampering Localization Via Enhanced Transformer and Co-Attention Fusion
- Effective Internal Language Model Training and Fusion for Factorized Transducer Model
- Efficient 3D Position Estimation in Badminton Scene
- Efficient Adapter Finetuning for Tail Languages in Streaming Multilingual ASR
- Efficient Adapter Tuning of Pre-Trained Speech Models for Automatic Speaker Verification
- Efficient Architecture Search for Real-Time Instance Segmentation
- Efficient Black-Box Speaker Verification Model Adaptation With Reprogramming And Backend Learning
- Efficient Content Reconstruction for High Dynamic Range Imaging
- Efficient Federated Learning with Smooth Aggregation for Non-IID Data from Multiple Edges
- Efficient Functional Link Adaptive Filters Based On Nearest Kronecker Product Decomposition
- Efficient Fusion of Depth Information for Defocus Deblurring
- Efficient Hierarchical Stripe Attention for Lightweight Image Super-Resolution
- Efficient High-Performance Bark-Scale Neural Network for Residual Echo and Noise Suppression
- Efficient Joint Rectification of Photometric and Geometric Distortions in Document Images
- Efficient Learned Image Compression with Selective Kernel Residual Module and Channel-Wise Causal Context Model
- Efficient Learning on Successive Test Time Augmentation
- Efficient Multi-Channel Speech Enhancement with Spherical Harmonics Injection for Directional Encoding
- Efficient Personal Voice Activity Detection with Wake Word Reference Speech
- Efficient Point Cloud Attribute Compression Framework using Attribute-Guided Graph Fourier Transform
- Efficient Point Cloud Attribute Compression Using Rich Parallelizable Context Model
- Efficient Polyp Segmentation via Integrity Learning
- Efficient Posenet with Coarse to Fine Transformer
- Efficient Quantum Recurrent Reinforcement Learning Via Quantum Reservoir Computing
- Efficient Scene Text Image Super-Resolution with Semantic Guidance
- Efficient Video and Audio Processing with Loihi 2
- EiffHDR: An Efficient Network for Multi-Exposure High Dynamic Range Imaging
- Eigendecomposition-Based Spatial-Temporal Attention for Brain Cognitive States Identification
- Electroencephalogram Helps Few-Shot Learning
- Electroencephalogram Sensor Data Compression Using an Asymmetrical Sparse Autoencoder with a Discrete Cosine Transform Layer
- Electrolaryngeal Speech Intelligibility Enhancement through Robust Linguistic Encoders
- Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision
- Elevating Visual Prompting in Transfer Learning Via Pruned Model Ensembles: No Retrain, No Pain
- Ellipse Detection Based On Structure-Preserving Anisotropic Edge Extraction
- Ellipse Detection Based on Contrast-Guided Arc Enhancement
- Embedded Feature Similarity Optimization with Specific Parameter Initialization for 2D/3D Medical Image Registration
- Embedded Graph Representation for Inter-Frame Coding of Dynamic Meshes
- EmoRED: A Dataset for Relation Extraction in Texts with Emoticons
- EmoTVR: A Hybrid Model to Estimate Continuous-Time and Continuous-Level Emotion from Electroencephalography
- EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model
- Emohrnet: High-Resolution Neural Network Based Speech Emotion Recognition
- Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
- Emotion-Aligned Contrastive Learning Between Images and Music
- Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
- Emphasized Non-Target Speaker Knowledge in Knowledge Distillation for Automatic Speaker Verification
- Employing Real Training Data for Deep Noise Suppression
- Empowering Vision-Language Models for Reasoning Ability through Large Language Models
- EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
- Enabling Device Control Planning Capabilities of Small Language Model
- Enabling Orientation-Free Mmwave-Based Vital Sign Sensing with Multi-Domain Signal Analysis
- Enabling Secure Wireless Communications via Movable Antennas
- Encoder-Minimal and Decoder-Minimal Framework for Remote Sensing Image Dehazing
- Encoding Seasonal Climate Predictions with Modular Neural Network
- Encoding Time and Energy Model for SVT-AV1 Based on Video Complexity
- End-To-End Personalized Cuff-Less Blood Pressure Monitoring Using ECG and PPG Signals
- End-To-End Real Time Tracking of Children's Reading with Pointer Network
- End-To-End Spatially-Constrained Multi-Perspective Fine-Grained Image Captioning
- End-to-End Learning of Gaussian Mixture Proposals Using Differentiable Particle Filters and Neural Networks
- End-to-End Speech Recognition Contextualization with Large Language Models
- End-to-End Speech Translation with Mutual Knowledge Distillation
- Energy Efficient Wake-Up Solution for Large-Scale Internet of Underwater Things Networks
- Energy-Aware Resolution Selection for Per-Title Encoding
- Energy-Based Models for Speech Synthesis
- Energy-Efficient Decentralized Learning Via Graph Sparsification
- Energy-Saving Cell-Free Massive MIMO Precoders with a per-AP Wideband Kronecker Channel Model
- Engineering the Neural Collapse Geometry of Supervised-Contrastive Loss
- Enhanced Axle-Based Vehicle Classification Using Angle-Based Micro-Doppler Signature
- Enhanced Channel Estimation in mm-Wave Mimo Systems Leveraging Integrated Communication and Sensing
- Enhanced Color Palette Modeling For Lossless Screen Content Compression
- Enhanced Deep Reinforcement Learning for Parcel Singulation in Non-Stationary Environments
- Enhanced KPI Anomaly Detection: An Unsupervised Hybrid Model with Dynamic Threshold
- Enhanced Low-Rank and Sparse Tucker Decomposition For Image Completion
- Enhanced Screen Shooting Resilient Document Watermarking
- Enhanced Transfer Learning with Efficient Modeling and Adaptive Fusion of Knowledge Via Prompt Tuning
- Enhanced Unsupervised Domain Adaptation with Dual-Attention Between Classification and Domain Alignment
- Enhancing Adversarial Robustness of DNNS Via Weight Decorrelation in Training
- Enhancing Adversarial Training with Prior Knowledge Distillation for Robust Image Compression
- Enhancing Adversarial Transferability in Object Detection with Bidirectional Feature Distortion
- Enhancing AoA Estimation Via Phase Modeling of Bluetooth 5 CTE Signals
- Enhancing Argumentative Relation Classification by Multi-Granularity Retrieval and Heterogeneous Graph Reasoning
- Enhancing Audio Generation Diversity with Visual Information
- Enhancing Audio-Visual Question Answering with Missing Modality via Trans-Modal Associative Learning
- Enhancing Code-Switching Speech Recognition With Interactive Language Biases
- Enhancing Conversation Smoothness in Language Learning Chatbots: An Evaluation of GPT4 for ASR Error Correction
- Enhancing Cross-Domain Detection: Adaptive Class-Aware Contrastive Transformer
- Enhancing Document-Level Event Extraction via Structure-Aware Heterogeneous Graph with Multi-Granularity Subsentences
- Enhancing End-to-End Conversational Speech Translation Through Target Language Context Utilization
- Enhancing Event Sequence Modeling with Contrastive Relational Inference
- Enhancing Expressiveness in Dance Generation Via Integrating Frequency and Music Style Information
- Enhancing GAN Performance Through Neural Architecture Search and Tensor Decomposition
- Enhancing Gender Privacy with Photo-Realistic Fusion of Disentangled Spatial Segments
- Enhancing Generalization Of Invisible Facial Privacy Cloak Via Gradient Accumulation
- Enhancing Generalization in Medical Visual Question Answering Tasks Via Gradient-Guided Model Perturbation
- Enhancing Generative Aspect-Based Sentiment Analysis with Relation-Level Supervision and Prompt
- Enhancing Healthcare with EOG: A Novel Approach to Sleep Stage Classification
- Enhancing Hyperspectral Anomaly Detection by Difference-of-Convex Sparse Anomaly Modeling
- Enhancing Image-Text Matching with Adaptive Feature Aggregation
- Enhancing Low-Latency Speaker Diarization with Spatial Dictionary Learning
- Enhancing Multi-Task Models For Recommendation with Tensor Trace Norm
- Enhancing Multilingual Speech Recognition through Language Prompt Tuning and Frame-Level Language Adapter
- Enhancing Multilingual TTS with Voice Conversion Based Data Augmentation and Posterior Embedding
- Enhancing Noisy Label Learning Via Unsupervised Contrastive Loss with Label Correction Based on Prior Knowledge
- Enhancing Note-Level Singing Transcription Model with Unlabeled and Weakly Labeled Data
- Enhancing Performance of Coarsened Graphs with Gradient-Matching
- Enhancing Pre-Trained ASR System Fine-Tuning for Dysarthric Speech Recognition Using Adversarial Data Augmentation
- Enhancing Quantised End-to-End ASR Models Via Personalisation
- Enhancing Realism in 3D Facial Animation Using Conformer-Based Generation and Automated Post-Processing
- Enhancing Reinforcement Learning via Causally Correct Input Identification and Targeted Intervention
- Enhancing Semantic Communication with Deep Generative Models: An Overview
- Enhancing Short-and Long-Term Sea Surface Temperature Forecasting with a Static and Dynamic Learnable Personalized Graph Convolution Network
- Enhancing Spatial Audio Generation with Source Separation and Channel Panning Loss
- Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach
- Enhancing Steganography of Generative Image Based on Image Retouching
- Enhancing Targeted Transferability VIA Feature Space Fine-Tuning
- Enhancing Two-Stage Finetuning for Speech Emotion Recognition Using Adapters
- Enhancing Violin Fingering Generation through Audio-Symbolic Fusion
- Enhancing the Domain Robustness of Self-Supervised pre-Training with Synthetic Images
- Enriching Music Descriptions with A Finetuned-LLM and Metadata for Text-to-Music Retrieval
- Entwined Inversion: Tune-Free Inversion For Real Image Faithful Reconstruction and Editing
- Environmental Sound Synthesis from Vocal Imitations and Sound Event Labels
- Esihgnn: Event-State Interactions Infused Heterogeneous Graph Neural Network for Conversational Emotion Recognition
- Estimating Directed Spectral Information Flow between Multi-Resolution Time Series
- Estimating Exercise-Induced Fatigue from Thermal Facial Images
- Estimating Symptoms and Clinical Signs Instead of Disorders: The Path Toward The Clinical Use of Voice and Speech Biomarkers In Psychiatry
- Estimation of Impulse Responses for a Moving Source Using Optimal Transport Regularization
- Estimation of Spectral Lines Using Expectation Propagation
- Evaluation of an Improved Ultrasonic Imaging Helmet for Observing Articulatory Data
- Evidence-Aware Multimodal Chinese Social Media Rumor Detection
- Evolution Backcasting of Edge Flows From Partial Observations Using Simplicial Vector Autoregressive Models
- Exact Classification of NMR Spectra from NMR Signals
- Exploiting A Quantum Multiple Kernel Learning Approach For Low-Resource Spoken Command Recognition
- Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
- Exploiting Modality-Specific Features for Multi-Modal Manipulation Detection and Grounding
- Exploiting Spatial-Temporal Data for Sleep Stage Classification via Hypergraph Learning
- Exploration of Visual Prompt in Grounded Pre-Trained Open-Set Detection
- Exploring Adapters with Conformers for Children's Automatic Speech Recognition
- Exploring Consistent Spatio-Temporal Distortion and Stable 3-D DCT Coefficients for Robust Blind Video Watermarking
- Exploring Label Hierarchy in Dialogue Intent Classification
- Exploring Large Scale Pre-Trained Models for Robust Machine Anomalous Sound Detection
- Exploring Latent Cross-Channel Embedding for Accurate 3d Human Pose Reconstruction in a Diffusion Framework
- Exploring Meta Information for Audio-Based Zero-Shot Bird Classification
- Exploring Multi-Modal Control in Music-Driven Dance Generation
- Exploring Object-Centered External Knowledge for Fine-Grained Video Paragraph Captioning
- Exploring Phonetic Context-Aware Lip-Sync for Talking Face Generation
- Exploring Self-Explainable Street-Level IP Geolocation with Graph Information Bottleneck
- Exploring Self-supervised Contrastive Learning of Spatial Sound Event Representation
- Exploring Soft Prompt Initialization Strategy for Few-Shot Continual Text Classification
- Exploring Spatio-Temporal Discriminative Cues for Group Activity Recognition Via Contrastive Learning
- Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study
- Exploring Targeted Universal Adversarial Attack for Deep Hashing
- Exploring the Utility of Clip Priors for Visual Relationship Prediction
- Expression Domain Translation Network for Cross-Domain Head Reenactment
- Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific Input Representation and Diffusion Outpainting
- Extending Implicit Neural Representations for Text-to-Image Generation
- Extending Large Language Models for Speech and Audio Captioning
- Extending Multilingual ASR to New Languages Using Supplementary Encoder and Decoder Components
- Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data
- Extending Whisper with Prompt Tuning to Target-Speaker ASR
- Extension of Clifford Data Regression Methods for Quantum Error Mitigation
- External Division of Two Proximity Operators: An Application to Signal Recovery with Structured Sparsity
- Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models
- Extremely Light-Weight Learning Based LDR to PQ HDR Conversion Using Bernstein Curves
- Extrinsic Versus App Information Feedback in Turbo Vep Mu-Mimo Receivers: Optimization Via Deep Unfolding
- Eye Motion Matters for 3D Face Reconstruction
- F1-EV score: Measuring The Likelihood of Estimating a Good Decision Threshold for Semi-Supervised Anomaly Detection
- F2GNN: An Adaptive Filter with Feature Segmentation for Graph-Based Fraud Detection
- FAMIM: A Novel Frequency-Domain Augmentation Masked Image Model Framework for Domain Generalizable Face Anti-Spoofing
- FAVANO: Federated Averaging with Asynchronous Nodes
- FCC-MF: Detecting Violence in Audio-Visual Context with Frame-Wise Cluster Contrast and Modality-Stage Flooding
- FDA-MIMO Radar Using Ambiguity Function for Target Two-Dimensional Localization
- FDC-NeRF: Learning Pose-Free Neural Radiance Fields with Flow-Depth Consistency
- FDIG: A Fine-Grained Data Integration Approach for Group Recommendation
- FDNet: A Novel Multivariate Time Series Classification Model Through Fusing Feature and Difference
- FED-SDS: Adaptive Structured Dynamic Sparsity for Federated Learning Under Heterogeneous Clients
- FEDKA: Federated Knowledge Augmentation for Multi-Center Medical Image Segmentation on non-IID Data
- FFT-Based Selection and Optimization of Statistics for Robust Recognition of Severely Corrupted Images
- FIBA: Federated Invisible Backdoor Attack
- FIRNet: Fundamental Frequency Controllable Fast Neural Vocoder With Trainable Finite Impulse Response Filter
- FPGNet: Single Image Deraining with High-Frequency Channel and Frequency Domain Prior Guidance
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.