ICASSP 2024 Accepted Papers
The full list of 2,679 papers accepted at ICASSP 2024 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- RL-LOGO: Deep Reinforcement Learning Localization for Logo Recognition
- RSED: Zero-Shot Relation Triplet Extraction via Relation Selection and Entity Boundary Detection
- RTLBP-AN Efficient Local Pattern For Facial Images Retrieval
- RVAE-EM: Generative Speech Dereverberation Based On Recurrent Variational Auto-Encoder And Convolutive Transfer Function
- RVDNet: A Two-Stage Network for Real-World Video Desnowing with Domain Adaptation
- Radar Perception with Scalable Connective Temporal Relations for Autonomous Driving
- Radar Recognition in the Wild: Enhancing Radar Emitter Recognition through Auto-Correlation Model-Agnostic Meta Learning
- Radardiff: Improving Sea Clutter Suppression Using Diffusion Models for Radar Images
- Rademacher Complexity Regularization for Correlation-Based Multiview Representation Learning
- Radio Slam with Hybrid Sensing for Mixed Reflection Type Environments
- Randomized Maximum Likelihood Via High-Dimensional Bayesian Optimization
- Ranking Enhanced Fine-Grained Contrastive Learning for Recommendation
- Ranking of Visual Trackers Using Robust Error Norms
- Rapid Change Localization in Dynamic Graphical Models
- Rapid Hybrid Modular Receive Beamforming Via Learned Optimization
- Rate-Quality Based Rate Control Model for Neural Video Compression
- Rating-Augmented No-Reference Point Cloud Quality Assessment Using Multi-Task Learning
- Read, Spell and Repeat: Scene Text Recognition with Vision-Language Circular Refinement
- Real-Oriented Object Detection Driven by Intelligent Stockbreeding
- Real-Time Low-Latency Music Source Separation Using Hybrid Spectrogram-Tasnet
- Real-Time Multi-Human Parsing on Embedded Devices
- Real-Time Privacy-Preserving Fall Risk Assessment with a Single Body-Worn Tracking Camera
- Real-Time Stereo Speech Enhancement with Spatial-Cue Preservation Based on Dual-Path Structure
- Recap: Retrieval-Augmented Audio Captioning
- Recent Advances in Scalable Energy-Efficient and Trustworthy Spiking Neural Networks: from Algorithms to Technology
- Recognition-Guided Diffusion Model for Scene Text Image Super-Resolution
- Reconstruction of Sound Field Through Diffusion Models
- Recovering Missing Node Features with Local Structure-Based Embeddings
- Recovering from Privacy-Preserving Masking with Large Language Models
- Recursive-Tail-Fista for Sparse Signal Recovery
- Redefining Night Vision: The Power of MSR-Driven Neural ISP
- Reduced-Dimensional Decomposition and Eigenspace Reconstruction of Coherent Sources with Arbitrary Rectangle Arrays
- Reducing the Complexity of Normalizing Flow Architectures for Point Cloud Attribute Compression
- Reference Line Network: On Simultaneous Gaussian Line Detection and Connection Graph Inference
- Refinement Bird's Eye View Feature for 3D Lane Detection with Dual-Branch View Transformation Module
- Refining 3D Human Mesh via Model-Free Offsets Estimation
- Refining Text Input For Augmentative and Alternative Communication (AAC) Devices: Analysing Language Model Layers For Optimisation
- Reflection Removal Using Recurrent Polarization-to-Polarization Network
- Reflow-TTS: A Rectified Flow Model for High-Fidelity Text-to-Speech
- Region-Adaptive Video Sharpening Via Rate-Perception Optimization
- Regularized Conditional Alignment for Multi-Domain Text Classification
- Reinforcement Learning Compensated Filter for Multi-Agents Cooperative Localization
- Reinforcement Learning-Guided Optogenetic Stimulation Policies for Robust Functional Network Discovery
- Relational Graph-Bridged Image-Text Interaction: A Novel Method for Multi-Modal Relation Extraction
- Remixed2remixed: Domain Adaptation for Speech Enhancement by Noise2noise Learning with Remixing
- Renyi Divergences Learning for explainable classification of SAR Image Pairs
- Reparameterization Head for Efficient Multi-Input Networks
- Representation Learning across Feature and Topology Views with Output Correction for Graph Convolutional Networks
- Representation and Boundary Enhancement for Action Segmentation Using Transformer
- Repurposing Mu-Mimo Downlink For Joint Wireless Communications And Imaging Via Virtual Users
- Residual Dense Swin Transformer for Continuous Depth-Independent Ultrasound Imaging
- Residualtransformer: Residual Low-Rank Learning With Weight-Sharing For Transformer Layers
- Resource-Constrained Stereo Singing Voice Cancellation
- Resource-Efficient Separation Transformer
- Retaining Informative Latent Variables in Probabilistic Segmentation
- Rethinking Normals: Direction Guided Point Cloud Recognition
- Rethinking Session Variability: Leveraging Session Embeddings for Session Robustness in Speaker Verification
- Rethinking Targeted Adversarial Attacks for Neural Machine Translation
- Retrieval Augmented End-to-End Spoken Dialog Models
- Retrieval-Augmented Text-to-Audio Generation
- Retrieval-Generation Synergy Augmented Large Language Models
- Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
- Reversible Jump Markov Chain Monte Carlo for Pulse Fitting
- Revise the NLU: A Prompting Strategy for Robust Dialogue System
- Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
- Revisiting the Equivalence of In-Context Learning and Gradient Descent: The Impact of Data Distribution
- Reweighted Atomic Norm Minimization for One-Bit Multichannel Spectral Compressed Sensing
- Riemannian Diffusion Adaptation over Graphs with Application to Online Distributed PCA
- Risk-Managed Sparse Index Tracking Via Market Graph Clustering
- RoFi: Robust WiFi Intrusion Detection via Distribution Matching
- Robust Beamforming for DFRC Systems in Complex Environments
- Robust Cross-Domain Speaker Verification with Multi-Level Domain Adapters
- Robust Decoding of the Auditory Attention from EEG Recordings Through Graph Convolutional Networks
- Robust DoA Estimation from Deep Acoustic Imaging
- Robust Face Recognition Based on an Angle-Aware Loss and Masked Autoencoder Pre-Training
- Robust Lightweight Depth Estimation Model via Data-Free Distillation
- Robust Localization of Key Fob Using Channel Impulse Response of Ultra Wide Band Sensors for Keyless Entry Systems
- Robust Low-Rank Correlation Fitting
- Robust Near-Field Beamforming for Millimeter Wave Communication System with Aperture Perturbations
- Robust Recovery of Joint Sparse Signals via Simultaneous Orthogonal Matching Pursuit
- Robust Regression Analysis Based on the K-Divergence
- Robust Self-Supervised Learning with Contrast Samples for Natural Language Understanding
- Robust Single-Particle Cryo-Em Image Denoising and Restoration
- Robust Speaker Personalisation Using Generalized Low-Rank Adaptation for Automatic Speech Recognition
- Robust Spoof Speech Detection Based on Multi-Scale Feature Aggregation and Dynamic Convolution
- Robust Symbol-Level Precoding via a Symbol-Perturbed Zero-Forcing Structure
- Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
- Robust and Imperceptible Commercial Camera-Screen Communication with 60Hz Refresh Rate
- RobustTSVar: A Robust Time Series Variance Estimation Algorithm
- Robustness Against Adversarial Attacks Via Learning Confined Adversarial Polytopes
- Robustness Evaluation of Machine Learning Models for Robot Arm Action Recognition in Noisy Environments
- Rényi Differential Privacy in the Shuffle Model: Enhanced Amplification Bounds
- S-Evaluator: Enhance Factual Consistency Evaluator with Adversarial Data Synthesized by Large Language Model
- S2E: Towards an End-to-End Entity Resolution Solution from Acoustic Signal
- SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR
- SADA: Saudi Audio Dataset for Arabic
- SADE: A Speaker-Aware Dual Encoding Model Based on Diagbert for Medical Triage and Pre-Diagnosis
- SALM: Speech-Augmented Language Model with in-Context Learning for Speech Recognition and Translation
- SAM-DEBLUR: Let Segment Anything Boost Image Deblurring
- SAM-GEBD: Zero-Cost Approach for Generic Event Boundary Detection
- SAM-OCTA: A Fine-Tuning Strategy for Applying Foundation Model OCTA Image Segmentation Tasks
- SAM: A Self-Adaptive Attention Module for Context-Aware Recommendation System
- SAMF: Small-Area-Aware Multi-Focus Image Fusion for Object Detection
- SAMVG: A Multi-Stage Image Vectorization Model with the Segment-Anything Model
- SAR2NDVI: Pre-Training for SAR-to-NDVI Image Translation
- SASA: Saliency-Aware Self-Adaptive Snapshot Compressive Imaging
- SBM: Smoothness-Based Minimization for Domain Generalization
- SC-MAD: Mixtures of Higher-Order Networks for Data Augmentation
- SCNet: Sparse Compression Network for Music Source Separation
- SCORE: Self-Supervised Correspondence Fine-Tuning for Improved Content Representations
- SCRN: A Spectrogram Convolutional Recurrent Network for AoA Estimation Using Bluetooth 5
- SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in Hubert
- SDEMG: Score-Based Diffusion Model for Surface Electromyographic Signal Denoising
- SDIF-DA: A Shallow-to-Deep Interaction Framework with Data Augmentation for Multi-Modal Intent Detection
- SDRNet: Saliency-Guided Dynamic Restoration Network for Rain and Haze Removal in Nighttime Images
- SE-SIS: Shadow-Embeddable Lossless Secret Image Sharing for Greyscale Images
- SEA-GNN: Sequence Extension Augmented Graph Neural Network for Sequential Recommendation
- SECP: A Speech Enhancement-Based Curation Pipeline for Scalable Acquisition of Clean Speech
- SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
- SEGLLM: Topic-Oriented Call Segmentation Via LLM-Based Conversation Synthesis
- SELM: Speech Enhancement using Discrete Tokens and Language Models
- SERC-GCN: Speech Emotion Recognition In Conversation Using Graph Convolutional Networks
- SG2SC: A Generative Semantic Communication Framework for Scene Understanding-Oriented Image Transmission
- SGM: A Dataset for 3D Garment Reconstruction from Single Hand-Drawn Sketch
- SGT: Self-Guided Transformer for Few-Shot Semantic Segmentation
- SIANet: Support Information-Aware Network for Category-Agnostic Pose Estimation
- SICRN: Advancing Speech Enhancement through State Space Model and Inplace Convolution Techniques
- SIMFALL: A Data Generator for RF-Based Fall Detection
- SIMMKD: Simple Mask-Flow Keypoint Detection for Both Typhoon Detection and Typhoon Eye Location
- SJTU-TMQA: A Quality Assessment Database for Static Mesh with Texture Map
- SMMA-Net: An Audio Clue-Based Target Speaker Extraction Network with Spectrogram Matching and Mutual Attention
- SO-Net: Model-Agnostic Sequential Hand Pose Optimization Framework
- SPASE: Spatial Saliency Explanation For Time Series Models
- SPATIALCODEC: Neural Spatial Speech Coding
- SPCL-MER: Supervised Prototypical Contrastive Learning for Micro-Expression Recognition
- SPDG-Net: Semantics Preserving Domain Augmentation through Style Interpolation for Multi-Source Domain Generalization
- SPEC-NERF: Multi-Spectral Neural Radiance Fields
- SPGFusion: A Semantic Prior Guided Infrared and Visible Image Fusion Network
- SPGM: Prioritizing Local Features for Enhanced Speech Separation Performance
- SPTESleepNet: Automatic Sleep Staging Model Based On Strip Patch Embeddings And Transformer Encoder
- SPY-Watermark: Robust Invisible Watermarking for Backdoor Attack
- SR-HuBERT : An Efficient Pre-Trained Model for Speaker Verification
- SR-VFA: Accurate Self-Refined Face Alignment in Videos
- SRECT: Machine-Specific Spatial-Resolution Enhancement in Computed Tomography
- SRP-UOD: Multi-Branch Hybrid Network Framework Based on Structural Re-Parameterization for Underwater Small Object Detection
- SSHNN: Semi-Supervised Hybrid NAS Network for Echocardiographic Image Segmentation
- SSL-Net: A Synergistic Spectral and Learning-Based Network for Efficient Bird Sound Classification
- SSR-GPCsT: Deep Learning Models Based on Functional Connectivity Maps in Autism Research
- SSTA: Salient Spatially Transformed Attack
- STEMGEN: A Music Generation Model That Listens
- STREAMVC: Real-Time Low-Latency Voice Conversion
- STS-CCL: Spatial-Temporal Synchronous Contextual Contrastive Learning for Urban Traffic Forecasting
- STYLECAP: Automatic Speaking-Style Captioning from Speech Based on Speech and Language Self-Supervised Learning Models
- STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
- SVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks
- SYNTHE-SEES: Face Based Text-to-Speech for Virtual Speaker
- Saliency Prediction of Sports Videos: A Large-Scale Database and a Self-Adaptive Approach
- Sam-Guided Enhanced Fine-Grained Encoding with Mixed Semantic Learning for Medical Image Captioning
- Sampling and Recovery of Signals Over Product Cell Structures
- Sandwiched Lo-Res Simulation for Scalable Flood Modeling
- Scalable Ensemble-Based Detection Method Against Adversarial Attacks For Speaker Verification
- Scalable Model-Based Gaussian Process Clustering
- Scalable and Efficient Speech Enhancement Using Modified Cold Diffusion: A Residual Learning Approach
- Scale-Aware Competition Network for Palmprint Recognition
- Scale-Free And Task-Generic Attack: Generating Photo-Realistic Adversarial Patterns With Patch Quilting Generator
- Scaling Results for Robust Distributed Estimation in Sensor Networks Using Order Statistics
- ScanPCGC: Learning-Based Lossless Point Cloud Geometry Compression using Sequential Slice Representation
- Scene Sketch-to-Image Synthesis Based on Multi-Object Control
- Score Calibration Based on Consistency Measure Factor for Speaker Verification
- Score-based Diffusion Models for Photoacoustic Tomography Image Reconstruction
- ScoreDec: A Phase-Preserving High-Fidelity Audio Codec with a Generalized Score-Based Diffusion Post-Filter
- SeACo-Paraformer: A Non-Autoregressive ASR System with Flexible and Effective Hotword Customization Ability
- Seam Mask Guided Partial Reconstruction with Quantum-Inspired Local Aggregation For Deep Image Stitching
- Search Robust and Adaptable Architecture
- Search for Gravitational Wave Probes - A Self-Supervised Learning for Pulsars Based on Signal Contexts
- Sec2Sec Co-Attention Transformer for Video-Based Apparent Affective Prediction
- Sector-Based Interference Cancellation for Robust Keyword Spotting Applications Using an Informed MPDR Beamformer
- Secure Energy Efficiency Fairness Maximization in Backscatter Throughput Constrained UAV-Assisted Data Collection
- Securely and Efficiently Outsourcing Neural Network Inference via Parallel MSB Extraction
- Security Equivalence Assessment between Cloud Standards by Mapping of Control Items
- Seeing Through The Conversation: Audio-Visual Speech Separation Based on Diffusion Model
- Seeking Similarities While Removing Differences: Graph Neural Networks Based on Node Correlation
- Segment Anything Model Guided Semantic Knowledge Learning For Remote Sensing Change Detection
- Segment Anything Model Meets Image Harmonization
- Segment then Match: Find the Carrier before Reasoning in Scene-Text VQA
- Segmentation-Driven Infrared and Visible Image Fusion Via Transformer-Enhanced Architecture Searching
- Segmented Error Minimisation (Semi) for Robust Training of Deep Learning Models with Non-Linear Shifts in Reference Data
- Selecting N-Lowest Scores for Training MOS Prediction Models
- Selective Domain-Invariant Feature for Generalizable Deepfake Detection
- Selective User Forwarded Cell-Free Massive Mimo with Quantized Symbols
- Self Knowledge Distillation Based On Layer-Wise Weighted Feature Imitation For Efficient Object Detection
- Self-Adaptive Scale Handling for Forecasting Time Series with Scale Heterogeneity
- Self-Distilled Dynamic Fusion Network for Language-Based Fashion Retrieval
- Self-Knowledge Distillation with Learning from Role-Model Samples
- Self-Motion As Supervision For Egocentric Audiovisual Localization
- Self-Supervised Adaptive AV Fusion Module for Pre-Trained ASR Models
- Self-Supervised Adaptive Pre-Training of Multilingual Speech Models for Language and Dialect Identification
- Self-Supervised Cross-Level Consistency Learning For Fundus Image Classification
- Self-Supervised Domain Exploration with an Optimal Transport Regularization for Open Set Cross-Domain Speech Emotion Recognition
- Self-Supervised Dual Generative Networks for Edge-Preserving Image Smoothing
- Self-Supervised Face Image Restoration with a One-Shot Reference
- Self-Supervised Learning for Anomalous Sound Detection
- Self-Supervised Learning for Sleep Stage Classification with Temporal Augmentation and False Negative Suppression
- Self-Supervised Models of Speech Infer Universal Articulatory Kinematics
- Self-Supervised Multi-Scale Hierarchical Refinement Method for Joint Learning of Optical Flow and Depth
- Self-Supervised Path Planning in UAV-Aided Wireless Networks Based on Active Inference
- Self-Supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
- Self-Supervised Pulse-Aware Interpretable Disentangled ECG Representation Learning
- Self-Supervised Reinforcement Learning for Out-of-Distribution Recovery via Auxiliary Reward
- Self-Supervised Spatially Variant PSF Estimation for Aberration-Aware Depth-from-Defocus
- Self-Supervised Speaker Verification Employing A Novel Clustering Algorithm
- Self-Supervised Speaker Verification with Adaptive Threshold and Hierarchical Training
- Self-Training Domain Adaptation Via Weight Transmission Between Generators
- SemDA: Communication-Efficient Data Aggregation Through Distributed Semantic Transmission
- Semantic Distillation and Structural Alignment Network for Fake News Detection
- Semantic Enrichment for Video Question Answering with Gated Graph Neural Networks
- Semantic Latent Decomposition with Normalizing Flows for Face Editing
- Semantic Proximity Alignment: Towards Human Perception-Consistent Audio Tagging by Aligning with Label Text Description
- Semantic Reconstruction of Continuous Language from Meg Signals
- Semantic Security: A Digital Watermark Method for Image Semantic Preservation
- Semantic Segmentation for Multi-Scene Remote Sensing Images with Noisy Labels Based on Uncertainty Perception
- Semantic-Enhanced Supervised Contrastive Learning
- Semantic-Guided Network with Contrastive Learning for Video Caption
- Semantic-Preserving Image Coding Based on Conditional Diffusion Models
- Semanticmapper: Region-Specific Domain Adaptation for 3D Shapes Through Lexical Delineation
- Semantics Driven Multi-View Knowledge Graph Embedding for Cross-Lingual Entity Alignment
- Semi-Autoregressive Streaming ASR with Label Context
- Semi-Blind Estimation of Direct-to-Reverberant Energy Ratio Using Residual Energy Test Statistics
- Semi-Decoupled 6D Pose Estimation via Multi-Modal Feature Fusion
- Semi-Supervised Domain Adaptation for Eeg-Based Sleep Stage Classification
- Semi-Supervised Metrics-Based Self-Training Root Cause Analysis for Cloud-Native Systems with Class-Imbalanced Data
- Semi-Supervised Sound Event Detection with Local and Global Consistency Regularization
- Semi-Supervised Volumetric Medical Image Segmentation via Class Prototype Guided Distribution-Aligned Representation Learning
- Sensi-Bert: Towards Sensitivity Driven Fine-Tuning for Parameter-Efficient Language Model
- Sensing with Random Signals
- Sensing-Aided Communication Channel Estimation with Tensor-Based Moving Target Localization
- Sensing-Assisted Distributed User Scheduling and Beamforming in Muli-Cell mmWave Networks
- Sequence of Linear Program for Robust Phase Retrieval
- Sequential Acquisition of Features and Experts for Datum-Wise Classification
- Sequential Detection of Anomalies in Noisy Outputs of an Unknown Function Using Gaussian and Yule-Simon Processes
- Sequential Monte Carlo Graph Convolutional Network for Dynamic Brain Connectivity
- Sequential Wasserstein Uncertainty Sets for Minimax Robust Online Change Detection
- Shapley Value Guided Extractive Text Summarization
- Shift Operator and Separation Filter for Different Period Mixed Signals Using Companion Matrix
- Shifted-Rectangle-Window Based Transformer for non-Displaced Femoral Neck Fracture Diagnosis
- Sifisinger: A High-Fidelity End-to-End Singing Voice Synthesizer Based on Source-Filter Model
- Signal Reconstruction from Nonideal Samples in Fractional Fourier Transform Domain
- Signal Transformer: Complex-Valued Attention and Meta-Learning for Signal Recognition
- Significant ASR Error Detection for Conversational Voice Assistants
- Similar but Faster: Manipulation of Tempo in Music Audio Embeddings for Tempo Prediction and Search
- Similarity Knowledge Distillation with Calibrated Mask
- Simple Contrastive Representation Learning for Time Series Forecasting
- Simultaneous Interior and Exterior Sound Field Synthesis Using Cylindrical and Spherical Loudspeaker Arrays
- Simultaneous Positioning and Tracking Using Dynamic Factor Graphs and Geometric Average Fusion
- SingFake: Singing Voice Deepfake Detection
- Single Image Reflection removal Using Feature Difference Enhancement
- Single and Few-Step Diffusion for Generative Speech Enhancement
- Single-Pixel Imaging Of Dynamic Flows Using Neural Ode Regularization
- Single-Source Domain Generalization in Fundus Image Segmentation Via Moderating and Interpolating Input Space Augmentation
- Situation-Aware Adaptive Transmit Beamforming for Automotive Radars
- Situational Signal Processing with Ecological Momentary Assessment: Leveraging Environmental Context for Cochlear Implant Users
- Sketch-Based 3D Shape Retrieval With Multi-View Fusion Transformer
- Sketched Column-Based Matrix Approximation With Side Information
- SkillNet-X: A Multilingual Multitask Model with Sparsely Activated Skills
- Skin Tone Disentanglement in 2D Makeup Transfer With Graph Neural Networks
- Skip-Step Contrastive Predictive Coding for Time Series Anomaly Detection
- SlideSpeech: A Large Scale Slide-Enriched Audio-Visual Corpus
- Slowfast Network for Continuous Sign Language Recognition
- Small Object Detection on the Water Surface Based on Radar and Camera Fusion
- Small-Footprint Automatic Speech Recognition System using Two-Stage Transfer Learning based Symmetrized Ternary Weight Network
- Small-Footprint Convolutional Neural Network with Reduced Feature Map for Voice Activity Detection
- Smooth Start: A Unified Approach for Gradual Transition from Cold to Old in Recommender Systems
- Snapshot Prompt Ensemble for Parameter-Efficient Soft Prompt Transfer
- Snore Sound Features Based on Percussive Enhancing and Positional Encoding Combined with Multi-Task Learning for Osahs Detection
- Social Learning with Adaptive Models
- Social Lode: Human Trajectory Prediction with Latent Odes
- Sod-Uav: Small Object Detection For Unmanned Aerial Vehicle Images Via Improved Yolov7
- Soft Alignment of Modality Space for End-to-End Speech Translation
- Soft Dynamic Time Warping with Variable Step Weights
- Soft Image Segmentation Using Gradient Graph Laplacian Regularizer
- Solution and Analysis For 3-D Localization In Closed-Form Integrating Sa and TDOA Measurements
- Sorting, Reasoning, and Extraction: An Easy-to-Hard Reasoning Framework for Document-Level Event Argument Extraction
- SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
- Source-Free Domain Adaptation for Millimeter Wave Radar Based Human Activity Recognition
- Source-Free Online Domain Adaptive Semantic Segmentation of Satellite Images Under Image Degradation
- SourceP: Detecting Ponzi Schemes on Ethereum with Source Code
- Space-Time Adaptive Processing for Radars in Connected and Automated Vehicular Platoons
- Sparse Bayesian Learning-Based Direct Localization for Distributed Sensor Arrays with Unknown Gain and Phase Errors
- Sparse Bayesian Synthetic Aperture Processing Based DOA Estimation with Deformed Towed Arrays
- Sparse Channel Representation and Estimation in Near Field Communications
- Sparse PCA with False Discovery Rate Controlled Variable Selection
- Sparse Regularization Based on Reverse Ordered Weighted L1-Norm and Its Application to Edge-Preserving Smoothing
- Sparse Sound Field Representation Using Complex Orthogonal Matching Pursuit
- Sparse, Weight-Constrained Arrays With O(N) Aperture for Reduced Mutual Coupling
- Sparsely Shared Lora on Whisper for Child Speech Recognition
- Sparsespikformer: A Co-Design Framework for Token and Weight Pruning in Spiking Transformer
- Spatial Formation-Guided Network for Group Activity Recognition
- Spatial Scaper: A Library to Simulate and Augment Soundscapes for Sound Event Localization and Detection in Realistic Rooms
- Spatial-Temporal Interaction Decoding Transformer for Unsupervised Multivariate Time Series Anomaly Detection
- Spatio-Temporal Action Detection with a Motion Sense and Semantic Correction Framework
- Spatio-Temporal Correlation Learning for Multiple Object Tracking
- Spatio-Temporal Data Mining with Information Integrity Protection: Graph Signal Based Air Quality Prediction
- Spatiotemporal Group Anomaly Detection via Graph Total Variation on Tensors
- Speak While You Think: Streaming Speech Synthesis During Text Generation
- Speaker Adaptation For Enhancement Of Bone-Conducted Speech
- Speaker Anonymization Using Neural Audio Codec Language Models
- Speaker-Adaptive Lipreading Via Spatio-Temporal Information Learning
- Speaker-Centric Multimodal Fusion Networks for Emotion Recognition in Conversations
- SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis
- Spectral Analysis of Vowels and Fricatives at Varied Levels of Dysarthria Severity for Amyotrophic Lateral Sclerosis
- Spectral Graph Neural Networks with Generalized Laguerre Approximation
- Spectro-Spatial Hyperspectral Image Reconstruction From Interferometric Acquisitions
- Spectrogram Smoothing for Estimation of the Evolutionary Spectra of Uniformly Modulated Processes
- SpectrumNet: Spectrum-Based Trajectory Encode Neural Network for Pedestrian Trajectory Prediction
- Speech Collage: Code-Switched Audio Generation by Collaging Monolingual Corpora
- Speech Emotion Recognition with Distilled Prosodic and Linguistic Affect Representations
- Speech Enhancement in Hearing Aids Using Target Speech Presence Estimation Based on a Delayed Remote Microphone Signal
- Speech Foundation Models on Intelligibility Prediction for Hearing-Impaired Listeners
- Speech Guided Masked Image Modeling for Visually Grounded Speech
- Speech Relationship Learning for Cross-Corpus Speech Emotion Recognition
- Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion Recognition
- Speech-Driven Emotional 3d Talking Face Animation Using Emotional Embeddings
- SpeechDPR: End-To-End Spoken Passage Retrieval For Open-Domain Spoken Question Answering
- Spiking Structured State Space Model for Monaural Speech Enhancement
- Spiking-Leaf: A Learnable Auditory Front-End for Spiking Neural Networks
- Spiral Shape Matters: Novel Bio-Inspired Cochlear Cepstrum
- Spontts: Modeling and Transferring Spontaneous Style for TTS
- Spoofing Attack Augmentation: Can Differently-Trained Attack Models Improve Generalisation?
- Srcodec: Split-Residual Vector Quantization for Neural Speech Codec
- Stability of Graph Convolutional Neural Networks Through The Lens of Small Perturbation Analysis
- Stable Distillation: Regularizing Continued Pre-Training for Low-Resource Automatic Speech Recognition
- Stable Knowledge Transfer for Contrastive Distillation
- Stable Optimization for Large Vision Model Based Deep Image Prior in Cone-Beam CT Reconstruction
- StableMiss+: Prediction with Incomplete Data Under Agnostic Mask Distribution Shift
- Stack-and-Delay: A New Codebook Pattern for Music Generation
- Stage-Regularized Neural Stein Critics For Testing Goodness-Of-Fit Of Generative Models
- State-Augmented Information Routing In Communication Systems With Graph Neural Networks
- Stateful Conformer with Cache-Based Inference for Streaming Automatic Speech Recognition
- Statistical and Computational Limits of Detecting and Recovering Hidden Submatrices
- Stealthy Backdoor Attack Towards Federated Automatic Speaker Verification
- Stein Variational Gradient Descent-Based Detection for Random Access with Preambles in MTC
- Stereo-Matching Knowledge Distilled Monocular Depth Estimation Filtered by Multiple Disparity Consistency
- Stereophonic Music Source Separation with Spatially-Informed Bridging Band-Split Network
- Stethoscope-Guided Supervised Contrastive Learning for Cross-Domain Adaptation on Respiratory Sound Classification
- Stochastic Configuration Networks for Laboratory Seismic Time-to-Failure Prediction
- StofNet: Super-Resolution Time of Flight Network
- StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
- Straightforward Adaptation of Particle Filter to Fish Eye Images for Top View Pedestrian Tracking
- Strategic Arms with Side Communication Prevail Over Low-Regret MAB Algorithms
- Streaming Active Learning for Regression Problems Using Regression via Classification
- Streaming Anchor Loss: Augmenting Supervision with Temporal Significance
- String Sound Synthesizer On Gpu-Accelerated Finite Difference Scheme
- Structure Matters: Analyzing Videos Via Graph Neural Networks for Social Media Platform Attribution
- Structure-Aware in-Air Handwritten Text Recognition with Graph-Guided Cross-Modality Translator
- Structure-Informed Positional Encoding for Music Generation
- Study of Abuse Detection in Continuous Speech for Indian Languages
- Style Adaptation for Domain-Adaptive Semantic Segmentation
- Stylespeech: Self-Supervised Style Enhancing with VQ-VAE-Based Pre-Training for Expressive Audiobook Speech Synthesis
- Subgroup Identification Through Multiplex Community Structure Within Functional Connectivity Networks
- Subnetwork-To-Go: Elastic Neural Network with Dynamic Training and Customizable Inference
- Subspace-Based Co-Array Processing For Nested Arrays without Eigendecomposition
- Subspace-Based Detection in OFDM ISAC Systems Under Different Constellations
- Subtype-Specific Biomarkers of Alzheimer's Disease from Anatomical and Functional Connectomes via Graph Neural Networks
- Summarizing Community-Based Question-Answer Pairs with Focus Rectification
- Sunflower Strategy for Bayesian Relational Data Analysis
- SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
- Supplementing Missing Visions Via Dialog for Scene Graph Generations
- Surface-Constrained Progressive Feature Preserving Point Cloud Compression
- SweepMM: A High-Quality Multimodal Dataset for Sweeping Robots in Home Scenarios for Vision-Language Model
- Syllable Level Features for Parkinson's Disease Detection from Speech
- Symmetric Consistency with Cross-Domain Mixup for Cross-Modality Cardiac Segmentation
- Symmetric VAR(1) Modelling with Guaranteed Stability
- Syncfusion: Multimodal Onset-Synchronized Video-to-Audio Foley Synthesis
- Synchformer: Efficient Synchronization From Sparse Cues
- Synonym Replacement and Generation Enhancement for Document Augmentation
- SynthTab: Leveraging Synthesized Data for Guitar Tablature Transcription
- Synthesizing Aβ-Pet Via An Image And Label Conditioning Latent Diffusion Model For Detecting Amyloid Status
- Synthesizing Black-Box Anti-Forensics Deepfakes With High Visual Quality
- Synthetic Conversations Improve Multi-Talker ASR
- Synthia's Melody: A Benchmark Framework for Unsupervised Domain Adaptation in Audio
- Synvox2: Towards A Privacy-Friendly Voxceleb2 Dataset
- T-EnFP: An Efficient Transformer Encoder-Based System for Driving Behavior Classification
- T-Foley: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis
- T-Pixel2Mesh: Combining Global and Local Transformer for 3D Mesh Generation from a Single Image
- T-SOT FNT: Streaming Multi-Talker ASR with Text-Only Domain Adaptation Capability
- TA2P: Task-Aware Adaptive Pruning Method for Image Classification on Edge Devices
- TACos: Learning Temporally Structured Embeddings for Few-Shot Keyword Spotting with Dynamic Time Warping
- TALDS-Net: Task-Aware Adaptive Local Descriptors Selection for Few-Shot Image Classification
- TAROT: A Hierarchical Framework with Multitask co-pretraining on Semi-Structured Data Towards Effective Person-Job fit
- TB-ResNet: Bridging the Gap from TDNN to ResNet in Automatic Speaker Verification with Temporal-Bottleneck Enhancement
- TCMP: End-to-End Topologically Consistent Magnitude Pruning for Miniaturized Graph Convolutional Networks
- TCNAS: Transformer Architecture Evolving in Code Clone Detection
- TD-GPT: Target Protein-Specific Drug Molecule Generation GPT
- TDT-KWS: Fast and Accurate Keyword Spotting Using Token-and-Duration Transducer
- TF-SepNet: An Efficient 1D Kernel Design in Cnns for Low-Complexity Acoustic Scene Classification
- TIA: A Teaching Intonation Assessment Dataset in Real Teaching Situations
- TNFormer: Single-Pass Multilingual Text Normalization with a Transformer Decoder Model
- TODM: Train Once Deploy Many Efficient Supernet-Based RNN-T Compression For On-Device ASR Models
- TRET: Two Stream-Based Regionally Enhanced Transformers for Person Re-Identification
- TRLS: A Time Series Representation Learning Framework Via Spectrogram for Medical Signal Processing
- TRUST-SER: On The Trustworthiness Of Fine-Tuning Pre-Trained Speech Embeddings For Speech Emotion Recognition
- Tackling Electrode Shift in Gesture Recognition with HD-EMG Electrode Subsets
- Tag Antenna Structure Calibrated Backscattering Signal Detection
- Tail Classes Matter: Long-Tailed Object Detection Revisited
- TalkNCE: Improving Active Speaker Detection with Talk-Aware Contrastive Learning
- Talking Face Generation for Impression Conversion Considering Speech Semantics
- Taming Prompt-Based Data Augmentation for Long-Tailed Extreme Multi-Label Text Classification
- Target Localization Based on Multistatic Mimo Radar via Double Coupled Canonical Polyadic Decomposition
- Target Optimization Direction Guided Transfer Learning for Image Classification
- Target Signal Power Improvement and Clutter Suppression via Beamforming for Integrated Sensing and Communication Systems
- Target Speaker Extraction by Directly Exploiting Contextual Information in the Time-Frequency Domain
- Target Speech Extraction with Pre-Trained Self-Supervised Learning Models
- Task Indicating Transformer for Task-Conditional Dense Predictions
- Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech Recognition
- Task Selection and Assignment for Multi-Modal Multi-Task Dialogue Act Classification with Non-Stationary Multi-Armed Bandits
- Task Vector Algebra for ASR Models
- Task-Wise Prompt Query Function for Rehearsal-Free Continual Learning
- Template-Guided Data Augmentation for Unbiased Scene Graph Generation
- Tempo Estimation as Fully Self-Supervised Binary Classification
- Temporal Conditional Coding for Dynamic Point Cloud Geometry Compression
- Temporal Convolution Shrinkage Network for Keyword Spotting
- Temporal Inconsistency-Based Active Learning
- Temporal Knowledge Graph Embedding using Householder Transformations
- Temporal Relational Context Learning for Extrapolation Reasoning on Temporal Knowledge Graphs
- Temporal-Spatial Prediction: Pre-Training on Diverse Datasets for EEG Classification
- Temporally-Guided Total Variation For Robust Spatiotemporal Fusion Of Satellite Images
- Ten-Guard: Tensor Decomposition for Backdoor Attack Detection in Deep Neural Networks
- Tensor Decomposition-Based Data Fusion for Biomarker Extraction from Multiple EEG Experiments
- Tensor Graph Decomposition for Temporal Networks
- Tensor Low-Rank Approximation of Finite-Horizon Value Functions
- Tensor Reconstruction-Based Sparse Array 2-D DOA Estimation of Mixed Coherent and Uncorrelated Signals
- Tensor-Guided Interpolation For Off-Grid Power Spectrum Map Construction
- Tensorial Convolutive Blind Source Separation
- Test-Time Distribution Learning Adapter for Cross-Modal Visual Reasoning
- Text Region Multiple Information Perception Network for Scene Text Detection
- Text-Driven Talking Face Synthesis by Reprogramming Audio-Driven Models
- Text-Only Unsupervised Domain Adaptation for Neural Transducer-Based ASR Personalization Using Synthesized Data
- Text-Video Completion Networks With Motion Compensation And Attention Aggregation
- Text2Avatar: Text to 3d Human Avatar Generation with Codebook-Driven Body Controllable Attribute
- TextrolSpeech: A Text Style Control Speech Corpus with Codec Language Text-to-Speech Models
- Textual Tokens Classification for Multi-Modal Alignment in Vision-Language Tracking
- Texture and Normal Map Estimation for 3D Face Reconstruction
- Texture-Unet: A Texture-Aware Network for Bone Marrow Smear Whole-Slide Image Region of Interest Segmentation
- The 2nd Clarity Prediction Challenge: A Machine Learning Challenge for Hearing Aid Intelligibility Prediction
- The Collaboration of 3D Convolutions and CRO-TSM in Lipreading
- The Devil is in Details: Delving Into Lite FFN Design for Vision Transformers
- The Double-Edged Sword Of Ai Safety: Balancing Anomaly Detection and OOD Generalization Via Model Anchoring
- The Effects of Loudness and Smiling on Timbre Features: Implications for Charismatic Voices in Mandarin, German and Danish
- The Joint Grid-Free DOA and Polarization Estimation Algorithm based on Atomic Norm Minimization
- The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker Extraction
- The Power of Few: Accelerating and Enhancing Data Reweighting with Coreset Selection
- The Rao, Wald, And Likelihood-Ratio Tests under Generalized Self-Concordance
- The Selectivity and Competition of the Mind's Eye in Visual Perception
- Theme-Enhanced Hard Negative Sample Mining for Open-Domain Question Answering
- Think as People: Context-Driven Multi-Image News Captioning with Adaptive Dual Attention
- Three-Dimensional Decoupled Atomic Norm Minimization
- Three-Dimensional Sound Wave Propagation Reproduction by CE-FDTD Simulation Applying Actual Radiation Characteristics
- Three-Dimensional Spatial-Temporal Near-Field Passive Localization Based on an Exact Spatial Propagation Model
- Through-The-Wall Radar Imaging With Wall Clutter Removal Via Riemannian Optimization On The Fixed-Rank Manifold
- Timbre-Trap: A Low-Resource Framework for Instrument-Agnostic Music Transcription
- Time Changed Normalizing Flows for Accurate SDE Modeling
- Time-Interval Visual Saliency Prediction in Mammogram Reading
- Time-Modulated Intelligent Reflecting Surface for Waveform Security
- Titan: Bringing the Deep Image Prior to Implicit Representations
- Token-Based Spatiotemporal Representation of the Events
- Tokenmotion: Motion-Guided Vision Transformer for Video Camouflaged Object Detection VIA Learnable Token Selection
- Topological Neural Networks over the Air
- Topology-Dependent Privacy Bound for Decentralized Federated Learning
- Topology-Regularized Self-Knowledge Distillation for Transductive-Inductive Learning of Brain Disorder Diagnosis
- Touring Sampling With Pushforward Maps
- Toward Quantifiable Face age Transformation
- Toward Sufficient Spatial-Frequency Interaction for Gradient-Aware Underwater Image Enhancement
- Towards 3D Computational Persicopy with an Ordinary Camera: a Separable Non-Linear Least Squares Formulation
- Towards A World-English Language Model for on-Device Virtual Assistants
- Towards ASR Robust Spoken Language Understanding Through in-Context Learning with Word Confusion Networks
- Towards Automatic Data Augmentation for Disordered Speech Recognition
- Towards Building The Federatedgpt: Federated Instruction Tuning
- Towards Controlled Table-to-Text Generation with Scientific Reasoning
- Towards Disease-Aware Self-Supervised Dynamic Brain Network Learning For Mental Diagnosis
- Towards Efficient Modeling and Inference in Multi-Dimensional Gaussian Process State-Space Models
- Towards Enabling DPOAE Estimation on Single-Speaker Earbuds
- Towards End-to-End Spoken Grammatical Error Correction
- Towards Faster End-to-End Data Transmission Over Voice Channels
- Towards Generic Deepfake Detection with Dynamic Curriculum
- Towards High Resolution Weather Monitoring With Sound Data
- Towards High-Performance and Low-Latency Feature-Based Speaker Adaptation of Conformer Speech Recognition Systems
- Towards Improving Speech Emotion Recognition Using Synthetic Data Augmentation from Emotion Conversion
- Towards Intelligent Design: A Self-Driven Framework for Collocated Clothing Synthesis Leveraging Fashion Styles and Textures
- Towards Interpretability of Automatic Phoneme Analysis in Cleft Lip and Palate Speech
- Towards Multi-Domain Face Landmark Detection with Synthetic Data from Diffusion Model
- Towards Omniscient Feature Alignment for Video Rescaling
- Towards Optimal Voice Disentanglement with Weak Supervision
- Towards Optimized Multi-Channel Modulo-ADCs: Moduli Selection Strategies and Bit Depth Analysis
- Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-Training and Multi-Modal Tokens
- Towards Resource-Efficient and Secure Federated Multimedia Recommendation
- Towards Robust Multimodal Prompting with Missing Modalities
- Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS
- Towards Video-Text Retrieval Adversarial Attack
- Towards a Unified View of Adversarial Training: A Contrastive Perspective
- Towards an Interpretable Representation of Speaker Identity via Perceptual Voice Qualities
- Towards an Objective Quality Metric for Interpolated Directional Room Impulse Responses
- Tracking Beyond the Unambiguous Range with Modulo Single-Photon Lidar
- Tracking of Multiple Spawning Targets with Heterogeneous Sensors for Seabed-To-Space Situational Awareness
- Trades++: Enhancing Multi-Object Tracking of Real Low Confidence Targets Using a Pyramid-Like Self-Attention Model
- Train Long and Test Long: Leveraging Full Document Contexts in Speech Processing
- Training Audio Captioning Models without Audio
- Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator
- Training Ultra-Low-Latency Spiking Neural Networks from Scratch
- Trajectory set Empowered Hypergraph Transformer for Mobile Sensor Based Traffic Prediction
- TranSentence: speech-to-speech Translation via Language-Agnostic Sentence-Level Speech Encoding without Language-Parallel Data
- TransAVS: End-to-End Audio-Visual Segmentation with Transformer
- TransCycle: A Data Augmentation Method for 3D Human Pose Estimation
- TransMUSIC: A Transformer-Aided Subspace Method for DOA Estimation with Low-Resolution ADCS
- Transducers with Pronunciation-Aware Embeddings for Automatic Speech Recognition
- Transfer the Linguistic Representations from TTS to Accent Conversion with Non-Parallel Data
- Transferable Models for Bioacoustics with Human Language Supervision
- Transferring Structure Knowledge: A New Task to Fake News Detection towards Cold-Start Propagation
- Transformer Model with Multi-Type Classification Decisions for Intrusion Attack Detection of Track Traffic and Vehicle
- Transformer-Inspired Lightweight Model for Efficient Time Series Forecasting
- Transforming Cardiovascular Health: a Transformer-Based Approach to Continuous, Non-Invasive Blood Pressure Estimation via Radar Sensing
- Translatotron 3: Speech to Speech Translation with Monolingual Data
- Transmit Beampattern Optimization for MIMO-ISAC Systems with Hybrid Beamforming
- Transmitting Data Through Reconfigurable Intelligent Surface: A Spatial Sigma-Delta Modulation Approach
- Tree Network Design for Faster Distributed Machine Learning Process with Distributed Dual Coordinate Ascent
- Tree of Uncertain Thoughts Reasoning for Large Language Models
- Treemil: A Multi-Instance Learning Framework for Time Series Anomaly Detection with Inexact Supervision
- Trend-Heuristic Reinforcement Learning Framework for News-Oriented Stock Portfolio Management
- Trusted Deep Domain Adaptation with Uncertainty Measure Based on Evidence Theory
- Turn-Taking and Backchannel Prediction with Acoustic and Large Language Model Fusion
- Two-Edge-Resolved 3d Non-Line-of-Sight Imaging: A Fisher Information Equalized Discretization
- Two-Stage Acoustic Echo Cancellation Network with Dual-Path Alignment
- Two-Stage Transfer Learning for Fusion and Classification of Airborne Hyperspectral Imagery
- Two-Step Knowledge Distillation for Tiny Speech Enhancement
- Type-Aware Decoding Via Explicitly Aggregating Event Information for Document-Level Event Extraction
- U2R: Underwater Ultrasonic Reflection Wave Dataset Toward Pose-Invariant Material Recognition
- UAV Operation Time Minimization for Wireless-Powered Data Collection
- UAV-Based Dynamic Object Tracking with Radio Map
- UNAD: Universal Anatomy-Initialized Noise Distribution Learning Framework Towards Low-Dose CT Denoising
- UNIDEAL: Curriculum Knowledge Distillation Federated Learning
- UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization
- UNeC: Unsupervised Exploring In Controllable Space
- USM-Lite: Quantization and Sparsity Aware Fine-Tuning for Speech Recognition with Universal Speech Models
- USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models
- Ultra Low Complexity Deep Learning Based Noise Suppression
- Ultra-Lightweight Neural Differential DSP Vocoder for High Quality Speech Synthesis
- Ultra-Low Delay Lossless Compression of Higher Order Ambisonics
- Uncertainty Quantification in Deep Learning Based Kalman Filters
- Uncertainty-Guided Contrastive Learning For Single Source Domain Generalisation
- Uncertainty-Guided Person Search Model with Auxiliary Shallow Feature Exploration
- Uncertainty-Guided Physics-Driven Deep Learning Reconstruction via Cyclic Measurement Consistency
- Uncovering Strong Ties: A Study of Indirect Sybil Attack on Signed Social Network
- Underlying-Complementarity and Surrounding-Correspondence for Multi-View Clustering
- Understanding Data Augmentation From A Robustness Perspective
- Understanding Gaussian Noise Mismatch: A Hellinger Distance Approach
- Understanding Probe Behaviors Through Variational Bounds of Mutual Information
- UniX-Encoder: A Universal X-Channel Speech Encoder for AD-HOC Microphone Array Speech Processing
- Unidirectional Brain-Computer Interface: Artificial Neural Network Encoding Natural Images to FMRI Response in the Visual Cortex
- Unified Analysis of Correlation-Aware Joint Sparse Support Recovery with ℓ0-Norm Constraint
- Unified Pretraining Target Based Video-Music Retrieval with Music Rhythm and Video Optical Flow Information
- Unified Probability Distributions of Generalized Composite Fading with Inverse-Type Distributions of Large-Scale Shadowing/Fluctuations
- Unified Speech and Gesture Synthesis Using Flow Matching
- Unified Srgb Real Noise Synthesizing with Adaptive Feature Modulation
- Unifying One-Shot Voice Conversion and Cloning with Disentangled Speech Representations
- Unimodal Aggregation for CTC-Based Speech Recognition
- Unintended Memorization in Large ASR Models, and How to Mitigate It
- Unitary Approximate Message Passing for Matrix Factorization
- Universal Adversarial Attack Against Speaker Recognition Models
- Unlabelled Sensing with Priors: Algorithm and Bounds
- Unleashing Trigger-Free Event Detection: Revealing Event Correlations Via a Contrastive Derangement Framework
- Unlocking Deep Learning: A BP-Free Approach for Parallel Block-Wise Training of Neural Networks
- Unravel Anomalies: an End-to-End Seasonal-Trend Decomposition Approach for Time Series Anomaly Detection
- Unraveling Explainable Reinforcement Learning Using Behavior Tree Structures
- Unrestricted Global Phase Bias-Aware Single-Channel Speech Enhancement with Conformer-Based Metric Gan
- Unrolled Proximal Gradient Descent Method for Non-Negative Least Squares Problem
- Unsupervised Accent Adaptation Through Masked Language Model Correction of Discrete Self-Supervised Speech Units
- Unsupervised Acoustic Scene Mapping Based on Acoustic Features and Dimensionality Reduction
- Unsupervised Anomaly Detection for Multivariate Time Series Using Diffusion Model
- Unsupervised Continual Learning of Image Representation Via Rememory-Based Simsiam
- Unsupervised Disparity Estimation for Light Field Videos
- Unsupervised Extractive Dialogue Summarization in Hyperdimensional Space
- Unsupervised Harmonic Parameter Estimation Using Differentiable DSP and Spectral Optimal Transport
- Unsupervised Human Activity Recognition Via Large Language Models and Iterative Evolution
- Unsupervised Learning Based End-to-End Delayless Generative Fixed-Filter Active Noise Control
- Unsupervised Learning of Facial Optical Flow via Occlusion-Aware Global-Local Matching
- Unsupervised Learning of Neural Semantic Mappings with the Hungarian Algorithm for Compositional Semantics
- Unsupervised Multi-Channel Separation And Adaptation
- Unsupervised Multi-Domain Data Selection for Asr Fine-Tuning
- Unsupervised Multiple Choices Question Answering Via Universal Corpus
- Unsupervised Optimal Power Flow Using Graph Neural Networks
- Unsupervised Pitch-Timbre Disentanglement of Musical Instruments Using a Jacobian Disentangled Sequential Autoencoder
- Unsupervised Remote Sensing Haze Removal Based on Saliency-Guided Transmission Refinement
- Unsupervised Speech Enhancement with Diffusion-Based Generative Models
- Unsupervised Speech Recognition with N-skipgram and Positional Unigram Matching
- Unsupervised Topic-Conditional Extractive Summarization
- Unsupervised multiple domain translation through controlled Disentanglement in variational autoencoder
- Updated Corpora and Benchmarks for Long-Form Speech Recognition
- Uplink Symbol Detection in Dynamic TDD Mimo Systems with AP-AP Interference
- Urban Traffic Flow Forecasting Based on Spatial-Temporal Graph Contrastive Learning
- User-Assisted Networked Sensing in OFDM Cellular Network with Erroneous Anchor Position Information
- Using Clustering to Improve the Performance of few-shot Learning
- Using Temporal Consistency for Compressed Sensing in High-Resolution mmWave Sounding
- Utilizing Second-Order Information in Noisy Information-Sharing Environments for Distributed Optimization
- V-DDPM: MRI Rician Noise Removal Model Based on VST and DDPM
- VCD: A Video Conferencing Dataset for Video Compression
- VFD-Net: Vocoder Fingerprints Detection for Fake Audio
- VGDIFFZERO: Text-To-Image Diffusion Models Can Be Zero-Shot Visual Grounders
- VIC-KD: Variance-Invariance-Covariance Knowledge Distillation to Make Keyword Spotting More Robust Against Adversarial Attacks
- VK-G2T: Vision and Context Knowledge Enhanced Gloss2text
- VL-FAS: Domain Generalization via Vision-Language Model For Face Anti-Spoofing
- VMCC-NET: Uncovering Challenging Regions in Semi-Supervised Medical Image Segmentation with Voxel Mask Based Cyclic-Consistency Network
- VRDMG: Vocal Restoration via Diffusion Posterior Sampling with Multiple Guidance
- VT-ReID: Learning Discriminative Visual-Text Representation for Polyp Re-Identification
- Variance Reduction Can Improve Trade-Off in Multi-Objective Learning
- Variational Analysis of Adversarial Regularization for Solving Inverse Problems
- Variational Connectionist Temporal Classification for Order-Preserving Sequence Modeling
- Vector Approximate Message Passing for Not So Large N.I.I.D. Generalized I/O Linear Models
- Vector Approximate message Passing with Arbitrary I.I.D. Noise Priors
- Vector Nonlinear Hawkes Model with Inhibition
- Vector Quantization Knowledge Transfer for End-to-End Text Image Machine Translation
- ViLaS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition
- Video Anomaly Prediction: Problem, Dataset and Method
- Video-Language Graph Convolutional Network for Human Action Recognition
- View Crafting For Instance-Level Representation from Scene Images
- Viewing Writing as Video: Optical Flow based Multi-Modal Handwritten Mathematical Expression Recognition
- Vision Transformer with 2D Explicit Position Encoding
- Vision-Sensor Attention Based Continual Multimodal Egocentric Activity Recognition
- Visual Adapt for RGBD Tracking
- Visual Prompt Tuning for Weakly Supervised Phrase Grounding
- Visual Speech Recognition for Languages with Limited Labeled Data Using Automatic Labels from Whisper
- Visual-Linguistic Representation Learning with Deep Cross-Modality Fusion for Referring Multi-Object Tracking
- Visually Dehallucinative Instruction Generation
- Visually Guided Binaural Audio Generation with Cross-Modal Consistency
- Vocal Fold Dynamics for Automatic Detection of Amyotrophic Lateral Sclerosis from Voice
- Voice Anonymization for All-Bias Evaluation of the Voice Privacy Challenge Baseline Systems
- Voice Toxicity Detection Using Multi-Task Learning
- VoiceFlow: Efficient Text-To-Speech with Rectified Flow Matching
- VoiceLDM: Text-to-Speech with Environmental Context
- Volumetric 3d Point Cloud Attribute Compression: Learned Polynomial Bilateral Filter for Prediction
- VoxMM: Rich Transcription of Conversations in the Wild
- Voxblink: A Large Scale Speaker Verification Dataset on Camera
- VoxtLM: Unified Decoder-Only Models for Consolidating Speech Recognition, Synthesis and Speech, Text Continuation Tasks
- Vulnerability of Face age Verification to Replay Attacks
- WAVER: Writing-Style Agnostic Text-Video Retrieval Via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
- WFTNet: Exploiting Global and Local Periodicity in Long-Term Time Series Forecasting
- WI-FI based Indoor Monitoring Enhanced by Multimodal Fusion
- WIFIACT: Enhancing Human Sensing Through Environment Robust Preprocessing And Bayesian Self-Supervised Learning
- Water Leak Detection via Domain Adaptation
- WaterDiff: Perceptual Image Watermarks Via Diffusion Model
- Wav2vec-VC: Voice Conversion via Hidden Representations of Wav2vec 2.0
- Wavelet-Decoupling Contrastive Enhancement Network for Fine-Grained Skeleton-Based Action Recognition
- Wavelet-Guided Acceleration of Text Inversion in Diffusion-Based Image Editing
- Wavelet-Inspired Multiscale Graph Convolutional Recurrent Network for Traffic Forecasting
- Weakly Semi-Supervised Tool Detection in Minimally Invasive Surgery Videos
- Weakly Supervised Few-Shot Segmentation Through Textual Prompt
- Weakly-Supervised Crowd Counting with Token Attention and Fusion: A Simple and Effective Baseline
- What Do Neural Networks Listen to? Exploring the Crucial Bands in Speech Enhancement Using SINC-Convolution
- What Do Self-Supervised Speech and Speaker Models Learn? New Findings from a Cross Model Layer-Wise Analysis
- When Green Learning Meets Federated Learning: Toward Distributed Learning with Low Complexity and Model Heterogeneity
- When Training-Free Nas Meets Vision Transformers: A Neural Tangent Kernel Perspective
- Which is the Better Teacher Action? A New Ranking Model and Dataset
- Whisper-Based Transfer Learning for Alzheimer Disease Classification: Leveraging Speech Segments with Full Transcripts as Prompts
- Widrow-Hoff LMS Adaline Demonstrator for Schools and Colleges
- Window-Based Convolutional Sparse Coding: Towards A Unified Framework
- X-CAUNET: Cross-Color Channel Attention with Underwater Image-Enhancing Transformer
- XMP: A Cross-Attention Multi-Scale Performer for File Fragment Classification
- YOLO-Med : Multi-Task Interaction Network for Biomedical Images
- ZE-FESG: A Zero-Shot Feature Extraction Method Based on Semantic Guidance for No-Reference Video Quality Assessment
- ZIV-Zakai Bound for DOA Estimation with Gain-Phase Error
- Zero Resource Code-Switched Speech Benchmark Using Speech Utterance Pairs for Multiple Spoken Languages
- Zero Shot Audio To Audio Emotion Transfer With Speaker Disentanglement
- Zero- and Few-Shot Sound Event Localization and Detection
- Zero-Shot Co-Salient Object Detection Framework
- Zero-Shot Imitation Policy Via Search In Demonstration Dataset
- Zero-Shot Intent Classification Using a Semantic Similarity Aware Contrastive Loss and Large Language Model
- Zero-Shot Object Detection with Partitioned Contrastive Feature Alignment
- Zigzag Attention: A Structural Aware Module For Lane Detection
- mmBaT: A Multi-Task Framework for Mmwave-Based Human Body Reconstruction and Translation Prediction
- uSee: Unified Speech Enhancement And Editing with Conditional Diffusion Models
- uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.