ICASSP 2025 Accepted Papers
The full list of 3,298 papers accepted at ICASSP 2025 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- "I've Heard of You!": Generate Spoken Named Entity Recognition Data for Unseen Entities
- 2.5D Top-K Ranked Multiple Instance Learning to Classify NSCLC PD-L1 Status on CT Images
- 30+ Years of Source Separation Research: Achievements and Future Challenges
- 3D Gaussian Splatting with Grouped Uncertainty for Unconstrained Images
- 3D Mesh Saliency Based on Dictionary Learning with Multi-Level Laplacian-Beltrami Operator
- 3D Shape Classification by Registration: Neural-Network-Free and Training-Free
- 3D TDOA-AOA Quaternion Based Acoustic SLAM for Drone Localization and Source Mapping
- 3D View Optimization for Improving Image Aesthetics
- 3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization
- 3DGCQA: A Quality Assessment Database for 3D AI-Generated Contents
- 3DSignDiff: Towards 3D Sign Language Gesture Generation
- 3GPP IVAS Codec - Perspectives on Development, Testing and Standardization
- A 3D Attenuation Coefficient based Degradation Estimation for Real Non-Homogeneous Dehazing
- A Bandwidth Efficient Dual Function Radar Communication System Based on a MIMO Radar Using OTFS Waveforms
- A Bayesian Interpretation of Adaptive Low-Rank Adaptation
- A Bayesian Perspective on Uncertainty Quantification for Estimated Graph Signals
- A Bilinear Source Separation, Dereverberation, and Background Noise Suppression Algorithm for Augmented Reality Applications
- A Blind Super-Resolution Method for Near-Field Channel Estimation with Angle-Range Recovery
- A Block Term Decomposition Model Based Algorithm for Tensor Completion of Multidimensional Harmonic Signals
- A CT-based Prediction System for Determining Respiratory Support Level in COVID-19 Patients
- A Chinese Expressive Long-dialogue Speech Dataset with Scripts
- A Clinical Knowledge-Driven Fine-Tuning Strategy for Applying Foundation Model to Fully Automatic Acute Ischemic Stroke Lesion Segmentation on Non-Contrast CT Scans
- A Comparative Analysis of Generalised Echo and Interference Cancelling and Extended Multichannel Wiener Filtering for Combined Noise Reduction and Acoustic Echo Cancellation
- A Comparative Study of Invariance-Aware Loss Functions for Deep Learning-based Gridless Direction-of-Arrival Estimation
- A Conditional KAN Diffusion Network for Human Activity Recognition with Missing Sensor Signal Series
- A Continual Learning Approach for Embodied Question Answering with Generative Adversarial Imitation Learning
- A Convolutional Recurrent Mixer Network For Radar Meteorological Image Super-Resolution
- A Cost-effective Solution for Remote Sensing Image Segmentation via Train/Test-Time Adaptation
- A Counterfactual Ultrasound Anti-Interference Self-Supervised Network for B-mode Ultrasound Tongue Extraction
- A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio
- A Cross-Modal Multi-Attitude Framework for the Generation of Space Target ISAR Images
- A Deformable-Based Source-Free Unsupervised Domain Adaptation Method for Cervical Cell Detection
- A Diffusion Model over Directed Acyclic Graphs for Event Schema Generation
- A Distillation-based Future-aware Graph Neural Network for Stock Trend Prediction
- A Divide-and-conquer Approach for Sparse Recovery in High Dimensions
- A Domain Adaptation Framework for Speech Recognition Systems with Only Synthetic data
- A Domain Adversarial Learning Framework for Major Depression Disorder Diagnosis
- A Domain-Specific Multilingual Speech Translation Corpus via Simultaneous Interpretation
- A Dual-Perspective Metaphor Detection Framework Using Large Language Models
- A Dual-Stream Network with Non-Stationary Characteristics-Enhanced for SST Image Prediction
- A Dynamic Edge-Selection Mechanism in HRV Hypergraph Learning for Improved Stress Detection
- A Dynamical Equation Approach For Quasi-Periodic Gaussian Processes
- A Fast Saturation Based Dehazing Framework with Accelerated Convolution and Attention Block
- A Federated Learning Network Intrusion Detection System for Multiple Imbalances
- A Federated Learning-Based Intrusion Detection System for Satellite-Terrestrial Integrated Networks
- A Framework Based on Data Augmentation for Knowledge Graph Entity Typing
- A Frequency-aware Augmentation Network for Mental Disorders Assessment from Audio
- A Fuzzy C-Means Clustering Algorithm for Real Medical Image Segmentation
- A GNSS-IR Aided Multispectral Satellite Data Fusion for Meter-Level Wide-Area Volumetric Soil Moisture Estimation
- A Generalized Graph Signal Processing Framework for Multiple Hypothesis Testing over Networks
- A Generative-Augmented Deep Matrix Factorization Model for POI Recommendations
- A Geometry-Based Node Activation Method for Relative Localization
- A Graph-Based Generative Adversarial Network Model for Inferring Task-State from Resting-State Functional Connectivity Networks
- A Grouping Strategy-Based Progressive Fusion Network for Hyperspectral Image Super-Resolution
- A Hierarchical Compression Technique for 3D Gaussian Splatting Compression
- A Hierarchical Flow for Few-shot Anomaly Detection via Global-local Aggregation Strategy
- A Hierarchical Reasoning Framework for Complex Question Answering over Knowledge Graph with Reinforcement Learning
- A Hierarchical Taxonomy For Deep State Space Models
- A High-Precision Character Cartoon Style Transfer Method Based on VToonify and Diffusion Models
- A Hybrid Model for Weakly-Supervised Speech Dereverberation
- A Hybrid Probabilistic-Deterministic Model Recursively Enhancing Speech
- A Joint Time-Frequency Attention for Leakage Detection in Water Distribution Networks Using Time Series Decomposition
- A Key to Effective Multi-task Learning: Separate Query Selection for Task-Synergized Handling and Node Utilization
- A Label Co-occurrence Transformation Network for Joint Empathy Detection and Empathy Intent Classification
- A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
- A Lowrate Variable-Bias Integrate-and-Fire Time Encoding Machine
- A Mamba-based Network for Semi-supervised Singing Melody Extraction Using Confidence Binary Regularization
- A Margin-Maximizing Fine-Grained Ensemble Method
- A Method for Removing Reflections from Water Surface Images Based on Pre-trained Image Restoration
- A Metric for Predicting the Quality of Ambisonic Spatial Audio Reproduced Using Spatially Interpolated or Extrapolated Room Impulse Responses
- A MoE Multimodal Graph Attention Network Framework for Multimodal Emotion Recognition
- A Model Stealing Attack Against Multi-Exit Networks
- A Modified Gain Normalized Step Size Adaptive Algorithm for Improved Online Secondary Path Modelling in Active Noise Control
- A Modified Nonlinear Matched Filter for Skewed Noise Based on the Gram-Charlier Expansion
- A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation
- A Multi-Agent Multi-Environment Mixed Q-Learning for Partially Decentralized Wireless Network Optimization
- A Multi-Label EEG Dataset for Mental Attention State Classification in Online Learning
- A Multi-Modal Information Fusion Model for Automatic Sleep Staging
- A Multi-Prior Fusion Network for Video-based Micro-Expression Recognition
- A Multi-Stage Feature Pipeline on Timestamped Speech Transcriptions for Dementia Assessment
- A Multi-Wavelength Optical Sensing Framework for Calibration-Free Wearable Blood Pressure Monitoring
- A Multi-annotated and Multi-modal Dataset for Wide-angle Video Quality Assessment
- A Multi-modal Approach to Dysarthria Detection and Severity Assessment Using Speech and Text Information
- A Multi-scenario Attention-based Generative Model for Personalized Blood Pressure Time Series Forecasting
- A Near-Field 3D Parameter Estimation Method Based on a Symmetric Enhanced Nested Array
- A New Model for Prototype-based Continual Learning in Hyperspherical Space
- A Noisy Label Filter based on GMM Binary Classification for Speaker Verification
- A Non-autoregressive Model for Joint STT and TTS
- A Novel Audio-Visual Multimodal Semi-Supervised Model Based on Graph Neural Networks for Depression Detection
- A Novel Compressive Compound Word Encoding and Independent Word Attention for Symbolic Music Generation
- A Novel Decision-Making Model for Playing Board Game Combining Planning and Opponent Behaviors
- A Novel Multimodal Method for Decoding Speech Perception from Brain Activities
- A Novel Network for Short-Term Wind Speed Prediction: Mitigating Distribution Shift and Feature Loss
- A Novel Self-Supervised Contrastive Learning Framework for Masked EEG Motor Imagery Modeling
- A Novel Single Continuous Shot Multiple Lesions Endoscopy Report Generation
- A Novel Split Deep Unfolding Transformer for Pan-Sharpening
- A Novel Underwater Acoustic Signal Denoising Model Based on Complex Convolution Dual-branch Multi-scale Attention Network
- A Novel Weighted Sparse Component Analysis for Underdetermined Blind Speech Separation
- A Parametric Non-Negative Coupled Canonical Polyadic Decomposition Algorithm for Hyperspectral Super-Resolution
- A Plug-and-Play Diffusion-Styled Conversion Model for Domain Discrepancies in Medical Image Segmentation
- A Practical Gated Recurrent Transformer Network Incorporating Multiple Fusions for Video Denoising
- A Pre-trained Plug-in Mixture-of-LoRAs Model for Transferable Sequential Recommendation
- A Pre-training Framework that Encodes Noise Information for Speech Quality Assessment
- A Privacy-Preserving Cross-Modal Retrieval Scheme Based on CLIP and Deep Hashing
- A Progressive Local Variance-guided Strategy for Improving Data Augmentation Reliability
- A Prompt Learning Framework with Large Language Model Augmentation for Few-shot Multi-label Intent Detection
- A Proximal Variable Smoothing for Nonsmooth Minimization Involving Weakly Convex Composite with MIMO Application
- A Quality-Aware Sampling Framework for Efficient 3D Point Cloud Transmission
- A Quantitative Metric Selection Approach for Time-series Forecasting Foundation Models
- A Ranking Scheme for Trust Region Multi-agent Reinforcement Learning
- A Reinforcement Learning Agent Controlled Multi-branch Small Object Detection Framework
- A Riemannian Approach to Ground Metric Learning for Optimal Transport
- A Risk Prediction Model for Real Estate Corporations Using High-Target Semantic BERT and Improved GRU
- A Robust Distributed Recurrent Neural Network for Multi-Agent Consensus Control
- A Robust Online Miscalibration Detection and Correction Method for LiDAR-Camera
- A Robust Quality Evaluator for Panoramic Videos
- A Scale-Adaptive and Background-Robust Method for Surface Defect Detection
- A Self-Evolving Framework for Multi-Agent Medical Consultation Based on Large Language Models
- A Self-supervised UAV Detection Method Based on Channel State Information
- A Singing Melody Extraction Network Via Self-Distillation and Multi-Level Supervision
- A Small-footprint Acoustic Echo Cancellation Solution for Mobile Full-Duplex Speech Interactions
- A Spherical-Harmonic Domain Selective Spatial Active Noise Control System Based on Sound Field Reproduction
- A Structured Neural Network Approach for Learning Improved Iterative Algorithms for SBL
- A Study of Improving The Privacy-Utility Trade-off of Task-specific Models with Learnable Privacy
- A Study of Multi-Scale Feature Learning From Pre-Trained Models on Speaker Verification
- A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
- A Systematic Evaluation of Machine Learning Methods for Fault Detection and Line Identification in Electrical Power Grids
- A Task-Oriented Real-Time and Robust Feature Compression and Selection Method in Collaborative Intelligence System
- A Teacher Action Quality Assessment Method Based on Label Constraint Strategy
- A Training-Free Correlation-Weighted Model for Zero-/Few-Shot Industrial Anomaly Detection with Retrieval Augmentation
- A Transmitter-Model Unaware Generative Image Compression Framework for Semantic Communication
- A Triangular Stable Node Network based on Self-supervised Learning for personalized prediction
- A Two-Stage AIGC Image Quality Assessment with T2I Correspondence and Visual Perception
- A Two-timescale Primal-dual Algorithm for Decentralized Optimization with Compression
- A Unified Hardware Accelerator for Fast Fourier Transform and Number Theoretic Transform
- A Unified Joint Contrastive Triplet Loss with Temporal and Frequency Signal Fusion for Diagnosing Heart Murmurs
- A Unified Metric for Simultaneous Evaluation of Error Rate and Annotation Cost
- A Unified Model for Oral Reading Fluency and Student Prosody
- A Unified Spatiotemporal Frequency Graph Neural Network for fMRI-based Brain Functional Connectivity Analysis
- A Weakly Supervised Semantic Segmentation Model with Enhanced CLIP Feature Extraction
- A Weighted Cross-entropy Loss for Mitigating LLM Hallucinations in Cross-lingual Continual Pretraining
- A Zero-Shot Physics-Informed Dictionary Learning Approach for Sound Field Reconstruction
- A decade of DCASE: Achievements, practices, evaluations and future challenges
- A first-order DirAC-based parametric Ambisonic coder for immersive communications
- A novel multimodal personality prediction method based on pretrained models and graph relational transformer network
- A spectrum-enhanced attention model for semantic segmentation of remote sensing images
- A-PeARCNN: a Physics-encoded AutoRegressive Convolutional Neural Network with AttentionNet for Solving Partial Differential Equations
- A2B: Neural Rendering of Ambisonic Recordings to Binaural
- A2GP-SF: Enhancing Few-shot Class Incremental Learning via Attribute Generative Prompting and Adaptive Sharpness Flattening
- AAD-DCE: An Aggregated Multimodal Attention Mechanism for Early and Late Dynamic Contrast Enhanced Prostate MRI Synthesis
- AC-Mix: Self-Supervised Adaptation for Low-Resource Automatic Speech Recognition using Agnostic Contrastive Mixup
- ACRL-10K: A Dataset for Air Conditioner Refrigerant Leak Smoke Detection
- AD2T: Adversarial Distortion Domain Translation for Robust Watermarking against Non-differentiable Distortions
- ADC-GS: Pose-Free 3D Gaussian Splatting with Adaptive Depth Consistency
- ADC: Enhancing Function Calling Via Adversarial Datasets and Code Line-Level Feedback
- ADD: A Detection Method for Image-Processing Adversarial Defenses
- AER-LLM: Ambiguity-aware Emotion Recognition Leveraging Large Language Models
- AGIAA-2K: A Fine-grained Dataset for Aesthetic and Alignment Evaluation of AI-Generated Images
- AGR: Age Group fairness Reward for Bias Mitigation in LLMs
- AI-Generated Music Detection and its Challenges
- AIDC: Benchmark for Analytical Learning in Incremental Disease Classification
- AKI360: Enabling Highly Interactive 360-degree Video Streaming by Adaptive Keyframe Interval
- ALIC: Adaptive Fusion Entropy Model for Learned Image Compression
- AMNS: Attention-Weighted Selective Mask and Noise Label Suppression for Text-to-Image Person Retrieval
- AMSER: Accelerate Mobile Speech Emotion Recognition with Signal Compression
- AMuSE: Attentive Multilingual Speech Encoding for Zero-Prior ASR
- ANASETC: Automatic Neural Architecture Search for Encrypted Traffic Classification
- AP-Net: Semi-Supervised Ultrasound Cardiac Segmentation Using Enhanced Anatomical Prior
- APLASE: Compression using Adaptive Piecewise Linear Approximation and Sparse Encoding
- APTSniffer: Detecting APT Attack Traffic Using Retrieval-Augmented Large Language Models
- ARIG-GCN: Anatomical Relationship and Isomorphic Graph Approximation Guided Graph Convolutional Network for Automated ASPECTS Scoring on Non-Contrast CT
- ARM : nnU-Net with Arena Mechanism for Medical Image Segmentation
- AS-Net: Adaptive Style-aware Network for Handwritten Text Generation
- ASANet: Scene Text Recognition With Alternate Self-Attention
- ASCDomain: Domain Invariant Device-Adversarial Isotropic Knowledge Distillation Convolutional Neural Architecture
- ASFC-NeRF: Large-Scale Scene Rendering with Adaptive Sampling and Feature-aware Compression
- ASR Benchmarking: Need for a More Representative Conversational Dataset
- ATGnet: Adaptive Temporal Graph Network for EEG-enabled Sound Source Tracking in Cocktail Party Scenarios
- ATP-TTS: Adaptive Thresholding Pseudo-Labeling for Low-Resource Multi-Speaker Text-to-Speech
- AUIED3K: A New Andaman Underwater Image Enhancement Dataset for Deep Learning-Driven Image Enhancement with Minimum Loss Dehazing
- AVS3P10 Standard for Real-time Speech Coding
- Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
- Accelerating Computation for Large-Scale Wide-Band RF Imaging
- Accelerating Convergence in Bounding Box Regression with a Refined IoU Loss Function
- Accelerometer-Based Person-in-Bed Detection Challenge
- AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
- Accompaniment Prompt Adherence: A measure for evaluating music accompaniment systems
- Accurate 3D Facial Paralysis Analysis Using Multi-View Infrared Structured Light System
- Accurate Hardware Trojan Detection for SGIN Device: A Prompt-Tuning and LangChain Approach
- AceParse: A Comprehensive Dataset with Diverse Structured Texts for Academic Literature Parsing
- Achieving Robustness in Blind Modulo Analog-to-Digital Conversion
- Acoustic Identification of Individual Animals with Hierarchical Contrastive Learning
- Acoustic Position Estimation of a Silent Listener
- Active Learning for Long-Tailed Annotation
- Active Listener: Continuous Generation of Listener's Head Motion Response in Dyadic Interactions
- Active Visual Learning for Robots with Dueling Deep Q-Networks and Transformer Encoders
- AdaBoost-Based Channel Estimation in One-Bit Millimeter-Wave MIMO
- AdaCS: Adaptive Normalization for Enhanced Code-Switching ASR
- AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
- AdapFed: Adaptive Devices Training Strategy for Heterogeneous Federated Learning
- Adapt and Feature Translation for Class-Incremental Learning with Pre-Trained Models
- AdaptVC: High Quality Voice Conversion with Adaptive Learning
- Adapter-Based Multi-Agent AVSR Extension for Pre-Trained ASR Models
- Adapting Large Language Model for Spatio-Temporal Understanding in Next Point-of-Interest Prediction
- Adapting Large Language Models to Forecast in Frequency Domain
- Adapting Single-Channel Pre-trained Transformer Models for Multi-Channel Sound Event Localization and Detection
- Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
- Adapting Without Seeing: Text-Aided Domain Adaptation for Adapting CLIP-like Models to Novel Domains
- Adaptive Acquisition in Bayesian Optimization with Agnostic Ensembles
- Adaptive Aspect Ratios with Patch-Mixup-ViT-based Vehicle ReID
- Adaptive Canonical Correlation Analysis With Application to Time Synchronization for Signal Alignment
- Adaptive Central Frequencies Locally Competitive Algorithm for Speech
- Adaptive Compression of Supervised and Self-Supervised Models for Green Speech Recognition
- Adaptive Contribution Modulation For Multi-Modal Manipulation Media Detection and Grounding
- Adaptive Decoding for Efficient Automatic Speech Recognition
- Adaptive Feature Aggregation for In-Air Handwritten Trajectory
- Adaptive Fine-Grained Feature Mining and RoI Feature Interaction Network for Small Object Detection in Aerial Images
- Adaptive Gradient-Based Timesurface for Event-based Detection
- Adaptive Large Language Models via Attention Shortcuts
- Adaptive Layered-Trust Robust Defense Mechanism for Personalized Federated Learning
- Adaptive Lossless Compression for Genomics Data by Multiple (s, k)-mer Encoding and XLSTM
- Adaptive Multi-Scale Local Correction for Semi-Supervised 3D Medical Image Segmentation
- Adaptive Password Guessing Framework Using Various Datasets
- Adaptive Prototype Learning for Anomalous Sound Detection with Partially Known Attributes
- Adaptive Receptive Field Convolution for Top-view Fisheye Images Segmentation
- Adaptive Skeleton Prompt Tuning for Cross-Dataset 3D Human Pose Estimation
- Adaptive Sparse Feature Location Activation Strategy for Sparse Detectors on Drone Images
- Adaptive Spatiotemporal Augmentation for Improving Dynamic Graph Learning
- Adaptive Time-Frequency Attention Network for Sleep Stage Classification Using Respiratory Signals
- Adaptive-Similarity-Based Brain Dynamic Functional Connectivity with Spatial-Temporal Attention and Domain Adaptation for Schizophrenia Diagnosis
- AdaptiveDrop: A Simple Adaptive Label Noise Filtering Scheme for Enhanced Self-supervised Speaker Verification
- Addressing Emotion Ambiguity and Annotator Subjectivity for Enhanced Speech Emotion Labeling
- Addressing Pilot Contamination in Channel Estimation with Variational Autoencoders
- Addressing Speed-Induced Dispersion in Stepped-Frequency PMCW Radar Systems
- Adopting Whisper for Confidence Estimation
- Advanced Graph-MLPs Distillation based on Global and Local Hyperbolic Geometry Learning
- Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
- Advances in Microphone Array Processing and Multichannel Speech Enhancement
- Advancing Active Speaker Detection for Egocentric Videos
- Advancing Dark Action Recognition via Modality Fusion and Dark-to-Light Diffusion Model
- Advancing Few-Shot Class-Incremental Learning with Virtual Prototype Guidance Prompting
- Advancing High-Resolution and Efficient Automotive Radar Imaging through Domain-Informed 1D Deep Learning
- Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset
- Advancing Non-intrusive Suppression on Enhancement Distortion for Noise Robust ASR
- Advancing Paired Image-Mask Synthesis for Automated Nanoparticle Phenotyping
- Advancing SAR Image Robustness: Integrating Diffusion Models for Adversarial Purification and Speckle Noise Suppression
- Advancing Single-Snapshot DOA Estimation with Siamese Neural Networks for Sparse Linear Arrays
- Advancing Streaming ASR with Chunk-wise Attention and Trans-chunk Selective State Spaces
- Adversarial Feature Disentanglement Framework for Voice Pathology Detection
- Adversarial Knowledge Transfer for Black-Box Model Inversion Attack
- Adversarial Learning For End-To-End Cochlear Speech Denoising Using Lightweight Deep Learning Models
- Adversarial Speech-Text Pre-Training for Speech Translation
- Adversarial Training and Cross-modal Feature Fusion in Multimodal Sentiment Analysis
- Adversarial Training and Gradient Optimization for Partially Deepfake Audio Localization
- Aesthetic Perception Prompting for Interpretable Image Aesthetics Assessment with MLLMs
- Age of Gossip with the Push-Pull Protocol
- AgentPose: Progressive Distribution Alignment via Feature Agent for Human Pose Distillation
- Agentic Copyright Watermarking against Adversarial Evidence Forgery with Purification-Agnostic Curriculum Proxy Learning
- Algorithm Design for Continual Learning in IoT Networks
- Aligned Contrastive Learning for Text-to-Music Retrieval
- Aligning Noisy-Clean Speech Pairs at Feature and Embedding Levels for Learning Noise-Invariant Speaker Representations
- Aligning Text-to-Image Diffusion Models without Human Feedback
- Alignment-Free Training for Transducer-based Multi-Talker ASR
- Ambisonics Binaural Rendering via Masked Magnitude Least Squares
- Ambisonics Coding in IVAS: A Hybrid SPAR and DirAC System
- Amplitude-Guidance Low-Light Image Enhancement with Frequency-based Channel Attention
- An Abnormal Audio Generation Method for Fault Diagnosis of Power Transformers
- An Adaptive Framework for Multi-View Clustering Leveraging Conditional Entropy Optimization
- An Adversarial Perturbation Generation Method for Image Anti-Forensics Based on Dual-Path Spatial Attention GAN
- An Attentive Dual-Encoder Framework Leveraging Multimodal Visual and Semantic Information for Automatic OSAHS Diagnosis
- An Attribute-Enriched Dataset and Auto-Annotated Pipeline for Open Detection
- An Automatic Extrinsic Calibration Method for LiDAR-Camera Fusion via Combining Semantic and Geometric Features
- An Efficient Hybrid Quantum Variational Classifier With Matrix Product State
- An Efficient Pore Annotation Framework for Tight Sandstone Images with Segment Anything Model
- An Efficient Residual-based Low-dose PET Reconstruction with Spatial-Frequency Integration
- An Efficient Sample Utilization Method for Deep Learning Based on Class Uncertainty
- An Efficient and Streaming Audio Visual Active Speaker Detection System
- An End-to-End Graph-Guided Spatiotemporal Model for Adaptive Frame-Level Facial Affect Analysis in the Wild
- An Ensemble Approach to Short-form Video Quality Assessment Using Multimodal LLM
- An Exceptional Dataset For Rare Pancreatic Tumor Segmentation
- An Experimental Study on Joint Modeling for Sound Event Localization and Detection with Source Distance Estimation
- An Explainable Probabilistic Attribute Embedding Approach for Spoofed Speech Characterization
- An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement
- An Improved Planar Approximation Localization Method in Distributed Airborne Radars
- An Information-Theoretic Analysis of Thompson Sampling with Infinite Action Spaces
- An Interactive Evaluation Framework for Empathetic Response Generation
- An Intra- and Cross-frame Topological Consistency Scheme for Semi-supervised Atherosclerotic Coronary Plaque Segmentation
- An LSTM Feature Imitation Network for Hand Movement Recognition from sEMG Signals
- An Optimized GPU-based Acceleration of CRYSTALS-Dilithium
- An Underwater Image Quality Dataset with Renewed Pairwise Voting
- AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder
- Analysis and Calibration of Nonlinear Power Amplifiers in Wideband OFDM-Based LEO Satellite Communication System
- Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization
- Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient Tuning
- Anchor-Prompt-based Segmentation and Embedding Model
- Anchored Monotonic Alignment and Representation Substitution for Rare Spontaneous Behaviors in Spontaneous Speech Synthesis
- Anima2: Cross-Species Animal Animation through Image-to-Video Synthesis with Subject Alignment
- AnimateSketches: Animate Sketches with Instance-Aware Mask
- Animation Anycolor: Enhancing Line Drawing Colorization with Keypoint Matching
- Annealing Distillation Algorithm for Transferring Unsupervised Clustering Knowledge to Supervised Student Models
- Apollo: Band-sequence Modeling for High-Quality Audio Restoration
- Appearance- and Orientation-aware Fine-grained Rotated Ship Detection in High-Resolution Satellite Imagery
- Appearance-adapter: A Self-supervised Pose-guided Human Image Synthesis Approach
- Approximation and Analysis of the One-Bit Hermite Law
- Archetypal Analysis for Binary Data
- ArtTwin: A Novel Concept of Developing Digital Twin of Human Arterial System
- Artistic Image Aesthetics Assessment Assisted by Photographic Visual Attributes
- Assessing Robustness of Multi-Modal Large Language Models in Image Classification through Hierarchical WordNet-Based Evaluation
- Asymptotic Behavior Analysis of Antenna Selection via Sparsity-Induced Precoder
- Atom-Constrained Maximum Likelihood Gridless DOA with Wirtinger Gradients
- Attacking Voice Anonymization Systems with Augmented Feature and Speaker Identity Difference
- Attention Augmented Structure-centric Bias Mitigation with Feature Disentanglement
- Attention Disentanglement for Semantic Diffusion Modeling in Text-to-Image Generation
- Attention Weighting and Conditional Entropy-driven Quantization Loss for Neural Audio Codecs
- Attention-Based Beamformer For Multi-Channel Speech Enhancement
- Attention-Driven Causal Discovery: From Transformer Matrices to Granger Causal Graphs for Non-Stationary Time-series Data
- Attention-Enhanced Feature Fusion Network for No-Reference Image Quality Assessment
- Attention-Enhanced Short-Time Wiener Solution for Acoustic Echo Cancellation
- Attribute Conditional Diffusion-Augmented Person Re-Identification
- Audio Array-Based 3D UAV Trajectory Estimation with LiDAR Pseudo-Labeling
- Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
- Audio Decoding by Inverse Problem Solving
- Audio Diffusion with Large Language Models
- Audio Explanation Synthesis with Generative Foundation Models
- Audio Features Investigation for Singing Voice Deepfake Detection
- Audio Sparse-Transformer for Speech Classification
- Audio Texture Manipulation by Exemplar-Based Analogy
- Audio-Driven Reinforcement Learning for Head-Orientation in Naturalistic Environments
- Audio-Faces Intra-Frame Alignment with Graph Attention Networks for Active Speaker Detection
- Audio-Visual Deepfake Detection With Local Temporal Inconsistencies
- Audio-Visual Representation Learning For Lip-Sync Estimation Through Ranking Augmented Contrastive Training
- AudioBERT: Audio Knowledge Augmented Language Model
- AudioCache: Accelerate Audio Generation With Training-Free Layer Caching
- AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
- AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework
- AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
- Audiogram-Informed End-to-End Noise Reduction and Wide Dynamic Range Compression for Hearing Aids
- Audiopedia: Audio QA with Knowledge
- Augmenting Short Enrollment Speech via Synthesis for Target Speaker Extraction
- AuscMLLM: Bridging Classification and Reasoning in Heart Sound Analysis with a Multimodal Large Language Model
- Automated Exposure Mapping for Networked Interference
- Automated Extraction of Spatio-Semantic Graphs for Identifying Cognitive Impairment
- Automated Graph Attention Network for Heterogeneous Entity Resolution
- Automatic Adaption of the Step Size in Gradient Descent Training
- Automatic Detection of Domain Shifts in Speech Enhancement Systems Using Confidence-Based Metrics
- Automatic Geometric Quantification and Rupture Risk Evaluation of 3D Intracranial Aneurysms
- Automatic Labelling & Semantic Segmentation with 4D Radar Tensors
- Automatic Numbering and Pathological Recognition of Pediatric Teeth Using CNN and Attention Mechanisms
- Automatic Parkinson's disease detection from speech: Layer selection vs adaptation of foundation models
- Automatic Speech Recognition and Spoken Language Understanding of Maritime Radio Communications: A case study with Singapore data
- Automatic Text Pronunciation Correlation Generation and Application for Contextual Biasing
- Automatic recognition of rodent call types using deep supervectors
- Automotive Radar Target Detection in Widely Separated and Distributed Aperture Radar Systems
- Autoregressive Density Estimation Transformers for Multivariate Time Series Anomaly Detection
- Autoregressive Language Model with Historical Context Re-encoding
- Auxiliary Tasks Benefit Skeleton-based Action Recognition
- Avoiding Domain Drift and Constant Predictions with Diffusion Enhanced Vector-Quantized Autoencoders for Temperature Predictions
- BAD: Bidirectional Auto-Regressive Diffusion for Text-to-Motion Generation
- BANC: Towards Efficient Binaural Audio Neural Codec for Overlapping Speech
- BCG data imputation via multimodal feature alignment and semantic sequence prediction
- BCS-Net: Multi-Task Breast Cancer Screening Network Enhanced by Multi-Modality Attention
- BDCKD: Unlocking the Power of Brownian Distance Covariance in Knowledge Distillation
- BDGAN: Boundary and Diversity-aware Generative Adversarial Network for Imbalanced Medical Image Augmentation
- BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection
- BIAWDiff: Enhancing Low-Light Images with Bio-Inspired Attention and Wavelet Diffusion
- BID-Net: Balanced Incremental Distillation Network for Fair Dermatological Disease Diagnosis
- BIF: A Biosignature Identification Framework for Model-agnostic Interpretation of MVI Diagnosis Models in HCC
- BIGFR: Bridging Individual and Group Fairness in Recommendation Systems
- BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR
- BP-GPT: Auditory Neural Decoding Using fMRI-prompted LLM
- BRDIA: Bidirectional Reasoning with Dynamic Instruction Adjustment for Multi-hop KGQA
- BS-Breath: Respiration Sensing with Cell-free Massive MIMO
- BadRefSR: Backdoor Attacks Against Reference-based Image Super Resolution
- Band Prompting Aided SAR and Multi-Spectral Data Fusion Framework for Local Climate Zone Classification
- Basis Function Learning for Variable-Length and Continuous-Indexed Signals
- Basket-Enhanced Heterogenous Hypergraph for Price-Sensitive Next Basket Recommendation
- Bayesian Filtering on Graphs
- Bayesian Nonparametric Clustering for Source Counting with a Small Aperture Microphone Array
- BeatKAN: An Efficient and Drum-Attuned Beat Tracking Method Using Kolmogorov-Arnold Networks
- Benchmarking Music Generation Models and Metrics via Human Preference Studies
- Bernoulli-Gaussian Scale Mixture Model and BP Method for Multi-Snapshot Sparse Signal Recovery
- Better Exploiting Spatial Separability in Multichannel Speech Enhancement with an Align-and-Filter Network
- Beyond Jensen's Inequality: Speeding Up ML Estimation of Generalized Hyperbolic Distributions
- Beyond Point Annotation: A Weakly Supervised Network Guided by Multi-Level Labels Generated from Four-Point Annotation for Thyroid Nodule Segmentation in Ultrasound Image
- Beyond Speaker Identity: Text Guided Target Speech Extraction
- Beyond Uniformity: Deblurring Images With Complex Noise Patterns Using Half Quadratic Splitting
- Bi-attention pyramid network for small defect with complex background in industrial detection
- BiCG: Binaural Cue Generation from Unified HRTF Datasets
- BiMA: Bidimensional multi-level attention embedded network for single-frame infrared small target detection
- Bidirectional Reference Image Quality Assessment via Content-Quality Correlation Modeling
- Big-Moe: Bypassing Isolated Gating For Generalized Multimodal Face Anti-Spoofing
- Bilevel Learning for Low-Light Image Enhancement and Detection
- Bilingual Dual-Head Deep Model for Parkinson's Disease Detection from Speech
- Binary Representation Learning for Discriminative Acoustic Unit Discovery
- Binary Stochastic Flip Optimization for Training Binary Neural Networks
- Biodenoising: Animal Vocalization Denoising without Access to Clean Data
- Birds of a Feather: Learning to Retrieve Dance Poses From Music Via Ground-Truth Annotation Lifting
- Black-Box Adversarial Defense Against Voice Conversion Using Latent Space Perturbation
- Blind Estimation of Sub-band Acoustic Parameters from Ambisonics Recordings using Spectro-Spatial Covariance Features
- Blind Spatial Impulse Response Generation from Separate Room- and Scene-Specific Information
- BloomCoreset: Fast Coreset Sampling using Bloom Filters for Fine-Grained Self-Supervised Learning
- BlurPaint: Image Inpainting using Blurring Diffusion Models
- Boli: A dataset for understanding stuttering experience and analyzing stuttered speech
- Bone Conducted Signal Guided Speech Enhancement For Voice Assistant on Earbuds
- Boolean Matrix Tri-Factorization
- Boolean matrix compressed sensing
- Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM
- Boosting Jailbreak Attack with Momentum
- Boosting Large Language Model for Speech Synthesis: An Empirical Study
- Boosting Lightweight Camouflaged Object Detection with Multi-Scale Context and Boundary Awareness
- Boosting Movie and TV Tag Accuracy with Knowledge Graphs
- Boosting Open-Vocabulary Object Detection Performance via Class-Agnostic Pseudo-Labels and MultiModal Hybrid Knowledge
- Boosting Stereo Image Noise Removal by Learning Uncertainty and Enriched Features
- Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models
- Boosting the Transferability of Adversarial Examples via Local Mixup and Adaptive Step Size
- Bootstrapping LLM-based Fact-checking via Iterative Rationalization Finetuning
- Bootstrapping Language-Audio Pre-training for Music Captioning
- Bottleneck-Constrained Contrastive Decoupled Network for Multimodal Aspect-based Sentiment Classification
- Boundary-Driven Table-Filling with Cross-Granularity Contrastive Learning for Aspect Sentiment Triplet Extraction
- Brain MRI Segmentation with Language-Driven Detection and Context-Aware Descriptions
- BrainChat: Interactive Semantic Information Decoding from fMRI Using Large-Scale Vision-Language Pretrained Models
- BrainVis: Exploring the Bridge between Brain and Visual Signals via Image Reconstruction
- Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition
- Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising
- Bridge-SR: Schrödinger Bridge for Efficient SR
- Bridging Modality Gap with Large Speech and Language Models for End-to-End Speech-to-Text Translation
- Bridging Neural and Symbolic Reasoning: A Dual-System Framework for Interpretable Question Answering
- Bridging Speech and Text Foundation Models with ReShape Attention
- Bridging Task Boundaries: Remote Sensing Image-Text Retrieval via Dictionary-Driven Adaptation
- Bridging the Fairness Gap: Enhancing Pre-trained Models with LLM-Generated Sentences
- Bridging the Modality Gap for Speech-image Retrieval with Text Supervision
- Build LLM-Based Zero-Shot Streaming TTS System with Cosyvoice
- C2AD: Dual Consistency Learning for Zero-Shot Anomaly Detection
- C3D-VIT: Consistency-Aware 3D Vision Transformer for Face Forgery Detection
- CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification
- CA-UAP: Content-Agnostic Universal Adversarial Perturbation for Enhanced Generalization
- CAAL-Unet: a Confusion Area Attention Lightweight Network for Brain Tumor Segmentation
- CAF-YOLO: A Robust Framework for Multi-Scale Lesion Detection in Biomedical Imagery
- CAMDet: Condition-Adaptive Multispectral Object Detection Using a Visible-Thermal Translation Model
- CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
- CAPAST: Content Affinity Preserved Arbitrary Style Transfer
- CASC-XVC: Zero-Shot Cross-Lingual Voice Conversion with Content Accordant and Speaker Contrastive Losses
- CASleepNet: A Cross Attention-based multimodal fusion approach for sleep staging with EEG and EOG
- CAT-Net: A Co-Adaptive Transfer Learning Network for BCI-Assisted Neurorehabilitation
- CAW-CL: Cascaded Adaptive Weighted Contrastive Learning for Unsupervised Ultrasound Plane-Wave Image Reconstruction
- CE-FFT: Communication-Efficient Federated Fine-Tuning for Large Language Models via Quantization and In-Context Learning
- CEMSSL: Conditional Embodied Self-Supervised Learning is All You Need for High-precision Multi-solution Inverse Kinematics of Robot Arms
- CFSum: A Transformer-Based Multi-Modal Video Summarization Framework With Coarse-Fine Fusion
- CGDD: Contrastive Gaussian-Dirac Diffusion Model
- CGEDN: Approximation of Graph Edit Distance with Path Generation via Learning Node Matching
- CGNet: Classification-Guided Multi-Task Interactive Network for Hyperspectral and Multispectral Image Fusion
- CHASE: Channel-Wise and Spatial Attention for Early Exiting in Image Classification
- CIEGCL: Counterfactual Intervention Enhancing Graph Contrastive Learning in Implicit Feedback
- CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
- CLAP-S: Support Set Based Adaptation for Downstream Fiber-optic Acoustic Recognition
- CLHi-MTS: A Contrastive Learning-Based Hierarchical Framework for Masked Medical Time-Series Modeling
- CLIPGaze: Zero-Shot Goal-Directed Scanpath Prediction Using CLIP
- CMFNThinker: A Novel Cross-source Multi-modal Fake News Detection Model
- CMGait: Enhancing Cross-Modality Gait Recognition between LiDAR and RGB through Contrastive Identity-consistent Feature Aggregation
- CMoS: Customizing Model Structures for Personalized Federated Learning
- COAST: Contrastive Learning with Augmented Spatio-Temporal Encoding for Next POI Recommendation
- COCO-OLAC: A Benchmark for Occluded Panoptic Segmentation and Image Understanding
- COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations
- COREMIL: Contextual Position Encoding-based Retrievable Multiple Instance Learning for Slide-level Classification
- COSMIC waveforms for Integrated Communication and Imaging
- CPA-Enhancer: Chain-of-Thought Prompted Adaptive Enhancer for Downstream Vision Tasks Under Unknown Degradations
- CPL: Curriculum Pseudo Labeling for Weakly Supervised Temporal Forgery Localization
- CPSNet: Comprehensive Enhancement Representation for Polyp Segmentation Task
- CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
- CR-CLIP: Image-Text Contrastive Regression for Generalized Gaze Estimation
- CSD: Weather forecasting with graph neural network based on cross-scale diffusivity
- CSMT: Combining Snoring and Metadata-based Text for Sleep Apnea Severity Classification
- CSS: Overcoming Pose and Scene Challenges in Crowd-Sourced 3D Gaussian Splatting
- CT Image Prediction Of PD-1 Gastric Cancer Patients Based On The PLSG Framework
- CTGDiff: A Conditional Diffusion Model for Cardiotocography Signal Synthesis
- CabiNet: A Deep Learning Framework for Multiclass Medical Image Segmentation from Multiple Single Class Datasets
- Calibration of Multiple Asynchronous Microphone Arrays using Hybrid TDOA
- Camouflaged Object Detection via Neural Architecture Search
- Camouflaged Object Detection with CNN-Transformer Harmonization and Calibration
- Can AI See What We Can't? Leveraging Deep Learning and Multi-Temporal Satellite Data to Revolutionize Crop Type Mapping and Yield Prediction
- Can Automated Speech Recognition Errors Provide Valuable Clues for Alzheimer's Disease Detection?
- Can Fairness and Robustness Be Simultaneously Achieved Under Byzantine Attacks?
- Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
- Can Large Language Models Grasp Event Signals? Exploring Pure Zero-Shot Event-based Recognition
- Can Quality Survive Scale? Toward an Equal-Quality Instance-Dependent Label Noise Model
- Can RAG-Driven Enhancements Amplify Audio LLMs for Low-Resource Languages?
- Can We "Cherry-Pick"? Investigating Multiple Renditions from a Generative Speech Synthesis Model
- CapT: A Hierarchical Capsule Representation Learning Approach for Class Continual Learning
- Capturing Rich Behavior Representations: A Dynamic Action Semantic-Aware Graph Transformer for Video Captioning
- CardioFlow: Learning to Generate ECG from PPG with Rectified Flow
- CardioRiskNet: Attention-based CVAE-enabled GCN for Risk Prediction in STEMI
- Carver: Learning to Reconstruct Right Ventricle from Sparse Multi-View 2D Echocardiograms
- CascadePAIE: Reallocating Relevance for Event Roles and Event Text in Event Argument Extraction
- Catch Causal Signals from Edges for Label Imbalance in Graph Classification
- Cauchy-Schwarz Divergence Transfer Entropy
- Causal Debiasing for Visual Commonsense Reasoning
- Causal Feature Supervision Decoupling: A Novel Method for Clothes-Changing Person Re-identification Algorithm
- Causal Speech Enhancement Based on a Two-Branch Nested U-Net Architecture Using Self-Supervised Speech Embeddings
- Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
- Causal fMRI-Mamba: Causal State Space Model for Neural Decoding and Brain Task States Recognition
- Causality-Guided Context-Aware Multimodal Public Speaking Anxiety Detection for Out-of-Distribution Generalization
- Certainty-guided Reasoning and Refinement Network for Camouflaged Object Detection
- Chain-of-Thought Prompting for Speech Translation
- Chained Motion Vector Prediction for Video Coding
- ChangeChat: An Interactive Model for Remote Sensing Change Analysis via Multimodal Instruction Tuning
- Channel and space-based joint rate allocation algorithm
- Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition
- ChannelMixer: A Hybrid CNN-Transformer Framework for Enhanced Multivariate Long-Term Time Series Forecasting
- Char-SAM: Turning Segment Anything Model into Scene Text Segmentation Annotator with Character-level Visual Prompts
- Chat-Driven 3D Human Pose and Shape Editing with Large Language Models
- ChatCAD: An MLLM-Guided Framework for Zero-shot CAD Drawing Restoration
- CheapNVS: Real-Time On-Device Narrow-Baseline Novel View Synthesis
- Chinese Speech Processing via Chinese Character Feature
- Chitrarth: Bridging Vision and Language for a Billion People
- ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
- CiGA: A Cross-Layer Fine-Grained Attention Correction Method for Large Language Model
- Class Relevance Learning for Out-of-Distribution Detection
- Class Semantic Prompts Enhanced Prototypical Fusion Method for Few-shot Named Entity Recognition
- Class-Difficulty Aware Hybrid Active Learning
- Class-wise Adaptive Logits Distillation with Meta-Learning
- Classification Error Bound for Low Bayes Error Conditions in Machine Learning
- Classification Inconsistency Alignment Network for Cross-corpus Speech Emotion Recognition
- Classification of Eye-Tracking Data Based on Spatiotemporal Attention Encoding
- Classification of Zhuang Dialect combined with Bert and SimAM
- Classifier-Guided Captioning Across Modalities
- Classifying Music-Induced Emotion Using Multi-Modal Ensembles of EEG and Audio Feature Models
- Climate Downscaling Using Neural Operator: Spatiotemporal Multimodal Fusion Operator with State-Query Coupled Kernel
- ClingTP: Curriculum Learning based Multi-style Title Prefix Generation
- Clinically Robust Polyp Segmentation: Enhanced Generalization and Perturbation Resistance
- Cloth-debiasing with Stable Diffusion in Cloth-changing Person Re-identification
- Cluster-Perceptive Graph Contrastive Learning for Community Detection
- Cluster-Refined Optimal Transport for Unsupervised Action Segmentation
- Clutter Resilient Occlusion Avoidance for Tightly-Coupled Motion-Assisted Detection
- Co-Attention Based Multi-Channel TF-GridNet for Speech Separation with Ad-Hoc Microphone Arrays
- Co-training with Progressive Distribution Alignment and Uncertainty-Interactive Relabeling for Semi-Supervised Domain Adaptive Semantic Segmentation
- CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models
- CoGAP: A Personalized Federated Learning Method Using Collaborative Optimization for Medical Image Classification
- CoLLAP: Contrastive Long-form Language-Audio Pretraining with Musical Temporal Structure Augmentation
- CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generation
- Coarse-to-Fine Text-to-Music Latent Diffusion
- Codar: Complex-valued Neural Network for Crossing-Floor Intrusion Detection via WiFi
- Code Drift: Towards Idempotent Neural Audio Codecs
- Codec-ASV: Exploring Neural Audio Codec For Speaker Representation Learning
- CogniDual Framework: Self-Training Large Language Models within a Dual-System Theoretical Framework for Improving Cognitive Tasks
- Cognitive Decline Detection using DLB Extraction Pipelines
- Cognitive Load Monitoring via Earable Acoustic Sensing
- Cognitive MIMO Radar Beamforming for Target Tracking Using a BCRB-based Criterion
- Cohort-Sensitive Labeling: An Effective Approach for Enhancing ASR Performance
- Col-OLHTR: A Novel Framework for Multimodal Online Handwritten Text Recognition
- Collaborative Association Network for Multi-view Multi-Human Association and Tracking using Constraint Optimization and Object Search
- Collaborative Automotive Radar Sensing via Mixed-Precision Distributed Array Completion
- Collaborative Dual-Branch Spatial-Frequency Enhancement Network for Low-Light Images
- Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
- Collaborative Personalized Federated Learning via Exponential Moving Average Optimization
- Collaborative Semantics-Assisted Large Language Models for Next POI Recommendation
- Collision-less and Balanced Sampling for Language-Queried Audio Source Separation
- Collusion-resistant Black-box Watermarking in Federated Learning through Weight Relevance Analysis
- Colored Point Cloud-based Mesh Registration for Enhancing Inter-frame Coding of Texture Video in V-DMC
- Colorization Network Watermarking in the CIE-Lab Domain
- Combined object-based audio and MASA format for enhanced spatial mobile communication
- Combining Loss-aware Curriculum Learning with Incomplete Graph Neural Networks
- Combining Spatio-Temporal Networks and Graph Attention Architectures for EEG-Based Workload Classification
- Commonality Augmented Disentanglement for Multimodal Crowdfunding Success Prediction
- Communication-efficient Exact Diffusion for Decentralized Learning
- Communication-efficient Verifiable and Oblivious Aggregation with Client Dropouts
- Community-entropy Based Graph Structure Learning for Topology-imbalance
- CompMTL: Layer-Wise Competitive Multi-Task Learning
- Compact Neural TTS Voices for Accessibility
- Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing
- Compgen: Synthesis and Generation of Faces From Edgemaps
- Complementary Graph Learning and Prompt-based Cross-modal Generation for Missing-modality Fake News Detection
- Complementary Learning System Theory-based Active Learning for Audio Classification
- Complete Reconstruction of the Tongue Contour Through Acoustic to Articulatory Inversion Using Real-Time MRI Data
- Completing Sets of Prototype Transfer Functions for Subspace-based Direction of Arrival Estimation of Multiple Speakers
- Complex Coprime Frequency Sum Based Signal Representation for Period Estimation
- Complex Open Information Extraction with Heterogeneous Syntax Forests
- ComplexDec: A Domain-robust High-fidelity Neural Audio Codec with Complex Spectrum Modeling
- Component-wise Self-Correction Network for Human Motion Prediction
- Compositional Audio Representation Learning
- Comprehensive Feature Processing Based on Attention Mechanism for Co-Salient Object Detection
- Comprehensive Perturbation Consistency for Semi-Supervised Change Detection in Remote Sensing Images
- Compressing a Flow-Based Privacy Protection Model via a Novel Joint Distilling and Pruning Method
- Compressive Imaging Reconstruction via Conditional Diffusion Model With Augmented Measurements
- ConPCO: Preserving Phoneme Characteristics For Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization
- ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps
- ConcealGS: Concealing Invisible Copyright Information in 3D Gaussian Splatting
- Concentrating Harder for Faster Audio Transformer
- Conditional Convolutions for End-to-End Single-Stage Video Text Detection
- Conditional Deep Canonical Time Warping
- Conditional Latent Diffusion-Based Speech Enhancement via Dual Context Learning
- Conditional-Balanced Adversarial Delta Tuning for Cross-Domain Implicit Discourse Relation Recognition
- Conformal Prediction for Manifold-based Source Localization with Gaussian Processes
- Confusion-Aware Prototypical Contrastive Learning for Open-Vocabulary Object Detection
- Consensus Graph Filter Learning for Multiple Graph Clustering
- Consensus Graph-Based Spectral Ensemble Clustering via Low-Rank Tensor Learning
- Conservative Offline Meta-Reinforcement Learning with Task Similarity Measurement
- Constraint-Awareness and Graph Reasoning for Temporal Question Answering
- Constructing Datasets From Public Police Body Camera Footage
- Contactless Nighttime Stress Monitoring with mmWave Radar
- Contactless Vital Sign Monitoring for Multiple People Using a Millimeter-wave MIMO Radar
- Content and Salient Semantics Collaboration for Cloth-Changing Person Re-Identification
- Content-Aware Dynamic Superpixel Segmentation
- Context-Aware Multi-Scale Polyp Segmentation Network
- Context-Guided Active Domain Adaptation for Blended Target Domain
- Contextual ASR with Retrieval Augmented Large Language Model
- Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction
- Contextual Value Alignment
- Contextualization of ASR with LLM using phonetic retrieval-based augmentation
- Continual Self-supervised Learning Considering Medical Domain Knowledge in Chest CT Images
- Continual Unsupervised Domain Adaptation for Audio Deepfake Detection
- Continuous-Discrete Differentiable Particle Filters for Irregular Time Series
- Continuously Learning New Words in Automatic Speech Recognition
- Continuously Learning Video-level Object Tokens for Robust UAV tracking
- Contrast Memory for Unsupervised Anomaly Detection
- Contrast-Unity for Partially-Supervised Temporal Sentence Grounding
- Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
- Contrastive Learning via Randomly Generated Deep Supervision
- Contrastive Lyrics Alignment with a Timestamp-Informed Loss
- Contrastive Pre-Training and Post-Tuning for Heterogeneous Graph Learning
- ControlMol: Adding Substructure Control To Molecule Diffusion Models
- Controllable Forgetting Mechanism for Few-Shot Class-Incremental Learning
- Controllable Generative Model for Brain Evolution
- Controlling the Number of Sample-Contributive Vertices in Generalized Sampling of Graph Signals
- Convergence Analysis of alpha-SVRG under Strong Convexity
- ConvexECG: Lightweight and Explainable Neural Networks for Personalized, Continuous Cardiac Monitoring
- Convolutional Retentive Network for EEG Decoding
- Convolutional Sparse Coding with Multipath Orthogonal Matching Pursuit
- Cooperative ISAC for Localization and Velocity Estimation Using OFDM Waveforms in Cell-Free MIMO Systems
- Cooperative Multi-Target Tracking Based on Multi-Detection TPHD in MIMO-OFDM Systems
- Cooperative Neural Radiance Field for Dynamic Scene Deblurring
- Cooperative and Competitive Functional Connectivity Based on Improved Ising Model
- CorrGAN: Simultaneous Learning of Speech Enhancement and Perceptual Quality Loss Functions
- Correlated Attention in Transformers for Multivariate Time Series
- Correlated Multiple IHC Virtual Staining for Breast Histopathological Images
- Correlative3D: Inter-Object Correlation-Aware 3D Scene Understanding
- Covariance Change Point Detection for Graph Signals
- Covert and Potent: A Weather-Camouflaged Backdoor Attacks on Self-Supervised Learning
- Cramér-Rao Bounds for Wideband Near-Field Sensing
- Credible and Detailed 3D Face Reconstruction in Large Pose
- CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs
- Critically-Damped Third-Order Langevin Dynamics
- CroPrompt: Cross-task Interactive Prompting for Zero-shot Spoken Language Understanding
- Cross-Channel Unlabeled Sensing over a Union of Signal Subspaces
- Cross-Component Residual Prediction for Geometry-Based Point Cloud Compression
- Cross-Domain Few-Shot Open-Set Keyword Spotting Using Keyword Adaptation and Prototype Reprojection
- Cross-Layer Cache Aggregation for Token Reduction in Ultra-Fine-Grained Image Recognition
- Cross-Layer Graph Knowledge Distillation for Image Recognition
- Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
- Cross-Modality Fusion Mamba for All-in-One Extreme Weather-Degraded Image Restoration
- Cross-Talk Detection in the IVAS Stereo Codec Based on GCC-PHAT
- Cross-Template-Based Hypergraph Transformer
- Cross-attention Inspired Selective State Space Models for Target Sound Extraction
- Cross-lingual Evaluation Of Hypernasality Using Wav2Vec2 Features
- Cross-modal Gaussian Localization Distillation for Optical Information guided SAR Object Detection
- CrossHash: Cross-scale Vision Transformer Hashing for Image Retrieval
- CrossSleep: Multi-Scale Attention with Cross-Time Learning for Single Channel EEG-Based Sleep Staging
- Crowdsourced Homophily Ties Based Graph Annotation Via Large Language Model
- Cued Speech Generation Leveraging a Pre-trained Audiovisual Text-to-Speech Model
- CurMIM: Curriculum Masked Image Modeling
- Curriculum Contrastive Learning for Aspect-based Sentiment Analysis
- Curriculum Learning aided Audio-Visual Speech Recognition with Arbitrary Speaker Number
- CycleFlow: Leveraging Cycle Consistency in Flow Matching for Speaker Style Adaptation
- D2-MLP: Dynamic Decomposed MLP Mixer for Medical Image Segmentation
- D2S: Towards Efficient Sparse 3D Object Detection via Dense to Sparse Knowledge Distillation
- D3RM: A Discrete Denoising Diffusion Refinement Model for Piano Transcription
- DA-LIF: Dual Adaptive Leaky Integrate-and-Fire Model for Deep Spiking Neural Networks
- DACAT: Dual-stream Adaptive Clip-aware Time Modeling for Robust Online Surgical Phase Recognition
- DAEF-VS: An Efficient Universal VoIP Steganalysis Framework Based on Domain-Aware Knowledge
- DAREK - Distance Aware Error for Kolmogorov Networks
- DARN: An Attention-Based Neural Network Using Residual Blocks for Sleep Micro-Events Detection
- DARNet: A Dual Attention Residual Network for Medical Image Classification
- DASSL: Domain Agnostic Self-Supervised Learning with Multiple Missing Information Reconstruction Branches
- DATA-VSR: Dynamic Trajectory Attention and Texture Adaptive Rooter for Video Super-Resolution
- DBCR: Exploiting Both Intra-cluster and Extra-cluster Relations for Compositional Reasoning
- DCASI: A Sequence-based Attack Investigation Method Using DTW Contrastive Learning
- DCCMamba: A Dual-stream Cross-time and Cross-feature with Mamba for Multivariate Time Series Forecasting
- DCCT-Net: A Network Combined Dynamic CNN and Transformer for Image Compressive Sensing
- DCD-MUSIC: Deep-Learning-Aided Cascaded Differentiable MUSIC Algorithm for Near-Field Localization of Multiple Sources
- DCFormer: Divide-and-Conquer in 3D Human Pose Estimation Tasks
- DCIM-AVSR: Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module
- DDA: Distillation-Driven Acceleration of the Reverse Diffusion Process for Stochastic Multi-Ship Trajectory Prediction
- DDNet: Deformable Convolution and Dense FPN for Surface Defect Detection in Recycled Books
- DDNet: Exploring Dual Dependencies for Long-Term Time Series Forecasting
- DDSP Guitar Amp: Interpretable Guitar Amplifier Modeling
- DEBT: Enhancing Entity Alignment in Knowledge Graphs through Description Enrichment and Bootstrap Training
- DEFormer: DCT-driven Enhancement Transformer for Low-light Image and Dark Vision
- DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face Synthesis
- DEP-SLAM: A Dynamic Environment Perception SLAM System with Large Language Models
- DETCP: Self-Detoxifying Language Models With Contrastive Pairs
- DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
- DFMA: Adaptive Dual Fusion for Multimodal Relation Extraction with Mutual Attention
- DFNeRF: Disentangled Facial Neural Radiance Fields for Text-based Editing of Free-view Talking Head
- DFT-Spread-Based OTFS Waveform Design With Good Peak-to-Average Power Ratio for Joint Sensing and Communications
- DFingerNet: Noise-Adaptive Speech Enhancement for Hearing Aids
- DGJA: Dependency Graph-enhanced Joint Attention Structure for Multimodal Sarcasm Detection
- DH-VTON: Deep Text-Driven Virtual Try-On via Hybrid Attention Learning
- DICS: Find Domain-Invariant and Class-Specific Features for Out-of-Distribution Generalization
- DKD2L: Dual Knowledge Distillation Dynamic Learning for sketch-based 3D shape retrieval
- DLM-VMTL: A Double LayerMapper For Heterogeneous Data Video Multi-Task Prompt Learning
- DMIBot: Dynamic Multimodal Interaction for Twitter Bot Detection
- DMKPN: Image Deblurring Under Multi-Factor Aliasing Diffusion Degradation
- DN-DR: Discriminative Network with Dual Reconstruction for Image Anomaly Detection
- DOA Estimation Based on Enhanced SRP-MVDR Using Kronecker Product Decomposition for Large Rectangular Microphone Arrays
- DOA Estimation of Coherent Sources Using Residual Network-based Subspace Reconstruction
- DOSE: Drum One-Shot Extraction from Music Mixture
- DPC: Large Model Alignment Method based on Decoding Probability Correction
- DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
- DPM-LVSN: A Diffusion Probabilistic Model-based Left Ventricular Segmentation Network
- DRANet: Dual-threshold Guided Reliability Aware Network for Semi-Supervised Image Semantic Segmentation
- DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
- DRDM: A Disentangled Representations Diffusion Model for Synthesizing Realistic Person Images
- DRSFANet: Dual-Path CNN with Residual and Frequency Attention for Image Denoising
- DS-BTIAN: A Novel Deep-Shallow Bidirectional Transformer Interactive Attention Network for Multimodal Emotion Recognition
- DSDIR: A Two-Stage Method for Addressing Noisy Long-Tailed Problems in Malicious Traffic Detection
- DSDN-Net: An Effective Network for Semantic Segmentation in Open-Pit Coal Mining Areas for Land Cover Recognition
- DSFormer: Deformable Pointformer for 3D Salient Object Detection
- DSINet: Towards Real-Time Target Speaker Extraction with Dynamic Speaker Information Fusion
- DSSM: Dual State Space Model For Human Motions Generation
- DTR: Dynamic Tree-Ring Watermarking Framework for Diffusion-Based Video Generation
- DU-PMVS: Learned Patchmatch Multi-View Stereo Based on Deformable Feature Pyramid and Uncertainty Awareness Modeling
- DULRTC-RME: A Deep Unrolled Low-rank Tensor Completion Network for Radio Map Estimation
- DUNE: Sim2Real Transfer for Depth-based Navigation in Unstructured Dynamic Indoor Environments
- DVM: Towards Controllable LLM Agents in Social Deduction Games
- DX2CT: Diffusion Model for 3D CT Reconstruction from Bi or Mono-planar 2D X-ray(s)
- DapPep: Domain Adaptive Peptide-agnostic Learning for Universal T-cell Receptor-antigen Binding Affinity Prediction
- Dark Experience for Incremental Keyword Spotting
- Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
- Data Glove-based Personalized Continuous Gesture Segmentation
- Data-Aided Regularization of Direct-Estimate Combiner in Distributed MIMO Systems
- Data-Driven Mispronunciation Pattern Discovery for Robust Speech Recognition
- Data-Driven White Noise Gain Constrained Robust Superdirective Beamformer for Speech Enhancement
- Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning
- Data-Free Post-Training Quantization with Block-wise Enhanced Sample Generation
- Data-driven Processing using Parametric Neural Network for Improved Bluetooth Channel Sounding Distance Estimation
- De-confusing Hard Samples for Text Semantic Hashing
- DeBeauty: A Joint Framework for Facial Beautification Removal Based on Spatial Collaborative Adaptation and Hyperplane Relocation
- DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
- Debiased Estimation for Cross-Domain Cold Start Recommendation
- Debiased Prototype Evolving for Point Cloud Domain Adaptation via 3D Foundation Models
- Debiased Training For Semi-supervised Sound Event Detection
- Decentralized Federated Dataset Dictionary Learning for Multi-Source Domain Adaptation
- Decentralized Online Ensembles of Gaussian Processes for Multi-Agent Systems
- Decentralized Stochastic Successive Convex Approximation for composite non-convex problems with non-linear functional constraints
- Decision-Aided Progressive Symbol Phase Equalizer in Sweep Spread Carrier Underwater Acoustic Communications
- Decoding Brain Structure and Gene Expression Interactions in Alzheimer's Disease Pathology
- Decoding the Unintelligible: Neural Speech Tracking in Low Signal-to-Noise Ratios
- Decoupled Feature Matching for Few-shot Counting and Localization
- DecoupledSynth: Enhancing Zero-Shot Text-to-Speech Via Factors Decoupling
- Decoupling While Coupling: Towards More Accurate Stereo Image Sand Removal Beyond Certainty
- Decreasing Word Error Rates in Paragraph Handwritten Text Recognition with Synthetic Data
- Deep Diffusion Gradients Leakage in Federated Learning
- Deep Dynamic Probabilistic Canonical Correlation Analysis
- Deep Enhancement Spotting Network for Low-complexity Keyword Spotting in Noisy Environments
- Deep Feedback Cancellation for Hearing Aids with Improved System Stability and Sound Quality
- Deep Generic Representations for Domain-Generalized Anomalous Sound Detection
- Deep Joint Source-Channel Coding for Wireless Point Cloud Transmission
- Deep Learning Amplified Early Stopping Bias: Overestimating Performance on Small Datasets
- Deep Learning for Modulo Sampling of FRI Signals
- Deep Learning-Based Perceptual Vibrotactile Codec with Rate Scalability
- Deep Metamorphic Registration for Tumor-Affected Medical Image Alignment
- Deep Model Pruning without Finetuning for Few Category Datasets
- Deep Receiver for Multi-Layer Data Transmission with Superimposed Pilots
- Deep Support Vein Machine for Lung Parcellation
- Deep Sylvester Posterior Inference for Adaptive Compressed Sensing in Ultrasound Imaging
- Deep Time Series Anomaly Detection with Local Temporal Pattern Learning
- Deep Transfer Regression for EEG-based Driving Fatigue Detection
- Deep Unfolded Approximate Message Passing for Quantitative Acoustic Microscopy Image Reconstruction
- Deep Unfolding Using Score-based Generative Networks for Automotive Radar Interference Mitigation
- Deep Unfolding of Full Waveform Inversion for Quantitative Ultrasound Imaging
- Deep Variational Sequential Monte Carlo for High-Dimensional Observations
- Deep-Relative-Trust-Based Diffusion for Decentralized Deep Learning
- DeepMatch: Navigating the Complexities of Underwater Textures for Enhanced Keypoint Matching
- DeepPEM-AFC: An Improved Prediction-Error-Method-based Adaptive Feedback Cancellation with Deep Learning for Hearing Aids
- DeepPreNet: A Deep Learning Pre-Processing Method for Speech Distortion Correction in Parametric Array Loudspeaker
- Deepfake Detection of Singing Voices With Whisper Encodings
- Deeply Coupling EEG Signals and Eye Movements for Multi-Modal and Region-Aware Emotion Recognition
- DeformAvatar: Point-Based Human Avatar Re-targeting and Rendering
- Deformable Attention-Based Edge-Aware Network for Single Image Super-Resolution
- Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
- Delving Into Coarse-Fine Feature Interaction Alignment for UAV Object Detection
- Delving into Transformer-based Network Architecture for Guided Depth Super-Resolution
- Denoising Student Features with Diffusion Models for Knowledge Distillation in Speaker Verification
- Denoising and Restoring Channel State Information for 5G Indoor Positioning in Low-SNR Scenarios
- Dense Point Clouds Matter: Dust-GS for Scene Reconstruction from Sparse Viewpoints
- Dense-Sparse Dynamic Time Warping for Customizing Piano Concerto Accompaniments
- Density-Adaptive Fuzzy Clustering with Isolation Kernel
- Density-aware and Depth-aware Visual Representation for Zero-Shot Object Counting
- DepMamba: Progressive Fusion Mamba for Multimodal Depression Detection
- Description-Based Controllable Text-to-Speech With Cross-Lingual Voice Control
- Design and Optimization of Superdirective Beamforming and Post-Filtering for Speech Enhancement
- Design of Multiple Binary Waveforms for the Joint MIMO Radar and Communications
- Design of Robust Differential Beamformers with Microphone Arrays of Arbitrary Planar Geometry
- DetailTTS: Learning Residual Detail Information for Zero-shot Text-to-speech
- Detecting Neurodegenerative Diseases using Frame-Level Handwriting Embeddings
- Detecting OOD Samples via Optimal Transport Scoring Function
- Detecting and Defending Against Adversarial Attacks on Automatic Speech Recognition via Diffusion Models
- Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
- Developing a Multilingual Dataset and Evaluation Metrics for Code-Switching: A Focus on Hong Kong's Polylingual Dynamics
- Device Selection for Resource-Efficient Edge Caching in a Federated Learning Framework
- Device-aware Optical Adversarial Attack for a Portable Projector-camera System
- DiGradPatch: Black-Box Patch Attacks via Diffusion-Based Double Gradient and Sensitive Distribution Guidance
- Diagram Formalization Enhanced Multi-Modal Geometry Problem Solver
- Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
- Diff4Steer: Steerable Diffusion Prior for Generative Music Retrieval with Semantic Guidance
- DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
- DiffAttack: Imperceptible and Transferable Audio Adversarial Attack via Diffusion Model
- DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models
- DiffDesign: A diffusion model using garment Knowledge-Enhanced for Fashion Design Synthesis
- DiffETM: Diffusion Process Enhanced Embedded Topic Model
- DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap
- DiffKillR: Killing and Recreating Diffeomorphisms for Cell Annotation in Dense Microscopy Images
- DiffListener: Discrete Diffusion Model for Listener Generation
- DiffMEL: A large-scale difficulty-graded dataset for Multimodal Entity Linking
- DiffRS: An Extensible Diffusion Model for Remote Sensing Image Generation
- DiffSR: Learning Radar Reflectivity Synthesis via Diffusion Model from Satellite Observations
- DiffSSD: A Diffusion-Based Dataset For Speech Forensics
- Difference Bonds Consistency and Complementarity to Enhance Multimodal Representation Learning
- Differentially Private Distribution Estimation Using Functional Approximation
- Differentially Private and Communication-efficient Decentralized Learning Using Deep Quantizers
- DiffuseFIST: A Fast Image-guided Style Transfer Method for Adapting Large-scale Diffusion Models
- Diffused Poses and Distilled Expressions for Controllable Audio-driven Talking Face Generation
- Diffusion Augmentation Sub-center Modeling for Unsupervised Anomalous Sound Detection with Partially Attribute-Unavailable Conditions
- Diffusion Counterfactual-Based Anomaly Detection in Class-Imbalanced Data
- Diffusion Features to Bridge Domain Gap for Semantic Segmentation
- Diffusion Learning Over Adaptive Competing Networks
- Diffusion Model Based Image Reconstruction in Lensless Imaging
- Diffusion Model with Multi-layer Wavelet Transform for Low-Light Image Enhancement
- Diffusion Models are Good Unsupervised Class-agnostic Shape Part Segmentators
- Diffusion Models are Zero-Shot Generative Text-Vision Retrievers
- Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning
- Diffusion-based Data Augmentation for Object Counting Problems
- Diffusion-based Identity-Preserving Facial Privacy Protection
- Diffusion-based Target Device Style Transfer for Robust Acoustic Scene Classification
- Diffusion-based Unsupervised Audio-visual Speech Enhancement
- Digital Operating Mode Classification of Real-World Amateur Radio Transmissions
- Digital Twin-Driven Bearing-Fault Detection in Induction Motor and Drives using Graph Sampling and Aggregation Network
- Dike: Enhancing Fairness and Efficiency in GPU Clusters for Deep Learning
- Dilated Convolution for Time Series Learning
- Dimensionality-Reduced Spatial Bipartite Graph Clustering for Hyperspectral and LiDAR Data
- Directional Source Separation for Robust Speech Recognition on Smart Glasses
- DirichNet Model for Detection of TMS-Induced Speech Errors in Patients Undergoing Epilepsy Surgery
- Discrete Unit-based Low-latency Multi-lingual Speech Synthesis for LIMMITS'25 Challenge
- Discriminating Mizo Hunting and War Chants using Acoustic Features
- Disentangle Heart Rate Signals for Improved Stress Detection
- Disentangled Representation Learning for Chinese Handwriting Recognition
- Disentanglement Analysis in Deep Latent Variable Models Matching Aggregate Posterior Distributions
- Disentangling Hierarchical Features for Anomalous Sound Detection Under Domain Shift
- Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
- Disparity-Guided Cross-View Transformer For Stereo Image Super-Resolution
- Distance Based Single-Channel Target Speech Extraction
- Distill To Detect: Amplifying Anomalies in Backdoor Models through Knowledge Distillation
- DistillW2N: A Lightweight One-Shot Whisper to Normal Voice Conversion Model Using Distillation of Self-Supervised Features
- Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment
- Distilling Generative-Discriminative Representations for Very Low-Resolution Face Recognition
- Distilling Knowledge from Large Video Models for Driver Visual Attention Prediction
- Distributed ATC Particle Filters for Cooperative Quaternion Tracking
- Distributed IRSs Mitigate Spatial Wideband & Beam Split Effects
- Distributed Interference Alignment Precoding and Detection for MU-MIMO OTSM Downlink in Time-Varying Channels
- Distributed Navigation with Dynamic Obstacles
- Distributed-Robust Source Localization in Wireless Acoustic Sensor Networks
- Distribution Alignment Informed Thresholding for Semi-Supervised Curvilinear Structure Segmentation
- Distributionally Robust Kalman Filtering over an Infinite-Horizon
- Diverse Collaboration in Multi-Agent Reinforcement Learning via Self-Adaptive Method
- Diversified Augmentation with Domain Adaptation for Debiased Video Temporal Grounding
- Diversity Matters: Co-training for Semi-Supervised Change Detection in Remote Sensing Images
- Diversity Seeking Techniques for Red-Teaming Large Language Models
- Divide-and-Conquer Variational Bayesian Inference for Multi-task Learning of High-resolution SAR Imagery
- Do Less and Achieve More: Free Condition Video Outpainting with Diffusion Model
- Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning
- DoA-Aided MMSE Channel Estimation for Wireless Communication Systems
- DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
- Domain Connection based Unsupervised Domain Adaptation for Semantic Segmentation
- Domain Obfuscation for Efficient Secure Aggregation of Sparse Feature Vectors
- Domain-Aware Knowledge Debiasing for Generalizable Video Understanding in CLIP
- Domain-Incremental Learning for Audio Classification
- Domain-Independent Automatic Generation of Descriptive Texts for Time-Series Data
- Domain-Specific Adaptation in Speech Emotion Recognition Using Emotional Distribution Alignment
- Domain-aware Node Representation Learning for Graph Out-of-Distribution Generalization
- Don't Lose Yourself: Boosting Multimodal Recommendation via Reducing Node-neighbor Discrepancy in Graph Convolutional Network
- Doppler Single-Photon Lidar
- Double Domain Converter Transformer For Improving EEG-Based Emotion Recognition from Video to Game Scenarios
- DrLLM: Prompt-Enhanced Distributed Denial-of-Service Resistance Method with Large Language Models
- DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
- DreamHA: Towards High-Quality Human Animation with Image-to-Video Diffusion Models
- DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance
- Driver Reaction Time Prediction Through Adaptive Evolutionary Synchrony Window and Convolutional-LSTM
- DuCol: Text-Tag Adaptive Colorization of Dual-Character Line Art
- DuPI: Dual-resolution Pseudo-label Integration for Semi-supervised Instance Segmentation
- Dual Attention for Space-Time Video Super-Resolution
- Dual Decoder for Fast Inference in Natural Language Generation
- Dual Encoders for Diffusion-based Image Inpainting
- Dual Multi-Scale GCN with Deformable Temporal Kernel for Skeleton-based Action Recognition
- Dual Path Unsupervised Real Image Denoising
- Dual Position Attention Time-Frequency Network for Binaural Audio Synthesis
- Dual Trajectory Revised Diffusion Model for Time Series Forecasting
- Dual-Domain Feature-Guided Task Alignment for Enhanced Small Object Detection
- Dual-Frequency Spatio-Temporal Phase Unwrapping
- Dual-Function Waveform Design in Wireless Sensor Networks via SoS Optimization
- Dual-Modality Guided Artistic Style Transfer with Pre-trained Diffusion Models
- Dual-PST: Dual-Branch SpatioTemporal-Planar Network for Video Forgery Detection
- Dual-Path Consistency Unsupervised Domain Adaptation for Nighttime Semantic Segmentation
- Dual-Path Contrastive Short Text Clustering with High-order Random Walk
- Dual-Path Model for Pulmonary Artery Segmentation
- Dual-Population Watermark Vaccine: Efficient and Imperceptible Adversarial Attack for Watermarked Image Protection
- Dual-Process Watermarked Diffusion: Integrating Watermarking With Denoising in Point Clouds
- Dual-Pyramid Attention Collaborative Network for Oracle Bone Inscription Classification
- Dual-Space Augmented Intrinsic-LoRA for Wind Turbine Segmentation
- Dual-Triple Transformer Networks for Accurate CT Pleural Effusion Segmentation
- Dual-energy CT metal artifact reduction by combined material decomposition and projection domain threshold segmentation
- Dual-level AMR Injection for Prompt-based Event Argument Extraction
- Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation
- Dynamic Category Queries Transformer for Generalized Few-shot Semantic Segmentation
- Dynamic Dictionary Design for Localization in Automotive Radar Systems
- Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
- Dynamic Graph Convolutional Networks with Spatiotemporal Missing Pattern Awareness
- Dynamic Graph Multi-granularity Attribute Scene Evolution Sequence Recommendation
- Dynamic Graph Recommendation via Sparse Augmentation and Singular Adaptation
- Dynamic Incentive Model for Federated Learning Model Trading via Evolutionary Game Theory
- Dynamic Language Group-based MoE: Enhancing Code-Switching Speech Recognition with Hierarchical Routing
- Dynamic Object Queries for Transformer-based Incremental Object Detection
- Dynamic Prototype Rehearsal for Continual ECG Arrhythmia Detection
- Dynamic ROI Adaptation for Accurate Non-Contact Heart Rate Estimation Using VGG-13 based Encoder-Decoder Model and Facial Landmarks
- Dynamic Routing and Calibration for Few-Shot Object Detection
- Dynamic SRM Curriculum for Trustworthy Multi-modal Classification
- Dynamic Soft Contrastive Learning for Time Series Anomaly Detection
- Dynamic Sparse Encoding and Cross-Temporal Attention for Remote Sensing Image Change Detection
- Dynamic Speech Generation to Enhance Intelligibility in Noisy Environments
- Dynamic SpikFormer: Low-Latency & Energy-Efficient Spiking Neural Networks with Dynamic Time Steps for Vision Transformers
- Dynamic Structure Hypergraph for Document-level Event Extraction
- Dynamic-static Feature Fusion with Multi-scale Attention for Continuous Blood Glucose Prediction
- DynamicAttention: Dynamic KV Cache for Disaggregate LLM Inference
- Dynamically Causal-Enhanced Exercise Representations for Adaptive Knowledge Tracing
- Dynamically Optimize MTD Strategy in Satellite Computing Systems Using A2C Reinforcement Learning
- Dysarthric Speech Conformer: Adaptation for Sequence-to-Sequence Dysarthric Speech Recognition
- E-RNS : Enhancing Negative Sample Quality from Gradient Perspective for Graph Recommendation
- E-URES 2.0: Efficient User-Centric Residual-Echo Suppression with a Lightweight Neural Network
- E1 TTS: Simple and Fast Non-Autoregressive TTS
- ECBANet: Exploiting Complementary Information for Efficient Burst Super-Resolution
- ECG-guided individual identification via PPG
- ECSNN: Spiking Neural Networks for Efficient Exposure Correction in Endoscopy Imaging
- EDSep: An Effective Diffusion-Based Method for Speech Source Separation
- EEG Correlation Analysis-guided Graph Local Enhanced Feature Learning For Emotion Recognition
- EEG Decoding and Visual Reconstruction via 3D Geometric with Nonstationarity Modelling
- EEG-Music Emotion Recognition: Challenge Overview
- EEG-ReMinD: Enhancing Neurodegenerative EEG Decoding through Self-Supervised State Reconstruction-Primed Riemannian Dynamics
- EFL-PEFT: A communication Efficient Federated Learning framework using PEFT sparsification for ASR
- EGAS: Enhanced Geometry-aware 3D Asset Generation Using Gaussian Splatting
- EGENN: An Efficient Graph-Enhanced Neural Network for Multivariate Time Series Forecasting
- EM-MIAs: Enhancing Membership Inference Attacks in Large Language Models through Ensemble Modeling
- EMMeTT: Efficient Multimodal Machine Translation Training
- EP-SAM: An Edge-Detection Prompt SAM Based Efficient Framework for Ultra-Low Light Video Segmentation
- EPCPE: A Real-time End-to-End Pipeline for RGB-based Category-level 6D Pose Estimation
- EPE-P: Evidence-based Parameter-efficient Prompting for Multimodal Learning with Missing Modalities
- EPI-Mamba: State Space Model for Semantic Segmentation from Light Fields
- EPIC: Error Pattern Informed Correction for Classroom ASR with Limited Labeled Data
- ERGNN: Spectral Graph Neural Network With Explicitly-Optimized Rational Graph Filters
- ES-NeRF: Enhancing Segmentation in NeRF with CLIP
- ETDE-Net: An End-to-End Time-Domain Enhancement Network for LPI Radar Signals
- EagerLog: Active Learning Enhanced Retrieval Augmented Generation for Log-based Anomaly Detection
- Earbuds Orientation Alignment Based on Markov Chain Monte Carlo Sampling
- Early Dementia Detection Using Multiple Spontaneous Speech Prompts: The PROCESS Challenge
- Easing Optimization Paths: a Circuit Perspective
- Easy, Interpretable, Effective: openSMILE for voice deepfake detection
- Easy-to-hard Instance-level Feature Fusion for Co-saliency Detection
- EasyControl: Adding Control to Video Diffusion for Controllable Video Generation and Interpolation
- Edge First: Edge-Guided Geometry for Superior 3D Roof Wireframe Reconstruction
- Edge-aware Laplacian Pyramid Network for Efficient Image Deblurring
- Edge-interaction Mamba Network for MRI Brain Tumor Segmentation
- Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer
- Effective Context Modeling Framework for Emotion Recognition in Conversations
- Effective Integration of KAN for Keyword Spotting
- Effective Pre-Training of Audio Transformers for Sound Event Detection
- Effective Techniques for Scaling Audio Encoder Pretraining
- Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
- EffectiveASR: A Single-Step Non-Autoregressive Mandarin Speech Recognition Architecture with High Accuracy and Inference Speed
- Efficient Anchor Graph Clustering Through Enhanced Within-Cluster Homogeneity
- Efficient Co-Approximate Parallel Compressive Depth Reconstruction on FPGA
- Efficient Co-clustering via Anchor-refined Label Spreading
- Efficient Data-Dependent Random Projection for Least Square Regressions
- Efficient Dataset Distillation through Low-Rank Space Sampling
- Efficient Defocus Deblurring Networks based on Diffusion Models
- Efficient Estimation of Kernel Matrix Spectral Norm using Random Features
- Efficient Extreme Large-Scale Speaker Verification: Dynamic Active Sub Fully-Connected Layers for Faster Training and Memory Optimization
- Efficient Fine-tuning Strategies for Enhancing Face Recognition Performance in Challenging Scenarios
- Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of Transformers
- Efficient Fusion of Computationally Diverse Modalities Using Chunking and Cross-Attention
- Efficient Global Attention and Correlation-Aware Fusion for Hyperspectral Image Classification
- Efficient Gridless Wideband Direction-of-Arrival Estimation From Many Frequencies
- Efficient Hierarchical Domain Adaptive Thermal Infrared Tracking
- Efficient Infrared Image Super-Resolution Reconstruction via Guided Filter Coefficients Estimation with Parallax Attention Mechanism
- Efficient Large-Scale Scene Point Cloud Upsampling with Implicit Neural Networks and Spatial Hashing
- Efficient Learning of Balanced Signed Graphs via Iterative Linear Programming
- Efficient Localized Perception for Resource-Constrained Vision Systems
- Efficient Long Document Ranking via Adaptive Token Pruning with Query-Document Alignment
- Efficient Long Speech Sequence Modelling for Time-Domain Depression Level Estimation
- Efficient Long-Form Speech Recognition for General Speech In-Context Learning
- Efficient MDCT-Based Multi-Channel Coding with Perceptual Whitening and Broadband ILD Compensation
- Efficient Modeling and Low Complexity Implementation of Rate Estimation in Versatile Video Coding
- Efficient Multi-branch Black-box Semantic-aware Targeted Attack Against Deep Hashing Retrieval
- Efficient Non-Sequential Relational Modeling for Temporal Knowledge Graph Link Predictions
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.