← Search

Yi Ding

38 accepted papers

2026

Deformba: Vision State Space Model with Adaptive State Fusion

ICML 2026poster

State Space Models (SSMs) have emerged as a powerful and efficient alternative to Transformers, demonstrating linear-time complexity and exceptional sequence modeling capabilities. However, their application to vision tasks remains challenging. First, existing vision SSMs largely depend on manually …

Cited by 0SourceScholar
2026

ECHO: Toward Contextual Seq2Seq Paradigms in Large EEG Models

ICLR 2026poster

Electroencephalography (EEG), with its broad range of applications, necessitates models that can generalize effectively across various tasks and datasets. Large EEG Models (LEMs) address this by pretraining encoder-centric architectures on large-scale unlabeled data to extract universal representati…

Cited by 0SourcecodeScholar
2026

EEG-Based Multimodal Learning via Hyperbolic Mixture-of-Curvature Experts

ICML 2026poster

Electroencephalography (EEG)-based multimodal learning integrates brain signals with complementary modalities to improve mental state assessment, providing great clinical potential. The effectiveness of such paradigms largely depends on the representation learning on heterogeneous modalities. For EE…

Cited by 0SourceScholar
2026

EEG-DLite: Dataset Distillation for Efficient Large EEG Model Training

AAAI 2026technical

Large-scale EEG foundation models have shown strong generalization across a range of downstream tasks, but their training remains resource-intensive due to the volume and variable quality of EEG data. In this work, we introduce EEG-DLite, a data distillation framework that enables more efficient pre

Cited by 0SourcePDFScholar
2026

EmBrace: A Collective Knowledge Fusion Framework Toward Unified EEG Foundation Models

ICML 2026poster

Electroencephalography (EEG) foundation models (EFMs) have achieved strong performance across a wide range of downstream EEG tasks via pretraining and fine-tuning. Through empirical analysis, we observe that (i) no single EFM consistently dominates all tasks, yet identifying the task-specific optima…

Cited by 0SourceScholar
2026

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips

CVPR 2026

Foley art plays a pivotal role in enhancing immersive auditory experiences in film, yet manual creation of spatio-temporal aligned audio remains labor-intensive. We propose FoleyDesigner, a novel framework inspired by professional Foley workflows, integrating film clip analysis, spatio-temporal cont

Cited by 0SourceScholar
2026

ForeDiffusion: Foresight-Conditioned Diffusion Policy via Future View Construction for Robot Manipulation

AAAI 2026technical

Diffusion strategies have advanced visual motor control by progressively denoising high-dimensional action sequences, providing a promising method for robot manipulation. However, as task complexity increases, the success rate of existing baseline models decreases considerably. Analysis indicates th

Cited by 0SourcePDFScholar
2026

Position: Modular Safety Guardrails Are Necessary for Foundation-Model-Enabled Robots in the Real World

ICML 2026poster

The integration of foundation models (FMs) into robotics has accelerated real-world deployment, while introducing new safety challenges arising from open-ended semantic reasoning and embodied physical action. These challenges require safety notions beyond physical constraint satisfaction. In this po…

Cited by 0SourceScholar
2026

Uni-NTFM: A Unified Foundation Model for EEG Signal Representation Learning

ICLR 2026poster

Current foundation models for electroencephalography (EEG) rely on architectures adapted from computer vision or natural language processing, typically treating neural signals as pixel grids or token sequences. This approach overlooks that the neural activity is activated by diverse sparse coding ac…

Cited by 0SourceScholar
2026

ViTSP: A Vision Language Models Guided Framework for Large-Scale Traveling Salesman Problems

ICLR 2026poster

Solving Traveling Salesman Problem (TSP) is NP-hard yet fundamental for wide real-world applications. Classical exact methods face challenges in scaling, and heuristic methods often require domain-specific parameter calibration. While learning-based approaches have shown promise, they suffer from po…

Cited by 0SourceScholar
2026

Where Signals Are Sparse, We Synthesize: Reinforcing Self-Corrective Reasoning in Vision–Language Models via Rollout Augmentation

ICML 2026poster

Self-correction is essential for solving complex reasoning problems in vision–language models (VLMs), yet existing reinforcement learning (RL) methods struggle to learn it. Effective self-correction behaviors emerge only rarely during RL, making learning signals sparse. To address this challenge, we…

Cited by 0SourceScholar
2025

CAT-Net: A Co-Adaptive Transfer Learning Network for BCI-Assisted Neurorehabilitation

ICASSP 2025accepted

Brain-computer interfaces (BCIs) hold great potential for motor recovery in post-stroke patients. However, the motor imagery decoding accuracy is limited by the non-stationarity of EEG signals across subjects and sessions. We propose CAT-Net: a Co-Adaptive Transfer learning network to simultaneously…

Cited by 0SourceScholar
2025

CPSNet: Comprehensive Enhancement Representation for Polyp Segmentation Task

ICASSP 2025accepted

Accurately segmenting polyp regions in colonoscopy images is crucial for the diagnosis and intervention of colorectal cancer. However, the task of polyp segmentation remains challenging due to the diverse size and shape variations among polyps, their extreme similarity to the background, and frequen…

Cited by 0SourceScholar
2025

DPM-LVSN: A Diffusion Probabilistic Model-based Left Ventricular Segmentation Network

ICASSP 2025accepted

To improve the accuracy and robustness of left ventricular segmentation, This paper proposes a diffusion probabilistic model-based left ventricular segmentation network (DPM-LVSN). DPM-LVSN integrates a U-shaped encoder-decoder for feature extraction, self-attention for semantic information, and a d…

Cited by 0SourceScholar
2025

ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time

ICLR 2025poster

Vision Language Models (VLMs) have become essential backbones for multi-modal intelligence, yet significant safety challenges limit their real-world application. While textual inputs can often be effectively safeguarded, adversarial visual inputs can often easily bypass VLM defense mechanisms. Exist…

2025

GDRIVE: Adaptive Object Detection in Autonomous Vehicles via Graph-Based Feature Learning

ICASSP 2025accepted

Navigating domain shifts in object detection is crucial for autonomous driving systems, particularly under varying weather conditions and diverse visual perspectives. Existing Cross-Domain Object Detection methods often struggle due to their reliance on broad semantic models, which can introduce bia…

Cited by 0SourceScholar
2025

MamBEV: Enabling State Space Models to Learn Birds-Eye-View Representations

ICLR 2025poster

3D visual perception tasks, such as 3D detection from multi-camera images, are essential components of autonomous driving and assistance systems. However, designing computationally efficient methods remains a significant challenge. In this paper, we propose a Mamba-based framework called MamBEV, whi…

2025

Multi-scale Graph Convolution with Corrective Contrastive Learning for Skeleton-based Action Recognition

ICASSP 2025accepted

For pursuing accurate skeleton-based action recognition, many existing graph-based approaches deploy the higher-order polynomials of the skeletal adjacency matrix to model the node correlations of distant neighbours. To further capture robust graphical patterns, a novel multi-scale graph convolution…

Cited by 0SourceScholar
2025

REFED: A Subject Real-time Dynamic Labeled EEG-fNIRS Synchronized Recorded Emotion Dataset

NeurIPS 2025poster

Affective brain-computer interfaces (aBCIs) play a crucial role in personalized human–computer interaction and neurofeedback modulation. To develop practical and effective aBCI paradigms and to investigate the spatial-temporal dynamics of brain activity under emotional inducement, portable electroen…

Cited by 0SourceScholar
2025

SelectiveFinetuning: Enhancing Transfer Learning In Sleep Staging Through Selective Domain Alignment

ICASSP 2025accepted

In practical sleep stage classification, a key challenge is the variability of EEG data across different subjects and environments. Differences in physiology, age, health status, and recording conditions can lead to domain shifts between data. These domain shifts often result in decreased model accu…

Cited by 0SourceScholar
2025

Ultrasound-Guided Registration Pseudo-Labels for Semi-Supervised Brachial Plexus Segmentation

ICASSP 2025accepted

In semi-supervised medical image segmentation, two main challenges arise. First, the quality of pseudo-labels generated by segmentation networks in data-limited scenarios is often poor, reducing segmentation accuracy. Second, many methods fail to effectively utilize the temporal context in video dat…

Cited by 0SourceScholar
2025

Unveiling Environmental Impacts of Large Language Model Serving: A Functional Unit View

ACL 2025long

Large language models (LLMs) offer powerful capabilities but come with significant environmental impact, particularly in carbon emissions. Existing studies benchmark carbon emissions but lack a standardized basis for comparison across different model configurations. To address this, we introduce the…

2025

Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection

EMNLP 2025

With the emergence of strong vision language capabilities, multimodal large language models (MLLMs) have demonstrated tremendous potential for real-world applications. However, the security vulnerabilities exhibited by the visual modality pose significant challenges to deploying such models in open-

2024

SparseSSP: 3D Subcellular Structure Prediction from Sparse-View Transmitted Light Images

ECCV 2024oral

"Traditional fluorescence staining is phototoxic to live cells, slow, and expensive; thus, the subcellular structure prediction (SSP) from transmitted light (TL) images is emerging as a label-free, faster, low-cost alternative. However, existing approaches utilize 3D networks for one-to-one voxel le…

2022

Learning-Enhanced Adaptive Robust GNSS Navigation in Challenging Environments

RA-L 2022

Global Navigation Satellite System (GNSS) is the widely used technology when it comes to outdoor positioning. But it has severe limitations with regard to safety-critical applications involving unmanned autonomous systems. Namely, the positioning performance degrades in harsh propagation environment

Cited by 16SourceScholar
2020

Dynamical Systems Theory for Causal Inference with Application to Synthetic Control Methods

AISTATS 2020poster

In this paper, we adopt results in nonlinear time series analysis for causal inference in dynamical settings. Our motivation is policy analysis with panel data, particularly through the use of “synthetic control" methods. These methods regress pre-intervention outcomes of the treated unit to outcom…

Cited by 5SourcePDFScholar
2020

Handling Missing Data with Graph Representation Learning

NeurIPS 2020poster

Machine learning with missing data has been approached in many different ways, including feature imputation where missing feature values are estimated based on observed values and label prediction where downstream labels are learned directly from incomplete data. However, existing imputation models…

2017

Multiresolution Kernel Approximation for Gaussian Process Regression

NeurIPS 2017spotlight

Gaussian process regression generally does not scale to beyond a few thousands data points without applying some sort of kernel approximation method. Most approximations focus on the high eigenvalue part of the spectrum of the kernel matrix, $K$, which leads to bad performance when the length scale…

Cited by 29SourcePDFScholar