← Search

Ying Sun

50 accepted papers

2026

Discovering Decoupled Functional Modules in Large Language Models

AAAI 2026technical

Understanding the internal functional organization of Large Language Models (LLMs) is crucial for improving their trustworthiness and performance. However, how LLMs organize different functions into modules remains highly unexplored. To bridge this gap, we formulate a function module discovery probl

Cited by 0SourcePDFScholar
2026

Forest-Based Graph Learning for Semi-Supervised Node Classification

ICLR 2026poster

Existing Graph Neural Networks usually learn long-distance knowledge via stacked layers or global attention, but struggle to balance cost-effectiveness and global receptive field. In this work, we break the dilemma by proposing a novel forest-based graph learning (FGL) paradigm that enables efficien…

Cited by 0SourceScholar
2026

LiDAR-to-4DRadar Diffusion Bridge via Cross-Modal Alignment and Translation in Latent Space

CVPR 2026

Millimeter-wave radar's all-weather capability makes it increasingly vital for autonomous perception. However, the high cost of radar data collection drives the need for data generation to augment radar datasets. Existing works mainly target partial radar representations, e.g., 2D or 3D slices, lead

Cited by 0SourceScholar
2026

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining

ICML 2026poster

Large language models (LLMs) have achieved remarkable breakthroughs across various applications. However, their architectures remain inefficient in pretraining due to two main limitations: (i) self-attention lacks an explicit inductive bias for locality, leading to redundant modeling of sequence-int…

Cited by 0SourceScholar
2026

Low-cost Full Fine-tuning: Learning What to Update for LLMs

ICML 2026poster

While Large language models (LLMs) have strong abilities, they generally rely on fine-tuning to supplement downstream task-specific knowledge. Due to the prohibitive memory overhead of full fine-tuning (FT), existing parameter-efficient fine-tuning techniques, e.g., LoRA and Adapters, update paramet…

Cited by 0SourceScholar
2026

MolEditRL: Structure-Preserving Molecular Editing via Discrete Diffusion and Reinforcement Learning

ICLR 2026poster

Molecular editing aims to modify a given molecule to optimize desired chemical properties while preserving structural similarity. However, current approaches typically rely on string-based or continuous representations, which fail to adequately capture the discrete, graph-structured nature of molecu…

Cited by 0SourceScholar
2026

NGTM: Substructure-based Neural Graph Topic Model for Interpretable Graph Generation

AAAI 2026technical

Graph generation plays a pivotal role across numerous domains, including molecular design and knowledge graph construction. Although existing methods achieve considerable success in generating realistic graphs, their interpretability remains limited, often obscuring the rationale behind structural d

Cited by 0SourcePDFScholar
2026

Representation-Aware Modularity: Efficient Cross-Task Generalization for LLMs

IJCAI 2026

Cross-task generalization (CTG) enables large language models (LLMs) to handle unseen tasks proficiently, enhancing their adaptability in real-world scenarios. However, existing methods relying on per-token dynamic routing to multiple trained LoRA adapters face high computational and GPU memory cost

Cited by 0Scholar
2026

Risk-Aware and Scalable Hierarchical Motion Planning for Large-Scale Robotic Swarms Via CVaR-Constrained MPC (I)

ICRA 2026poster

Motion planning for large-scale robotic swarms presents significant challenges in terms of scalability and safety assurance in cluttered environments. To address these issues, this manuscript proposes a Closed-loop hierarchical Risk-aware swarm mOtion planner using Conditional ValuE at Risk (C-ROVER…

Cited by 0Scholar
2026

Towards Illumination-Aware Restoration of Metalens-Captured Images: A New Dataset and a Strong Baseline

AAAI 2026technical

Metalenses offer compelling advantages such as lightweight and ultra-thin design, making them promising alternatives to conventional lenses. However, their widespread adoption is hindered by image quality degradation caused by chromatic and angular aberrations. To mitigate this, restoration processe

Cited by 0SourcePDFScholar
2026

Your AI-Generated Image Detector Can Secretly Achieve SOTA Accuracy, If Calibrated

AAAI 2026technical

Despite being trained on balanced datasets, existing AI-generated image detectors often exhibit systematic bias at test time, frequently misclassifying fake images as real. We hypothesize that this behavior stems from distributional shift in fake samples and implicit priors learned during training.

Cited by 0SourcePDFScholar
2025

EFSkip: A New Error Feedback with Linear Speedup for Compressed Federated Learning with Arbitrary Data Heterogeneity

AAAI 2025technical

Due to the communication bottleneck in distributed and decentralized federated learning applications, algorithms using compressed communication have attracted significant attention. The Error Feedback (EF) is a widely-studied compression framework for convergence with biased compressors such as top-…

Cited by 0SourcePDFScholar
2025

LLMs Can Simulate Standardized Patients via Agent Coevolution

ACL 2025long

Training medical personnel using standardized patients (SPs) remains a complex challenge, requiring extensive domain expertise and role-specific practice. Most research on Large Language Model (LLM)-based simulated patients focuses on improving data retrieval accuracy or adjusting prompts through hu…

2025

Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video Understanding

NeurIPS 2025poster

Understanding long video content is a complex endeavor that often relies on densely sampled frame captions or end-to-end feature selectors, yet these techniques commonly overlook the logical relationships between textual queries and visual elements. In practice, computational constraints necessitate…

Cited by 0SourceScholar
2025

OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects

NeurIPS 2025spotlight

Free-moving object reconstruction from monocular video remains challenging, particularly without reliable pose or depth cues and under arbitrary object motion. We introduce OnlineSplatter, a novel online feed-forward framework generating high-quality, object-centric 3D Gaussians directly from RGB fr…

Cited by 0SourceScholar
2025

Revisiting Noise Resilience Strategies in Gesture Recognition: Short-Term Enhancement in sEMG Analysis

ICML 2025poster

Gesture recognition based on surface electromyography (sEMG) has been gaining importance in many 3D Interactive Scenes. However, sEMG is easily influenced by various forms of noise in real-world environments, leading to challenges in providing long-term stable interactions through sEMG. Existing met…

Cited by 0SourcePDFScholar
2025

Understanding the Statistical Accuracy-Communication Trade-off in Personalized Federated Learning with Minimax Guarantees

ICML 2025poster

Personalized federated learning (PFL) offers a flexible framework for aggregating information across distributed clients with heterogeneous data. This work considers a personalized federated learning setting that simultaneously learns global and local models. While purely local training has no commu…

Cited by 0SourcePDFScholar
2025

Unifying Knowledge from Diverse Datasets to Enhance Spatial-Temporal Modeling: A Granularity-Adaptive Geographical Embedding Approach

ICML 2025poster

Spatio-temporal forecasting provides potential for discovering evolutionary patterns in geographical scientific data. However, geographical scientific datasets are often manually collected across studies, resulting in limited time spans and data scales. This hinders existing methods that rely on ric…

Cited by 0SourcePDFScholar
2024

Anomaly Subgraph Detection through High-Order Sampling Contrastive Learning

IJCAI 2024poster

Anomaly subgraph detection is a crucial task in various real-world applications, including identifying high-risk areas, detecting river pollution, and monitoring disease outbreaks. Early traditional graph-based methods can obtain high-precision detection results in scenes with small-scale graphs and…

Cited by 0SourcePDFScholar
2024

Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs

NeurIPS 2024poster

Large Language Models (LLMs) have shown remarkable reasoning capabilities on complex tasks, but they still suffer from out-of-date knowledge, hallucinations, and opaque decision-making. In contrast, Knowledge Graphs (KGs) can provide explicit and editable knowledge for LLMs to alleviate these issues…

2024

Risk-Aware Non-Myopic Motion Planner for Large-Scale Robotic Swarm Using CVaR Constraints

IROS 2024poster

Swarm robotics has garnered significant attention due to its ability to accomplish elaborate and synchronized tasks. Existing methodologies for motion planning of swarm robotic systems mainly encounter difficulties in scalability and safety guarantee. To address these limitations, we propose a Risk-…

Cited by 1SourceScholar
2024

SpGesture: Source-Free Domain-adaptive sEMG-based Gesture Recognition with Jaccard Attentive Spiking Neural Network

NeurIPS 2024poster

Surface electromyography (sEMG) based gesture recognition offers a natural and intuitive interaction modality for wearable devices. Despite significant advancements in sEMG-based gesture recognition models, existing methods often suffer from high computational latency and increased energy consumptio…

2024

Tackling Uncertain Correspondences for Multi-Modal Entity Alignment

NeurIPS 2024poster

Recently, multi-modal entity alignment has emerged as a pivotal endeavor for the integration of Multi-Modal Knowledge Graphs (MMKGs) originating from diverse data sources. Existing works primarily focus on fully depicting entity features by designing various modality encoders or fusion approaches. H…

Cited by 5SourcePDFScholar
2024

TransFusion: Covariate-Shift Robust Transfer Learning for High-Dimensional Regression

AISTATS 2024poster

The main challenge that sets transfer learning apart from traditional supervised learning is the distribution shift, reflected as the shift between the source and target models and that between the marginal covariate distributions. In this work, we tackle model shifts in the presence of covariate sh…

Cited by 19SourcePDFScholar
2023

A General Regret Bound of Preconditioned Gradient Method for DNN Training

CVPR 2023highlight

While adaptive learning rate methods, such as Adam, have achieved remarkable improvement in optimizing Deep Neural Networks (DNNs), they consider only the diagonal elements of the full preconditioned matrix. Though the full-matrix preconditioned gradient methods theoretically have a lower regret bou…

2023

Beyond Homophily: Robust Graph Anomaly Detection via Neural Sparsification

IJCAI 2023poster

Recently, graph-based anomaly detection (GAD) has attracted rising attention due to its effectiveness in identifying anomalies in relational and structured data. Unfortunately, the performance of most existing GAD methods suffers from the inherent structural noises of graphs induced by hidden anomal…

2023

Counterfactual Dynamics Forecasting – a New Setting of Quantitative Reasoning

AAAI 2023technical

Rethinking and introspection are important elements of human intelligence. To mimic these capabilities, counterfactual reasoning has attracted attention of AI researchers recently, which aims to forecast the alternative outcomes for hypothetical scenarios (“what-if”). However, most existing approach…

2023

Learning to Learn: How to Continuously Teach Humans and Machines

ICCV 2023poster

Curriculum design is a fundamental component of education. For example, when we learn mathematics at school, we build upon our knowledge of addition to learn multiplication. These and other concepts must be mastered before our first algebra lesson, which also reinforces our addition and multiplicati…

Cited by 5PDFScholar
2023

Meta Compositional Referring Expression Segmentation

CVPR 2023poster

Referring expression segmentation aims to segment an object described by a language expression from an image. Despite the recent progress on this task, existing models tackling this task may not be able to fully capture semantics and visual representations of individual concepts, which limits their…

Cited by 34SourcePDFScholar
2023

Visuo-Tactile Feedback-Based Robot Manipulation for Object Packing

RA-L 2023

Robots are increasingly expected to manipulate objects, of which properties have high perceptual uncertainty from any single sensory modality. This directly impacts successful object manipulation. Object packing is one of the challenging tasks in robot manipulation. In this work, a new visuo-tactile

Cited by 23SourceScholar
2022

Hybrid Local SGD for Federated Learning with Heterogeneous Communications

ICLR 2022spotlight

Communication is a key bottleneck in federated learning where a large number of edge devices collaboratively learn a model under the orchestration of a central server without sharing their own training data. While local SGD has been proposed to reduce the number of FL rounds and become the algorithm…

Cited by 60SourcePDFScholar
2022

Monitoring Vegetation From Space at Extremely Fine Resolutions via Coarsely-Supervised Smooth U-Net

IJCAI 2022poster

Monitoring vegetation productivity at extremely fine resolutions is valuable for real-world agricultural applications, such as detecting crop stress and providing early warning of food insecurity. Solar-Induced Chlorophyll Fluorescence (SIF) provides a promising way to directly measure plant product…

Cited by 2SourcePDFScholar
2022

Self-Supervised Global-Local Structure Modeling for Point Cloud Domain Adaptation With Reliable Voted Pseudo Labels

CVPR 2022poster

In this paper, we propose an unsupervised domain adaptation method for deep point cloud representation learning. To model the internal structures in target point clouds, we first propose to learn the global representations of unlabeled data by scaling up or down point clouds and then predicting the…

Cited by 68PDFScholar
2022

Tackling Data Heterogeneity: A New Unified Framework for Decentralized SGD with Sample-induced Topology

ICML 2022spotlight

We develop a general framework unifying several gradient-based stochastic optimization methods for empirical risk minimization problems both in centralized and distributed scenarios. The framework hinges on the introduction of an augmented graph consisting of nodes modeling the samples and edges mod…

Cited by 19SourcePDFScholar
2021

Discerning Decision-Making Process of Deep Neural Networks with Hierarchical Voting Transformation

NeurIPS 2021poster

Neural network based deep learning techniques have shown great success for numerous applications. While it is expected to understand their intrinsic decision-making processes, these deep neural networks often work in a black-box way. To this end, in this paper, we aim to discern the decision-making…

2020

Accelerated Primal-Dual Algorithms for Distributed Smooth Convex Optimization over Networks

AISTATS 2020poster

This paper proposes a novel family of primal-dual-based distributed algorithms for smooth, convex, multi-agent optimization over networks that uses only gradient information and gossip communications. The algorithms can also employ acceleration on the computation and communications. We provide a u…

2017

D2L: Decentralized dictionary learning over dynamic networks

ICASSP 2017accepted

The paper studies a general class of distributed dictionary learning (DL) problems where the learning task is distributed over a multi-agent network with (possibly) time-varying (non-symmetric) connectivity. This setting is relevant, for instance, in scenarios where massive amounts of data are not c…

Cited by 0SourceScholar
2016

Orthogonal sparse eigenvectors: A procrustes problem

ICASSP 2016accepted

The problem of estimating sparse eigenvectors of a symmetric matrix attracts a lot of attention in many applications, especially those with high dimensional data set. While classical eigenvectors can be obtained as the solution of a maximization problem, existing approaches formulated this problem b…

Cited by 0SourceScholar
2015

Robust estimation of structured covariance matrix for heavy-tailed distributions

ICASSP 2015accepted

In this paper, we consider the robust covariance estimation problem in the non-Gaussian set-up. In particular, Tyler's M-estimator is adopted for samples drawn from a heavy-tailed elliptical distribution. For some applications, the covariance matrix naturally possesses certain structure. Therefore,…

Cited by 0SourceScholar