← Search

Yixuan Sun

12 accepted papers

2026

Beyond Buffer Limits: Energy-Based Data Reassembly for Continual Learning

ICML 2026poster

Continual learning (CL) aims to acquire new knowledge from a non-stationary data stream while retaining performance on previously learned tasks. Memory-based replay methods mitigate catastrophic forgetting by storing and revisiting past samples, but their effectiveness is fundamentally constrained b…

Cited by 0SourceScholar
2026

Beyond Instance-Level Self-Supervision in 3D Multi-Modal Medical Imaging

ICML 2026poster

Self-supervised pre-training methods in medical imaging typically treat each individual as an isolated instance, learning representations through augmentation-based objectives or masked reconstruction. They often do not adequately capitalize on a key characteristic of physiological features: anatomi…

Cited by 0SourceScholar
2026

Tracing the Heart’s Pathways: ECG Representation Learning from a Cardiac Conduction Perspective

AAAI 2026technical

The multi-lead electrocardiogram (ECG) stands as a cornerstone of cardiac diagnosis. Recent strides in electrocardiogram self-supervised learning (eSSL) have brightened prospects for enhancing representation learning without relying on high-quality annotations. Yet earlier eSSL methods suffer a key

Cited by 0SourcePDFScholar
2025

QA-MDT: Quality-aware Masked Diffusion Transformer for Enhanced Music Generation

IJCAI 2025

Text-to-music (TTM) generation, which converts textual descriptions into audio, opens up innovative avenues for multimedia creation. Achieving high quality and diversity in this process demands extensive, high-quality data, which are often scarce in available datasets. Most open-source datasets freq

2024

Pixel-level Semantic Correspondence through Layout-aware Representation Learning and Multi-scale Matching Integration

CVPR 2024poster

Establishing precise semantic correspondence across object instances in different images is a fundamental and challenging task in computer vision. In this task difficulty arises often due to three challenges: confusing regions with similar appearance inconsistent object scale and indistinguishable n…

2023

Collision-free Coverage Path Planning for the Variable-speed Curvature-constrained Robot

ICRA 2023poster

Dubins coverage has been extensively researched to address the coverage path planning (CPP) problem of a known environment for the curvature-constrained robot. However, its fixed-speed assumption prevents the robot from accelerating to reduce the time and limits its flexibility to avoid obstacles. T…

Cited by 2SourceScholar
2023

Correspondence Transformers With Asymmetric Feature Learning and Matching Flow Super-Resolution

CVPR 2023poster

This paper solves the problem of learning dense visual correspondences between different object instances of the same category with only sparse annotations. We decompose this pixel-level semantic matching problem into two easier ones: (i) First, local feature descriptors of source and target images…

2023

Hierarchical Visual Categories Modeling: A Joint Representation Learning and Density Estimation Framework for Out-of-Distribution Detection

ICCV 2023poster

Detecting out-of-distribution inputs for visual recognition models has become critical in safe deep learning. This paper proposes a novel hierarchical visual category modeling scheme to separate out-of-distribution data from in-distribution data through joint representation learning and statistical…

Cited by 3PDFScholar
2023

MISC210K: A Large-Scale Dataset for Multi-Instance Semantic Correspondence

CVPR 2023poster

Semantic correspondence have built up a new way for object recognition. However current single-object matching schema can be hard for discovering commonalities for a category and far from the real-world recognition tasks. To fill this gap, we design the multi-instance semantic correspondence task wh…

2023

Weakly Supervised Learning of Semantic Correspondence through Cascaded Online Correspondence Refinement

ICCV 2023poster

In this paper, we develop a weakly supervised learning algorithm to learn robust semantic correspondences from large-scale datasets with only image-level labels. Following the spirit of multiple instance learning (MIL), we decompose the weakly supervised correspondence learning problem into three st…

Cited by 1PDFcodeScholar
2022

FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos

CVPR 2022poster

Current benchmarks for facial expression recognition (FER) mainly focus on static images, while there are limited datasets for FER in videos. It is still ambiguous to evaluate whether performances of existing methods remain satisfactory in real-world application-oriented scenes. For example, the "Ha…

Cited by 110PDFcodeScholar