← Search

Yi Pan

10 accepted papers

2026

SIGMA-PPG: Statistical-prior Informed Generative Masking Architecture for PPG Foundation Model

ICML 2026poster

Current foundation model for photoplethysmography (PPG) signals is challenged by the intrinsic redundancy and noise of the signal. Standard masked modeling often yields trivial solutions while contrastive methods lack morphological precision. To address these limitations, we propose a Statistical-pr…

Cited by 0SourceScholar
2025

AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs

ACL 2025long

When aligning large language models (LLMs), their performance across various tasks (such as being helpful, harmless, and honest) is heavily influenced by the composition of the training data. However, it is difficult to determine what mixture of data should be used to produce a model with strong per…

2025

ECHOPulse: ECG Controlled Echocardio-gram Video Generation

ICLR 2025poster

Echocardiography (ECHO) is essential for cardiac assessments, but its video quality and interpretation heavily relies on manual expertise, leading to inconsistent results from clinical and portable devices. ECHO video generation offers a solution by improving automated monitoring through synthetic d…

Cited by 5SourcePDFScholar
2025

FAWL: Weakly-Supervised Video Corpus Moment Retrieval with Frame-Wise Auxiliary Alignment and Weighted Contrastive Learning

ICASSP 2025accepted

Video Corpus Moment Retrieval (VCMR) is a challenging task that aims to localize query-specified moments from a collection of untrimmed videos. The recent state-of-the-art method, JSG, tries to tackle this task using only video-level annotations in a weakly-supervised setting. However, the late fusi…

Cited by 0SourceScholar
2025

HELENE: Hessian Layer-wise Clipping and Gradient Annealing for Accelerating Fine-tuning LLM with Zeroth-order Optimization

EMNLP 2025

Fine-tuning large language models (LLMs) faces significant memory challenges due to the high cost of back-propagation. MeZO addresses this using zeroth-order (ZO) optimization, matching memory usage to inference but suffering from slow convergence due to varying curvatures across model parameters. T

Cited by 0SourcePDFScholar
2025

ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval

EMNLP 2025

Partially Relevant Video Retrieval (PRVR) is a practical yet challenging task that involves retrieving videos based on queries relevant to only specific segments. While existing works follow the paradigm of developing models to process unimodal features, powerful pretrained vision-language models li

2025

RefCap: Zero-shot Video Corpus Moment Retrieval Based on Refined Dense Video Captioning

ICASSP 2025accepted

Video corpus moment retrieval (VCMR) is a challenging task aimed at localizing specific segments from untrimmed videos within a vast video collection. It has long been addressed using end-to-end supervised or weakly-supervised methods, which often lack explainability and rely on laborious annotation…

Cited by 0SourceScholar
2023

Deconfounded Opponent Intention Inference for Football Multi-Player Policy Learning

IROS 2023poster

Due to the high complexity of a football match, the opponents' strategies are variable and unknown. Thus predicting the opponents' future intentions accurately based on current situation is crucial for football players' decision-making. To better anticipate the opponents and learn more effective str…

Cited by 2SourceScholar
2023

Lazy Agents: A New Perspective on Solving Sparse Reward Problem in Multi-agent Reinforcement Learning

ICML 2023poster

Sparse reward remains a valuable and challenging problem in multi-agent reinforcement learning (MARL). This paper addresses this issue from a new perspective, i.e., lazy agents. We empirically illustrate how lazy agents damage learning from both exploration and exploitation. Then, we propose a novel…

2021

Alexa Conversations: An Extensible Data-driven Approach for Building Task-oriented Dialogue Systems

NAACL 2021system demonstrations

Traditional goal-oriented dialogue systems rely on various components such as natural language understanding, dialogue state tracking, policy learning and response generation. Training each component requires annotations which are hard to obtain for every new domain, limiting scalability of such sys…

Cited by 22SourcePDFScholar