← Search

Yifan Hu

28 accepted papers

2026

Bridging Past and Future: Distribution-Aware Alignment for Time Series Forecasting

ICLR 2026poster

Although contrastive and other representation-learning methods have long been explored in vision and NLP, their adoption in modern time series forecasters remains limited. We believe they hold strong promise for this domain. To unlock this potential, we explicitly align past and future representatio…

Cited by 0SourcecodeScholar
2026

Erosion Attack for Adversarial Training to Enhance Semantic Segmentation Robustness

ICASSP 2026poster

Existing segmentation models exhibit significant vulnerability to adversarial attacks.To improve robustness, adversarial training incorporates adversarial examples into model training. However, existing attack methods consider only global semantic information and ignore contextual semantic relations…

Cited by 0SourcePDFScholar
2026

From Observations to States: Latent Time Series Forecasting

ICML 2026poster

Deep learning has achieved strong performance in Time Series Forecasting (TSF). However, we identify a critical representation paradox, termed Latent Chaos: models with accurate predictions often learn latent representations that are temporally disordered and lack continuity. We attribute this pheno…

Cited by 0SourceScholar
2026

Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy Optimization

ICLR 2026poster

Multimodal large language models (MLLMs) are well suited to image aesthetic assessment, as they can capture high-level aesthetic features leveraging their cross-modal understanding capacity. However, the scarcity of multimodal aesthetic reasoning data and the inherently subjective nature of aestheti…

Cited by 0SourcecodeScholar
2025

Adaptive Multi-Scale Decomposition Framework for Time Series Forecasting

AAAI 2025technical

Transformer-based and MLP-based methods have emerged as leading approaches in time series forecasting (TSF). However, real-world time series often show different patterns at different scales, and future changes are shaped by the interplay of these overlapping scales, requiring high-capacity models.…

2025

Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis

ACL 2025finding

Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CSS models face interpretability limitations due to insufficient emotional perception and redundant discrete speech coding…

2025

Efficiently Serving Large Multimodal Models Using EPD Disaggregation

ICML 2025poster

Large Multimodal Models (LMMs) extend Large Language Models (LLMs) by handling diverse inputs such as images, audio, and video, but at the cost of adding a multimodal encoding stage that increases both computational and memory overhead. This step negatively affects key Service Level Objectives (SLOs…

2025

MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference Alignment

ICML 2025poster

Reinforcement Learning from Human Feedback (RLHF) has shown promise in aligning large language models (LLMs). Yet its reliance on a singular reward model often overlooks the diversity of human preferences. Recent approaches address this limitation by leveraging multi-dimensional feedback to fine-tun…

Cited by 0SourcePDFScholar
2025

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

AAAI 2025technical

Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in understanding the spatial environment from the image. Many attempts have been made to extract global spatial visual informat…

2025

TimeBridge: Non-Stationarity Matters for Long-term Time Series Forecasting

ICML 2025poster

Non-stationarity poses significant challenges for multivariate time series forecasting due to the inherent short-term fluctuations and long-term trends that can lead to spurious regressions or obscure essential long-term relationships. Most existing methods either eliminate or retain non-stationarit…

2025

TimeFilter: Patch-Specific Spatial-Temporal Graph Filtration for Time Series Forecasting

ICML 2025poster

Time series forecasting methods generally fall into two main categories: Channel Independent (CI) and Channel Dependent (CD) strategies. While CI overlooks important covariate relationships, CD captures all dependencies without distinction, introducing noise and reducing generalization. Recent advan…

2025

Unraveling Interwoven Roles of Large Language Models in Authorship Privacy: Obfuscation, Mimicking, and Verification

EMNLP 2025

Recent advancements in large language models (LLMs) have been fueled by large-scale training corpora drawn from diverse sources such as websites, news articles, and books. These datasets often contain explicit user information, such as person names, addresses, that LLMs may unintentionally reproduce

2025

Vanish into Thin Air: Cross-prompt Universal Adversarial Attacks for SAM2

NeurIPS 2025spotlight

Recent studies reveal the vulnerability of the image segmentation foundation model SAM to adversarial examples. Its successor, SAM2, has attracted significant attention due to its strong generalization capability in video segmentation. However, its robustness remains unexplored, and it is unclear wh…

Cited by 0SourceScholar
2024

Contextual Bilevel Reinforcement Learning for Incentive Alignment

NeurIPS 2024poster

The optimal policy in various real-world strategic decision-making problems depends both on the environmental configuration and exogenous events. For these settings, we introduce Contextual Bilevel Reinforcement Learning (CB-RL), a stochastic bilevel decision-making model, where the lower level cons…

2024

Distributionally Robust Model-based Reinforcement Learning with Large State Spaces

AISTATS 2024poster

Three major challenges in reinforcement learning are the complex dynamical systems with large state spaces, the costly data acquisition processes, and the deviation of real-world dynamics from the training environment deployment. To overcome these issues, we study distributionally robust Markov deci…

2024

Emotion Rendering for Conversational Speech Synthesis with Heterogeneous Graph-Based Context Modeling

AAAI 2024technical

Conversational Speech Synthesis (CSS) aims to accurately express an utterance with the appropriate prosody and emotional inflection within a conversational setting. While recognising the significance of CSS task, the prior studies have not thoroughly investigated the emotional expressiveness problem…

2024

Generalization Bounds of Nonconvex-(Strongly)-Concave Stochastic Minimax Optimization

AISTATS 2024poster

This paper studies the generalization performance of algorithms for solving nonconvex-(strongly)-concave (NC-SC/NC-C) stochastic minimax optimization measured by the stationarity of primal functions. We first establish algorithm-agnostic generalization bounds via uniform convergence between the empi…

Cited by 5SourcePDFScholar
2024

Group Robust Preference Optimization in Reward-free RLHF

NeurIPS 2024poster

Adapting large language models (LLMs) for specific tasks usually involves fine-tuning through reinforcement learning with human feedback (RLHF) on preference data. While these data often come from diverse labelers' groups (e.g., different demographics, ethnicities, company teams, etc.), traditional…

Cited by 20SourcePDFScholar
2024

Stochastic Optimization Algorithms for Instrumental Variable Regression with Streaming Data

NeurIPS 2024poster

We develop and analyze algorithms for instrumental variable regression by viewing the problem as a conditional stochastic optimization problem. In the context of least-squares instrumental variable regression, our algorithms neither require matrix inversions nor mini-batches thereby providing a full…

Cited by 3SourcePDFScholar
2023

Deep Directly-Trained Spiking Neural Networks for Object Detection

ICCV 2023poster

Spiking neural networks (SNNs) are brain-inspired energy-efficient models that encode information in spatiotemporal dynamics. Recently, deep SNNs trained directly have shown great success in achieving high performance on classification tasks with very few time steps. However, how to design a directl…

Cited by 102PDFcodeScholar
2023

Multi-Modal Food Classification in a Diet Tracking System with Spoken and Visual Inputs

ICASSP 2023accepted

In this paper, we present multi-modal approaches to diet tracking. As health and well-being become increasingly important, mobile applications for diet tracking attract much interest. However, these applications often require users to log their meals based on relatively unreliable memory recall, the…

Cited by 0SourceScholar
2022

Perturbations in the Wild: Leveraging Human-Written Text Perturbations for Realistic Adversarial Attack and Defense

ACL 2022findings

We proposes a novel algorithm, ANTHRO, that inductively extracts over 600K human-written text perturbations in the wild and leverages them for realistic adversarial attack. Unlike existing character-based attacks which often deductively hypothesize a set of manipulation strategies, our work is groun…

2021

Going Deeper With Directly-Trained Larger Spiking Neural Networks

AAAI 2021technical

Spiking neural networks (SNNs) are promising in a bio-plausible coding for spatio-temporal information and event-driven signal processing, which is very suited for energy-efficient implementation in neuromorphic hardware. However, the unique working mode of SNNs makes them more difficult to train th…

Cited by 590SourcePDFScholar
2020

Biased Stochastic First-Order Methods for Conditional Stochastic Optimization and Applications in Meta Learning

NeurIPS 2020poster

Conditional stochastic optimization covers a variety of applications ranging from invariant learning and causal inference to meta-learning. However, constructing unbiased gradient estimators for such problems is challenging due to the composition structure. As an alternative, we propose a biased sto…

Cited by 73SourcePDFScholar