← Search

Yi He

37 accepted papers

2026

CARD: Coarse-to-fine Autoregressive Modeling with Radix-based Decomposition for Transferable Free Energy Estimation

ICML 2026poster

Estimating free energy differences quantifies thermodynamic preferences in molecular interactions, which is central to chemistry and drug discovery. Despite fruitful progress, existing methods still face key limitations: classical computational approaches remain prohibitively expensive due to their …

Cited by 0SourceScholar
2026

Enhancing the Security of Visual Speaker Authentication Based on Dynamic Lip-Print Analysis

CVPR 2026

In recent years, face-based authentication methods are gradually replacing traditional methods across various applications, offering enhanced security and user convenience. However, these methods are threatened by the continuously evolving DeepFake techniques. In this paper, a novel Visual Speaker A

Cited by 0SourceScholar
2026

FinMathBench: A Formula-Driven Benchmark for Evaluating LLMs’ Math Reasoning Capabilities in Finance

AAAI 2026technical

Many existing financial math reasoning benchmarks suffer from data contamination and high manual construction costs. To address this, we propose a novel formula-driven approach to dynamically construct math reasoning benchmarks in finance. Our two-stage approach: (1) generates single-formula questio

Cited by 0SourcePDFScholar
2026

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

AAAI 2026technical

Existing autoregressive large-scale text-to-speech (TTS) models have advantages in speech naturalness, but their token-by-token generation mechanism makes it difficult to precisely control the duration of synthesized speech. This becomes a significant limitation in applications requiring strict audi

Cited by 0SourcePDFScholar
2026

Information Shapes Koopman Representation

ICLR 2026oral

The Koopman operator provides a powerful framework for modeling dynamical systems and has attracted growing interest from the machine learning community. However, its infinite-dimensional nature makes identifying suitable finite-dimensional subspaces challenging, especially for deep architectures. W…

Cited by 0SourcecodeScholar
2026

MARLIN: Multi-Agent Reinforcement Learning for Incremental DAG Discovery

AAAI 2026technical

Uncovering causal structures from observational data is crucial for understanding complex systems and making informed decisions. While reinforcement learning (RL) has shown promise in identifying these structures in the form of a directed acyclic graph (DAG), existing methods often lack efficiency,

Cited by 0SourcePDFScholar
2026

MMPD-Bench: Bridging Multimodal Fission with Multi-Polarimetric Modalities Decomposition

ICML 2026poster

Recovering multiple physical parameters from high-dimensional optical measurements remains challenging in computational optics. We present *MMPD-Bench*, a pioneering benchmark that reframes multi-polarimetric modalities decomposition from Mueller matrix observations as a *modality fission* problem u…

Cited by 0SourceScholar
2026

Relink: Constructing Query-Driven Evidence Graph On-the-Fly for GraphRAG

AAAI 2026technical

Graph-based Retrieval-Augmented Generation (GraphRAG) mitigates hallucinations in Large Language Models (LLMs) by grounding them in structured knowledge. However, current GraphRAG methods are constrained by a prevailing build-then-reason paradigm, which relies on a static, pre-constructed Knowledge

Cited by 0SourcePDFScholar
2025

A Novel Sparse Active Online Learning Framework for Fast and Accurate Streaming Anomaly Detection Over Data Streams

IJCAI 2025

Online Anomaly Detection (OAD) is critical for identifying rare yet important data points in large, dynamic, and complex data streams. A key challenge lies in achieving accurate and consistent detection of anomalies while maintaining computational and memory efficiency. Conventional OAD approaches,

Cited by 0SourcePDFScholar
2025

Chaos Meets Attention: Transformers for Large-Scale Dynamical Prediction

ICML 2025poster

Generating long-term trajectories of dissipative chaotic systems autoregressively is a highly challenging task. The inherent positive Lyapunov exponents amplify prediction errors over time. Many chaotic systems possess a crucial property — ergodicity on their attractors, which makes long-term predic…

2025

EmotiCrafter: Text-to-Emotional-Image Generation based on Valence-Arousal Model

ICCV 2025poster

Recent research shows that emotions can enhance users' cognition and influence information communication. While research on visual emotion analysis is extensive, limited work has been done on helping users generate emotionally rich image content. Existing work on emotional image generation relies on…

2025

Enhancing Visual Forced Alignment with Local Context-Aware Feature Extraction and Multi-Task Learning

ICASSP 2025accepted

This paper introduces a novel approach to Visual Forced Alignment (VFA), aiming to accurately synchronize utterances with corresponding lip movements, without relying on audio cues. We propose a novel VFA approach that integrates a local context-aware feature extractor and employs multitask learning…

Cited by 0SourceScholar
2025

How to Mitigate Information Loss in Knowledge Graphs for GraphRAG: Leveraging Triple Context Restoration and Query-Driven Feedback

IJCAI 2025

Knowledge Graph (KG)-augmented Large Language Models (LLMs) have recently propelled significant advances in complex reasoning tasks, thanks to their broad domain knowledge and contextual awareness. Unfortunately, current methods often assume KGs to be complete, which is impractical given the inheren

2025

Metric-Agnostic Continual Learning for Sustainable Group Fairness

AAAI 2025technical

Group Fairness-aware Continual Learning (GFCL) aims to eradicate discriminatory predictions against certain demographic groups in a sequence of diverse learning tasks. This paper explores an even more challenging GFCL problem – how to sustain a fair classifier across a sequence of tasks with covaria…

2025

Query-Driven Multimodal GraphRAG: Dynamic Local Knowledge Graph Construction for Online Reasoning

ACL 2025finding

An increasing adoption of Large Language Models (LLMs) in complex reasoning tasks necessitates their interpretability and reliability. Recent advances to that end include retrieval-augmented generation (RAG) and knowledge graph-enhanced RAG (GraphRAG), whereas they are constrained by static knowledg…

Cited by 0SourcePDFScholar
2025

Tensor-Var: Efficient Four-Dimensional Variational Data Assimilation

ICML 2025poster

Variational data assimilation estimates the dynamical system states by minimizing a cost function that fits the numerical models with the observational data. Although four-dimensional variational assimilation (4D-Var) is widely used, it faces high computational costs in complex nonlinear systems and…

Cited by 0SourcePDFScholar
2024

MKG-FENN: A Multimodal Knowledge Graph Fused End-to-End Neural Network for Accurate Drug–Drug Interaction Prediction

AAAI 2024technical

Taking incompatible multiple drugs together may cause adverse interactions and side effects on the body. Accurate prediction of drug-drug interaction (DDI) events is essential for avoiding this issue. Recently, various artificial intelligence-based approaches have been proposed for predicting DDI ev…

2024

Speaker-Adaptive Lipreading Via Spatio-Temporal Information Learning

ICASSP 2024accepted

Lipreading has been rapidly developed recently with the help of large-scale datasets and large models. Despite the significant progress made, the performance of lipreading models still falls short when dealing with unseen speakers. Therefore, it is necessary to utilize the speaker’s videos for fine-…

Cited by 0SourceScholar
2023

Exploring Hypergraph of Earnings Call for Risk Prediction (Student Abstract)

AAAI 2023technical

In financial economics, studies have shown that the textual content in the earnings conference call transcript has predictive power for a firm's future risk. However, the conference call transcript is very long and contains diverse non-relevant content, which poses challenges for the text-based risk…

Cited by 2SourcePDFScholar
2023

Online Random Feature Forests for Learning in Varying Feature Spaces

AAAI 2023technical

In this paper, we propose a new online learning algorithm tailored for data streams described by varying feature spaces (VFS), wherein new features constantly emerge and old features may stop to be observed over various time spans. Our proposed algorithm, named Online Random Feature Forests for Feat…

Cited by 14SourcePDFScholar
2023

Online Semi-supervised Learning with Mix-Typed Streaming Features

AAAI 2023technical

Online learning with feature spaces that are not fixed but can vary over time renders a seemingly flexible learning paradigm thus has drawn much attention. Unfortunately, two restrictions prohibit a ubiquitous application of this learning paradigm in practice. First, whereas prior studies mainly ass…

2023

Reducing Language Confusion for Code-Switching Speech Recognition with Token-Level Language Diarization

ICASSP 2023accepted

Code-switching (CS) occurs when languages switch within a speech signal and leads to language confusion for automatic speech recognition (ASR). We address the problem of language confusion for improving CS-ASR from two perspectives: incorporating and disentangling language information. We incorporat…

Cited by 0SourceScholar
2023

Towards Utilitarian Online Learning -- A Review of Online Algorithms in Open Feature Space

IJCAI 2023poster

Human intelligence comes from the capability to describe and make sense of the world surrounding us, often in a lifelong manner. Online Learning (OL) allows a model to simulate this capability, which involves processing data in sequence, making predictions, and learning from predictive errors. Howev…

Cited by 7SourcePDFScholar
2021

Online Learning in Variable Feature Spaces under Incomplete Supervision

AAAI 2021technical

This paper explores a new online learning problem where the input sequence lives in an over-time varying feature space and the ground-truth label of any input point is given only occasionally, making online learners less restrictive and more applicable. The crux in this setting lies in how to exploi…

Cited by 34SourcePDFScholar
2019

SVD: A Large-Scale Short Video Dataset for Near-Duplicate Video Retrieval

ICCV 2019poster

With the explosive growth of video data in real applications, near-duplicate video retrieval (NDVR) has become indispensable and challenging, especially for short videos. However, all existing NDVR datasets are introduced for long videos. Furthermore, most of them are small-scale and lack of diversi…

Cited by 61PDFcodeScholar
2019

Semi-Supervised Skin Detection by Network With Mutual Guidance

ICCV 2019poster

We present a new data-driven method for robust skin detection from a single human portrait image. Unlike previous methods, we incorporate human body as a weak semantic guidance into this task, considering acquiring large-scale of human labeled skin data is commonly expensive and time-consuming. To b…

Cited by 32PDFScholar