← Search

Yanna Ding

4 accepted papers

2025

Architecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential Equation

AAAI 2025technical

Learning curve extrapolation predicts neural network performance from early training epochs and has been applied to accelerate AutoML, facilitating hyperparameter tuning and neural architecture search. However, existing methods typically model the evolution of learning curves in isolation, neglectin…

2025

Epigraph Based Multilevel Optimization (EMO) for Enhancing Chain-of-Thought Reasoning Capabilities

ICASSP 2025accepted

Chain-of-thought (CoT) reasoning applies to complex tasks with multiple intermediate steps, a key feature of large language models. Recent studies have revealed CoT as a composition of in-context filtering and learning. This paper proposes a unified framework for CoT optimization that exploits the n…

Cited by 0SourceScholar
2025

Inferring from Logits: Exploring Best Practices for Decoding-Free Generative Candidate Selection

ACL 2025long

Generative Language Models rely on autoregressive decoding to produce the output sequence token by token. Many tasks such as preference optimization, require the model to produce task-level output consisting of multiple tokens directly by selecting candidates from a pool as predictions. Determining…

Cited by 0SourcePDFScholar
2025

Optimality and NP-Hardness of Transformers in Learning Markovian Dynamical Functions

NeurIPS 2025poster

Transformer architectures can solve unseen tasks based on input-output pairs in a given prompt due to in-context learning (ICL). Existing theoretical studies on ICL have mainly focused on linear regression tasks, often with i.i.d. inputs. To understand how transformers express in-context learning wh…

Cited by 0SourceScholar