← Search

Shunta Akiyama

9 accepted papers

2026

A Strictly Proper Scoring Rule and a Calibration Metric for Interval-Censored Data Analysis

ICML 2026poster

Interval-censored data present unique challenges in statistical analysis due to the partial observability of event times within known intervals, requiring assumptions about the censoring mechanism. This paper explores the theoretical relationship between two foundational assumptions: independent mon…

Cited by 0SourceScholar
2026

Why Agentic Theorem Prover Works: A Statistical Provability Theory of Mathematical Reasoning Models

ICML 2026poster

Agentic theorem provers---pipelines that couple a mathematical reasoning model with library retrieval, decomposition/search, and a proof assistant verifier---have recently achieved striking empirical success, yet it remains unclear which components drive performance and why such systems work at all …

Cited by 0SourceScholar
2024

SILVER: Single-loop variance reduction and application to federated learning

ICML 2024poster

Most variance reduction methods require multiple times of full gradient computation, which is time-consuming and hence a bottleneck in application to distributed optimization. We present a single-loop variance-reduced gradient estimator named SILVER (SIngle-Loop VariancE-Reduction) for the finite-su…

Cited by 0SourcePDFScholar
2023

Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods

ICLR 2023poster

While deep learning has outperformed other methods for various tasks, theoretical frameworks that explain its reason have not been fully established. We investigate the excess risk of two-layer ReLU neural networks in a teacher-student regression model, in which a student network learns an unknown t…

Cited by 10SourcePDFScholar
2021

Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods

ICLR 2021spotlight

Establishing a theoretical analysis that explains why deep learning can outperform shallow learning such as kernel methods is one of the biggest issues in the deep learning literature. Towards answering this question, we evaluate excess risk of a deep learning estimator trained by a noisy gradient d…

Cited by 22SourcePDFScholar
2021

On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting

ICML 2021spotlight

Deep learning empirically achieves high performance in many applications, but its training dynamics has not been fully understood theoretically. In this paper, we explore theoretical analysis on training two-layer ReLU neural networks in a teacher-student regression model, in which a student network…

Cited by 16SourcePDFScholar