← Search

Satoki Ishikawa

4 accepted papers

2026

Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks

ICLR 2026oral

Empirical scaling laws have driven the evolution of large language models (LLMs), yet their coefficients shift whenever the model architecture or data pipeline changes. Mixture‑of‑Experts (MoE) models, now standard in state‑of‑the‑art systems, introduce a new sparsity dimension that current dense‑mo…

Cited by 0SourcecodeScholar
2025

Local Loss Optimization in the Infinite Width: Stable Parameterization of Predictive Coding Networks and Target Propagation

ICLR 2025poster

Local learning, which trains a network through layer-wise local targets and losses, has been studied as an alternative to backpropagation (BP) in neural computation. However, its algorithms often become more complex or require additional hyperparameters due to the locality, making it challenging to…

Cited by 0SourcePDFScholar
2025

PhiNets: Brain-inspired Non-contrastive Learning Based on Temporal Prediction Hypothesis

ICLR 2025poster

Predictive coding has been established as a promising neuroscientific theory to describe the mechanism of information processing in the retina or cortex. This theory hypothesises that cortex predicts sensory inputs at various levels of abstraction to minimise prediction errors. Inspired by predict…

Cited by 1SourcePDFScholar
2024

On the Parameterization of Second-Order Optimization Effective towards the Infinite Width

ICLR 2024poster

Second-order optimization has been developed to accelerate the training of deep neural networks and it is being applied to increasingly larger-scale models. In this study, towards training on further larger scales, we identify a specific parameterization for second-order optimization that promotes f…

Cited by 5SourcePDFScholar