← Search

Yasuyuki Okoshi

5 accepted papers

2026

The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms

AAAI 2026technical

The strong lottery ticket hypothesis (SLTH) conjectures that high-performing subnetworks, called strong lottery tickets (SLTs), are hidden in randomly initialized neural networks. Although recent theoretical studies have established the SLTH across various neural architectures, the SLTH for transfor

Cited by 0SourcePDFScholar
2026

Towards Quantization-Aware Training for Ultra-Low-Bit Reasoning LLMs

ICLR 2026poster

Large language models (LLMs) have achieved remarkable performance across diverse reasoning tasks, yet their deployment is hindered by prohibitive computational and memory costs. Quantization-aware training (QAT) enables ultra-low-bit compression (<4 bits per weight), but existing QAT methods often d…

Cited by 0SourcecodeScholar
2025

Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression

NeurIPS 2025poster

This paper proposes a novel matrix quantization method, Binary Quadratic Quan- tization (BQQ). In contrast to conventional first-order quantization approaches— such as uniform quantization and binary coding quantization—that approximate real-valued matrices via linear combinations of binary bases, B…

Cited by 0SourceScholar
2025

Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling

NeurIPS 2025poster

Test-time scaling (TTS) has proven effective in enhancing the reasoning capabilities of large language models (LLMs). Verification plays a key role in TTS, simultaneously influencing (1) reasoning performance and (2) compute efficiency, due to the quality and computational cost of verification. In t…

Cited by 0SourceScholar
2022

Multicoated Supermasks Enhance Hidden Networks

ICML 2022spotlight

Hidden Networks (Ramanujan et al., 2020) showed the possibility of finding accurate subnetworks within a randomly weighted neural network by training a connectivity mask, referred to as supermask. We show that the supermask stops improving even though gradients are not zero, thus underutilizing back…

Cited by 13SourcePDFScholar