← Search

Rares Dolga

3 accepted papers

2026

wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models

ICLR 2026poster

Improving the reasoning capabilities of diffusion-based large language models (dLLMs) through reinforcement learning (RL) remains an open problem. The intractability of dLLMs likelihood function necessitates approximating the current, old, and reference policy likelihoods at each policy optimization…

Cited by 62SourceScholar
2025

From Characters to Tokens: Dynamic Grouping with Hierarchical BPE

EMNLP 2025

Subword tokenization methods like Byte Pair Encoding (BPE) are widely used in large language models due to their balance of vocabulary compactness and representational power. However, they suffer from inefficiencies in representing rare words and require large embedding matrices. Character-level mod

Cited by 0SourcePDFScholar
2025

Incremental Sequence Classification with Temporal Consistency

NeurIPS 2025spotlight

We address the problem of incremental sequence classification, where predictions are updated as new elements in the sequence are revealed. Drawing on temporal-difference learning from reinforcement learning, we identify a temporal-consistency condition that successive predictions should satisfy. We…

Cited by 0SourceScholar