← Search

Anh T Nguyen

2 accepted papers

2026

Push, Pop, Parallelize: Stack-Augmented Linear Attention via the Delta Rule

ICML 2026poster

Linear attention architectures based on the Delta rule, such as DeltaNet and RWKV-7, combine Transformers' training scalability with RNNs' inference efficiency and can provably solve regular language tasks. However, due to their fixed-size state, these models fundamentally struggle to capture the re…

Cited by 0SourceScholar
2025

CASUAL: Conditional Support Alignment for Domain Adaptation with Label Shift

AAAI 2025technical

Unsupervised domain adaptation (UDA) refers to a domain adaptation framework in which a learning model is trained based on the labeled samples on the source domain and unlabelled ones in the target domain. The dominant existing methods in the field that rely on the classical covariate shift assumpti…

Cited by 0SourcePDFScholar