← Search

Jack Xin

6 accepted papers

2026

Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark

ICML 2026poster

Grading in large undergraduate STEM courses often yields minimal feedback due to heavy instructional workloads. We present a large-scale empirical study of AI grading on real, handwritten single-variable calculus work from a major U.S. public research university. Using OCR-conditioned large language…

Cited by 0SourceScholar
2026

SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging

ICML 2026poster

Attention is the critical component of a transformer. Yet the quadratic computational complexity of vanilla full attention in the input size and the inability of its linear attention variant to focus have been challenges for computer vision tasks. We provide a mathematical definition of generalized …

Cited by 0SourceScholar
2025

Global Well-posedness and Convergence Analysis of Score-based Generative Models via Sharp Lipschitz Estimates

ICLR 2025poster

We establish global well-posedness and convergence of the score-based generative models (SGM) under minimal general assumptions of initial data for score estimation. For the smooth case, we start from a Lipschitz bound of the score function with optimal time length. The optimality is validated by an…

Cited by 1SourcePDFScholar
2024

Rethinking the Benefits of Steerable Features in 3D Equivariant Graph Neural Networks

ICLR 2024poster

Theoretical and empirical comparisons have been made to assess the expressive power and performance of invariant and equivariant GNNs. However, there is currently no theoretical result comparing the expressive power of $k$-hop invariant GNNs and equivariant GNNs. Additionally, little is understood a…

Cited by 7SourcePDFScholar
2022

Glassoformer: A Query-Sparse Transformer for Post-Fault Power Grid Voltage Prediction

ICASSP 2022accepted

We propose GLassoformer, a novel and efficient transformer architecture leveraging group Lasso regularization to reduce the number of queries of the standard self-attention mechanism. Due to the sparsified queries, GLassoformer is more computationally efficient than the standard transformers. On the…

Cited by 0SourceScholar
2019

Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets

ICLR 2019poster

Training activation quantized neural networks involves minimizing a piecewise constant training loss whose gradient vanishes almost everywhere, which is undesirable for the standard back-propagation or chain rule. An empirical way around this issue is to use a straight-through estimator (STE) (Bengi…

Cited by 382SourcePDFScholar