← Search

Muquan Li

6 accepted papers

2026

Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset Distillation

CVPR 2026

Decoupled dataset distillation (DD) compresses large corpora into a few synthetic images by matching a frozen teacher's statistics. However, current residual-matching pipelines rely on static real patches, creating a fit-complexity gap and a pull-to-anchor effect that reduce intra-class diversity an

Cited by 0SourceScholar
2026

Mind Your Margin and Boundary: Are Your Distilled Datasets Truly Robust?

ICML 2026oral

Dataset distillation (DD) compresses a large training set into a small synthetic set for efficient training, but most DD methods optimize only clean accuracy and leave robustness uncontrolled. Recent robust DD methods improve robustness, yet they often suffer from a poor accuracy–robustness trade-of…

Cited by 0SourceScholar
2026

PADA-Coder: Improving Plan-Following Code Generation via Perturbation-Verified Attention Distillation and Dynamic Alignment

ICML 2026poster

The Plan-then-Code paradigm effectively enhances Large Language Models (LLMs) in complex code generation by decomposing reasoning into explicit, interpretable steps. However, introducing the plan and verification report substantially enlarges the context, which in turn misdirects the model’s attenti…

Cited by 0SourceScholar
2026

Stop When Further Reasoning Won’t Help: Attention-State Adaptive Generation in Reasoning Models

ICML 2026spotlight

By incorporating test-time compute scaling, large reasoning models (LRMs) are able to solve complex problems by generating explicit chain-of-thought (CoT) reasoning processes. However, they often suffer from overthinking during generation, resulting in redundant token outputs and degraded accuracy. …

Cited by 0SourceScholar
2025

Beyond Random: Automatic Inner-loop Optimization in Dataset Distillation

NeurIPS 2025poster

The growing demand for efficient deep learning has positioned dataset distillation as a pivotal technique for compressing training dataset while preserving model performance. However, existing inner-loop optimization methods for dataset distillation typically rely on random truncation strategies, wh…

Cited by 0SourceScholar