← Search

Su Lu

6 accepted papers

2025

Maximizing the Effectiveness of Larger BERT Models for Compression

ACL 2025long

Knowledge distillation (KD) is a widely used approach for BERT compression, where a larger BERT model serves as a teacher to transfer knowledge to a smaller student model. Prior works have found that distilling a larger BERT with superior performance may degrade student’s performance than a smaller…

2024

Revisit the Essence of Distilling Knowledge through Calibration

ICML 2024poster

Knowledge Distillation (KD) has evolved into a practical technology for transferring knowledge from a well-performing model (teacher) to a weak model (student). A counter-intuitive phenomenon known as capacity mismatch has been identified, wherein KD performance may not be good when a better teacher…

Cited by 1SourcePDFScholar
2021

Dual-Arm Needle Manipulation with the da Vinci® Surgical Robot Under Uncertainty

ICRA 2021poster

This paper proposes a path correction method for surgical robotic systems performing needle handoff manipulations as part of autonomous execution of surgical suturing. During handoff motions, the position and orientation of the needle is subject to perturbations from the idealized planned pose due t…

Cited by 2SourceScholar
2021

Tailoring Embedding Function to Heterogeneous Few-Shot Tasks by Global and Local Feature Adaptors

AAAI 2021technical

Few-Shot Learning (FSL) is essential for visual recognition. Many methods tackle this challenging problem via learning an embedding function from seen classes and transfer it to unseen classes with a few labeled instances. Researchers recently found it beneficial to incorporate task-specific feature…

Cited by 29SourcePDFScholar