← Search

Swanand Kadhe

3 accepted papers

2026

Entropy-Aware On-Policy Distillation of Language Models

ICML 2026poster

On-policy distillation is a promising approach for transferring knowledge between language models, where a student learns from dense token-level signals along its own trajectories. This framework typically uses reverse KL divergence, encouraging the student to match the teacher's high-confidence pre…

Cited by 0SourceScholar
2023

LESS-VFL: Communication-Efficient Feature Selection for Vertical Federated Learning

ICML 2023poster

We propose LESS-VFL, a communication-efficient feature selection method for distributed systems with vertically partitioned data. We consider a system of a server and several parties with local datasets that share a sample ID space but have different feature sets. The parties wish to collaboratively…

Cited by 30SourcePDFScholar
2021

Leveraging Spatial and Temporal Correlations in Sparsified Mean Estimation

NeurIPS 2021poster

We study the problem of estimating at a central server the mean of a set of vectors distributed across several nodes (one vector per node). When the vectors are high-dimensional, the communication cost of sending entire vectors may be prohibitive, and it may be imperative for them to use sparsificat…

Cited by 18SourcePDFScholar