← Search

Kai Zhong

16 accepted papers

2026

SFT Doesn’t Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs

ICLR 2026poster

Supervised Fine-Tuning (SFT) on domain-specific datasets is a common approach to adapt Large Language Models (LLMs) to specialized tasks but is often believed to degrade their general capabilities. In this work, we revisit this trade-off and present both empirical and theoretical insights. First, we…

Cited by 0SourceScholar
2025

Secure Analog Beamforming Design for Wireless Communication Systems With Movable Antennas

ICASSP 2025accepted

Movable antennas (MA) allow flexible positioning within a specified region, enhancing wireless communication performance. This paper explores leveraging MA to improve physical layer security in analog beamforming (AB) systems. Specifically, we aim to maximize the secrecy rate by jointly optimizing t…

Cited by 0SourceScholar
2023

ReAugKD: Retrieval-Augmented Knowledge Distillation For Pre-trained Language Models

ACL 2023short

Knowledge Distillation (KD) is one of the most effective approaches to deploying large-scale pre-trained language models in low-latency environments by transferring the knowledge contained in the large-scale models to smaller student models. Prior KD approaches use the soft labels and intermediate a…

Cited by 23SourcePDFScholar
2018

Binary Classification with Karmic, Threshold-Quasi-Concave Metrics

ICML 2018oral

Complex performance measures, beyond the popular measure of accuracy, are increasingly being used in the context of binary classification. These complex performance measures are typically not even decomposable, that is, the loss evaluated on a batch of samples cannot typically be expressed as a sum…

Cited by 36SourcePDFScholar
2018

MixLasso: Generalized Mixed Regression via Convex Atomic-Norm Regularization

NeurIPS 2018poster

We consider a generalization of mixed regression where the response is an additive combination of several mixture components. Standard mixed regression is a special case where each response is generated from exactly one component. Typical approaches to the mixture regression problem employ local sea…

Cited by 3SourcePDFScholar
2017

Fast Classification with Binary Prototypes

AISTATS 2017poster

In this work, we propose a new technique for \emphfast k-nearest neighbor (k-NN) classification in which the original database is represented via a small set of learned binary prototypes. The training phase simultaneously learns a hash function which maps the data points to binary codes, and a set o…

2016

Dual Decomposed Learning with Factorwise Oracle for Structural SVM of Large Output Domain

NeurIPS 2016poster

Many applications of machine learning involve structured output with large domain, where learning of structured predictor is prohibitive due to repetitive calls to expensive inference oracle. In this work, we show that, by decomposing training of Structural Support Vector Machine (SVM) into a series…

Cited by 10SourcePDFScholar
2016

PD-Sparse : A Primal and Dual Sparse Approach to Extreme Multiclass and Multilabel Classification

ICML 2016poster

We consider Multiclass and Multilabel classification with extremely large number of classes, of which only few are labeled to each instance. In such setting, standard methods that have training, prediction cost linear to the number of classes become intractable. State-of-the-art methods thus aim to…

Cited by 233SourcePDFScholar
2015

A Convex Exemplar-based Approach to MAD-Bayes Dirichlet Process Mixture Models

ICML 2015poster

MAD-Bayes (MAP-based Asymptotic Derivations) has been recently proposed as a general technique to derive scalable algorithm for Bayesian Nonparametric models. However, the combinatorial nature of objective functions derived from MAD-Bayes results in hard optimization problem, for which current pract…

Cited by 8SourcePDFScholar
2015

Sparse Linear Programming via Primal and Dual Augmented Coordinate Descent

NeurIPS 2015poster

Over the past decades, Linear Programming (LP) has been widely used in different areas and considered as one of the mature technologies in numerical optimization. However, the complexity offered by state-of-the-art algorithms (i.e. interior-point method and primal, dual simplex methods) is still uns…

Cited by 40SourcePDFScholar