← Search

Wang-Cheng Kang

4 accepted papers

2026

How to train data-efficient LLMs

ICLR 2026poster

The training of large language models (LLMs) is expensive. In this paper, we study data-efficient approaches for pre-training LLMs, \ie, techniques that aim to optimize the Pareto frontier of model quality and training resource/data consumption. We seek to understand the tradeoffs associated with da…

Cited by 0SourceScholar
2025

ActionPiece: Contextually Tokenizing Action Sequences for Generative Recommendation

ICML 2025spotlight

Generative recommendation (GR) is an emerging paradigm where user actions are tokenized into discrete token patterns and autoregressively generated as predictions. However, existing GR models tokenize each action independently, assigning the same fixed tokens to identical actions across all sequence…

2023

Unified Embedding: Battle-Tested Feature Representations for Web-Scale ML Systems

NeurIPS 2023spotlight

Learning high-quality feature embeddings efficiently and effectively is critical for the performance of web-scale machine learning systems. A typical model ingests hundreds of features with vocabularies on the order of millions to billions of tokens. The standard approach is to represent each featur…

Cited by 11SourcePDFScholar
2019

Complete the Look: Scene-Based Complementary Product Recommendation

CVPR 2019poster

Modeling fashion compatibility is challenging due to its complexity and subjectivity. Existing work focuses on predicting compatibility between product images (e.g. an image containing a t-shirt and an image containing a pair of jeans). However, these approaches ignore real-world 'scene' images (e.g…

Cited by 96PDFcodeScholar