← Search

Hanyu Peng

6 accepted papers

2023

Dataset Pruning: Reducing Training Data by Examining Generalization Influence

ICLR 2023poster

The great success of deep learning heavily relies on increasingly larger training data, which comes at a price of huge computational and infrastructural costs. This poses crucial questions that, do all training data contribute to model's performance? How much does each individual training sample or…

Cited by 135SourcePDFScholar
2023

MetaTPTrans: A Meta Learning Approach for Multilingual Code Representation Learning

AAAI 2023technical

Representation learning of source code is essential for applying machine learning to software engineering tasks. Learning code representation from a multilingual source code dataset has been shown to be more effective than learning from single-language datasets separately, since more training data f…

2022

Cross-Lingual Cross-Modal Consolidation for Effective Multilingual Video Corpus Moment Retrieval

NAACL 2022findings

Existing multilingual video corpus moment retrieval (mVCMR) methods are mainly based on a two-stream structure. The visual stream utilizes the visual content in the video to estimate the query-visual similarity, and the subtitle stream exploits the query-subtitle similarity. The final query-video si…

Cited by 23SourcePDFScholar