IJCAI 2022poster69 citations

Data Valuation in Machine Learning: "Ingredients", Strategies, and Open Challenges

Rachael Hwee Ling Sim, Xinyi Xu, Bryan Kian Hsiang Low

Abstract

Data valuation in machine learning (ML) is an emerging research area that studies the worth of data in ML. Data valuation is used in collaborative ML to determine a fair compensation for every data owner and in interpretable ML to identify the most responsible, noisy, or misleading training examples. This paper presents a comprehensive technical survey that provides a new formal study of data valuation in ML through its “ingredients” and the corresponding properties, grounds the discussion of common desiderata satisfied by existing data valuation strategies on our proposed ingredients, and identifies open research challenges for designing new ingredients, data valuation strategies, and cost reduction techniques.

Survey Track: Machine LearningSurvey Track: AI Ethics, Trust, FairnessSurvey Track: Multidisciplinary Topics and Applications
BibTeX
@inproceedings{ijcai2022p782,
  title     = {Data Valuation in Machine Learning: "Ingredients", Strategies, and Open Challenges},
  author    = {Sim, Rachael Hwee Ling and Xu, Xinyi and Low, Bryan Kian Hsiang},
  booktitle = {Proceedings of the Thirty-First International Joint Conference on
               Artificial Intelligence, {IJCAI-22}},
  publisher = {International Joint Conferences on Artificial Intelligence Organization},
  editor    = {Lud De Raedt},
  pages     = {5607--5614},
  year      = {2022},
  month     = {7},
  note      = {Survey Track},
  doi       = {10.24963/ijcai.2022/782},
  url       = {https://doi.org/10.24963/ijcai.2022/782},
}
Data Valuation in Machine Learning: "Ingredients", Strategies, and Open Challenges · IJCAI 2022