← Search

David Dao

6 accepted papers

2026

Field Deployment of BiodivX Drones in the Amazon Rainforest for Biodiversity Monitoring (I)

ICRA 2026poster

Tropical rainforests are among the most biodiverse ecosystems on Earth and also among the most threatened by anthropogenic pressures such as deforestation and climate change. Understanding human impact and the efficacy of conservation and preservation efforts requires scalable and comprehensive biod…

Cited by 0Scholar
2024

Data Debugging with Shapley Importance over Machine Learning Pipelines

ICLR 2024poster

When a machine learning (ML) model exhibits poor quality (e.g., poor accuracy or fairness), the problem can often be traced back to errors in the training data. Being able to discover the data examples that are the most likely culprits is a fundamental concern that has received a lot of attention re…

2024

OAM-TCD: A globally diverse dataset of high-resolution tree cover maps

NeurIPS 2024poster

Accurately quantifying tree cover is an important metric for ecosystem monitoring and for assessing progress in restored sites. Recent works have shown that deep learning-based segmentation algorithms are capable of accurately mapping trees at country and continental scales using high-resolution aer…

2023

GEO-Bench: Toward Foundation Models for Earth Monitoring

NeurIPS 2023poster

Recent progress in self-supervision has shown that pre-training large neural networks on vast amounts of unsupervised data can lead to substantial increases in generalization to downstream tasks. Such models, recently coined foundation models, have been transformational to the field of natural lang…

2021

Scalability vs. Utility: Do We Have To Sacrifice One for the Other in Data Importance Quantification?

CVPR 2021poster

Quantifying the importance of each training point to a learning task is a fundamental problem in machine learning and the estimated importance scores have been leveraged to guide a range of data workflows such as data summarization and domain adaption. One simple idea is to use the leave-one-out err…

Cited by 79PDFcodeScholar
2019

Towards Efficient Data Valuation Based on the Shapley Value

AISTATS 2019poster

{\em “How much is my data worth?”} is an increasingly common question posed by organizations and individuals alike. An answer to this question could allow, for instance, fairly distributing profits among multiple data contributors and determining prospective compensation when data breaches happen. I…

Cited by 570SourcePDFScholar