← Search

Xinyi Xu

21 accepted papers

2026

HiTeA: Hierarchical Temporal Alignment for Training-Free Long-Video Temporal Grounding

ICLR 2026poster

Temporal grounding in long, untrimmed videos is critical for real-world video understanding, yet it remains a challenging task owing to complex temporal structures and pervasive visual redundancy. Existing methods rely heavily on supervised training with task-specific annotations, which inherently l…

Cited by 0SourceScholar
2026

Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning

ICML 2026oral

Collaborative machine learning involves training high-quality models using datasets from a number of sources. To incentivize sources to share data, existing data valuation methods fairly reward each source based on its data submitted as is. However, as these methods do not verify nor incentivize dat…

Cited by 0SourceScholar
2026

TD-VAD: Breaking Visual Dependence in Video Anomaly Detection with Text-Driven Learning

ICML 2026poster

Visual data is typically a prerequisite for training existing video anomaly detection (VAD) methods. However, obtaining sufficient annotated anomaly data for training is challenging and not scalable due to the rarity of anomaly data and the wide variety of abnormal events. In this work, we advocate …

Cited by 0SourceScholar
2025

Efficient Top-m Data Values Identification for Data Selection

ICLR 2025poster

Data valuation has found many real-world applications, e.g., data pricing and data selection. However, the most adopted approach -- Shapley value (SV) -- is computationally expensive due to the large number of model trainings required. Fortunately, most applications (e.g., data selection) require on…

Cited by 0SourcePDFScholar
2025

Uncovering Scaling Laws for Large Language Models via Inverse Problems

EMNLP 2025

Large Language Models (LLMs) are large-scale pretrained models that have achieved remarkable success across diverse domains. These successes have been driven by unprecedented complexity and scale in both data and computations. However, due to the high costs of training such models, brute-force trial

Cited by 0SourcePDFScholar
2024

DETAIL: Task DEmonsTration Attribution for Interpretable In-context Learning

NeurIPS 2024poster

In-context learning (ICL) allows transformer-based language models that are pre-trained on general text to quickly learn a specific task with a few "task demonstrations" without updating their parameters, significantly boosting their flexibility and generality. ICL possesses many distinct character…

2024

Data Distribution Valuation

NeurIPS 2024poster

Data valuation is a class of techniques for quantitatively assessing the value of data for applications like pricing in data marketplaces. Existing data valuation methods define a value for a discrete dataset. However, in many use cases, users are interested in not only the value of the dataset, but…

2024

Position Paper: Data-Centric AI in the Age of Large Language Models

EMNLP 2024finding

This position paper proposes a data-centric viewpoint of AI research, focusing on large language models (LLMs). We start by making a key observation that data is instrumental in the developmental (e.g., pretraining and fine-tuning) and inferential stages (e.g., in-context learning) of LLMs, and advo…

Cited by 1SourcePDFScholar
2023

FAIR: Fair Collaborative Active Learning with Individual Rationality for Scientific Discovery

AISTATS 2023poster

Scientific discovery aims to find new patterns and test specific hypotheses by analysing large-scale experimental data. However, various practical limitations (e.g., high experimental costs or the inability to perform some experiments) make it challenging for researchers to collect sufficient experi…

Cited by 15SourcePDFScholar
2023

Fair yet Asymptotically Equal Collaborative Learning

ICML 2023poster

In collaborative learning with streaming data, nodes (e.g., organizations) jointly and continuously learn a machine learning (ML) model by sharing the latest model updates computed from their latest streaming data. For the more resourceful nodes to be willing to share their model updates, they need…

2023

Incentives in Private Collaborative Machine Learning

NeurIPS 2023poster

Collaborative machine learning involves training models on data from multiple parties but must incentivize their participation. Existing data valuation methods fairly value and reward each party based on shared data or model parameters but neglect the privacy risks involved. To address this, we int…

Cited by 7SourcePDFScholar
2023

Model Shapley: Equitable Model Valuation with Black-box Access

NeurIPS 2023poster

Valuation methods of data and machine learning (ML) models are essential to the establishment of AI marketplaces. Importantly, certain practical considerations (e.g., operational constraints, legal restrictions) favor the use of model valuation over data valuation. Also, existing marketplaces that i…

2023

Probably Approximate Shapley Fairness with Applications in Machine Learning

AAAI 2023technical

The Shapley value (SV) is adopted in various scenarios in machine learning (ML), including data valuation, agent valuation, and feature attribution, as it satisfies their fairness requirements. However, as exact SVs are infeasible to compute in practice, SV estimates are approximated instead. This a…

2022

Data Valuation in Machine Learning: "Ingredients", Strategies, and Open Challenges

IJCAI 2022poster

Data valuation in machine learning (ML) is an emerging research area that studies the worth of data in ML. Data valuation is used in collaborative ML to determine a fair compensation for every data owner and in interpretable ML to identify the most responsible, noisy, or misleading training examples…

Cited by 69SourcePDFScholar
2022

Incentivizing Collaboration in Machine Learning via Synthetic Data Rewards

AAAI 2022technical

This paper presents a novel collaborative generative modeling (CGM) framework that incentivizes collaboration among self-interested parties to contribute data to a pool for training a generative model (e.g., GAN), from which synthetic data are drawn and distributed to the parties as rewards commensu…

2022

On the Convergence of the Shapley Value in Parametric Bayesian Learning Games

ICML 2022spotlight

Measuring contributions is a classical problem in cooperative game theory where the Shapley value is the most well-known solution concept. In this paper, we establish the convergence property of the Shapley value in parametric Bayesian learning games where players perform a Bayesian inference using…

2021

Gradient Driven Rewards to Guarantee Fairness in Collaborative Machine Learning

NeurIPS 2021poster

In collaborative machine learning(CML), multiple agents pool their resources(e.g., data) together for a common learning task. In realistic CML settings where the agents are self-interested and not altruistic, they may be unwilling to share data or model information without adequate rewards. Furtherm…

Cited by 95SourcePDFScholar
2021

Validation Free and Replication Robust Volume-based Data Valuation

NeurIPS 2021poster

Data valuation arises as a non-trivial challenge in real-world use cases such as collaborative machine learning, federated learning, trusted data sharing, data marketplaces. The value of data is often associated with the learning performance (e.g., validation accuracy) of a model trained on the data…

Cited by 81SourcePDFScholar