← Search

Syed Zawad

5 accepted papers

2026

GneissWeb: Preparing High Quality Data for LLMs at Scale

ICLR 2026poster

Data quantity and quality play a vital role in determining the performance of Large Language Models (LLMs). High-quality data, in particular, can significantly boost the LLM's ability to generalize on a wide range of downstream tasks. In this paper, we introduce **GneissWeb**, a large dataset of aro…

Cited by 0SourceScholar
2026

When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets

ICLR 2026poster

Aligning large language models (LLMs) is a central objective of post-training, often achieved through reward modeling and reinforcement learning methods. Among these, direct preference optimization (DPO) has emerged as a widely adopted technique that fine-tunes LLMs on preferred completions over les…

Cited by 0SourceScholar
2025

Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance

NeurIPS 2025spotlight

Recent work on large language models (LLMs) has increasingly focused on post-training and alignment with datasets curated to enhance instruction following, world knowledge, and specialized skills. However, most post-training datasets used in leading open- and closed-source LLMs remain inaccessible t…

Cited by 0SourceScholar
2023

DySR: Adaptive Super-Resolution via Algorithm and System Co-design

ICLR 2023poster

Super resolution (SR) is a promising approach for improving the quality of low resolution steaming services on mobile devices. On mobile devices, the available computing and memory resources change dynamically depending on other running applications. Due to the high computation and memory demands of…

Cited by 1SourcePDFScholar
2021

Curse or Redemption? How Data Heterogeneity Affects the Robustness of Federated Learning

AAAI 2021technical

Data heterogeneity has been identified as one of the key features in federated learning but often overlooked in the lens of robustness to adversarial attacks. This paper focuses on characterizing and understanding its impact on backdooring attacks in federated learning through comprehensive experime…