← Search

Muhammad Faaiz Taufiq

6 accepted papers

2026

D-FUSEr: Diverse Failure, Unified Success via Error-Distribution Shaping in LLM Reasoning

ICML 2026poster

Test-time scaling methods such as majority vote aggregation and iterative refinement (e.g., self-reflection or multi-agent inference) improve reasoning performance by leveraging multiple solution samples. However, their efficacy depends not only on raw performance, but critically on the distribution…

Cited by 0SourceScholar
2025

Understanding Chain-of-Thought in LLMs through Information Theory

ICML 2025poster

Large Language Models (LLMs) have shown impressive performance in complex reasoning tasks through the use of Chain-of-Thought (CoT) reasoning, allowing models to break down problems into manageable sub-tasks. However, existing CoT evaluation techniques either require annotated CoT data or fall short…

Cited by 6SourcePDFScholar
2024

Achievable Fairness on Your Data With Utility Guarantees

NeurIPS 2024poster

In machine learning fairness, training models that minimize disparity across different sensitive groups often leads to diminished accuracy, a phenomenon known as the fairness-accuracy trade-off. The severity of this trade-off inherently depends on dataset characteristics such as dataset imbalances o…

2023

Manifold Restricted Interventional Shapley Values

AISTATS 2023poster

Shapley values are model-agnostic methods for explaining model predictions. Many commonly used methods of computing Shapley values, known as off-manifold methods, rely on model evaluations on out-of-distribution input samples. Consequently, explanations obtained are sensitive to model behaviour outs…

2023

Marginal Density Ratio for Off-Policy Evaluation in Contextual Bandits

NeurIPS 2023poster

Off-Policy Evaluation (OPE) in contextual bandits is crucial for assessing new policies using existing data without costly experimentation. However, current OPE methods, such as Inverse Probability Weighting (IPW) and Doubly Robust (DR) estimators, suffer from high variance, particularly in cases of…

2022

Conformal Off-Policy Prediction in Contextual Bandits

NeurIPS 2022accept

Most off-policy evaluation methods for contextual bandits have focused on the expected outcome of a policy, which is estimated via methods that at best provide only asymptotic guarantees. However, in many applications, the expectation may not be the best measure of performance as it does not capture…

Cited by 20SourcePDFScholar