← Search

Fangru Lin

8 accepted papers

2026

Can Large Language Models Generalize Procedures Across Representations?

ICML 2026poster

Large language models (LLMs) are trained and tested extensively on symbolic representations such as code and graphs, yet real-world user tasks are often specified in natural language. To what extent can LLMs generalize across these representations? Here, we approach this question by studying isomorp…

Cited by 0SourceScholar
2025

Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks

ACL 2025long

Language is not monolithic. While benchmarks, including those designed for multiple languages, are often used as proxies to evaluate the performance of Large Language Models (LLMs), they tend to overlook the nuances of within-language variation and thus fail to model the experience of speakers of no…

Cited by 0SourcePDFScholar
2025

Designing Specialized Two-Dimensional Graph Spectral Filters for Spatial-Temporal Graph Modeling

AAAI 2025technical

Spatial-temporal graph modeling is challenging due to the diverse node interactions across spatial and temporal dimensions. Recent studies typically adopt Graph Neural Networks (GNNs) to perform node-level aggregation at different time steps, acting as a series of low-pass graph spectral filters, fo…

2025

Leveraging Heterophily in Spatial-Temporal Graphs for Multivariate Time-Series Forecasting

ICASSP 2025accepted

Multivariate Time-Series (MTS) forecasting is challenging due to the complex spatial-temporal dependencies inherent in MTS data. Recent studies typically adopt spatial-temporal graph models to leverage this information. However, most of these approaches assume homophily in graphs and perform only im…

Cited by 0SourceScholar
2025

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

NeurIPS 2025poster

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as `safety' and `robustness' requires strong construct validity, that is, having measures t…

Cited by 0SourceScholar
2025

TCP: a Benchmark for Temporal Constraint-Based Planning

EMNLP 2025

Temporal reasoning and planning are essential capabilities for large language models (LLMs), yet most existing benchmarks evaluate them in isolation and under limited forms of complexity. To address this gap, we introduce the Temporal Constraint-based Planning (TCP) benchmark, that jointly assesses

2024

Graph-enhanced Large Language Models in Asynchronous Plan Reasoning

ICML 2024poster

Planning is a fundamental property of human intelligence. Reasoning about asynchronous plans is challenging since it requires sequential and parallel planning to optimize time costs. Can large language models (LLMs) succeed at this task? Here, we present the first large-scale study investigating thi…

2024

Probing Large Language Models for Scalar Adjective Lexical Semantics and Scalar Diversity Pragmatics

COLING 2024main

Scalar adjectives pertain to various domain scales and vary in intensity within each scale (e.g. certain is more intense than likely on the likelihood scale). Scalar implicatures arise from the consideration of alternative statements which could have been made. They can be triggered by scalar adject…