← Search

Linjing Li

16 accepted papers

2025

A Novel Decision-Making Model for Playing Board Game Combining Planning and Opponent Behaviors

ICASSP 2025accepted

Board game offers a unique platform for exploring the capabilities of artificial intelligence in decision-making. It demands long-term strategic planning and opponent behaviors to refine decision-making. Since the success of AlphaGo family, learning agents have become pivotal methods for board game.…

Cited by 0SourceScholar
2025

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

EMNLP 2025

Many studies focus on data annotation techniques for training effective PRMs. However, current methods encounter a significant issue when applied to long CoT reasoning processes: they tend to focus solely on the first incorrect step and all preceding steps, assuming that all subsequent steps are inc

Cited by 0SourcePDFScholar
2025

Conservative Offline Meta-Reinforcement Learning with Task Similarity Measurement

ICASSP 2025accepted

Offline meta-reinforcement learning (OMRL) enables reinforcement learning (RL) agents to adapt to unseen tasks without interacting with the environment. However, OMRL faces challenges such as Q-function overestimation and difficulties in inferring tasks correctly and robustly due to distribution dis…

Cited by 0SourceScholar
2025

Evaluating Generalization Capability of Language Models across Abductive, Deductive and Inductive Logical Reasoning

COLING 2025main

Transformer-based language models (LMs) have demonstrated remarkable performance on many natural language tasks, yet to what extent LMs possess the capability of generalizing to unseen logical rules remains not explored sufficiently. In classical logic category, abductive, deductive and inductive (A…

2025

Learning Dynamics in Continual Pre-Training for Large Language Models

ICML 2025oral

Continual Pre-Training (CPT) has become a popular and effective method to apply strong foundation models to specific downstream tasks. In this work, we explore the **learning dynamics** throughout the CPT process for large language models (LLMs). We specifically focus on how general and downstream…

Cited by 0SourcePDFScholar
2025

Learning Strategy Representation for Imitation Learning in Multi-Agent Games

AAAI 2025technical

The offline datasets for imitation learning (IL) in multi-agent games typically contain player trajectories exhibiting diverse strategies, which necessitate measures to prevent learning algorithms from acquiring undesirable behaviors. Learning representations for these trajectories is an effective a…

Cited by 0SourcePDFScholar
2025

Learning Theorem Rationale for Improving the Mathematical Reasoning Capability of Large Language Models

AAAI 2025technical

Large language models (LLMs) have achieved significant progress in mathematical reasoning, especially in elementary math. However, they remain indisposed on tackling complex questions at high-school or college levels, which put forward a more advanced requirement of mastering relevant mathematical t…

2025

Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL

NeurIPS 2025poster

Large reasoning models (LRMs) are proficient at generating explicit, step-by-step reasoning sequences before producing final answers. However, such detailed reasoning can introduce substantial computational overhead and latency, particularly for simple problems. To address this over-thinking problem…

Cited by 0SourcecodeScholar
2025

POSITION BIAS MITIGATES POSITION BIAS: Mitigate Position Bias Through Inter-Position Knowledge Distillation

EMNLP 2025

Positional bias (PB), manifesting as non-uniform sensitivity across different contextual locations, significantly impairs long-context comprehension and processing capabilities. Previous studies have addressed PB either by modifying the underlying architectures or by employing extensive contextual a

2025

Sociologically-Informed Graph Neural Network for Opinion Prediction

ICASSP 2025accepted

Social media platforms has long served as open arenas where individuals discuss and change their opinions on various events, subsequently influencing the progression of these events. Public opinion, recognized as an important social signal, is instrumental in understanding the developmental patterns…

Cited by 0SourceScholar
2025

Uncertainty Unveiled: Can Exposure to More In-context Examples Mitigate Uncertainty for Large Language Models?

ACL 2025finding

Recent advances in handling long sequences have unlocked new possibilities for long-context in-context learning (ICL). While existing research predominantly focuses on performance gains driven by additional in-context examples, the impact on the trustworthiness of generated responses remains underex…

Cited by 0SourcePDFScholar
2025

Unearthing Gems from Stones: Policy Optimization with Negative Sample Augmentation for LLM Reasoning

EMNLP 2025

Recent advances in reasoning language models have witnessed a paradigm shift from short to long CoT pattern. Given the substantial computational cost of rollouts in long CoT models, maximizing the utility of fixed training datasets becomes crucial. Our analysis reveals that negative responses contai

Cited by 0SourcePDFScholar
2024

Integrating Language Models with Symbolic Formulas for First-Order Logic Reasoning

ICASSP 2024accepted

Performing logical reasoning based on prior knowledge is a crucial human cognitive ability and has been a long-standing objective in the field of artificial intelligence. Large language models based on transformer architecture have been a common approach for logical reasoning over text. However, the…

Cited by 0SourceScholar
2024

Unveiling Factual Recall Behaviors of Large Language Models through Knowledge Neurons

EMNLP 2024main

In this paper, we investigate whether Large Language Models (LLMs) actively recall or retrieve their internal repositories of factual knowledge when faced with reasoning tasks. Through an analysis of LLMs’ internal factual recall at each reasoning step via Knowledge Neurons, we reveal that LLMs fail…

2023

LDM$^2$: A Large Decision Model Imitating Human Cognition with Dynamic Memory Enhancement

EMNLP 2023long findings

With the rapid development of large language models (LLMs), it is highly demanded that LLMs can be adopted to make decisions to enable the artificial general intelligence. Most approaches leverage manually crafted examples to prompt the LLMs to imitate the decision process of human. However, design…

Cited by 0SourceScholar