← Search

Junshu Pan

3 accepted papers

2026

Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs

ICLR 2026poster

Enhancing the complex reasoning capabilities of Large Language Models (LLMs) attracts widespread attention. While reinforcement learning (RL) has shown superior performance for improving complex reasoning, its impact on cross-lingual generalization compared to Supervised Fine-Tuning (SFT) remains un…

Cited by 0SourceScholar
2026

Pre-DPO: Improving Data Utilization in Direct Preference Optimization Using a Guiding Reference Model

AAAI 2026technical

Direct Preference Optimization (DPO) simplifies reinforcement learning from human feedback (RLHF) for large language models (LLMs) by directly training on offline preference data to align with human preferences. During DPO training, the reference model serves as a data weight adjuster. However, the

Cited by 0SourcePDFScholar
2022

Hierarchical Cross-Modality Semantic Correlation Learning Model for Multimodal Summarization

AAAI 2022technical

Multimodal summarization with multimodal output (MSMO) generates a summary with both textual and visual content. Multimodal news report contains heterogeneous contents, which makes MSMO nontrivial. Moreover, it is observed that different modalities of data in the news report correlate hierarchically…