← Search

Shuo Sun

27 accepted papers

2026

Generalization of RLVR Using Causal Reasoning as a Testbed

ICLR 2026poster

Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for post-training large language models (LLMs) on complex reasoning tasks. Yet, the conditions under which RLVR yields robust generalization remain poorly understood. This paper provides an empirical study of R…

Cited by 0SourceScholar
2026

IMPACT: Behavioral Intention-Aware Multimodal Trajectory Prediction With Adaptive Context Trimming

RA-L 2026

This paper presents a unified framework that jointly predicts behavioral intentions and vectorized occupancy, leveraging them as priors to dynamically prune context information during trajectory decoding, thereby enhancing prediction accuracy, interpretability, and efficiency. While most prior work

Cited by 4SourceScholar
2026

IMPACT: Behavioral Intention-Aware Multimodal Trajectory Prediction with Adaptive Context Trimming

ICRA 2026poster

This paper presents a unified framework that jointly predicts behavioral intentions and vectorized occupancy, leveraging them as priors to dynamically prune context information during trajectory decoding, thereby enhancing prediction accuracy, interpretability, and efficiency. While most prior work …

2026

NPRIP: Nucleus-to-Periphery Retrieval-Iterative Prompting for Improved Abstractive Summarization in Low-Resource Mongolian

IJCAI 2026

Large language models often face challenges in low-resource agglutinative language text summarization tasks due to poorly designed prompts, leading to core information dilution, reduced fidelity, and critical information loss caused by the complex grammatical structures of agglutinative languages. F

Cited by 0Scholar
2025

AGI-Elo: How Far Are We From Mastering A Task?

NeurIPS 2025poster

As the field progresses toward Artificial General Intelligence (AGI), there is a pressing need for more comprehensive and insightful evaluation frameworks that go beyond aggregate performance metrics. This paper introduces a unified rating system that jointly models the difficulty of individual test…

Cited by 0SourcecodeScholar
2025

AudioBench: A Universal Benchmark for Audio Large Language Models

NAACL 2025long

We introduce AudioBench, a universal benchmark designed to evaluate Audio Large Language Models (AudioLLMs). It encompasses 8 distinct tasks and 26 datasets, among which, 7 are newly proposed datasets. The evaluation targets three main aspects: speech understanding, audio scene understanding, and vo…

2025

Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data

EMNLP 2025

Recent speech-LLMs have shown impressive performance in tasks like transcription and translation, yet they remain limited in understanding the paralinguistic aspects of speech crucial for social and emotional intelligence. We propose CP-Bench, a benchmark for evaluating speech-LLMs on contextual par

2025

Can Large Language Models Translate Unseen Languages in Underrepresented Scripts?

EMNLP 2025

Large language models (LLMs) have demonstrated impressive performance in machine translation, but still struggle with unseen low-resource languages, especially those written in underrepresented scripts. To investigate whether LLMs can translate such languages with the help of linguistic resources, w

2025

MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders

ICASSP 2025accepted

The rapid advancements in large language models (LLMs) have significantly enhanced natural language processing capabilities, facilitating the development of AudioLLMs that process and understand speech and audio inputs alongside text. Existing AudioLLMs typically combine a pre-trained audio encoder…

Cited by 0SourceScholar
2025

RMP-YOLO: A Robust Motion Predictor for Partially Observable Scenarios Even if You Only Look Once

ICRA 2025

We introduce RMP-YOLO, a unified framework designed to provide robust motion predictions even with incomplete input data. Our key insight stems from the observation that complete and reliable historical trajectory data plays a pivotal role in ensuring accurate motion prediction. Therefore, we propos

Cited by 7SourcecodeScholar
2024

3QFP: Efficient neural implicit surface reconstruction using Tri-Quadtrees and Fourier feature Positional encoding

ICRA 2024poster

Neural implicit surface representations are currently receiving a lot of interest as a means to achieve high-fidelity surface reconstruction at a low memory cost, compared to traditional explicit representations. However, state-of-the-art methods still struggle with excessive memory usage and non-sm…

Cited by 1SourcecodeScholar
2024

DriveSceneGen: Generating Diverse and Realistic Driving Scenarios From Scratch

RA-L 2024

Realistic and diverse traffic scenarios in large quantities are crucial for the development and validation of autonomous driving systems. However, owing to numerous difficulties in the data collection process and the reliance on intensive annotations, real-world datasets lack sufficient quantity and

Cited by 32SourceScholar
2024

EarnHFT: Efficient Hierarchical Reinforcement Learning for High Frequency Trading

AAAI 2024technical

High-frequency trading (HFT) is using computer algorithms to make trading decisions in short time scales (e.g., second-level), which is widely used in the Cryptocurrency (Crypto) market, (e.g., Bitcoin). Reinforcement learning (RL) in financial research has shown stellar performance on many quantita…

2024

High-Fidelity SLAM Using Gaussian Splatting with Rendering-Guided Densification and Regularized Optimization

IROS 2024poster

We propose a dense RGBD SLAM system based on 3D Gaussian Splatting that provides metrically accurate pose tracking and visually realistic reconstruction. To this end, we first propose a Gaussian densification strategy based on the rendering loss to map unobserved areas and refine reobserved areas. S…

Cited by 13SourcecodeScholar
2024

Market-GAN: Adding Control to Financial Market Data Generation with Semantic Context

AAAI 2024technical

Financial simulators play an important role in enhancing forecasting accuracy, managing risks, and fostering strategic financial decision-making. Despite the development of financial market simulation methodologies, existing frameworks often struggle with adapting to specialized simulation context.…

Cited by 10SourcePDFScholar
2024

Mitigating Linguistic Artifacts in Emotion Recognition for Conversations from TV Scripts to Daily Conversations

COLING 2024main

Emotion Recognition in Conversations (ERC) is a well-studied task with numerous potential real-world applications. However, existing ERC models trained on the MELD dataset derived from TV series, struggle when applied to daily conversation datasets. A closer examination of the datasets unveils the p…

Cited by 0SourcePDFScholar
2024

SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages

EMNLP 2024main

Southeast Asia (SEA) is a region rich in linguistic diversity and cultural variety, with over 1,300 indigenous languages and a population of 671 million people. However, prevailing AI models suffer from a significant lack of representation of texts, images, and audio datasets from SEA, compromising…

2023

An Exploratory Study on Model Compression for Text-to-SQL

ACL 2023findings

Text-to-SQL translates user queries into SQL statements that can retrieve relevant answers from relational databases. Recent approaches to Text-to-SQL rely on pre-trained language models that are computationally expensive and technically challenging to deploy in real-world applications that require…

2023

Battle of the Large Language Models: Dolly vs LLaMA vs Vicuna vs Guanaco vs Bard vs ChatGPT - A Text-to-SQL Parsing Comparison

EMNLP 2023long findings

The success of ChatGPT has ignited an AI race, with researchers striving to develop new large language models (LLMs) that can match or surpass the language understanding and generation abilities of commercial ones. In recent times, a number of models have emerged, claiming performance near that of…

Cited by 0SourceScholar
2023

FISS+: Efficient and Focused Trajectory Generation and Refinement Using Fast Iterative Search and Sampling Strategy

IROS 2023poster

Trajectory planning plays a crucial role in autonomous driving systems, as it is tasked to generate feasible trajectories under highly dynamic scenarios within the time constraint. This paper proposes a novel two-stage coarse-to-fine framework for efficient sampling-based trajectory planning. The pr…

Cited by 7SourceScholar
2023

TradeMaster: A Holistic Quantitative Trading Platform Empowered by Reinforcement Learning

NeurIPS 2023poster

The financial markets, which involve over \$90 trillion market capitals, attract the attention of innumerable profit-seeking investors globally. Recent explosion of reinforcement learning in financial trading (RLFT) research has shown stellar performance on many quantitative trading tasks. However,…

2022

AfriCLIRMatrix: Enabling Cross-Lingual Information Retrieval for African Languages

EMNLP 2022main

Language diversity in NLP is critical in enabling the development of tools for a wide range of users.However, there are limited resources for building such tools for many languages, particularly those spoken in Africa.For search, most existing datasets feature few or no African languages, directly i…

2022

FISS: A Trajectory Planning Framework Using Fast Iterative Search and Sampling Strategy for Autonomous Driving

RA-L 2022

Trajectory planning is a critical component in autonomous vehicles directly responsible for driving safety and efficiency during deployment. The ability to find the optimal trajectory in real-time is critical for autonomous driving. This paper presents a novel general framework using the Fast Iterat

Cited by 20SourceScholar
2021

Classification-based Quality Estimation: Small and Efficient Models for Real-world Applications

EMNLP 2021main

Sentence-level Quality estimation (QE) of machine translation is traditionally formulated as a regression task, and the performance of QE models is typically measured by Pearson correlation with human labels. Recent QE models have achieved previously-unseen levels of correlation with human judgments…

Cited by 3SourcePDFScholar