← Search

Nuo Chen

58 accepted papers

2026

ChemEval: A Multi-level and Fine-grained Chemical Capability Evaluation for Large Language Models

ICLR 2026poster

The emergence of Large Language Models (LLMs) in chemistry marks a significant advancement in applying artificial intelligence to chemical sciences. While these models show promising potential, their effective application in chemistry demands sophisticated evaluation protocols that address the field…

Cited by 0SourcecodeScholar
2026

CoGenSAM: Codebook-Interactive Generative Labeling for Adapting SAM to Crack Segmentation

AAAI 2026technical

The goal of this work is to adapt Segment Anything Models (SAM) into crack segmentation tasks via automatic label generation, thus eliminating manual annotation cost. In this regard, an intuitive approach is to extract edges of crack samples and generate labels via the dilation and erosion processes

Cited by 0SourcePDFScholar
2026

Exposing Weaknesses of Large Reasoning Models through Graph Algorithm Problems

ICLR 2026poster

Large Reasoning Models (LRMs) have advanced rapidly, yet existing benchmarks on mathematics, code, and common-sense reasoning remain limited: they lack long-context evaluation, offer insufficient challenge, and provide answers that are difficult to verify programmatically. We introduce GrAlgoBench,…

Cited by 0SourceScholar
2026

ExtendAttack: Attacking Servers of LRMs via Extending Reasoning

AAAI 2026technical

Large Reasoning Models (LRMs) have demonstrated promising performance in complex tasks. However, the resource-consuming reasoning processes may be exploited by attackers to maliciously occupy the resources of the servers, leading to a crash, like the DDoS attack in cyber. To this end, we propose a n

Cited by 0SourcePDFScholar
2026

MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs

IJCAI 2026

Multimodal Large Language Models (MLLMs) perform strongly in high-resource languages, yet often produce fluent but culturally "thin" descriptions in low-resource settings. We argue that this failure is not merely a linguistic limitation: culture-specific visual knowledge depends on native visual-tex

Cited by 0Scholar
2026

Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization

ICLR 2026poster

As Large Language Models (LLMs) are increasingly deployed in real-world applications, it is important to ensure their behaviors align with human values, societal norms, and ethical principles. However, safety alignment under Reinforcement Learning (RL) often suffers from forgetting learned general a…

Cited by 0SourcecodeScholar
2026

Towards Persistence: Learning Topological Constraints for Event-based Small Object Detection

CVPR 2026

Small object detection (SOD) plays a vital role in applications such as anti-UAV tasks, yet conventional image-based methods struggle in high-speed scenarios due to the limited frame rate. Event cameras offer a promising alternative by capturing spatiotemporal event streams with microsecond-level te

Cited by 0SourceScholar
2026

Trust3R: Unifying Feed-Forward Pointmap Prediction and Evidential Learning for Trust-Aware 3D Reconstruction

ICML 2026poster

Geometric foundation models hold promise for unconstrained dense geometry prediction from uncalibrated images; however, in current feed-forward designs, their predicted confidence scores are heuristic, lack probabilistic interpretation, and often fail to indicate where and how much the predicted geo…

Cited by 0SourceScholar
2026

WeatherSyn: An Instruction Tuning MLLM For Weather Forecasting Report Generation

ICML 2026poster

Accurate weather forecast reporting enables individuals and communities to better plan daily activities, agricultural operations, and transportation. However, the current reporting process primarily relies on manual analysis of multi-source data, which often leads to information overload and reduced…

Cited by 0SourceScholar
2025

CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

ACL 2025finding

Image captioning has been a longstanding challenge in vision-language research. With the rise of LLMs, modern Vision-Language Models (VLMs) generate detailed and comprehensive image descriptions. However, benchmarking the quality of such captions remains unresolved. This paper addresses two key ques…

Cited by 0SourcePDFScholar
2025

Chain of Execution Supervision Promotes General Reasoning in Large Language Models

NeurIPS 2025poster

Building robust and general reasoning ability is a central goal in the development of large language models (LLMs). Recent efforts increasingly turn to code as a rich training source, given its inherent logical structure and diverse reasoning paradigms—such as divide-and-conquer, topological orderin…

Cited by 0SourceScholar
2025

Compress to Impress: Unleashing the Potential of Compressive Memory in Real-World Long-Term Conversations

COLING 2025main

Existing retrieval-based methods have made significant strides in maintaining long-term conversations. However, these approaches face challenges in memory database management and accurate memory retrieval, hindering their efficacy in dynamic, real-world interactions. This study introduces a novel fr…

2025

DRBO: Mitigating the Bottleneck Effect via Dynamic Reward Balancing in Multi-reward LLM Optimization

EMNLP 2025

In the current landscape of large language models (LLMs), many evaluation metrics have been developed and used as rewards during training to improve specific metrics. However, balancing these metrics and dynamically adjusting reward weights remains challenging, as current approaches often fail to en

2025

Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family Experts

ICLR 2025poster

Adapting medical Large Language Models to local languages can reduce barriers to accessing healthcare services, but data scarcity remains a significant challenge, particularly for low-resource languages. To address this, we first construct a high-quality medical dataset and conduct analysis to ensu…

2025

Event-based Tiny Object Detection: A Benchmark Dataset and Baseline

ICCV 2025poster

Small object detection (SOD) in anti-UAV task is a challenging problem due to the small size of UAVs and complex backgrounds. Traditional frame-based cameras struggle to detect small objects in complex environments due to their low frame rates, limited dynamic range, and data redundancy. Event camer…

2025

GraphArena: Evaluating and Exploring Large Language Models on Graph Computation

ICLR 2025poster

The ``arms race'' of Large Language Models (LLMs) demands new benchmarks to examine their progresses. In this paper, we introduce GraphArena, a benchmarking tool designed to evaluate LLMs on real-world graph computational problems. It offers a suite of four polynomial-time tasks (e.g., Shortest Dist…

2025

How does Misinformation Affect Large Language Model Behaviors and Preferences?

ACL 2025long

Large Language Models (LLMs) have shown remarkable capabilities in knowledge-intensive tasks, while they remain vulnerable when encountering misinformation. Existing studies have explored the role of LLMs in combating misinformation, but there is still a lack of fine-grained analysis on the specific…

2025

Is Your LLM Outdated? A Deep Look at Temporal Generalization

NAACL 2025long

The rapid advancement of Large Language Models (LLMs) has led to the development of benchmarks that consider temporal dynamics, however, there remains a gap in understanding how well these models can generalize across temporal contexts due to the inherent dynamic nature of language and information.…

2025

MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria

NAACL 2025long

Multimodal large language models (MLLMs) have broadened the scope of AI applications. Existing automatic evaluation methodologies for MLLMs are mainly limited in evaluating objective queries without considering real-world user experiences, inadequately addressing the nuances of creative and associat…

2025

MegaAgent: A Large-Scale Autonomous LLM-based Multi-Agent System Without Predefined SOPs

ACL 2025finding

LLM-based multi-agent systems (MAS) have shown promise in tackling complex tasks. However, existing solutions often suffer from limited agent coordination and heavy reliance on predefined Standard Operating Procedures (SOPs), which demand extensive human input. To address these limitations, we propo…

2025

MotionGrasp: Long-Term Grasp Motion Tracking for Dynamic Grasping

RA-L 2025

Dynamic grasping, which aims to grasp moving objects in unstructured environment, is crucial for robotics community. Previous methods propose to track the initial grasps or objects by matching between the latest two frames. However, this neighbour-frame matching strategy ignores the long-term histor

Cited by 6SourceScholar
2025

PHMamba: Preheating State Space Models with Context-Augmented Features for Medical Image Segmentation

ICASSP 2025accepted

The recent Mamba model has demonstrated the competitive potential of State Space Models (SSMs) on various image benchmarks, particularly in modeling long-range sequences. However, most of the improvements in Mamba methods focus on scanning strategies, and lack an effective means of aggregating conte…

Cited by 0SourceScholar
2025

RelEdit: Evaluating Conceptual Knowledge Editing in Language Models via Relational Reasoning

ACL 2025finding

The conceptual knowledge in Large Language Models (LLMs) can become outdated over time, and concept editing is often an option. Current evaluations on conceptual knowledge editing primarily focus on whether the definitions of concepts are successfully edited, neglecting the impact on the model’s rel…

2025

Revisiting LoRA through the Lens of Parameter Redundancy: Spectral Encoding Helps

ACL 2025finding

Low-Rank Adaptation (LoRA) has emerged as a prominent technique for fine-tuning large foundation models. Despite its successes, the substantial parameter redundancy, which limits the capacity and efficiency of LoRA, has been recognized as a bottleneck. In this work, we systematically investigate the…

Cited by 0SourcePDFScholar
2024

Are Your Models Still Fair? Fairness Attacks on Graph Neural Networks via Node Injections

NeurIPS 2024poster

Despite the remarkable capabilities demonstrated by Graph Neural Networks (GNNs) in graph-related tasks, recent research has revealed the fairness vulnerabilities in GNNs when facing malicious adversarial attacks. However, all existing fairness attacks require manipulating the connectivity between e…

2024

Breaking Language Barriers in Multilingual Mathematical Reasoning: Insights and Observations

EMNLP 2024finding

Existing research predominantly focuses on developing powerful large language models (LLMs) for mathematical reasoning within monolingual languages, with few explorations in preserving efficacy in a multilingual context. To bridge this gap, this paper pioneers exploring and training powerful Multili…

2024

ChatDev: Communicative Agents for Software Development

ACL 2024long

Software development is a complex task that necessitates cooperation among multiple members with diverse skills. Numerous studies used deep learning to improve specific phases in a waterfall model, such as design, coding, and testing. However, the deep learning model in each phase requires unique de…

2024

ControlMath: Controllable Data Generation Promotes Math Generalist Models

EMNLP 2024main

Utilizing large language models (LLMs) for data augmentation has yielded encouraging results in mathematical reasoning. However, these approaches face constraints in problem diversity, potentially restricting them to in-domain/distribution data generation. To this end, we propose **ControlMath**, an…

2024

CryptoTrade: A Reflective LLM-based Agent to Guide Zero-shot Cryptocurrency Trading

EMNLP 2024main

The utilization of Large Language Models (LLMs) in financial trading has primarily been concentrated within the stock market, aiding in economic and financial decisions. Yet, the unique opportunities presented by the cryptocurrency market, noted for its on-chain data’s transparency and the critical…

2024

Make Prompt-based Black-Box Tuning Colorful: Boosting Model Generalization from Three Orthogonal Perspectives

COLING 2024main

Large language models (LLMs) have shown increasing power on various natural language processing (NLP) tasks. However, tuning these models for downstream tasks usually needs exorbitant costs or is unavailable due to commercial considerations. Recently, black-box tuning has been proposed to address th…

2024

Multiagent Multitraversal Multimodal Self-Driving: Open MARS Dataset

CVPR 2024poster

Large-scale datasets have fueled recent advancements in AI-based autonomous vehicle research. However these datasets are usually collected from a single vehicle's one-time pass of a certain location lacking multiagent interactions or repeated traversals of the same place. Such information could lead…

2024

Retrieving, Rethinking and Revising: The Chain-of-Verification Can Improve Retrieval Augmented Generation

EMNLP 2024finding

Recent Retrieval Augmented Generation (RAG) aims to enhance Large Language Models (LLMs) by incorporating extensive knowledge retrieved from external sources. However, such approach encounters some challenges: Firstly, the original queries may not be suitable for precise retrieval, resulting in erro…

Cited by 4SourcePDFScholar
2024

SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous Driving

IROS 2024

Monocular scene understanding is a foundational component of autonomous systems. Within the spectrum of monocular perception topics, one crucial and useful task for holistic 3D scene understanding is semantic scene completion (SSC), which jointly completes semantic information and geometric details

Cited by 90SourcecodeScholar
2024

Structure-aware Fine-tuning for Code Pre-trained Models

COLING 2024main

Over the past few years, we have witnessed remarkable advancements in Code Pre-trained Models (CodePTMs). These models achieved excellent representation capabilities by designing structure-based pre-training tasks for code. However, how to enhance the absorption of structural knowledge when fine-tun…

Cited by 1SourcePDFScholar
2024

TransCoder: Towards Unified Transferable Code Representation Learning Inspired by Human Skills

COLING 2024main

Code pre-trained models (CodePTMs) have recently demonstrated a solid capacity to process various code intelligence tasks, e.g., code clone detection, code translation, and code summarization. The current mainstream method that deploys these models to downstream tasks is to fine-tune them on individ…

2023

Alleviating Over-smoothing for Unsupervised Sentence Representation

ACL 2023long

Currently, learning better unsupervised sentence representations is the pursuit of many natural language processing communities. Lots of approaches based on pre-trained language models (PLMs) and contrastive learning have achieved promising results on this task. Experimentally, we observe that the o…

2023

Evaluating and Enhancing the Robustness of Code Pre-trained Models through Structure-Aware Adversarial Samples Generation

EMNLP 2023long findings

Code pre-trained models (CodePTMs) have significantly advanced the field of neural code intelligence. Despite their capabilities, these models are susceptible to adversarial attacks that subtly modify the model inputs, resulting in incorrect outputs or predictions. Previous methods of robustness ev…

Cited by 0SourceScholar
2023

FiTs: Fine-Grained Two-Stage Training for Knowledge-Aware Question Answering

AAAI 2023technical

Knowledge-aware question answering (KAQA) requires the model to answer questions over a knowledge base, which is essential for both open-domain QA and domain-specific QA, especially when language models alone cannot provide all the knowledge needed. Despite the promising result of recent KAQA system…

2023

Human Mobility Modeling during the COVID-19 Pandemic via Deep Graph Diffusion Infomax

AAAI 2023technical

Non-Pharmaceutical Interventions (NPIs), such as social gathering restrictions, have shown effectiveness to slow the transmission of COVID-19 by reducing the contact of people. To support policy-makers, multiple studies have first modelled human mobility via macro indicators (e.g., average daily tra…

2023

Improving Retrieval-Based Dialogue System Via Syntax-Informed Attention

ICASSP 2023accepted

Multi-turn response selection is a challenging task due to its high demands on efficient extraction of the matching features from abundant information provided by context utterances. Since incorporating syntactic information like dependency structures into neural models can promote a better understa…

Cited by 0SourceScholar
2023

Large Language Models Meet Harry Potter: A Dataset for Aligning Dialogue Agents with Characters

EMNLP 2023long findings

In recent years, Dialogue-style Large Language Models (LLMs) such as ChatGPT and GPT4 have demonstrated immense potential in constructing open-domain dialogue agents. However, aligning these agents with specific characters or individuals remains a considerable challenge due to the complexities of ch…

Cited by 0SourceScholar
2023

Mapping Degeneration Meets Label Evolution: Learning Infrared Small Target Detection With Single Point Supervision

CVPR 2023poster

Training a convolutional neural network (CNN) to detect infrared small targets in a fully supervised manner has gained remarkable research interests in recent years, but is highly labor expensive since a large number of per-pixel annotations are required. To handle this problem, in this paper, we ma…

2023

Natural Response Generation for Chinese Reading Comprehension

EMNLP 2023long findings

Machine reading comprehension (MRC) is an important area of conversation agents and draws a lot of attention. However, there is a notable limitation to current MRC benchmarks: The labeled answers are mostly either spans extracted from the target corpus or the choices of the given candidates, ignorin…

Cited by 0SourcecodeScholar
2023

Orca: A Few-shot Benchmark for Chinese Conversational Machine Reading Comprehension

EMNLP 2023long findings

The conversational machine reading comprehension (CMRC) task aims to answer questions in conversations, which has been a hot research topic in recent years because of its wide applications. However, existing CMRC benchmarks in which each conversation is assigned a static passage are inconsistent wit…

Cited by 0SourcecodeScholar
2023

Structural Contrastive Pretraining for Cross-Lingual Comprehension

ACL 2023findings

To present, multilingual language models trained using various pre-training tasks like mask language modeling (MLM) have yielded encouraging results on a wide range of downstream tasks. Despite the promising performances, structural knowledge in cross-lingual corpus is less explored in current works…

2023

Uncertainty-aware Parameter-Efficient Self-training for Semi-supervised Language Understanding

EMNLP 2023long findings

The recent success of large pre-trained language models (PLMs) heavily hinges on massive labeled data, which typically produces inferior performance in low-resource scenarios. To remedy this dilemma, we study self-training as one of the predominant semi-supervised learning (SSL) approaches, which ut…

Cited by 0SourcecodeScholar
2023

When Gradient Descent Meets Derivative-Free Optimization: A Match Made in Black-Box Scenario

ACL 2023findings

Large pre-trained language models (PLMs) have garnered significant attention for their versatility and potential for solving a wide spectrum of natural language processing (NLP) tasks. However, the cost of running these PLMs may be prohibitive. Furthermore, PLMs may not be open-sourced due to commer…

Cited by 8SourcePDFScholar
2022

A Transformer-based Threshold-Free Framework for Multi-Intent NLU

COLING 2022main

Multi-intent natural language understanding (NLU) has recently gained attention. It detects multiple intents in an utterance, which is better suited to real-world scenarios. However, the state-of-the-art joint NLU models mainly detect multiple intents on threshold-based strategy, resulting in one ma…

2022

Bridging the Gap between Language Models and Cross-Lingual Sequence Labeling

NAACL 2022long

Large-scale cross-lingual pre-trained language models (xPLMs) have shown effective in cross-lingual sequence labeling tasks (xSL), such as machine reading comprehension (xMRC) by transferring knowledge from a high-resource language to low-resource languages. Despite the great success, we draw an emp…

2022

CAT-probing: A Metric-based Approach to Interpret How Pre-trained Models for Programming Language Attend Code Structure

EMNLP 2022finding

Code pre-trained models (CodePTMs) have recently demonstrated significant success in code intelligence. To interpret these models, some probing methods have been applied. However, these methods fail to consider the inherent characteristics of codes. In this paper, to address the problem, we propose…

2022

End-to-end Spoken Conversational Question Answering: Task, Dataset and Model

NAACL 2022findings

In spoken question answering, the systems are designed to answer questions from contiguous text spans within the related speech transcripts. However, the most natural way that human seek or test their knowledge is via human conversations. Therefore, we propose a new Spoken Conversational Question An…

Cited by 37SourcePDFScholar
2022

From Good to Best: Two-Stage Training for Cross-Lingual Machine Reading Comprehension

AAAI 2022technical

Cross-lingual Machine Reading Comprehension (xMRC) is a challenging task due to the lack of training data in low-resource languages. Recent approaches use training data only in a resource-rich language (such as English) to fine-tune large-scale cross-lingual pre-trained language models, which transf…

2021

Adaptive Bi-Directional Attention: Exploring Multi-Granularity Representations for Machine Reading Comprehension

ICASSP 2021accepted

Recently, the attention-enhanced multi-layer encoder, such as Transformer, has been extensively studied in Machine Reading Comprehension (MRC). To predict the answer, it is common practice to employ a predictor to draw information only from the final encoder layer which generates the coarse-grained…

Cited by 0SourceScholar
2021

MRD-Net: Multi-Modal Residual Knowledge Distillation for Spoken Question Answering

IJCAI 2021poster

Spoken question answering (SQA) has recently drawn considerable attention in the speech community. It requires systems to find correct answers from the given spoken passages simultaneously. The common SQA systems consist of the automatic speech recognition (ASR) module and text-based question answer…

Cited by 37SourcePDFScholar
2021

Self-supervised Contrastive Cross-Modality Representation Learning for Spoken Question Answering

EMNLP 2021finding

Spoken question answering (SQA) requires fine-grained understanding of both spoken documents and questions for the optimal answer prediction. In this paper, we propose novel training schemes for spoken question answering with a self-supervised training stage and a contrastive representation learning…

Cited by 61SourcePDFScholar