← Search

Linlin Wang

18 accepted papers

2026

Pseudocode-Guided Structured Reasoning for Automating Reliable Inference in Vision-Language Models

ICRA 2026poster

Vision-Language Models (VLMs) are becoming the cornerstone of high-level reasoning for robotic automation, enabling robots to parse natural language commands and perceive their environments. However, their susceptibility to hallucinations introduces critical failures in decision-making, posing signi…

2025

ACE-M3: Automatic Capability Evaluator for Multimodal Medical Models

COLING 2025main

As multimodal large language models (MLLMs) gain prominence in the medical field, the need for precise evaluation methods to assess their effectiveness has become critical. While benchmarks provide a reliable means to evaluate the capabilities of MLLMs, traditional metrics like ROUGE and BLEU employ…

Cited by 0SourcePDFScholar
2025

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation

ACL 2025long

With the proliferation of large language models (LLMs) in the medical domain, there is increasing demand for improved evaluation techniques to assess their capabilities. However, traditional metrics like F1 and ROUGE, which rely on token overlaps to measure quality, significantly overlook the import…

Cited by 0SourcePDFScholar
2025

Hierarchical Divide-and-Conquer for Fine-Grained Alignment in LLM-Based Medical Evaluation

AAAI 2025technical

In the rapidly evolving landscape of large language models (LLMs) for medical applications, ensuring the reliability and accuracy of these models in clinical settings is paramount. Existing benchmarks often focus on fixed-format tasks like multiple-choice QA, which fail to capture the complexity of…

2025

Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models

ACL 2025finding

With the increasing size of Large Vision-Language Models (LVLMs), network pruning techniques aimed at compressing models for deployment in resource-constrained environments have garnered significant attention. However, we observe that pruning often leads to a degradation in safety performance. To ad…

2025

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning

AAAI 2025technical

The emergence of fine-tuning-as-a-service has revealed a new vulnerability in large language models (LLMs). A mere handful of malicious data uploaded by users can subtly manipulate the fine-tuning process, leading to a compromised alignment state. Existing methods to counteract fine-tuning attacks t…

2024

CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical Scenarios

EMNLP 2024main

With the proliferation of Large Language Models (LLMs) in diverse domains, there is a particular need for unified evaluation standards in clinical medical scenarios, where models need to be examined very thoroughly. We present CliMedBench, a comprehensive benchmark with 14 expert-guided core clinica…

2024

Fine-Grained Scene Graph Generation via Sample-Level Bias Prediction

ECCV 2024poster

"Scene Graph Generation (SGG) aims to explore the relationships between objects in images and obtain scene summary graphs, thereby better serving downstream tasks. However, the long-tailed problem has adversely affected the scene graph’s quality. The predictions are dominated by coarse-grained relat…

2024

Generating Persona-Aware Empathetic Responses with Retrieval-Augmented Prompt Learning

ICASSP 2024accepted

Empathetic response generation requires perceiving and understanding the user’s emotion to deliver suitable responses. However, existing models generally lack an ability to respond in a persona-specific way, which has been shown to play a vital role in expressing appropriate empathy. To address this…

Cited by 0SourceScholar
2024

How Do Humans Write Code? Large Models Do It the Same Way Too

EMNLP 2024main

Program-of-Thought (PoT) replaces natural language-based Chain-of-Thought (CoT) as the most popular method in Large Language Models (LLMs) mathematical reasoning tasks by utilizing external tool calls to circumvent computational errors. However, our evaluation of the GPT-4 and Llama series reveals t…

2024

Hypernetwork-Assisted Parameter-Efficient Fine-Tuning with Meta-Knowledge Distillation for Domain Knowledge Disentanglement

NAACL 2024findings

Domain adaptation from labeled source domains to the target domain is important in practical summarization scenarios. However, the key challenge is domain knowledge disentanglement. In this work, we explore how to disentangle domain-invariant knowledge from source domains while learning specific kno…

Cited by 1SourcePDFScholar
2024

MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models

AAAI 2024technical

The emergence of various medical large language models (LLMs) in the medical domain has highlighted the need for unified evaluation standards, as manual evaluation of LLMs proves to be time-consuming and labor-intensive. To address this issue, we introduce MedBench, a comprehensive benchmark for the…

2023

A Disentangled-Attention Based Framework with Persona-Aware Prompt Learning for Dialogue Generation

AAAI 2023technical

Endowing dialogue agents with personas is the key to delivering more human-like conversations. However, existing persona-grounded dialogue systems still lack informative details of human conversations and tend to reply with inconsistent and generic responses. One of the main underlying causes is tha…

Cited by 5SourcePDFScholar
2023

Disentangled CVAEs with Contrastive Learning for Explainable Recommendation

AAAI 2023technical

Modern recommender systems are increasingly expected to provide informative explanations that enable users to understand the reason for particular recommendations. However, previous methods struggle to interpret the input IDs of user--item pairs in real-world datasets, failing to extract adequate ch…

Cited by 7SourcePDFScholar
2023

UMRSpell: Unifying the Detection and Correction Parts of Pre-trained Models towards Chinese Missing, Redundant, and Spelling Correction

ACL 2023long

Chinese Spelling Correction (CSC) is the task of detecting and correcting misspelled charac- ters in Chinese texts. As an important step for various downstream tasks, CSC confronts two challenges: 1) Character-level errors consist not only of spelling errors but also of missing and redundant ones th…

2022

Curriculum Prompt Learning with Self-Training for Abstractive Dialogue Summarization

EMNLP 2022main

Succinctly summarizing dialogue is a task of growing interest, but inherent challenges, such as insufficient training data and low information density impede our ability to train abstractive models. In this work, we propose a novel curriculum-based prompt learning method with self-training to addres…

2022

Multi-Scale Distribution Deep Variational Autoencoder for Explanation Generation

ACL 2022findings

Generating explanations for recommender systems is essential for improving their transparency, as users often wish to understand the reason for receiving a specified recommendation. Previous methods mainly focus on improving the generation quality, but often produce generic explanations that fail to…

Cited by 5SourcePDFScholar
2016

Improved DNN-based segmentation for multi-genre broadcast audio

ICASSP 2016accepted

Automatic segmentation is a crucial initial processing step for processing multi-genre broadcast (MGB) audio. It is very challenging since the data exhibits a wide range of both speech types and background conditions with many types of non-speech audio. This paper describes a segmentation system for…

Cited by 0SourceScholar