← Search

Shoushan Li

22 accepted papers

2026

GRASP: Awakening Latent Spatial Reasoning in LVLMs via Training-free Geometric Rectification

ICML 2026poster

Large Vision-Language Models (LVLMs) exhibit remarkable general capabilities but struggle significantly with spatial reasoning tasks. In this paper, we uncover a critical representation-output misalignment via linear probing: LVLMs correctly encode spatial features internally, but generate incorrect…

Cited by 0SourceScholar
2026

Mis: Light Response Agent for Video Comment with Multimodal Informative Seeking

ICRA 2026poster

Automatic response generation of video comments (RGVC) aims to generate a target reply to the content of the target comment based on the video context. Existing works for RGVC normally rely on large language models (LLMs), and mostly neglect the importance of extracting key information from both lin…

Cited by 0Scholar
2026

PDD-RRG: Posterior Diagnostic Decision for Study-level Radiology Report Generation

IJCAI 2026

Automatic radiology report generation (RRG) aims to simulate the workflow of radiologists, assisting them in clinical diagnosis. However, existing methods often fall short in utilizing all information relevant to the examination, as is typically done in clinical practice. Although some works attempt

Cited by 0Scholar
2025

A Comprehensive Graph Framework for Question Answering with Mode-Seeking Preference Alignment

ACL 2025finding

Recent advancements in retrieval-augmented generation (RAG) have enhanced large language models in question answering by integrating external knowledge. However, challenges persist in achieving global understanding and aligning responses with human ethical and quality preferences. To address these i…

2025

Exploring Knowledge Filtering for Retrieval-Augmented Discriminative Tasks

ACL 2025finding

Retrieval-augmented methods have achieved remarkable advancements in alleviating the hallucination of large language models.Nevertheless, the introduction of external knowledge does not always lead to the expected improvement in model performance, as irrelevant or harmful information present in the…

Cited by 0SourcePDFScholar
2025

Exploring Unified Training Framework for Multimodal User Profiling

COLING 2025main

With the emergence of social media and e-commerce platforms, accurate user profiling has become increasingly vital for recommendation systems and personalized services. Recent studies have focused on generating detailed user profiles by extracting various aspects of user attributes from textual revi…

Cited by 0SourcePDFScholar
2025

Pathological Section Staining Transferring with Tailored Metric-based Model Selection

ICASSP 2025accepted

As the important pathological section staining, Immunohistochemistry (IHC) staining uses labeled antibodies to highlight specific antigens, providing clearer results for malignancy identification compared with Hematoxylin and Eosin (H&E) staining. However, obtaining IHC manually is labor-intensive a…

Cited by 0SourceScholar
2025

Vision-aided Unsupervised Constituency Parsing with Multi-MLLM Debating

ACL 2025finding

This paper presents a novel framework for vision-aided unsupervised constituency parsing (VUCP), leveraging multimodal large language models (MLLMs) pre-trained on diverse image-text or video-text data. Unlike previous methods requiring explicit cross-modal alignment, our approach eliminates this ne…

Cited by 0SourcePDFScholar
2025

Zero-shot Cross-lingual NER via Mitigating Language Difference: An Entity-aligned Translation Perspective

EMNLP 2025

Cross-lingual Named Entity Recognition (CL-NER) aims to transfer knowledge from high-resource languages to low-resource languages. However, existing zero-shot CL-NER (ZCL-NER) approaches primarily focus on Latin script language (LSL), where shared linguistic features facilitate effective knowledge t

Cited by 0SourcePDFScholar
2024

Cross-domain NER with Generated Task-Oriented Knowledge: An Empirical Study from Information Density Perspective

EMNLP 2024main

Cross-domain Named Entity Recognition (CDNER) is crucial for Knowledge Graph (KG) construction and natural language processing (NLP), enabling learning from source to target domains with limited data. Previous studies often rely on manually collected entity-relevant sentences from the web or attempt…

2023

Word-level Prefix/Suffix Sense Detection: A Case Study on Negation Sense with Few-shot Learning

ACL 2023findings

Morphological analysis is an important research issue in the field of natural language processing. In this study, we propose a context-free morphological analysis task, namely word-level prefix/suffix sense detection, which deals with the ambiguity of sense expressed by prefix/suffix. To research th…

2022

Aspect-based Sentiment Analysis with Opinion Tree Generation

IJCAI 2022poster

Existing studies usually extract these sentiment elements by decomposing the complex structure prediction task into multiple subtasks. Despite their effectiveness, these methods ignore the semantic structure in ABSA problems and require extensive task-specific designs. In this study, we introduce an…

2022

KATG: Keyword-Bias-Aware Adversarial Text Generation for Text Classification

AAAI 2022technical

Recent work has shown that current text classification models are vulnerable to small adversarial perturbation to inputs, and adversarial training that re-trains the models with the support of adversarial examples is the most popular way to alleviate the impact of the perturbation. However, current…

Cited by 6SourcePDFScholar
2022

One-Teacher and Multiple-Student Knowledge Distillation on Sentiment Classification

COLING 2022main

Knowledge distillation is an effective method to transfer knowledge from a large pre-trained teacher model to a compacted student model. However, in previous studies, the distilled student models are still large and remain impractical in highly speed-sensitive systems (e.g., an IR system). In this s…

2021

Joint Multi-modal Aspect-Sentiment Analysis with Auxiliary Cross-modal Relation Detection

EMNLP 2021main

Aspect terms extraction (ATE) and aspect sentiment classification (ASC) are two fundamental and fine-grained sub-tasks in aspect-level sentiment analysis (ALSA). In the textual analysis, joint extracting both aspect terms and sentiment polarities has been drawn much attention due to the better appli…

2021

More than Text: Multi-modal Chinese Word Segmentation

ACL 2021short

Chinese word segmentation (CWS) is undoubtedly an important basic task in natural language processing. Previous works only focus on the textual modality, but there are often audio and video utterances (such as news broadcast and face-to-face dialogues), where textual, acoustic and visual modalities…

2021

Multi-modal Graph Fusion for Named Entity Recognition with Targeted Visual Guidance

AAAI 2021technical

Multi-modal named entity recognition (MNER) aims to discover named entities in free text and classify them into pre-defined types with images. However, dominant MNER models do not fully exploit fine-grained semantic correspondences between semantic units of different modalities, which have the poten…

2021

Multi-modal Multi-label Emotion Recognition with Heterogeneous Hierarchical Message Passing

AAAI 2021technical

As an important research issue in affective computing community, multi-modal emotion recognition has become a hot topic in the last few years. However, almost all existing studies perform multiple binary classification for each emotion with focus on complete time series data. In this paper, we focus…

2020

End-to-End Emotion-Cause Pair Extraction with Graph Convolutional Network

COLING 2020main

Emotion-cause pair extraction (ECPE), which aims at simultaneously extracting emotion-cause pairs that express emotions and their corresponding causes in a document, plays a vital role in understanding natural languages. Considering that most emotions usually have few causes mentioned in their conte…

2020

Multimodal Topic-Enriched Auxiliary Learning for Depression Detection

COLING 2020main

From the perspective of health psychology, human beings with long-term and sustained negativity are highly possible to be diagnosed with depression. Inspired by this, we argue that the global topic information derived from user-generated contents (e.g., texts and images) is crucial to boost the perf…