← Search

Qingcai Chen

26 accepted papers

2026

D2Dewarp: Dual Dimensions Geometric Representation Learning Based Document Image Dewarping

CVPR 2026

Document image dewarping remains a challenging task in the deep learning era. While existing methods have improved by leveraging text line awareness, they typically focus only on a single horizontal dimension. In this paper, we propose a fine-grained deformation perception model that focuses on Dual

Cited by 0SourcecodeScholar
2026

MMDIR: Multimodal Instruction-Driven Framework for Mixed-Degradation Document Image Restoration

CVPR 2026

Restoring degraded document image is essential for both improving visual quality and optimizing performance in downstream document analysis tasks. Although existing methods have demonstrated substantial improvements in restoration outcomes, they primarily address single-type degradation scenarios. C

Cited by 0SourcecodeScholar
2025

Reasoning Graph Enhanced Exemplars Retrieval for In-Context Learning

COLING 2025main

Large language models (LLMs) have exhibited remarkable few-shot learning capabilities and unified the paradigm of NLP tasks through the in-context learning (ICL) technique. Despite the success of ICL, the quality of the exemplar demonstrations can significantly influence the LLM’s performance. Exist…

2024

Discriminative Forests Improve Generative Diversity for Generative Adversarial Networks

AAAI 2024technical

Improving the diversity of Artificial Intelligence Generated Content (AIGC) is one of the fundamental problems in the theory of generative models such as generative adversarial networks (GANs). Previous studies have demonstrated that the discriminator in GANs should have high capacity and robustness…

2024

Linguistic Rule Induction Improves Adversarial and OOD Robustness in Large Language Models

COLING 2024main

Ensuring robustness is especially important when AI is deployed in responsible or safety-critical environments. ChatGPT can perform brilliantly in both adversarial and out-of-distribution (OOD) robustness, while other popular large language models (LLMs), like LLaMA-2, ERNIE and ChatGLM, do not perf…

Cited by 1SourcePDFScholar
2024

TDeLTA: A Light-Weight and Robust Table Detection Method Based on Learning Text Arrangement

AAAI 2024technical

The diversity of tables makes table detection a great challenge, leading to existing models becoming more tedious and complex. Despite achieving high performance, they often overfit to the table style in training set, and suffer from significant performance degradation when encountering out-of-distr…

2024

ZO-AdaMU Optimizer: Adapting Perturbation by the Momentum and Uncertainty in Zeroth-Order Optimization

AAAI 2024technical

Lowering the memory requirement in full-parameter training on large models has become a hot research area. MeZO fine-tunes the large language models (LLMs) by just forward passes in a zeroth-order SGD optimizer (ZO-SGD), demonstrating excellent performance with the same GPU memory usage as inference…

2023

Controllable Contrastive Generation for Multilingual Biomedical Entity Linking

EMNLP 2023long main

Multilingual biomedical entity linking (MBEL) aims to map language-specific mentions in the biomedical text to standardized concepts in a multilingual knowledge base (KB) such as Unified Medical Language System (UMLS). In this paper, we propose Con2GEN, a prompt-based controllable contrastive genera…

Cited by 0SourceScholar
2023

EARA: Improving Biomedical Semantic Textual Similarity with Entity-Aligned Attention and Retrieval Augmentation

EMNLP 2023long findings

Measuring Semantic Textual Similarity (STS) is a fundamental task in biomedical text processing, which aims at quantifying the similarity between two input biomedical sentences. Unfortunately, the STS datasets in the biomedical domain are relatively smaller but more complex in semantics than common…

Cited by 0SourcecodeScholar
2023

FashionSAP: Symbols and Attributes Prompt for Fine-Grained Fashion Vision-Language Pre-Training

CVPR 2023poster

Fashion vision-language pre-training models have shown efficacy for a wide range of downstream tasks. However, general vision-language pre-training models pay less attention to fine-grained domain features, while these features are important in distinguishing the specific domain tasks from general t…

2022

CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

ACL 2022long

Artificial Intelligence (AI), along with the recent progress in biomedical language understanding, is gradually offering great promise for medical practice. With the development of biomedical language understanding benchmarks, AI applications are widely used in the medical field. However, most bench…

2022

Calibration Meets Explanation: A Simple and Effective Approach for Model Confidence Estimates

EMNLP 2022main

Calibration strengthens the trustworthiness of black-box models by producing better accurate confidence estimates on given examples. However, little is known about if model explanations can help confidence calibration. Intuitively, humans look at important features attributions and decide whether th…

2022

Diaformer: Automatic Diagnosis via Symptoms Sequence Generation

AAAI 2022technical

Automatic diagnosis has attracted increasing attention but remains challenging due to multi-step reasoning. Recent works usually address it by reinforcement learning methods. However, these methods show low efficiency and require task-specific reward functions. Considering the conversation between d…

2022

End-to-End ASR-Enhanced Neural Network for Alzheimer's Disease Diagnosis

ICASSP 2022accepted

This paper presents an approach to Alzheimer’s disease (AD) diagnosis from spontaneous speech using an end-to-end ASR-enhanced neural network. Under the condition that only audio data are provided and accurate transcripts are unavailable, this paper proposes a system that can analyze utterances to d…

Cited by 0SourceScholar
2022

Enhancing Entity Representations with Prompt Learning for Biomedical Entity Linking

IJCAI 2022poster

Biomedical entity linking aims to map mentions in biomedical text to standardized concepts or entities in a curated knowledge base (KB) such as Unified Medical Language System (UMLS). The latest research tends to solve this problem in a unified framework solely based on surface form matching between…

2022

Unifying Model Explainability and Robustness for Joint Text Classification and Rationale Extraction

AAAI 2022technical

Recent works have shown explainability and robustness are two crucial ingredients of trustworthy and reliable text classification. However, previous works usually address one of two aspects: i) how to extract accurate rationales for explainability while being beneficial to prediction; ii) how to mak…

2021

HCAG: A Hierarchical Context-Aware Graph Attention Model for Depression Detection

ICASSP 2021accepted

Depression is one of the most common mental health disorders, it’s crucial to design an effective and robust model for automatic depression detection (ADD). Although current approaches rely on extra topic models or manually topic-selection procedures which is time-consuming, they still haven’t thoro…

Cited by 0SourceScholar
2021

Leveraging Capsule Routing to Associate Knowledge with Medical Literature Hierarchically

EMNLP 2021main

Integrating knowledge into text is a promising way to enrich text representation, especially in the medical field. However, undifferentiated knowledge not only confuses the text representation but also imports unexpected noises. In this paper, to alleviate this problem, we propose leveraging capsule…

2021

Multi-hop Graph Convolutional Network with High-order Chebyshev Approximation for Text Reasoning

ACL 2021long

Graph convolutional network (GCN) has become popular in various natural language processing (NLP) tasks with its superiority in long-term and non-consecutive word interactions. However, existing single-hop graph reasoning in GCN may miss some important non-consecutive dependencies. In this study, we…

2020

Learning to Generate Diverse Questions from Keywords

ICASSP 2020accepted

Diverse text generation has been emerging as an important topic of natural language generation. Traditional studies on question generation mainly investigate how to generate one question based on a given input (one-to-one). In this paper, we focus on a more complex question generation task, i.e., ge…

Cited by 0SourceScholar
2020

MedWriter: Knowledge-Aware Medical Text Generation

COLING 2020main

To exploit the domain knowledge to guarantee the correctness of generated text has been a hot topic in recent years, especially for high professional domains such as medical. However, most of recent works only consider the information of unstructured text rather than structured information of the kn…

Cited by 7SourcePDFScholar
2020

SED-MDD: Towards Sentence Dependent End-To-End Mispronunciation Detection and Diagnosis

ICASSP 2020accepted

A mispronunciation detection and diagnosis (MD&D) system typically consists of multiple stages, such as an acoustic model, a language model and a Viterbi decoder. In order to integrate these stages, we propose SED-MDD, an end-to-end model for sentence dependent mispronunciation detection and diagnos…

Cited by 0SourceScholar