← Search

Deqiang Jiang

19 accepted papers

2026

DenseMLLM: Standard Multimodal LLMs are Intrinsic Dense Predictors

ICML 2026poster

Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in high-level visual understanding. However, extending these models to fine-grained dense prediction tasks, such as semantic segmentation and depth estimation, typically necessitates the incorporation of complex, tas…

Cited by 0SourceScholar
2026

RAR: Reversing Visual Attention Re-Sinking for Unlocking Potential in Multimodal Large Language Models

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks, yet they frequently exhibit suboptimal output layers, where intermediate decoder layers outperform the final ones, signaling underutilized model capacity. In this work, we delve into the root causes a…

Cited by 0SourceScholar
2026

UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm

CVPR 2026

Despite remarkable advancements in multimodal large language models (MLLMs), their fine-grained visual understanding is constrained by a primary reliance on sparse textual supervision. Existing efforts to introduce visual supervision typically do so during post-training, when visual representations

Cited by 0SourceScholar
2024

Enhancing Visual Document Understanding with Contrastive Learning in Large Visual-Language Models

CVPR 2024poster

Recently the advent of Large Visual-Language Models (LVLMs) has received increasing attention across various domains particularly in the field of visual document understanding (VDU). Different from conventional vision-language tasks VDU is specifically concerned with text-rich scenarios containing a…

Cited by 14SourcePDFScholar
2024

Few-shot Temporal Pruning Accelerates Diffusion Models for Text Generation

COLING 2024main

Diffusion models have achieved significant success in computer vision and shown immense potential in natural language processing applications, particularly for text generation tasks. However, generating high-quality text using these models often necessitates thousands of iterations, leading to slow…

Cited by 1SourcePDFScholar
2024

Grab What You Need: Rethinking Complex Table Structure Recognition with Flexible Components Deliberation

AAAI 2024technical

Recently, Table Structure Recognition (TSR) task, aiming at identifying table structure into machine readable formats, has received increasing interest in the community. While impressive success, most single table component-based methods can not perform well on unregularized table cases distracted b…

Cited by 9SourcePDFScholar
2024

HRVDA: High-Resolution Visual Document Assistant

CVPR 2024poster

Leveraging vast training data multimodal large language models (MLLMs) have demonstrated formidable general visual comprehension capabilities and achieved remarkable performance across various tasks. However their performance in visual document understanding still leaves much room for improvement. T…

Cited by 20SourcePDFScholar
2024

Talk With Human-like Agents: Empathetic Dialogue Through Perceptible Acoustic Reception and Reaction

ACL 2024long

Large Language Model (LLM)-enhanced agents become increasingly prevalent in Human-AI communication, offering vast potential from entertainment to professional domains. However, current multi-modal dialogue systems overlook the acoustic information present in speech, which is crucial for understandin…

2023

Attention Where It Matters: Rethinking Visual Document Understanding with Selective Region Concentration

ICCV 2023poster

We propose a novel end-to-end document understanding model called SeRum (SElective Region Understanding Model) for extracting meaningful information from document images, including document analysis, retrieval, and office automation. Unlike state-of-the-art approaches that rely on multi-stage techni…

Cited by 15PDFScholar
2023

TaCo: Textual Attribute Recognition via Contrastive Learning

AAAI 2023technical

As textual attributes like font are core design elements of document format and page style, automatic attributes recognition favor comprehensive practical applications. Existing approaches already yield satisfactory performance in differentiating disparate attributes, but they still suffer in distin…

Cited by 0SourcePDFScholar
2023

The Devil Is in the Frequency: Geminated Gestalt Autoencoder for Self-Supervised Visual Pre-training

AAAI 2023technical

The self-supervised Masked Image Modeling (MIM) schema, following "mask-and-reconstruct" pipeline of recovering contents from masked image, has recently captured the increasing interest in the community, owing to the excellent ability of learning visual representation from unlabeled data. Aiming at…

Cited by 38SourcePDFScholar
2023

Turning a CLIP Model Into a Scene Text Detector

CVPR 2023poster

The recent large-scale Contrastive Language-Image Pretraining (CLIP) model has shown great potential in various downstream tasks via leveraging the pretrained vision and language knowledge. Scene text, which contains rich textual and visual information, has an inherent connection with a model like C…

2022

GMN: Generative Multi-modal Network for Practical Document Information Extraction

NAACL 2022long

Document Information Extraction (DIE) has attracted increasing attention due to its various advanced applications in the real world. Although recent literature has already achieved competitive results, these approaches usually fail when dealing with complex documents with noisy OCR results or mutati…

Cited by 12SourcePDFScholar
2022

Neural Collaborative Graph Machines for Table Structure Recognition

CVPR 2022poster

Recently, table structure recognition has achieved impressive progress with the help of deep graph models. Most of them exploit single visual cues of tabular elements or simply combine visual cues with other modalities via early fusion to reason their graph relationships. However, neither early fusi…

Cited by 41PDFScholar
2022

NomMer: Nominate Synergistic Context in Vision Transformer for Visual Recognition

CVPR 2022poster

Recently, Vision Transformers (ViT), with the self-attention (SA) as the de facto ingredients, have demonstrated great potential in the computer vision community. For the sake of trade-off between efficiency and performance, a group of works merely perform SA operation within local patches, whereas…

Cited by 20PDFcodeScholar
2022

Perceiving Stroke-Semantic Context: Hierarchical Contrastive Learning for Robust Scene Text Recognition

AAAI 2022technical

We introduce Perceiving Stroke-Semantic Context (PerSec), a new approach to self-supervised representation learning tailored for Scene Text Recognition (STR) task. Considering scene text images carry both visual and semantic properties, we equip our PerSec with dual context perceivers which can cont…

Cited by 50SourcePDFScholar
2022

Semantic-Preserving Abstractive Text Summarization with Siamese Generative Adversarial Net

NAACL 2022findings

We propose a novel siamese generative adversarial net for abstractive text summarization (SSPGAN), which can preserve the main semantics of the source text. Different from previous generative adversarial net based methods, SSPGAN is equipped with a siamese semantic-preserving discriminator, which ca…

2022

Sequence-to-Action: Grammatical Error Correction with Action Guided Sequence Generation

AAAI 2022technical

The task of Grammatical Error Correction (GEC) has received remarkable attention with wide applications in Natural Language Processing (NLP) in recent years. While one of the key principles of GEC is to keep the correct parts unchanged and avoid over-correction, previous sequence-to-sequence (seq2se…

2021

Hierarchical Multi-label Text Classification with Horizontal and Vertical Category Correlations

EMNLP 2021main

Hierarchical multi-label text classification (HMTC) deals with the challenging task where an instance can be assigned to multiple hierarchically structured categories at the same time. The majority of prior studies either focus on reducing the HMTC task into a flat multi-label problem ignoring the v…

Cited by 31SourcePDFScholar