← Search

Zhendong Mao

85 accepted papers

2026

Cat-PO: Cross-modal Adaptive Token-rewards for Preference Optimization in Truthful Multimodal LLMs

ICLR 2026poster

Multi-modal Large Language Models (MLLMs) have shown remarkable generative capabilities across multi-modal tasks, yet remain plagued by hallucinations where generated textual contents are semantically inconsistent with the input images. This work reveals that existing multi-modal preference optimiza…

Cited by 0SourceScholar
2026

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

ICLR 2026poster

Deep Research Agents (DRAs) are emerging as one of the most practical classes of LLM-based agents. Given an open-ended research task, they find, analyze, and synthesize large numbers of online sources to produce a comprehensive report at the level of a research analyst. This can compress hours of ma…

Cited by 0SourcecodeScholar
2026

LayerEdit: Disentangled Multi-Object Editing via Conflict-Aware Multi-Layer Learning

AAAI 2026technical

Text-driven multi-object image editing which aims to precisely modify multiple objects within an image based on text descriptions, has recently attracted considerable interest. Existing works primarily follow the localize-editing paradigm, focusing on independent object localization and editing whil

Cited by 0SourcePDFScholar
2026

MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools

AAAI 2026technical

The Model Context Protocol (MCP) is rapidly emerging as a pivotal open standard, designed to enhance agent-tool integration and interoperability, and is positioned to unlock a new era of powerful, interconnected, and genuinely utilitarian agentic AI. However, despite MCP

Cited by 0SourcePDFScholar
2026

MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation

ICLR 2026poster

Long video generation with Diffusion Transformers (DiTs) is bottlenecked by the quadratic scaling of full attention with sequence length. Since attention is highly redundant, outputs are dominated by a small subset of query–key pairs. Existing sparse methods rely on blockwise coarse estimation, whos…

Cited by 0SourcecodeScholar
2026

Self-guided Semantic Inspection for Zero-Shot Composed Image Retrieval

CVPR 2026

Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve target images using a composed query of a reference image and a textual modification, without relying on triplet-based supervision. As the two inputs describe related but semantically unaligned information, the key challenge lies in interp

Cited by 0SourcecodeScholar
2026

SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder

AAAI 2026technical

Reward models (RMs) are a core component in the post-training of large language models (LLMs), serving as proxies for human preference evaluation and guiding model alignment. However, training reliable RMs under limited resources remains challenging due to the reliance on large-scale preference anno

Cited by 0SourcePDFScholar
2026

UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation

ICML 2026poster

Multi-modal image generation, particularly subject-driven customization, has garnered growing attention in recent years. Despite the rapid advancement of generative models, their evaluation remains largely lagging. Existing methods, whether embedding-based or Multi-modal Large Language Model (MLLM)-…

Cited by 0SourceScholar
2026

Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models

CVPR 2026

Direct Preference Optimization (DPO) has proven to be an effective solution for mitigating hallucination in Multimodal Large Language Models (MLLMs) by learning from preference pairs. One of its key challenges lies in how to transfer the sequence-level preference into fine-grained supervision on vis

Cited by 0SourcecodeScholar
2026

Video-LevelGauge: Investigating Contextual Positional Bias in Video Language Models.

ICLR 2026poster

Large video language models (LVLMs) have made notable progress in video understanding, spurring the development of corresponding evaluation benchmarks. However, existing benchmarks generally assess overall performance across entire video sequences, overlooking nuanced behaviors such as contextual po…

Cited by 0SourceScholar
2025

A4A: Adapter for Adapter Transfer via All-for-All Mapping for Cross-Architecture Models

CVPR 2025poster

Large-scale text-to-image models evolve rapidly in size and architecture. The existing adapters struggle to keep pace with these models, requiring extensive retraining. This paper proposes a novel adapter transfer framework, A4A (Adapter for Adapter), which uses an all-for-all mapping approach to se…

Cited by 0SourcePDFScholar
2025

Alleviating Hallucinations in Large Language Models via Truthfulness-driven Rank-adaptive LoRA

ACL 2025finding

Improving the truthfulness of LLMs to alleviate hallucinations has become critical for promoting the practical deployment of LLMs. Current fine-tuning-based methods ignore the intrinsic discrepancy in the truthfulness correlations across LLM internal modules, and instead treat them equally, which ma…

2025

Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach

EMNLP 2025

Creative writing is a key capability of Large Language Models (LLMs), with potential applications in literature, storytelling, and various creative domains. However, evaluating the creativity of machine-generated texts remains a significant challenge, as existing methods either rely on costly manual

Cited by 0SourcePDFScholar
2025

CustomContrast: A Multilevel Contrastive Perspective for Subject-Driven Text-to-Image Customization

AAAI 2025technical

Subject-driven text-to-image (T2I) customization has drawn significant interest in academia and industry. This task enables pre-trained models to generate novel images based on unique subjects. Existing studies adopt a self-reconstructive perspective, focusing on capturing all details of a single im…

Cited by 6SourcePDFScholar
2025

D^2iT: Dynamic Diffusion Transformer for Accurate Image Generation

CVPR 2025poster

Diffusion models are widely recognized for their ability to generate high-fidelity images. Despite the excellent performance and scalability of the Diffusion Transformer (DiT) architecture, it applies fixed compression across different image regions during the diffusion process, disregarding the nat…

2025

DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video Customization

ICCV 2025poster

Customized text-to-video generation with pre-trained large-scale models has recently garnered significant attention by focusing on identity and motion consistency. Existing works typically follow the isolated customized paradigm, where the subject identity or motion dynamics are customized exclusive…

2025

ELDER: Enhancing Lifelong Model Editing with Mixture-of-LoRA

AAAI 2025technical

Large language models (LLMs) require model editing to efficiently update specific knowledge within them and avoid factual errors. Most model editing methods are solely designed for single-time use and result in a significant forgetting effect in lifelong editing scenarios, where sequential edits ar…

2025

FeedEdit: Text-Based Image Editing with Dynamic Feedback Regulation

CVPR 2025poster

Text-based image editing which aims at generating rigid or non-rigid changes to images conditioned on the given text, has recently attracted considerable interest. Previous works mainly follow the multi-step denoising diffusion paradigm, which adopts a fixed text guidance intensity (i.e., editing in…

Cited by 0SourcePDFScholar
2025

Fine-grained Knowledge Enhancement for Retrieval-Augmented Generation

ACL 2025finding

Retrieval-augmented generation (RAG) effectively mitigates hallucinations in large language models (LLMs) by filling knowledge gaps with retrieved external information. Most existing studies primarily retrieve knowledge documents based on semantic similarity to assist in answering questions but igno…

Cited by 0SourcePDFScholar
2025

From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding

ACL 2025long

The pursuit of diverse, complex, and large-scale instruction data is crucial for automatically aligning large language models (LLMs). While there are methods capable of generating synthetic instructions at scale, they either suffer from limited grounding sources, leading to a narrow distribution, or…

2025

Graph Mixture of Experts and Memory-augmented Routers for Multivariate Time Series Anomaly Detection

AAAI 2025technical

Multivariate time series (MTS) anomaly detection is a critical task that involves identifying abnormal patterns or events in data that consist of multiple interrelated time series. In order to better model the complex interdependence between entities and the various inherent characteristics of each…

2025

Hierarchy-Aware Pseudo Word Learning with Text Adaptation for Zero-Shot Composed Image Retrieval

ICCV 2025poster

Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve the target image based on a reference image and a text describing the user's intention without training on the triplet datasets. The key to this task is to make specified changes to specific objects in the reference image based on the text…

Cited by 0SourcePDFScholar
2025

Improve Safety Training of Large Language Models with Safety-Critical Singular Vectors Localization

ACL 2025long

The rapid advancement of large language models (LLMs) has brought about increased concerns regarding their safety, especially as adversaries develop jailbreak techniques to bypass LLMs’ safety mechanism. Although recent work on safety training with modules such as low-rank adaptation (LoRA) to resis…

2025

Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models

NeurIPS 2025poster

The widespread adoption of large language models (LLMs) across industries has increased the demand for high-quality and customizable outputs. However, traditional alignment methods often require retraining large pretrained models, making it difficult to quickly adapt and optimize LLMs for diverse ap…

Cited by 0SourceScholar
2025

Leveraging robust optimization for llm alignment under distribution shifts

NeurIPS 2025poster

Preference alignment methods are increasingly critical for steering large language models (LLMs) to generate outputs consistent with human values. While recent approaches often rely on synthetic data generated by LLMs for scalability and cost-efficiency reasons, this reliance can introduce distribut…

Cited by 0SourceScholar
2025

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory

ICCV 2025poster

Animation colorization is a crucial part of real animation industry production. Long animation colorization has high labor costs. Therefore, automated long animation colorization based on the video generation model has significant research value. Existing studies are limited to short-term colorizati…

Cited by 0SourcePDFScholar
2025

M-RangeDetector: Enhancing Generalization in Machine-Generated Text Detection through Multi-Range Attention Masks

ACL 2025finding

The increasing capability and widespread usage of large language models (LLMs) highlight the desirability of automatic detection of machine-generated text. Existing supervised detectors often overfit within their training domains, as they have primarily learned domain-specific textual features, such…

Cited by 0SourcePDFScholar
2025

MIRROR: Multi-agent Intra- and Inter-Reflection for Optimized Reasoning in Tool Learning

IJCAI 2025

Complex tasks involving tool integration pose significant challenges for Large Language Models (LLMs), leading to the emergence of multi-agent workflows as a promising solution. Reflection has emerged as an effective strategy for correcting erroneous trajectories in agentic workflows. However, exist

Cited by 0SourcePDFScholar
2025

Multi-Prototype Grouping for Continual Learning in Visual Question Answering

ICASSP 2025accepted

Visual Question Answering (VQA) aims to answer questions utilizing information from both textual and visual modalities. New data categories and novel combinations of the two modalities will continuously emerge in practical applications, necessitating continual learning. For this unique compositional…

Cited by 0SourceScholar
2025

On-the-fly Preference Alignment via Principle-Guided Decoding

ICLR 2025poster

With the rapidly expanding landscape of large language models, aligning model generations with human values and preferences is becoming increasingly important. Popular alignment methods, such as Reinforcement Learning from Human Feedback, have shown significant success in guiding models with greater…

2025

Pro3D-Editor: A Progressive Framework for Consistent and Precise 3D Editing

NeurIPS 2025poster

Text-guided 3D editing aims to locally modify 3D objects based on editing prompts, which has significant potential for applications in 3D game and film domains. Existing methods typically follow a view-agnostic paradigm: editing 2D view images indiscriminately and projecting them back into 3D space.…

Cited by 0SourceScholar
2025

Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability

ACL 2025finding

Training language models with rationales augmentation has been shown to be beneficial in many existing works. In this paper, we identify that such a prevailing view does not hold consistently. We conduct comprehensive investigations to thoroughly inspect the impact of rationales on model performance…

2025

RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models

ICCV 2025poster

Unifying diverse image generation tasks within a single framework remains a fundamental challenge in visual generation. While large language models (LLMs) achieve unification through task-agnostic data and generation, existing visual generation models fail to meet these principles. Current approache…

2025

SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation

CVPR 2025poster

Vision-language temporal alignment is a crucial capability for human dynamic recognition and cognition in real-world scenarios. While existing research focuses on capturing vision-language relevance, it faces limitations due to biased temporal distributions, imprecise annotations, and insufficient c…

Cited by 0SourcePDFScholar
2024

Benchmarking and Improving Compositional Generalization of Multi-aspect Controllable Text Generation

ACL 2024long

Compositional generalization, representing the model’s ability to generate text with new attribute combinations obtained by recombining single attributes from the training data, is a crucial property for multi-aspect controllable text generation (MCTG) methods. Nonetheless, a comprehensive compositi…

2024

Chain-of-Question: A Progressive Question Decomposition Approach for Complex Knowledge Base Question Answering

ACL 2024findings

Complex KBQA leverages the knowledge base (KB) to answer complex natural questions involving complicated semantics like multi-hop reasoning. Existing methods involve a question decomposition process, i.e., breaking a complex question into several simpler sub-questions, to assist obtaining logical fo…

Cited by 0SourcePDFScholar
2024

Disentangled Learning with Synthetic Parallel Data for Text Style Transfer

ACL 2024long

Text style transfer (TST) is an important task in natural language generation, which aims to transfer the text style (e.g., sentiment) while keeping its semantic information. Due to the absence of parallel datasets for supervision, most existing studies have been conducted in an unsupervised manner,…

2024

DreamIdentity: Enhanced Editability for Efficient Face-Identity Preserved Image Generation

AAAI 2024technical

While large-scale pre-trained text-to-image models can synthesize diverse and high-quality human-centric images, an intractable problem is how to preserve the face identity and follow the text prompts simultaneously for conditioned input face images and texts. Despite existing encoder-based methods…

Cited by 34SourcePDFScholar
2024

Feature-Adaptive and Data-Scalable In-Context Learning

ACL 2024long

In-context learning (ICL), which promotes inference with several demonstrations, has become a widespread paradigm to stimulate LLM capabilities for downstream tasks. Due to context length constraints, it cannot be further improved in spite of more training data, and general features directly from LL…

2024

Gradual Residuals Alignment: A Dual-Stream Framework for GAN Inversion and Image Attribute Editing

AAAI 2024technical

GAN-based image attribute editing firstly leverages GAN Inversion to project real images into the latent space of GAN and then manipulates corresponding latent codes. Recent inversion methods mainly utilize additional high-bit features to improve image details preservation, as low-bit codes cannot f…

Cited by 3SourcePDFScholar
2024

Homology Consistency Constrained Efficient Tuning for Vision-Language Models

NeurIPS 2024poster

Efficient transfer learning has shown remarkable performance in tuning large-scale vision-language models (VLMs) toward downstream tasks with limited data resources. The key challenge of efficient transfer lies in adjusting image-text alignment to be task-specific while preserving pre-trained genera…

Cited by 0SourcePDFScholar
2024

IDEATE: Detecting AI-Generated Text Using Internal and External Factual Structures

COLING 2024main

The effective detection of AI-generated text is a vital principle to ensure responsible use of large language models (LLMs). Previous studies mainly focused on discovering and utilizing internal evidences contained in the text itself to perform the detection, while ignoring external evidences implic…

2024

Identification of Necessary Semantic Undertakers in the Causal View for Image-Text Matching

AAAI 2024technical

Image-text matching bridges vision and language, which is a fundamental task in multimodal intelligence. Its key challenge lies in how to capture visual-semantic relevance. Fine-grained semantic interactions come from fragment alignments between image regions and text words. However, not all fragmen…

2024

Improving Radiology Report Generation with D2-Net: When Diffusion Meets Discriminator

ICASSP 2024accepted

Radiology report generation (RRG) aims to automatically provide observations and insight into a patient’s condition based on radiology images, which is able to greatly reduce the workload of physicians on the premise of ensuring the quality of medical treatment. Existing works leverage the Transform…

Cited by 0SourceScholar
2024

KNN-Instruct: Automatic Instruction Construction with K Nearest Neighbor Deduction

EMNLP 2024main

Supervised fine-tuning (SFT) is a critical procedure for aligning large language models. Despite its efficiency, the construction of SFT data often struggles with issues of quality, diversity, and scalability. Many existing methods, inspired by the Self-Instruct framework, typically generate synthet…

2024

Knowledge Context Modeling with Pre-trained Language Models for Contrastive Knowledge Graph Completion

ACL 2024findings

Text-based knowledge graph completion (KGC) methods utilize pre-trained language models for triple encoding and further fine-tune the model to achieve completion. Despite their excellent performance, they neglect the knowledge context in inferring process. Intuitively, knowledge contexts, which refe…

Cited by 5SourcePDFScholar
2024

LIRE: listwise reward enhancement for preference alignment

ACL 2024findings

Recently, tremendous strides have been made to align the generation of Large Language Models (LLMs) with human values to mitigate toxic or unhelpful content. Leveraging Reinforcement Learning from Human Feedback (RLHF) proves effective and is widely adopted by researchers. However, implementing RLHF…

2024

Linguistic-Aware Patch Slimming Framework for Fine-grained Cross-Modal Alignment

CVPR 2024poster

Cross-modal alignment aims to build a bridge connecting vision and language. It is an important multi-modal task that efficiently learns the semantic similarities between images and texts. Traditional fine-grained alignment methods heavily rely on pre-trained object detectors to extract region featu…

2024

RESEMO: A Benchmark Chinese Dataset for Studying Responsive Emotion from Social Media Content

ACL 2024findings

On social media platforms, users’ emotions are triggered when they encounter particular content from other users,where such emotions are different from those that spontaneously emerged, owing to the “responsive” nature. Analyzing the aforementioned responsive emotions from user interactions is a tas…

Cited by 0SourcePDFScholar
2024

RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization

CVPR 2024poster

Text-to-image customization which aims to synthesize text-driven images for the given subjects has recently revolutionized content creation. Existing works follow the pseudo-word paradigm i.e. represent the given subjects as pseudo-words and then compose them with the given text. However the inheren…

2024

Visual-Linguistic Dependency Encoding for Image-Text Retrieval

COLING 2024main

Image-text retrieval is a fundamental task to bridge the semantic gap between natural language and vision. Recent works primarily focus on aligning textual meanings with visual appearance. However, they often overlook the semantic discrepancy caused by syntactic structure in natural language express…

2023

$k$NN Prompting: Beyond-Context Learning with Calibration-Free Nearest Neighbor Inference

ICLR 2023poster

In-Context Learning (ICL), which formulates target tasks as prompt completion conditioned on in-context demonstrations, has become the prevailing utilization of LLMs. In this paper, we first disclose an actual predicament for this typical usage that it can not scale up with training data due to cont…

2023

Air-Decoding: Attribute Distribution Reconstruction for Decoding-Time Controllable Text Generation

EMNLP 2023long main

Controllable text generation (CTG) aims to generate text with desired attributes, and decoding-time-based methods have shown promising performance on this task. However, in this paper, we identify the phenomenon of Attribute Collapse for the first time. It causes the fluency of generated text to rap…

Cited by 0SourcecodeScholar
2023

Crossing the Gap: Domain Generalization for Image Captioning

CVPR 2023poster

Existing image captioning methods are under the assumption that the training and testing data are from the same domain or that the data from the target domain (i.e., the domain that testing data lie in) are accessible. However, this assumption is invalid in real-world applications where the data fro…

Cited by 19SourcePDFScholar
2023

Grammatical Error Correction via Mixed-Grained Weighted Training

EMNLP 2023long findings

The task of Grammatical Error Correction (GEC) aims to automatically correct grammatical errors in natural texts. Almost all previous works treat annotated training data equally, but inherent discrepancies in data are neglected. In this paper, the inherent discrepancies are manifested in two aspect…

Cited by 0SourceScholar
2023

IAEval: A Comprehensive Evaluation of Instance Attribution on Natural Language Understanding

EMNLP 2023long findings

Instance attribution (IA) aims to identify the training instances leading to the prediction of a test example, helping researchers understand the dataset better and optimize data processing. While many IA methods have been proposed recently, how to evaluate them still remains open. Previous evaluati…

Cited by 0SourceScholar
2023

Improving Image Captioning via Predicting Structured Concepts

EMNLP 2023long main

Having the difficulty of solving the semantic gap between images and texts for the image captioning task, conventional studies in this area paid some attention to treating semantic concepts as a bridge between the two modalities and improved captioning performance accordingly. Although promising res…

Cited by 0SourceScholar
2023

Inductive Relation Prediction from Relational Paths and Context with Hierarchical Transformers

ICASSP 2023accepted

Relation prediction on knowledge graphs (KGs) is a key research topic. Dominant embedding-based methods mainly focus on the transductive setting and lack the inductive ability to generalize to new entities for inference. Existing methods for inductive reasoning mostly mine the connections between en…

Cited by 0SourceScholar
2023

Learning Semantic Relationship Among Instances for Image-Text Matching

CVPR 2023poster

Image-text matching, a bridge connecting image and language, is an important task, which generally learns a holistic cross-modal embedding to achieve a high-quality semantic alignment between the two modalities. However, previous studies only focus on capturing fragment-level relation within a sampl…

2023

Not All Image Regions Matter: Masked Vector Quantization for Autoregressive Image Generation

CVPR 2023poster

Existing autoregressive models follow the two-stage generation paradigm that first learns a codebook in the latent space for image reconstruction and then completes the image generation autoregressively based on the learned codebook. However, existing codebook learning simply models all local region…

2023

On the Calibration of Large Language Models and Alignment

EMNLP 2023long findings

As large language models attract increasing attention and find widespread application, concurrent challenges of reliability also arise at the same time. Confidence calibration, an effective analysis method for gauging the reliability of deep models, serves as a crucial tool for assessing and improvi…

Cited by 0SourceScholar
2023

Random Entity Quantization for Parameter-Efficient Compositional Knowledge Graph Representation

EMNLP 2023long main

Representation Learning on Knowledge Graphs (KGs) is essential for downstream tasks. The dominant approach, KG Embedding (KGE), represents entities with independent vectors and faces the scalability challenge. Recent studies propose an alternative way for parameter efficiency, which represents ent…

Cited by 0SourcecodeScholar
2023

S2ynRE: Two-stage Self-training with Synthetic data for Low-resource Relation Extraction

ACL 2023long

Current relation extraction methods suffer from the inadequacy of large-scale annotated data. While distant supervision alleviates the problem of data quantities, there still exists domain disparity in data qualities due to its reliance on domain-restrained knowledge bases. In this work, we propose…

2023

SADE: A Self-Adaptive Expert for Multi-Dataset Question Answering

ICASSP 2023accepted

Multi-dataset question answering (QA) aims to combine multiple QA datasets to build models that not only perform well on training distributions, but also transfer and generalize well to new distributions. Some prior work considered building a collection of dataset-specific experts upon a shared Tran…

Cited by 0SourceScholar
2023

Text Style Transfer with Contrastive Transfer Pattern Mining

ACL 2023long

Text style transfer (TST) is an important task in natural language generation, which aims to alter the stylistic attributes (e.g., sentiment) of a sentence and keep its semantic meaning unchanged. Most existing studies mainly focus on the transformation between styles, yet ignore that this transform…

2023

Towards Accurate Image Coding: Improved Autoregressive Image Generation With Dynamic Vector Quantization

CVPR 2023highlight

Existing vector quantization (VQ) based autoregressive models follow a two-stage generation paradigm that first learns a codebook to encode images as discrete codes, and then completes generation based on the learned codebook. However, they encode fixed-size image regions into fixed-length codes and…

2022

ER-SAN: Enhanced-Adaptive Relation Self-Attention Network for Image Captioning

IJCAI 2022poster

Image captioning (IC), bringing vision to language, has drawn extensive attention. Precisely describing visual relations between image objects is a key challenge in IC. We argue that the visual relations, that is geometric positions (i.e., distance and size) and semantic interactions (i.e., actions…

2022

EmRel: Joint Representation of Entities and Embedded Relations for Multi-triple Extraction

NAACL 2022long

Multi-triple extraction is a challenging task due to the existence of informative inter-triple correlations, and consequently rich interactions across the constituent entities and relations. While existing works only explore entity representations, we propose to explicitly introduce relation represe…

2022

Improving Chinese Spelling Check by Character Pronunciation Prediction: The Effects of Adaptivity and Granularity

EMNLP 2022main

Chinese spelling check (CSC) is a fundamental NLP task that detects and corrects spelling errors in Chinese texts. As most of these spelling errors are caused by phonetic similarity, effectively modeling the pronunciation of Chinese characters is a key factor for CSC. In this paper, we consider intr…

2022

Show Your Faith: Cross-Modal Confidence-Aware Network for Image-Text Matching

AAAI 2022technical

Image-text matching bridges vision and language, which is a crucial task in the field of multi-modal intelligence. The key challenge lies in how to measure image-text relevance accurately as matching evidence. Most existing works aggregate the local semantic similarities of matched region-word pairs…

2022

UniRel: Unified Representation and Interaction for Joint Relational Triple Extraction

EMNLP 2022main

Relational triple extraction is challenging for its difficulty in capturing rich correlations between entities and relations. Existing works suffer from 1) heterogeneous representations of entities and relations, and 2) heterogeneous modeling of entity-entity interactions and entity-relation interac…

2021

Entity Structure Within and Throughout: Modeling Mention Dependencies for Document-Level Relation Extraction

AAAI 2021technical

Entities, as the essential elements in relation extraction tasks, exhibit certain structure. In this work, we formulate such entity structure as distinctive dependencies between mention pairs. We then propose SSAN, which incorporates these structural dependencies within the standard self-attention m…

2021

Image Captioning with Context-Aware Auxiliary Guidance

AAAI 2021technical

Image captioning is a challenging computer vision task, which aims to generate a natural language description of an image. Most recent researches follow the encoder-decoder framework which depends heavily on the previous generated words for the current prediction. Such methods can not effectively ta…

2021

Lesion-Aware Transformers for Diabetic Retinopathy Grading

CVPR 2021poster

Diabetic retinopathy (DR) is the leading cause of permanent blindness in the working-age population. And automatic DR diagnosis can assist ophthalmologists to design tailored treatments for patients, including DR grading and lesion discovery. However, most of existing methods treat DR grading and le…

Cited by 137PDFScholar
2021

Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Recognition

CVPR 2021poster

Linguistic knowledge is of great benefit to scene text recognition. However, how to effectively model linguistic rules in end-to-end deep networks remains a research challenge. In this paper, we argue that the limited capacity of language models comes from: 1) implicitly language modeling; 2) unidir…

Cited by 461PDFcodeScholar
2020

Graph Structured Network for Image-Text Matching

CVPR 2020poster

Image-text matching has received growing interest since it bridges vision and language. The key challenge lies in how to learn correspondence between image and text. Existing works learn coarse correspondence based on object co-occurrence statistics, while failing to learn fine-grained phrase corres…

Cited by 306PDFcodeScholar
2020

Overcoming Language Priors with Self-supervised Learning for Visual Question Answering

IJCAI 2020poster

Most Visual Question Answering (VQA) models suffer from the language prior problem, which is caused by inherent data biases. Specifically, VQA models tend to answer questions (e.g., what color is the banana?) based on the high-frequency answers (e.g., yellow) ignoring image contents. Existing approa…

2017

Double-bit quantization and weighting for nearest neighbor search

ICASSP 2017accepted

Binary embedding is an effective way for nearest neighbor (NN) search as binary code is storage efficient and fast to compute. It tries to convert real-value signatures into binary codes while preserving similarity of the original data. However, it greatly decreases the discriminability of original…

Cited by 0SourceScholar