← Search

Hongtao Xie

59 accepted papers

2026

From Evaluation to Defense: Advancing Safety in Video Large Language Models

ICLR 2026poster

While the safety risks of image-based large language models (Image LLMs) have been extensively studied, their video-based counterparts (Video LLMs) remain critically under-examined. To systematically study this problem, we introduce \textbf{VideoSafetyEval} - the first large-scale, real-world benchm…

Cited by 0SourceScholar
2026

Latents-Inv:Robust Semantic Watermark via Dual-Path Mutual Information Redundancy for Diffusion Models

IJCAI 2026

Semantic watermarking methods, embedding identity into the initial latent noise, provide an imperceptible identity traceability for diffusion models in copyright protection and source verification. However, existing methods are highly vulnerable to adversarial attacks, especially geometric transform

Cited by 0Scholar
2026

Meerkat-VL: Implicit Risk Safety Alignment in Multimodal LLMs via Perceptual Reasoning and Self-Verification

ICML 2026poster

Multimodal LLMs (MLLMs) are increasingly deployed across diverse applications, but they pose significant safety concerns due to cross-modal interactions. To improve model safety awareness, existing methods rely on explicit-risk preference datasets and reinforcement learning guided by safety rewards.…

Cited by 0SourceScholar
2026

Orthogonal Concept Erasure for Diffusion Models

ICML 2026oral

Concept erasure has emerged as a promising approach to mitigate undesired or unsafe content in diffusion models, yet existing methods still face significant limitations. While training-based methods are effective, their high computational cost limits scalability. Editing-based methods are more effic…

Cited by 0SourceScholar
2026

RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding

AAAI 2026technical

Multi-modal Retrieval-Augmented Generation (RAG) has become a critical method for empowering LLMs by leveraging candidate visual documents. However, current methods consider the entire document as the basic retrieval unit, introducing substantial irrelevant visual content in two ways: 1) Relevant do

Cited by 0SourcePDFScholar
2026

SDErasure: Concept-Specific Trajectory Shifting for Concept Erasure via Adaptive Diffusion Classifier

ICLR 2026poster

Concept erasure methods have proven effective in mitigating the potential for text‑to‑image diffusion models to produce harmful content. Nevertheless, prevailing methods based on post fine-tuning introduce substantial disruption to the original model’s parameter distribution and suffer from excessiv…

Cited by 0SourceScholar
2026

Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement

CVPR 2026

Recent advances in Multimodal Large Language Models (MLLMs) have enabled automated generation of structured layouts from natural language descriptions. Existing methods typically follow a code-only paradigm that generates code to represent layouts, which are then rendered by graphic engines to produ

Cited by 0SourcecodeScholar
2026

ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement

CVPR 2026

While existing generation and unified models excel at general image generation, they struggle with tasks requiring deep reasoning, planning, and precise data-to-visual mapping abilities beyond general scenarios. To push beyond the existing limitations, we introduce a new and challenging task: creati

Cited by 0SourcecodeScholar
2026

SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability

AAAI 2026technical

Multimodal Large Language Models (MLLMs) have shown remarkable progress in temporal or spatial localization tasks, but struggle with joint spatio-temporal video grounding (STVG). We identify two key bottlenecks hindering this capability: (1) the sheer number of visual tokens makes long-range and fin

Cited by 0SourcePDFScholar
2026

Test-Time Scaling with Reflective Generative Model

ICLR 2026poster

We introduce a new Reflective Generative Model (RGM), which obtains OpenAI o3-mini's performance via a novel Reflective Generative Form. This form focuses on high-quality reasoning trajectory selection and contains two novelties: 1) A unified interface for policy and process reward model: we share t…

Cited by 0SourcecodeScholar
2025

CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness

NeurIPS 2025poster

Visual captioning benchmarks have become outdated with the emergence of modern multimodal large language models (MLLMs), as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effectively. While recent benchmarks attempt to address this by focusing on keyword ex…

Cited by 0SourceScholar
2025

CLIP-Adapted Region-to-Text Learning for Generative Open-Vocabulary Semantic Segmentation

ICCV 2025poster

In recent years, Open-Vocabulary Semantic Segmentation (OVSS) has been largely advanced. However, existing methods mostly rely on a pre-trained vision-language model (e.g., CLIP) and require a predefined set of classes to guide the semantic segmentation process during the inference. This not only na…

Cited by 0SourcePDFScholar
2025

CPL: Curriculum Pseudo Labeling for Weakly Supervised Temporal Forgery Localization

ICASSP 2025accepted

In forgery detection, temporal forgery localization offers a more nuanced perspective than binary detection by providing more precise temporal boundaries of manipulations. However, its need for frame-wise annotations limits real-world practicality. Therefore, we present the task of Weakly Supervised…

Cited by 0SourceScholar
2025

Forensic-MoE: Exploring Comprehensive Synthetic Image Detection Traces with Mixture of Experts

ICCV 2025poster

Recently, synthetic images have evolved incredibly realistic with the development of generative techniques. To avoid the spread of misinformation and identify synthetic content, research on synthetic image detection becomes urgent. Unfortunately, limited to the singular forensic perspective, existin…

2025

GRIP: A Graph-Based Reasoning Instruction Producer

NeurIPS 2025poster

Large-scale, high-quality data is essential for advancing the reasoning capabilities of large language models (LLMs). As publicly available Internet data becomes increasingly scarce, synthetic data has emerged as a crucial research direction. However, existing data synthesis methods often suffer fro…

Cited by 0SourceScholar
2025

GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation

ICCV 2025poster

While increasing attention has been paid to co-speech gesture synthesis, most previous works neglect to investigate hand gestures with explicit and essential semantics. In this paper, we study co-speech gesture generation with an emphasis on specific hand gesture activation, which can deliver more i…

Cited by 0SourcePDFScholar
2025

Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models

CVPR 2025poster

Recent Multi-modal Large Language Models (MLLMs) have been challenged by the computational overhead resulting from massive video frames, often alleviated through compression strategies. However, the visual content is not equally contributed to user instructions, existing strategies (e.g., average po…

2025

IDseq: Decoupled and Sequentially Detecting and Grounding Multi-Modal Media Manipulation

AAAI 2025technical

Detecting and grounding multi-modal media manipulation aims to categorize the type and localize the region of manipulation for image-text pairs in both two modalities. Existing methods have not sufficiently explored the intrinsic properties of the manipulated images, which contain both forgery and c…

Cited by 0SourcePDFScholar
2025

IGD: Instructional Graphic Design with Multimodal Layer Generation

ICCV 2025poster

Graphic design visually conveys information and data by creating and combining text, images and graphics. Two-stage methods that rely primarily on layout generation lack creativity and intelligence, making graphic design still labor-intensive. Existing diffusion-based methods generate non-editable g…

2025

Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design

ICCV 2025poster

With the increasing demand for the right to be forgotten, machine unlearning (MU) has emerged as a vital tool for enhancing trust and regulatory compliance by enabling the removal of sensitive data influences from machine learning (ML) models. However, most MU algorithms primarily rely on in-trainin…

Cited by 0SourcePDFScholar
2025

IterMeme: Expert-Guided Multimodal LLM for Interactive Meme Creation with Layout-Aware Generation

IJCAI 2025

Meme creation is a creative process that blends images and text. However, existing methods lack critical components, failing to support intent-driven caption-layout generation and personalized generation, making it difficult to generate high-quality memes. To address this limitation, we propose Iter

2025

Mask^2DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation

CVPR 2025poster

Sora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challenging task of multi-scene video generation, which offers broader applications, remains relatively underexplored. To bridge this gap, we propose Mask^2DiT,…

2025

SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text Recognition

ICCV 2025poster

Connectionist temporal classification (CTC)-based scene text recognition (STR) methods, e.g., SVTR, are widely employed in OCR applications, mainly due to their simple architecture, which only contains a visual model and a CTC-aligned linear classifier, and therefore fast inference. However, they ge…

2025

SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled Synthesis

CVPR 2025poster

Due to the limited scale of multimodal table understanding (MTU) data, model performance is constrained. A straightforward approach is to use multimodal large language models to obtain more samples, but this may cause hallucinations, generate incorrect sample pairs, and cost significantly.To address…

2024

AlignZeg: Mitigating Objective Misalignment for Zero-shot Semantic Segmentation

ECCV 2024poster

"A serious issue that harms the performance of zero-shot visual recognition is named objective misalignment, i.e., the learning objective prioritizes improving the recognition accuracy of seen classes rather than unseen classes, while the latter is the true target to pursue. This issue becomes more…

Cited by 4SourcePDFScholar
2024

Boosting Semi-Supervised Scene Text Recognition via Viewing and Summarizing

NeurIPS 2024poster

Existing scene text recognition (STR) methods struggle to recognize challenging texts, especially for artistic and severely distorted characters. The limitation lies in the insufficient exploration of character morphologies, including the monotonousness of widely used synthetic training data and the…

2024

DEADiff: An Efficient Stylization Diffusion Model with Disentangled Representations

CVPR 2024highlight

The diffusion-based text-to-image model harbors immense potential in transferring reference style. However current encoder-based approaches significantly impair the text controllability of text-to-image models while transferring styles. In this paper we introduce DEADiff to address this issue using…

2024

DiffAM: Diffusion-based Adversarial Makeup Transfer for Facial Privacy Protection

CVPR 2024poster

With the rapid development of face recognition (FR) systems the privacy of face images on social media is facing severe challenges due to the abuse of unauthorized FR systems. Some studies utilize adversarial attack techniques to defend against malicious FR systems by generating adversarial examples…

2024

Focus on the Whole Character: Discriminative Character Modeling for Scene Text Recognition

IJCAI 2024poster

Recently, scene text recognition (STR) models have shown significant performance improvements. However, existing models still encounter difficulties in recognizing challenging texts that involve factors such as severely distorted and perspective characters. These challenging texts mainly cause two…

2024

How Control Information Influences Multilingual Text Image Generation and Editing?

NeurIPS 2024poster

Visual text generation has significantly advanced through diffusion models aimed at producing images with readable and realistic text. Recent works primarily use a ControlNet-based framework, employing standard font text images to control diffusion models. Recognizing the critical role of control in…

2024

Knowledge Context Modeling with Pre-trained Language Models for Contrastive Knowledge Graph Completion

ACL 2024findings

Text-based knowledge graph completion (KGC) methods utilize pre-trained language models for triple encoding and further fine-tune the model to achieve completion. Despite their excellent performance, they neglect the knowledge context in inferring process. Intuitively, knowledge contexts, which refe…

Cited by 5SourcePDFScholar
2024

Leveraging Text Localization for Scene Text Removal via Text-aware Masked Image Modeling

ECCV 2024poster

"Existing scene text removal (STR) task suffers from insufficient training data due to the expensive pixel-level labeling. In this paper, we aim to address this issue by introducing a Text-aware Masked Image Modeling algorithm (TMIM), which can pretrain STR models with low-cost text detection labels…

2024

OTE: Exploring Accurate Scene Text Recognition Using One Token

CVPR 2024poster

In this paper we propose a novel framework to fully exploit the potential of a single vector for scene text recognition (STR). Different from previous sequence-to-sequence methods that rely on a sequence of visual tokens to represent scene text images we prove that just one token is enough to charac…

2024

Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition

IJCAI 2024poster

In text recognition, self-supervised pre-training emerges as a good solution to reduce dependence on expansive annotated real data. Previous studies primarily focus on local visual representation by leveraging mask image modeling or sequence contrastive learning. However, they omit modeling the ling…

2024

ShowMaker: Creating High-Fidelity 2D Human Video via Fine-Grained Diffusion Modeling

NeurIPS 2024poster

Although significant progress has been made in human video generation, most previous studies focus on either human facial animation or full-body animation, which cannot be directly applied to produce realistic conversational human videos with frequent hand gestures and various facial movements simul…

Cited by 4SourcePDFScholar
2024

Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment Retrieval

AAAI 2024technical

Video Moment Retrieval (VMR) aims to retrieve temporal segments in untrimmed videos corresponding to a given language query by constructing cross-modal alignment strategies. However, these existing strategies are often sub-optimal since they ignore the modality imbalance problem, i.e., the semantic…

2023

Exploring Stroke-Level Modifications for Scene Text Editing

AAAI 2023technical

Scene text editing (STE) aims to replace text with the desired one while preserving background and styles of the original text. However, due to the complicated background textures and various text styles, existing methods fall short in generating clear and legible edited text images. In this study,…

2023

Learning Orthogonal Prototypes for Generalized Few-Shot Semantic Segmentation

CVPR 2023poster

Generalized few-shot semantic segmentation (GFSS) distinguishes pixels of base and novel classes from the background simultaneously, conditioning on sufficient data of base classes and a few examples from novel class. A typical GFSS approach has two training phases: base class learning and novel cla…

2023

Linguistic More: Taking a Further Step toward Efficient and Accurate Scene Text Recognition

IJCAI 2023poster

Vision model have gained increasing attention due to their simplicity and efficiency in Scene Text Recognition (STR) task. However, due to lacking the perception of linguistic knowledge and information, recent vision models suffer from two problems: (1) the pure vision-based query results in attenti…

2023

MomentDiff: Generative Video Moment Retrieval from Random to Real

NeurIPS 2023poster

Video moment retrieval pursues an efficient and generalized solution to identify the specific temporal segments within an untrimmed video that correspond to a given language description. To achieve this goal, we provide a generative diffusion-based framework called MomentDiff, which simulates a typi…

2023

Progressive Spatio-Temporal Prototype Matching for Text-Video Retrieval

ICCV 2023oral

The performance of text-video retrieval has been significantly improved by vision-language cross-modal learning schemes. The typical solution is to directly align the global video-level and sentence-level features during learning, which would ignore the intrinsic video-text relations, i.e., a text…

Cited by 46PDFcodeScholar
2023

TPS++: Attention-Enhanced Thin-Plate Spline for Scene Text Recognition

IJCAI 2023poster

Text irregularities pose significant challenges to scene text recognizers. Thin-Plate Spline (TPS)-based rectification is widely regarded as an effective means to deal with them. Currently, the calculation of TPS transformation parameters purely depends on the quality of regressed text borders. It i…

2022

Bridging the Gap Between Vision Transformers and Convolutional Neural Networks on Small Datasets

NeurIPS 2022accept

There still remains an extreme performance gap between Vision Transformers (ViTs) and Convolutional Neural Networks (CNNs) when training from scratch on small datasets, which is concluded to the lack of inductive bias. In this paper, we further consider this problem and point out two weaknesses of V…

2022

Detecting Tampered Scene Text in the Wild

ECCV 2022poster

"Text manipulation technologies cause serious worries in recent years, however, corresponding tampering detection methods have not been well explored. In this paper, we introduce a new task, named Tampered Scene Text Detection (TSTD), to localize text instances and recognize the texture authenticity…

2022

Dual-Stream Knowledge-Preserving Hashing for Unsupervised Video Retrieval

ECCV 2022poster

"Unsupervised video hashing usually optimizes binary codes by learning to reconstruct input videos. Such reconstruction constraint spends much effort on frame-level temporal context changes without focusing on video-level global semantics that are more useful for retrieval. Hence, we address this pr…

Cited by 24SourcePDFScholar
2022

Partial Class Activation Attention for Semantic Segmentation

CVPR 2022poster

Current attention-based methods for semantic segmentation mainly model pixel relation through pairwise affinity and coarse segmentation. For the first time, this paper explores modeling pixel relation via Class Activation Map (CAM). Beyond the previous CAM generated from image-level classification,…

Cited by 51PDFcodeScholar
2021

Dynamic Inconsistency-aware DeepFake Video Detection

IJCAI 2021poster

The spread of DeepFake videos causes a serious threat to information security, calling for effective detection methods to distinguish them. However, the performance of recent frame-based detection methods become limited due to their ignorance of the inter-frame inconsistency of fake videos. In this…

Cited by 0SourcePDFScholar
2021

Frequency-Aware Discriminative Feature Learning Supervised by Single-Center Loss for Face Forgery Detection

CVPR 2021poster

Face forgery detection is raising ever-increasing interest in computer vision since facial manipulation technologies cause serious worries. Though recent works have reached sound achievements, there are still unignorable problems: a) learned features supervised by softmax loss are separable but not…

Cited by 333PDFScholar
2021

From Two to One: A New Scene Text Recognizer With Visual Language Modeling Network

ICCV 2021poster

In this paper, we abandon the dominant complex language model and rethink the linguistic learning process in the scene text recognition. Different from previous methods considering the visual and linguistic information in two separate structures, we propose a Visual Language Modeling Network (Vision…

Cited by 184PDFcodeScholar
2021

Query-Memory Re-Aggregation for Weakly-supervised Video Object Segmentation

AAAI 2021technical

Weakly-supervised video object segmentation (WVOS) is an emerging video task that can track and segment the target given a simple bounding box label. However, existing WVOS methods are still unsatisfied in either speed or accuracy, since they only use the exemplar frame to guide the prediction while…

Cited by 26SourcePDFScholar
2021

Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Recognition

CVPR 2021poster

Linguistic knowledge is of great benefit to scene text recognition. However, how to effectively model linguistic rules in end-to-end deep networks remains a research challenge. In this paper, we argue that the limited capacity of language models comes from: 1) implicitly language modeling; 2) unidir…

Cited by 461PDFcodeScholar
2021

Semantic-guided Reinforced Region Embedding for Generalized Zero-Shot Learning

AAAI 2021technical

Generalized zero-shot Learning (GZSL) aims to recognize images from either seen or unseen domain, mainly by learning a joint embedding space to associate image features with the corresponding category descriptions. Recent methods have proved that localizing important object regions can effectively b…

Cited by 37SourcePDFScholar
2020

ContourNet: Taking a Further Step Toward Accurate Arbitrary-Shaped Scene Text Detection

CVPR 2020poster

Scene text detection has witnessed rapid development in recent years. However, there still exists two main challenges: 1) many methods suffer from false positives in their text representations; 2) the large scale variance of scene texts makes it hard for network to learn samples. In this paper, we p…

Cited by 273PDFcodeScholar
2020

Domain-Aware Visual Bias Eliminating for Generalized Zero-Shot Learning

CVPR 2020poster

Generalized zero-shot learning aims to recognize images from seen and unseen domains. Recent methods focus on learning a unified semantic-aligned visual representation to transfer knowledge between two domains, while ignoring the effect of semantic-free visual representation in alleviating the biase…

Cited by 201PDFcodeScholar
2020

Graph Structured Network for Image-Text Matching

CVPR 2020poster

Image-text matching has received growing interest since it bridges vision and language. The key challenge lies in how to learn correspondence between image and text. Existing works learn coarse correspondence based on object co-occurrence statistics, while failing to learn fine-grained phrase corres…

Cited by 306PDFcodeScholar
2020

Hierarchical Granularity Transfer Learning

NeurIPS 2020poster

In the real world, object categories usually have a hierarchical granularity tree. Nowadays, most researchers focus on recognizing categories in a specific granularity, \emph{e.g.,} basic-level or sub(ordinate)-level. Compared with basic-level categories, the sub-level categories provide more valuab…

Cited by 5SourcePDFScholar
2020

Real-World Automatic Makeup via Identity Preservation Makeup Net

IJCAI 2020poster

This paper focuses on the real-world automatic makeup problem. Given one non-makeup target image and one reference image, the automatic makeup is to generate one face image, which maintains the original identity with the makeup style in the reference image. In the real-world scenario, face makeup ta…

Cited by 0SourcePDFScholar
2017

Double-bit quantization and weighting for nearest neighbor search

ICASSP 2017accepted

Binary embedding is an effective way for nearest neighbor (NN) search as binary code is storage efficient and fast to compute. It tries to convert real-value signatures into binary codes while preserving similarity of the original data. However, it greatly decreases the discriminability of original…

Cited by 0SourceScholar