← Search

Zhi Tang

15 accepted papers

2026

Uni-DocRobust: Universal Plug-and-Play Robustness Enhancement for Multi-modal LLMs via Feature Restoration

ICML 2026poster

Real-world degradations, such as noise, blur, and low resolution, significantly impair the performance of Multi-modal Large Language Models (MLLMs) in document understanding tasks. Despite recent advancements, progress in this field remains stifled by two critical bottlenecks: the scarcity of large-…

Cited by 0SourceScholar
2025

GlyphSR: A Simple Glyph-Aware Framework for Scene Text Image Super-Resolution

AAAI 2025technical

The goal of scene text image super-resolution (STISR) is to enhance the clarity of text within line images, thereby improving readability and enabling more accurate text recognition. However, existing STISR methods often rely heavily on Text Prior (TP) derived from trained recognizers, which can be…

Cited by 0SourcePDFScholar
2024

Maskstr: Guide Scene Text Recognition Models with Masking

ICASSP 2024accepted

Text recognition in information loss scenarios like blurriness, occlusion, and perspective distortion is challenging in real-world applications. To enhance robustness, some studies use extra unlabeled data for encoder pretraining. Others focus on improving decoder context reasoning. However, pretrai…

Cited by 0SourceScholar
2024

Recognition-Guided Diffusion Model for Scene Text Image Super-Resolution

ICASSP 2024accepted

Scene Text Image Super-Resolution (STISR) aims to enhance the resolution and legibility of text within low-resolution (LR) images, consequently elevating recognition accuracy in Scene Text Recognition (STR). Previous methods predominantly employ discriminative Convolutional Neural Networks (CNNs) au…

Cited by 0SourceScholar
2024

Training-Free Open-Ended Object Detection and Segmentation via Attention as Prompts

NeurIPS 2024poster

Existing perception models achieve great success by learning from large amounts of labeled data, but they still struggle with open-world scenarios. To alleviate this issue, researchers introduce open-set perception tasks to detect or segment unseen objects in the training set. However, these models…

Cited by 3SourcePDFScholar
2023

GeoDRL: A Self-Learning Framework for Geometry Problem Solving using Reinforcement Learning in Deductive Reasoning

ACL 2023findings

Ensuring both interpretability and correctness is a great challenge in automated geometry problem solving (GPS), and the scarcity of labeled data hinders learning mathematical reasoning from samples. Therefore, we present GeoDRL, a self-learning geometry problem solving framework that integrates log…

2022

BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework

NeurIPS 2022accept

Fusing the camera and LiDAR information has become a de-facto standard for 3D object detection tasks. Current methods rely on point clouds from the LiDAR sensor as queries to leverage the feature from the image space. However, people discovered that this underlying assumption makes the current fusio…

2022

CMUA-Watermark: A Cross-Model Universal Adversarial Watermark for Combating Deepfakes

AAAI 2022technical

Malicious applications of deepfakes (i.e., technologies generating target facial attributes or entire faces from facial images) have posed a huge threat to individuals' reputation and security. To mitigate these threats, recent studies have proposed adversarial watermarks to combat deepfake models,…

2022

Cycle Representation Learning for Inductive Relation Prediction

ICML 2022spotlight

In recent years, algebraic topology and its modern development, the theory of persistent homology, has shown great potential in graph representation learning. In this paper, based on the mathematics of algebraic topology, we propose a novel solution for inductive relation prediction, an important le…

2022

Neural Approximation of Graph Topological Features

NeurIPS 2022accept

Topological features based on persistent homology capture high-order structural information so as to augment graph neural network methods. However, computing extended persistent homology summaries remains slow for large and dense graphs and can be a serious bottleneck for the learning pipeline. Insp…

2021

Adaptive Edge Attention for Graph Matching with Outliers

IJCAI 2021poster

Graph matching aims at establishing correspondence between node sets of given graphs while keeping the consistency between their edge sets. However, outliers in practical scenarios and equivalent learning of edge representations in deep learning methods are still challenging. To address these issues…

2021

Link Prediction with Persistent Homology: An Interactive View

ICML 2021spotlight

Link prediction is an important learning task for graph-structured data. In this paper, we propose a novel topological approach to characterize interactions between two nodes. Our topological feature, based on the extended persistent homology, encodes rich structural information regarding the multi-…

2021

OPANAS: One-Shot Path Aggregation Network Architecture Search for Object Detection

CVPR 2021poster

Recently, neural architecture search (NAS) has been exploited to design feature pyramid networks (FPNs) and achieved promising results for visual object detection. Encouraged by the success, we propose a novel One-Shot Path Aggregation Network Architecture Search (OPANAS) algorithm, which significan…

Cited by 72PDFcodeScholar
2017

Mutual Enhancement for Detection of Multiple Logos in Sports Videos

ICCV 2017poster

Detecting logo frequency and duration in sports videos provides sponsors an effective way to evaluate their advertising efforts. However, general-purposed object detection methods cannot address all the challenges in sports videos. In this paper, we propose a mutual-enhanced approach that can improv…

Cited by 40PDFScholar