← Search

Qi Zheng

29 accepted papers

2026

AssemMate: Graph-Based LLM for Robotic Assembly Assistance

ICRA 2026poster

Large Language Model (LLM)-based robotic assembly assistance has gained significant research attention. It requires the injection of domain-specific knowledge to guide the assembly process through natural language interaction with humans. Despite some progress, existing methods represent knowledge i…

2026

Towards a Unified Generative Model for Scarce Time Series with Domain Experts

ICML 2026poster

Synthesizing realistic time series with generative models has wide-ranging applications in real-world scenarios. Despite recent progress, most existing methods are trained under the assumption of abundant training data, which substantially limits their effectiveness in data-scarce settings. In this …

Cited by 0SourceScholar
2025

4KAgent: Agentic Any Image to 4K Super-Resolution

NeurIPS 2025poster

We present 4KAgent, a unified agentic super-resolution generalist system designed to universally upscale any image to 4K resolution (and even higher, if applied iteratively). Our system can transform images from extremely low resolutions with severe degradations, for example, highly distorted inputs…

Cited by 0SourcecodeScholar
2025

A Robust and Efficient Visual-Inertial Initialization With Probabilistic Normal Epipolar Constraint

RA-L 2025

Accurate and robust initialization is essential for Visual-Inertial Odometry (VIO), as poor initialization can severely degrade pose accuracy. During initialization, it is crucial to estimate parameters such as accelerometer bias, gyroscope bias, initial velocity, gravity, etc. Most existing VIO ini

Cited by 4SourcecodeScholar
2025

A Simple yet Effective Layout Token in Large Language Models for Document Understanding

CVPR 2025poster

Recent methods that integrate spatial layouts with text for document understanding in large language models (LLMs) have shown promising results. A commonly used method is to represent layout information as text tokens and interleave them with text content as inputs to the LLMs. However, such a metho…

Cited by 0SourcePDFScholar
2025

End-to-End HOI Reconstruction Transformer with Graph-based Encoding

CVPR 2025highlight

Human-object interaction (HOI) reconstruction has garnered significant attention due to its diverse applications and the success of capturing human meshes. Existing HOI reconstruction methods often rely on explicitly modeling interactions between humans and objects. However, such a way leads to a na…

Cited by 0SourcePDFScholar
2025

Frequency-Biased Synergistic Design for Image Compression and Compensation

CVPR 2025poster

Compression artifacts removal (CAR), an effective post-processing method to reduce compression distortion in edge-side codecs, demonstrates remarkable results by utilizing convolutional neural networks (CNNs) on high computational power cloud side. Traditional image compression reduces redundancy in…

Cited by 0SourcePDFScholar
2025

Intelligent Document Parsing: Towards End-to-end Document Parsing via Decoupled Content Parsing and Layout Grounding

EMNLP 2025

In the daily work, vast amounts of documents are stored in pixel-based formats such as images and scanned PDFs, posing challenges for efficient database management and data processing. Existing methods often fragment the parsing process into the pipeline of separated subtasks on the layout element l

Cited by 0SourcePDFScholar
2025

Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding

EMNLP 2025

Multimodal large language models (MLLMs) have shown impressive capabilities in document understanding, a rapidly growing research area with significant industrial demand. As a multimodal task, document understanding requires models to possess both perceptual and cognitive abilities. However, due to

Cited by 0SourcePDFScholar
2025

ProcTag: Process Tagging for Assessing the Efficacy of Document Instruction Data

AAAI 2025technical

Recently, large language models (LLMs) and multimodal large language models (MLLMs) have demonstrated promising results on document visual question answering (VQA) task, particularly after training on document instruction datasets. An effective evaluation method for document instruction data is cruc…

2025

ST-ReP: Learning Predictive Representations Efficiently for Spatial-Temporal Forecasting

AAAI 2025technical

Spatial-temporal forecasting is crucial and widely applicable in various domains such as traffic, energy, and climate. Benefiting from the abundance of unlabeled spatial-temporal data, self-supervised methods are increasingly adapted to learn spatial-temporal representations. However, it encounters…

2024

DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing

EMNLP 2024main

Parsing documents from pixels, such as pictures and scanned PDFs, into hierarchical structures is extensively demanded in the daily routines of data storage, retrieval and understanding. However, previously the research on this topic has been largely hindered since most existing datasets are small-s…

2024

LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding

CVPR 2024poster

Recently leveraging large language models (LLMs) or multimodal large language models (MLLMs) for document understanding has been proven very promising. However previous works that employ LLMs/MLLMs for document understanding have not fully explored and utilized the document layout information which…

2024

MuEP: A Multimodal Benchmark for Embodied Planning with Foundation Models

IJCAI 2024poster

Foundation models have demonstrated significant emergent abilities, holding great promise for enhancing embodied agents' reasoning and planning capacities. However, the absence of a comprehensive benchmark for evaluating embodied agents with multimodal observations in complex environments remains a…

2024

WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation

ECCV 2024poster

"In the era of content creation revolution propelled by advancements in generative models, the field of web design remains unexplored despite its critical role in modern digital communication. The web design process is complex and often time-consuming, especially for those with limited expertise. In…

2023

GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render Tree

EMNLP 2023long main

Inexhaustible web content carries abundant perceptible information beyond text. Unfortunately, most prior efforts in pre-trained Language Models (LMs) ignore such cyber-richness, while few of them only employ plain HTMLs, and crucial information in the rendered web, such as visual, layout, and style…

Cited by 0SourceScholar
2023

GeoLayoutLM: Geometric Pre-Training for Visual Information Extraction

CVPR 2023highlight

Visual information extraction (VIE) plays an important role in Document Intelligence. Generally, it is divided into two tasks: semantic entity recognition (SER) and relation extraction (RE). Recently, pre-trained models for documents have achieved substantial progress in VIE, particularly in SER. Ho…

2023

LISTER: Neighbor Decoding for Length-Insensitive Scene Text Recognition

ICCV 2023poster

The diversity in length constitutes a significant characteristic of text. Due to the long-tail distribution of text lengths, most existing methods for scene text recognition (STR) only work well on short or seen-length text, lacking the capability of recognizing longer text or performing length extr…

Cited by 29PDFcodeScholar
2023

LORE: Logical Location Regression Network for Table Structure Recognition

AAAI 2023technical

Table structure recognition (TSR) aims at extracting tables in images into machine-understandable formats. Recent methods solve this problem by predicting the adjacency relations of detected cell boxes, or learning to generate the corresponding markup sequences from the table images. However, they e…

2023

Modeling Video As Stochastic Processes for Fine-Grained Video Representation Learning

CVPR 2023highlight

A meaningful video is semantically coherent and changes smoothly. However, most existing fine-grained video representation learning methods learn frame-wise features by aligning frames across videos or exploring relevance between multiple views, neglecting the inherent dynamic process of each video.…

2022

No-Reference Quality Assessment of Variable Frame-Rate Videos Using Temporal Bandpass Statistics

ICASSP 2022accepted

Recent advances in mobile devices and cloud computing techniques have made it possible to capture, process, and share high resolution, high frame rate (HFR) videos across the Internet nearly instantaneously. Being able to monitor and control the quality of these streamed videos can enable the de-liv…

Cited by 0SourceScholar
2022

Understanding Gender Bias in Knowledge Base Embeddings

ACL 2022long

Knowledge base (KB) embeddings have been shown to contain gender biases. In this paper, we study two questions regarding these biases: how to quantify them, and how to trace their origins in KB? Specifically, first, we develop two novel bias measures respectively for a group of person entities and a…

Cited by 10SourcePDFScholar
2020

An End-to-End OCR Text Re-organization Sequence Learning for Rich-text Detail Image Comprehension

ECCV 2020poster

Nowadays rich description on detail images help users know more about the commodities. With the help of OCR technology, the description text can be detected and recognized as auxiliary information to remove the comprehending barriers among the visual impaired users. However, for lack of proper logic…

Cited by 29SourcePDFScholar
2019

Rain Streak Removal via Multi-scale Mixture Exponential Power Model

ICASSP 2019accepted

Rain streaks severely hamper the visible performance of the outdoor surveillance videos, which becomes an attractive issue in recent computer vision research. Existing methods usually encode rain streaks into Gaussian Mixture Model (GM-M). However, the limited number of Gaussian components in the GM…

Cited by 0SourceScholar
2018

Hierarchical Bilinear Pooling for Fine-Grained Visual Recognition

ECCV 2018poster

Fine-grained visual recognition is challenging because it highly relies on the modeling of various semantic parts and fine-grained feature learning. Bilinear pooling based models have been shown to be effective at fine-grained recognition, while most previous approaches neglect the fact that inter-l…

2017

Transferring clothing parsing from fashion dataset to surveillance

ICASSP 2017accepted

In this paper we address the problem of automatic clothing parsing in surveillance video with the information from user-generated tags such as “jeans” and “T-shirt”. Although clothing parsing has achieved great success in fashion clothing, it is quite challenging to parse clothing in practical surve…

Cited by 0SourceScholar