← Search

Ze Liu

17 accepted papers

2026

From Pixel to Precision: Enhancing Handwritten Mathematical Expression Recognition with Image-Level Reward

CVPR 2026

Handwritten mathematical expression recognition is hindered by a fundamental misalignment between the dual representations of LaTeX formulas: the symbolic text and the rendered visual image. This discrepancy means that textually distinct LaTeX sequences can produce visually identical outputs, while

Cited by 0SourceScholar
2026

OmniGen2: Towards Instruction-Aligned Multimodal Generation

CVPR 2026

Multimodal generative models can process instructions in various modalities and demonstrate outstanding performance across a wide range of image generation tasks. However, their robustness in complex real-world scenarios remains limited due to insufficient generalized instruction alignment. We intro

Cited by 0SourcecodeScholar
2026

Trust, but Verify: Uncertainty-Driven Evidential Multimodal Representation Learning

IJCAI 2026

Effective multimodal learning in real-world scenarios depends on a nuanced treatment of uncertainty, which arises at three levels: (1) Intrinsic Uncertainty from modality-specific noise or ambiguity; (2) Relational Uncertainty due to cross-modal conflicts or redundancy; and (3) Aggregated Uncertaint

Cited by 0Scholar
2025

Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval

ACL 2025long

With the popularity of multimodal techniques, it receives growing interests to acquire useful information in visual forms. In this work, we formally define an emerging IR paradigm called Visualized Information Retrieval, or Vis-IR, where multimodal information, such as texts, images, tables and char…

2025

MegaPairs: Massive Data Synthesis for Universal Multimodal Retrieval

ACL 2025long

Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs, a novel data synthesis method that leverages vision language models (VLMs) and open-domain images, together with a massi…

2024

Enhancing Leg Odometry in Legged Robots with Learned Contact Bias: An LSTM Recurrent Neural Network Approach

IROS 2024poster

To address the leg odometry drift caused by the non-stationary foot contact, this paper introduces a novel data-driven based leg odometry technique for legged robots. By leveraging a Long Short-Term Memory (LSTM) Recurrent Neural Network (RNN), the method learns the biases in the robot’s foot contac…

Cited by 0SourceScholar
2024

Generalization Error Bounds for Two-stage Recommender Systems with Tree Structure

NeurIPS 2024oral

Two-stage recommender systems play a crucial role in efficiently identifying relevant items and personalizing recommendations from a vast array of options. This paper, based on an error decomposition framework, analyzes the generalization error for two-stage recommender systems with a tree structure…

Cited by 0SourcePDFScholar
2024

V-DETR: DETR with Vertex Relative Position Encoding for 3D Object Detection

ICLR 2024poster

We introduce a highly performant 3D object detector for point clouds using the DETR framework. The prior attempts all end up with suboptimal results because they fail to learn accurate inductive biases from the limited scale of training data. In particular, the queries often attend to points that ar…

2022

Could Giant Pre-trained Image Models Extract Universal Representations?

NeurIPS 2022accept

Frozen pretrained models have become a viable alternative to the pretraining-then-finetuning paradigm for transfer learning. However, with frozen models there are relatively few parameters available for adapting to downstream tasks, which is problematic in computer vision where tasks vary significan…

Cited by 11SourcePDFScholar
2022

Swin Transformer V2: Scaling Up Capacity and Resolution

CVPR 2022poster

We present techniques for scaling Swin Transformer [??] up to 3 billion parameters and making it capable of training with images of up to 1,536x1,536 resolution. By scaling up capacity and resolution, Swin Transformer sets new records on four representative vision benchmarks: 84.0% top-1 accuracy on…

Cited by 2410PDFcodeScholar
2021

RGB-D Salient Object Detection via 3D Convolutional Neural Networks

AAAI 2021technical

RGB-D salient object detection (SOD) recently has attracted increasing research interest and many deep learning methods based on encoder-decoder architectures have emerged. However, most existing RGB-D SOD models conduct feature fusion either in the single encoder or the decoder stage, which hardly…

2021

Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows

ICCV 2021poster

This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision. Challenges in adapting Transformer from language to vision arise from differences between the two domains, such as large variations in the scale of visual ent…

Cited by 29925PDFcodeScholar
2021

Syntactically Diverse Adversarial Network for Knowledge-Grounded Conversation Generation

EMNLP 2021finding

Generative conversation systems tend to produce meaningless and generic responses, which significantly reduce the user experience. In order to generate informative and diverse responses, recent studies proposed to fuse knowledge to improve informativeness and adopt latent variables to enhance the di…

Cited by 3SourcePDFScholar