← Search

Zixiao Wang

14 accepted papers

2026

General Process Reward Modeling for Robotic Reinforcement Learning

CVPR 2026

The primary obstacle for applying reinforcement learning (RL) to real-world robotics is the design of effective reward functions. While recently learning-based Process Reward Models (PRMs) are a promising direction, they are often hindered by two fundamental limitations: their reward models lack ste

Cited by 0SourcecodeScholar
2026

OLion: Approaching the Hadamard Ideal by Intersecting Spectral and L inf Implicit Biases

ICML 2026poster

Many optimizers can be interpreted as steepest-descent methods under norm-induced geometries, and thus inherit corresponding implicit biases. We introduce Orthogonal Lion which combines spectral control from orthogonalized update directions with $\ell_\infty$-style coordinate control from sign updat…

Cited by 0SourceScholar
2026

Test-Time Scaling with Reflective Generative Model

ICLR 2026poster

We introduce a new Reflective Generative Model (RGM), which obtains OpenAI o3-mini's performance via a novel Reflective Generative Form. This form focuses on high-quality reasoning trajectory selection and contains two novelties: 1) A unified interface for policy and process reward model: we share t…

Cited by 0SourcecodeScholar
2025

Beyond Profile: From Surface-Level Facts to Deep Persona Simulation in LLMs

ACL 2025finding

Previous approaches to persona simulation large language models (LLMs) have typically relied on learning basic biographical information, or using limited role-play dialogue datasets to capture a character’s responses. However, a holistic representation of an individual goes beyond surface-level fact…

2025

Optimal Nuisance Function Tuning for Estimating a Doubly Robust Functional under Proportional Asymptotics

NeurIPS 2025spotlight

In this paper, we explore the asymptotically optimal tuning parameter choice in ridge regression for estimating nuisance functions of a statistical functional that has recently gained prominence in conditional independence testing and causal inference. Given a sample of size $n$, we study estimat…

Cited by 0SourceScholar
2025

SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled Synthesis

CVPR 2025poster

Due to the limited scale of multimodal table understanding (MTU) data, model performance is constrained. A straightforward approach is to use multimodal large language models to obtain more samples, but this may cause hallucinations, generate incorrect sample pairs, and cost significantly.To address…

2024

Boosting Semi-Supervised Scene Text Recognition via Viewing and Summarizing

NeurIPS 2024poster

Existing scene text recognition (STR) methods struggle to recognize challenging texts, especially for artistic and severely distorted characters. The limitation lies in the insufficient exploration of character morphologies, including the monotonousness of widely used synthetic training data and the…

2024

Focus on the Whole Character: Discriminative Character Modeling for Scene Text Recognition

IJCAI 2024poster

Recently, scene text recognition (STR) models have shown significant performance improvements. However, existing models still encounter difficulties in recognizing challenging texts that involve factors such as severely distorted and perspective characters. These challenging texts mainly cause two…

2024

Identification and Estimation for Nonignorable Missing Data: A Data Fusion Approach

ICML 2024poster

We consider the task of identifying and estimating a parameter of interest in settings where data is missing not at random (MNAR). In general, such parameters are not identified without strong assumptions on the missing data model. In this paper, we take an alternative approach and introduce a metho…

Cited by 1SourcePDFScholar
2024

Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition

IJCAI 2024poster

In text recognition, self-supervised pre-training emerges as a good solution to reduce dependence on expansive annotated real data. Previous studies primarily focus on local visual representation by leveraging mask image modeling or sequence contrastive learning. However, they omit modeling the ling…

2023

ATFormer: A Learned Performance Model with Transfer Learning Across Devices for Deep Learning Tensor Programs

EMNLP 2023long main

The training and inference efficiency of ever-larger deep neural networks highly rely on the performance of tensor operators on specific hardware platforms. Therefore, a compilation-based optimization flow with automatic tensor generation and parameter tuning is necessary for efficient model deploym…

Cited by 0SourceScholar
2023

Truncate-Split-Contrast: A Framework for Learning from Mislabeled Videos

AAAI 2023technical

Learning with noisy label is a classic problem that has been extensively studied for image tasks, but much less for video in the literature. A straightforward migration from images to videos without considering temporal semantics and computational cost is not a sound choice. In this paper, we propos…

Cited by 3SourcePDFScholar