← Search

Ze Wang

34 accepted papers

2026

Hierarchical Attention Network with Correction for Cross-Domain User Association

AAAI 2026technical

Despite the rich spatiotemporal patterns contained in trajectory data from multiple Location-Based Social Network (LBSN) platforms, heterogeneous formats, semantic inconsistencies, and unequal user scales across platforms create substantial barriers to reliable identity mapping. Furthermore, GPS dri

Cited by 0SourcePDFScholar
2026

ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning

ICLR 2026poster

The rapid advancement of text-to-image (T2I) models has increased the need for reliable human preference modeling, a demand further amplified by recent progress in reinforcement learning for preference alignment. However, existing approaches typically quantify the quality of a generated image using…

Cited by 0SourceScholar
2026

Multi-Priority Reactive Motion Control for Safe and Coordinated Dual-Arm Manipulation in Dynamic Environments

RA-L 2026

Reactive motion generation for dual-arm robotic systems is challenging due to their high degrees of freedom, nonlinear characteristics as well as the presence of multiple constraints, including kinematic limits, collision avoidance, dual-arm coordination, and other task-specific requirements. These

Cited by 0SourceScholar
2026

OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera

CVPR 2026

Robust 3D semantic occupancy is essential for legged and humanoid robots, yet most Semantic Scene Completion (SSC) systems are built for wheeled platforms with forward-facing sensors. We present OneOcc, a vision-only panoramic SSC framework tailored to severe body jitter and 360deg continuity. OneOc

Cited by 0SourcecodeScholar
2026

VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking

CVPR 2026

Video agentic models have advanced challenging video-language tasks. However, most agentic approaches still heavily rely on greedy parsing over densely sampled video frames, resulting in high computational cost. We present VideoSeek, a long-horizon video agent that leverages video logic flow to acti

Cited by 3SourcecodeScholar
2026

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models

ICLR 2026poster

Omni-modal large language models (OLLMs) aim to unify audio, vision, and text understanding within a single framework. While existing benchmarks have advanced multimodal evaluation, it remains unclear whether OLLMs achieve modality-invariant reasoning or inherit modality-specific biases. We introduc…

Cited by 0SourceScholar
2025

Agent Laboratory: Using LLM Agents as Research Assistants

EMNLP 2025

Historically, scientific discovery has been a lengthy and costly process, demanding substantial time and resources from initial conception to final results. To accelerate scientific discovery, reduce research costs, and improve research quality, we introduce Agent Laboratory, an autonomous LLM-based

Cited by 0SourcePDFScholar
2025

Cross-Domain Trajectory Association Based on Hierarchical Spatiotemporal Enhanced Attention Hypergraph

AAAI 2025technical

Identifying and linking the same users across different social platforms is crucial for understanding user behavior and preferences. However, cross-domain datasets exhibit diverse characteristics, such as varying check-in frequencies, significant disparities in data precision, and distinct distribut…

Cited by 0SourcePDFScholar
2025

Masked Autoencoders Are Effective Tokenizers for Diffusion Models

ICML 2025spotlight

Recent advances in latent diffusion models have demonstrated their effectiveness for high-resolution image synthesis. However, the properties of the latent space from tokenizer for better learning and generation of diffusion models remain under-explored. Theoretically and empirically, we find that i…

Cited by 8SourcePDFScholar
2025

SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration

COLING 2025main

The development of unbiased large language models is widely recognized as crucial, yet existing benchmarks fall short in detecting biases due to limited scope, contamination, and lack of a fairness baseline. SAGED(bias) is the first holistic benchmarking pipeline to address these problems. The pipel…

2025

SF-TIM: A Simple Framework for Enhancing Quadrupedal Robot Jumping Agility by Combining Terrain Imagination and Measurement

IROS 2025

Dynamic jumping on high platforms and over gaps differentiates legged robots from wheeled counterparts. Compared to walking on rough terrains, dynamic locomotion on abrupt surfaces requires fusing proprioceptive and exteroceptive perception for explosive movements. In this paper, we propose SF-TIM (

Cited by 3SourcecodeScholar
2025

Self-Taught Agentic Long Context Understanding

ACL 2025long

Answering complex, long-context questions remains a major challenge for large language models (LLMs) as it requires effective question clarifications and context retrieval. We propose Agentic Long-Context Understanding (AgenticLU), a framework designed to enhance an LLM’s understanding of such queri…

2025

SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer

CVPR 2025poster

Efficient image tokenization with high compression ratios remains a critical challenge for training generative models.We present SoftVQ-VAE, a continuous image tokenizer that leverages soft categorical posteriors to aggregate multiple codewords into each latent token, substantially increasing the re…

2025

Tuning Timestep-Distilled Diffusion Model Using Pairwise Sample Optimization

ICLR 2025poster

Recent advancements in timestep-distilled diffusion models have enabled high-quality image generation that rivals non-distilled multi-step models, but with significantly fewer inference steps. While such models are attractive for applications due to the low inference cost and latency, fine-tuning th…

Cited by 2SourcePDFScholar
2025

Unleashing Hour-Scale Video Training for Long Video-Language Understanding

NeurIPS 2025spotlight

Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has left the training of hour-long Video-LMMs underexplored. To close this gap, we present VideoMarathon, a large-scale hou…

Cited by 0SourceScholar
2024

Constructing Concept-based Models to Mitigate Spurious Correlations with Minimal Human Effort

ECCV 2024poster

"Enhancing model interpretability can address spurious correlations by revealing how models draw their predictions. Concept Bottleneck Models (CBMs) can provide a principled way of disclosing and guiding model behaviors through human-understandable concepts, albeit at a high cost of human efforts in…

2024

JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models

EMNLP 2024finding

The use of Large Language Models (LLMs) in hiring has led to legislative actions to protect vulnerable demographic groups. This paper presents a novel framework for benchmarking hierarchical gender hiring bias in Large Language Models (LLMs) for resume scoring, revealing significant issues of revers…

Cited by 8SourcePDFScholar
2024

Training Diffusion Models Towards Diverse Image Generation with Reinforcement Learning

CVPR 2024poster

Diffusion models have demonstrated unprecedented capabilities in image generation. Yet they incorporate and amplify the data bias (e.g. gender age) from the original training set limiting the diversity of generated images. In this paper we propose a diversity-oriented fine-tuning method using reinfo…

Cited by 10SourcePDFScholar
2023

Ring-Rotor: A Novel Retractable Ring-Shaped Quadrotor With Aerial Grasping and Transportation Capability

RA-L 2023

This letter presents a novel and retractable ring-shaped quadrotor called Ring-Rotor that can adjust the vehicle's length and width simultaneously. Unlike other morphing quadrotors with high platform complexity and poor controllability, Ring-Rotor uses only one servo motor for morphing but reduces t

Cited by 29SourceScholar
2022

Few-Shot Fast-Adaptive Anomaly Detection

NeurIPS 2022accept

The ability to detect anomaly has long been recognized as an inherent human ability, yet to date, practical AI solutions to mimic such capability have been lacking. This lack of progress can be attributed to several factors. To begin with, the distribution of ``abnormalities'' is intractable. Anythi…

Cited by 30SourcePDFScholar
2022

LF-VIO: A Visual-Inertial-Odometry Framework for Large Field-of-View Cameras with Negative Plane

IROS 2022poster

Visual-inertial-odometry has attracted extensive attention in the field of autonomous driving and robotics. The size of Field of View (FoV) plays an important role in Visual-Odometry (VO) and Visual-Inertial-Odometry (VIO), as a large FoV enables to perceive a wide range of surrounding scene element…

Cited by 21SourcecodeScholar
2022

Task Space Contouring Error Estimation and Precision Iterative Control of Robotic Manipulators

RA-L 2022

The task space contouring performance is significant for the machining accuracy of industrial robotic manipulators, but the contouring control of end-effector which is important in the industry has received scant attention. In this letter, a novel task space contouring error estimation and control s

Cited by 19SourceScholar
2021

Cirrus: A Long-range Bi-pattern LiDAR Dataset

ICRA 2021poster

In this paper, we introduce Cirrus, a new long-range bi-pattern LiDAR public dataset for autonomous driving tasks such as 3D object detection, critical to highway driving and timely decision making. Our platform is equipped with a high-resolution video camera and a pair of LiDAR sensors with a 250-m…

Cited by 41SourceScholar
2021

Learning to Learn Dense Gaussian Processes for Few-Shot Learning

NeurIPS 2021poster

Gaussian processes with deep neural networks demonstrate to be a strong learner for few-shot learning since they combine the strength of deep learning and kernels while being able to well capture uncertainty. However, it remains an open problem to leverage the shared knowledge provided by related ta…

Cited by 32SourcePDFScholar
2021

Meaningful Centroidal Frame Orientation of Multi-body Floating Locomotion Systems

ICRA 2021poster

In this paper, we propose a meaningful definition of rotational centroidal orientation which is somewhat missed in the state-of-the-art centroidal momentum and dynamics theory for locomotion robots with one floating base. This centroidal instantaneous orientation rotates as the robot runs, and it is…

Cited by 6SourceScholar
2021

Spatiotemporal Joint Filter Decomposition in 3D Convolutional Neural Networks

NeurIPS 2021poster

In this paper, we introduce spatiotemporal joint filter decomposition to decouple spatial and temporal learning, while preserving spatiotemporal dependency in a video. A 3D convolutional filter is now jointly decomposed over a set of spatial and temporal filter atoms respectively. In this way, a 3D…

Cited by 8SourcePDFScholar
2020

A Dictionary Approach to Domain-Invariant Learning in Deep Networks

NeurIPS 2020poster

In this paper, we consider domain-invariant deep learning by explicitly modeling domain shifts with only a small amount of domain-specific parameters in a Convolutional Neural Network (CNN). By exploiting the observation that a convolutional filter can be well approximated as a linear combination o…

Cited by 11SourcePDFScholar