← Search

Zixin Zhang

14 accepted papers

2026

EvDiff3D: Event-Aware Diffusion Repair for High-Fidelity Event-Based 3D Reconstruction

AAAI 2026technical

Event cameras are bio-inspired sensors that capture visual information through asynchronous brightness changes, offering distinct advantages including high temporal resolution and wide dynamic range. While prior research has investigated event-based 3D reconstruction for extreme scenarios, existing

Cited by 0SourcePDFScholar
2026

ImpText: A Benchmark and Tool-Augmented Framework for Implicit Text Reasoning

ICML 2026poster

Multimodal Large Language Models (MLLMs) have demonstrated exceptional proficiency in standard text extraction, but they encounter significant challenges when confronting real-world implicit text. Such content typically contains malicious information, intentionally concealed through physical deforma…

Cited by 0SourceScholar
2026

Show, Don't Tell: Morphing Latent Reasoning into Image Generation

ICML 2026poster

Text-to-image (T2I) generation has achieved remarkable progress, yet existing methods often lack the ability to dynamically reason and refine during generation--a hallmark of human creativity. Current reasoning-augmented paradigms mostly rely on explicit thought processes, where intermediate reasoni…

Cited by 0SourceScholar
2026

TiViBench: Benchmarking Think-in-Video Reasoning for Video Generation

CVPR 2026

The rapid evolution of video generative models has shifted their focus from producing visually plausible outputs to tackling tasks requiring physical plausibility and logical consistency. However, despite recent breakthroughs such as Veo 3's chain-of-frames reasoning, it remains unclear whether thes

Cited by 0SourcecodeScholar
2025

ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback

NeurIPS 2025poster

With the rapid advancement of generative models, general-purpose generation has gained increasing attention as a promising approach to unify diverse tasks across modalities within a single system. Despite this progress, existing open-source frameworks often remain fragile and struggle to support com…

Cited by 0SourceScholar
2025

Event-Guided Consistent Video Enhancement with Modality-Adaptive Diffusion Pipeline

NeurIPS 2025poster

Recent advancements in low-light video enhancement (LLVE) have increasingly leveraged both RGB and event cameras to improve video quality under challenging conditions. However, existing approaches share two key drawbacks. First, they are tuned for steady low-light scenes, so their performance drops…

Cited by 0SourceScholar
2025

Robots with Attitude: Singularity-Free Quaternion-Based Model-Predictive Control for Agile Legged Robots

ICRA 2025

We present a model-predictive control (MPC) framework for legged robots that avoids the singularities associated with common three-parameter attitude representations like Euler angles during large-angle rotations. Our method parameterizes the robot's attitude with singularity-free unit quaternions a

Cited by 2SourcecodeScholar
2025

Sample-Efficient Online Control Policy Learning with Real-Time Recursive Model Updates

CoRL 2025poster

Data-driven control methods need to be sample-efficient and lightweight, especially when data acquisition and computational resources are limited---such as during learning on hardware. Most modern data-driven methods require large datasets and struggle with real-time updates of models, limiting thei…

Cited by 0SourceScholar
2024

Enhancing Storage and Computational Efficiency in Federated Multimodal Learning for Large-Scale Models

ICML 2024poster

The remarkable generalization of large-scale models has recently gained significant attention in multimodal research. However, deploying heterogeneous large-scale models with different modalities under Federated Learning (FL) to protect data privacy imposes tremendous challenges on clients' limited…

2023

Cerberus: Low-Drift Visual-Inertial-Leg Odometry For Agile Locomotion

ICRA 2023poster

We present an open-source Visual-Inertial-Leg Odometry (VILO) state estimation solution for legged robots, called Cerberus, which precisely estimates position on various terrains in real-time using a set of standard sensors, including stereo cameras, IMU, joint encoders, and contact sensors. In addi…

Cited by 37SourcecodeScholar
2023

PlankAssembly: Robust 3D Reconstruction from Three Orthographic Views with Learnt Shape Programs

ICCV 2023poster

In this paper, we develop a new method to automatically convert 2D line drawings from three orthographic views into 3D CAD models. Existing methods for this problem reconstruct 3D models by back-projecting the 2D observations into 3D space while maintaining explicit correspondence between the input…

Cited by 5PDFcodeScholar