← Search

Zhixiang Wang

29 accepted papers

2026

Continuous Gaussian Process Pre-Optimization for Asynchronous Event-Inertial Odometry

RA-L 2026

Event cameras, as bio-inspired sensors, are asynchronously triggered with high-temporal resolution compared to intensity cameras. Recent work has focused on fusing the event measurements with inertial measurements to enable ego-motion estimation in high-speed and HDR environments. However, existing

Cited by 5SourcecodeScholar
2026

Continuous Gaussian Process Pre-Optimization for Asynchronous Event-Inertial Odometry

ICRA 2026poster

Event cameras, as bio-inspired sensors, are asynchronously triggered with high-temporal resolution compared to intensity cameras. Recent work has focused on fusing the event measurements with inertial measurements to enable ego-motion estimation in high-speed and HDR environments. However, existing …

2026

One Tool Is Enough: Reinforcement Learning of LLM Agents for Repository-Level Code Navigation

ICML 2026poster

Locating files and functions requiring modification in large software repositories is challenging due to their scale and structural complexity. Existing LLM-based methods typically treat this as a repository-level retrieval task and rely on multiple auxiliary tools, which often overlook code executi…

Cited by 0SourceScholar
2026

Position: The Systemic Lack of Agency in Visual Reasoning

ICML 2026poster

This paper argues that a systemic lack of Agency constrains the implicit reasoning capabilities of current Vision-Language Models (VLMs). Implicit reasoning refers to the ability to autonomously discover and utilize hidden visual evidence to bridge information gaps, rather than merely relying on exp…

Cited by 0SourceScholar
2026

Reflection Separation from a Single Image via Joint Latent Diffusion

CVPR 2026

Single-image reflection separation is highly challenging under extreme conditions like glare or weak reflections. Existing methods often struggle to recover both layers in glare or weak-reflection scenarios because of insufficient information. This paper presents a diffusion model explicitly fine-tu

Cited by 0SourcecodeScholar
2025

3DM: Distill, Dynamic Drop, and Merge for Debiasing Multi-modal Large Language Models

ACL 2025finding

The rapid advancement of Multi-modal Language Models (MLLMs) has significantly enhanced performance in multimodal tasks, yet these models often exhibit inherent biases that compromise their reliability and fairness. Traditional debiasing methods face a trade-off between the need for extensive labele…

2025

Beyond Demonstrations: Dynamic Vector Construction from Latent Representations

EMNLP 2025

In-Context derived Vector (ICV) methods extract task-relevant representations from large language models (LLMs) and reinject them during inference, achieving comparable performance to few-shot In-Context Learning (ICL) without repeated demonstration processing. However, existing ICV methods remain s

Cited by 0SourcePDFScholar
2025

Fault Joint Detection and Adaptive Fault-Tolerant Control of Legged Robots Under Joint Partial Failures

RA-L 2025

Legged robots employ multiple joint actuators that are susceptible to abrupt partial failures during prolonged operation. Joint failures take multiple forms. They are not limited to complete lockout failures, which are the main focus of existing literature. They also include partial torque tracking

Cited by 3SourceScholar
2025

Foundation Model Insights and a Multi-Model Approach for Superior Fine-Grained One-shot Subset Selection

ICML 2025oral

One-shot subset selection serves as an effective tool to reduce deep learning training costs by identifying an informative data subset based on the information extracted by an information extractor (IE). Traditional IEs, typically pre-trained on the target dataset, are inherently dataset-dependent.…

2025

Large Language Models Might Not Care What You Are Saying: Prompt Format Beats Descriptions

EMNLP 2025

With the help of in-context learning (ICL), large language models (LLMs) have achieved impressive performance across various tasks. However, the function of descriptive instructions during ICL remains under-explored. In this work, we propose an ensemble prompt framework to describe the selection cri

2025

Sekai: A Video Dataset towards World Exploration

NeurIPS 2025poster

Video generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration. However, existing video generation datasets are not well-suited for world exploration training as they suffer from some limitations: limited locations, short duration, static s…

Cited by 0SourceScholar
2025

Subjective Camera 1.0: Bridging Human Cognition and Visual Reconstruction through Sequence-Aware Sketch-Guided Diffusion

ICCV 2025poster

We introduce the concept of a subjective camera to reconstruct meaningful moments that physical cameras fail to capture. We propose Subjective Camera 1.0, a framework for reconstructing real-world scenes from readily accessible subjective readouts, i.e., textual descriptions and progressively drawn…

2024

Asynchronous Event-Inertial Odometry using a Unified Gaussian Process Regression Framework

IROS 2024poster

Recent works have combined monocular event camera and inertial measurement unit to estimate the SE(3) trajectory. However, the asynchronicity of event cameras brings a great challenge to conventional fusion algorithms. In this paper, we present an asynchronous event-inertial odometry under a unified…

Cited by 2SourceScholar
2024

Contributing Dimension Structure of Deep Feature for Coreset Selection

AAAI 2024technical

Coreset selection seeks to choose a subset of crucial training samples for efficient learning. It has gained traction in deep learning, particularly with the surge in training dataset sizes. Sample selection hinges on two main aspects: a sample's representation in enhancing performance and the role…

2024

HumanNeRF-SE: A Simple yet Effective Approach to Animate HumanNeRF with Diverse Poses

CVPR 2024poster

We present HumanNeRF-SE a simple yet effective method that synthesizes diverse novel pose images with simple input. Previous HumanNeRF works require a large number of optimizable parameters to fit the human images. Instead we reload these approaches by combining explicit and implicit human represent…

Cited by 4SourcePDFScholar
2024

Revisiting Adversarial Patches for Designing Camera-Agnostic Attacks against Person Detection

NeurIPS 2024poster

Physical adversarial attacks can deceive deep neural networks (DNNs), leading to erroneous predictions in real-world scenarios. To uncover potential security risks, attacking the safety-critical task of person detection has garnered significant attention. However, we observe that existing attack met…

Cited by 1SourcePDFScholar
2024

SCOI: Syntax-augmented Coverage-based In-context Example Selection for Machine Translation

EMNLP 2024main

In-context learning (ICL) greatly improves the performance of large language models (LLMs) on various down-stream tasks, where the improvement highly depends on the quality of demonstrations. In this work, we introduce syntactic knowledge to select better in-context examples for machine translation…

2024

Self-reconfiguration Strategies for Space-distributed Spacecraft

IROS 2024poster

This paper proposes a distributed on-orbit spacecraft assembly algorithm, where future spacecraft can assemble modules with different functions on orbit to form a spacecraft structure with specific functions. This form of spacecraft organization has the advantages of reconfigurability, fast mission…

Cited by 0SourceScholar
2023

HOTCOLD Block: Fooling Thermal Infrared Detectors with a Novel Wearable Design

AAAI 2023technical

Adversarial attacks on thermal infrared imaging expose the risk of related applications. Estimating the security of these systems is essential for safely deploying them in the real world. In many cases, realizing the attacks in the physical space requires elaborate special perturbations. These solut…

2023

Rethinking Video Frame Interpolation from Shutter Mode Induced Degradation

ICCV 2023poster

Image restoration from various motion-related degradations, like blurry effects recorded by a global shutter (GS) and jello effects caused by a rolling shutter (RS), has been extensively studied. It has been recently recognized that such degradations encode temporal information, which can be exploit…

Cited by 8PDFScholar
2022

Both Style and Fog Matter: Cumulative Domain Adaptation for Semantic Foggy Scene Understanding

CVPR 2022oral

Although considerable progress has been made in semantic scene understanding under clear weather, it is still a tough problem under adverse weather conditions, such as dense fog, due to the uncertainty caused by imperfect observations. Besides, difficulties in collecting and labeling foggy images hi…

Cited by 64PDFScholar
2022

Neural Global Shutter: Learn To Restore Video From a Rolling Shutter Camera With Global Reset Feature

CVPR 2022poster

Most computer vision systems assume distortion-free images as inputs. The widely used rolling-shutter (RS) image sensors, however, suffer from geometric distortion when the camera and object undergo motion during capture. Extensive researches have been conducted on correcting RS distortions. However…

Cited by 15PDFcodeScholar
2020

Beyond Intra-modality: A Survey of Heterogeneous Person Re-identification

IJCAI 2020poster

An efficient and effective person re-identification (ReID) system relieves the users from painful and boring video watching and accelerates the process of video analysis. Recently, with the explosive demands of practical applications, a lot of research efforts have been dedicated to heterogeneous pe…

2020

Domain-Specific Mappings for Generative Adversarial Style Transfer

ECCV 2020poster

Style transfer generates an image whose content comes from one image and style from the other. Image-to-image translation approaches with disentangled representations have been shown effective for style transfer between two image categories. However, previous methods often assume a shared domain-inv…

2019

Learning to Reduce Dual-Level Discrepancy for Infrared-Visible Person Re-Identification

CVPR 2019poster

Infrared-Visible person RE-IDentification (IV-REID) is a rising task. Compared to conventional person re-identification (re-ID), IV-REID concerns the additional modality discrepancy originated from the different imaging processes of spectrum cameras, in addition to the person's appearance discrepanc…

Cited by 522PDFcodeScholar
2019

Non-local Self-attention Structure for Function Approximation in Deep Reinforcement Learning

ICASSP 2019accepted

Reinforcement learning is a framework to make sequential decisions. The combination with deep neural networks further improves the ability of this framework. Convolutional nerual networks make it possible to make sequential decisions based on raw pixels information directly and make reinforcement le…

Cited by 0SourceScholar