← Search

Jiyuan Zhang

17 accepted papers

2026

OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs

AAAI 2026technical

Existing sparse attention methods primarily target inference-time acceleration by selecting critical tokens under predefined sparsity patterns. However, they often fail to bridge the training–inference gap and lack the capacity for fine-grained token selection across multiple dimensions—such as quer

Cited by 0SourcePDFScholar
2026

Rapid Robot Manipulation Policy Learning Via Hierarchical Foundation-Model Prior Distillation

ICRA 2026poster

In robotic skill acquisition, rapid policy learning remains challenging due to high-dimensional state-action spaces and inefficient exploration in the early stage of training cite{p1}. Although the pre-trained OpenVLA model exhibits cross-task generalization and can generate goal-directed actions fo…

Cited by 0Scholar
2025

ACT as Human: Multimodal Large Language Model Data Annotation with Critical Thinking

NeurIPS 2025poster

Supervised learning relies on high-quality labeled data, but obtaining such data through human annotation is both expensive and time-consuming. Recent work explores using large language models (LLMs) for annotation, but LLM-generated labels still fall short of human-level quality. To address this pr…

Cited by 0SourceScholar
2025

USP-Gaussian: Unifying Spike-based Image Reconstruction, Pose Correction and Gaussian Splatting

CVPR 2025highlight

Spike camera, as an innovative type of neuromorphic camera that captures scenes with 0-1 bit stream at 40 kHz, is increasingly being employed for the novel view synthesis task building on the techniques such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). Previous spike-based appr…

2024

Continuous Spatiotemporal Events Decoupling through Spike-based Bayesian Computation

NeurIPS 2024poster

Numerous studies have demonstrated that the cognitive processes of the human brain can be modeled using the Bayesian theorem for probabilistic inference of the external world. Spiking neural networks (SNNs), capable of performing Bayesian computation with greater physiological interpretability, offe…

Cited by 0SourcePDFScholar
2024

Exploring Efficient Asymmetric Blind-Spots for Self-Supervised Denoising in Real-World Scenarios

CVPR 2024poster

Self-supervised denoising has attracted widespread attention due to its ability to train without clean images. However noise in real-world scenarios is often spatially correlated which causes many self-supervised algorithms that assume pixel-wise independent noise to perform poorly. Recent works hav…

Cited by 10SourcePDFScholar
2024

Spike-guided Motion Deblurring with Unknown Modal Spatiotemporal Alignment

CVPR 2024poster

The traditional frame-based cameras that rely on exposure windows for imaging experience motion blur in high-speed scenarios. Frame-based deblurring methods lack reliable motion cues to restore sharp images under extreme blur conditions. The spike camera is a novel neuromorphic visual sensor that ou…

2024

SpikeReveal: Unlocking Temporal Sequences from Real Blurry Inputs with Spike Streams

NeurIPS 2024spotlight

Reconstructing a sequence of sharp images from the blurry input is crucial for enhancing our insights into the captured scene and poses a significant challenge due to the limited temporal features embedded in the image. Spike cameras, sampling at rates up to 40,000 Hz, have proven effective in captu…

2024

Transient Glimpses: Unveiling Occluded Backgrounds through the Spike Camera

AAAI 2024technical

The de-occlusion problem, involving extracting clear background images by removing foreground occlusions, holds significant practical importance but poses considerable challenges. Most current research predominantly focuses on generating discrete images from calibrated camera arrays, but this approa…

2023

Enhancing Motion Deblurring in High-Speed Scenes with Spike Streams

NeurIPS 2023poster

Traditional cameras produce desirable vision results but struggle with motion blur in high-speed scenes due to long exposure windows. Existing frame-based deblurring algorithms face challenges in extracting useful motion cues from severely blurred images. Recently, an emerging bio-inspired vision se…

Cited by 13SourcePDFScholar
2023

Learning Temporal-Ordered Representation for Spike Streams Based on Discrete Wavelet Transforms

AAAI 2023technical

Spike camera, a new type of neuromorphic visual sensor that imitates the sampling mechanism of the primate fovea, can capture photons and output 40000 Hz binary spike streams. Benefiting from the asynchronous sampling mechanism, the spike camera can record fast-moving objects and clear images can be…

2023

Unsupervised Optical Flow Estimation with Dynamic Timing Representation for Spike Camera

NeurIPS 2023poster

Efficiently selecting an appropriate spike stream data length to extract precise information is the key to the spike vision tasks. To address this issue, we propose a dynamic timing representation for spike streams. Based on multi-layers architecture, it applies dilated convolutions on temporal dime…

2022

Hierarchical Representation-based Dynamic Reasoning Network for Biomedical Question Answering

COLING 2022main

Recently, Biomedical Question Answering (BQA) has attracted growing attention due to its application value and technical challenges. Most existing works treat it as a semantic matching task that predicts answers by computing confidence among questions, options and evidence sentences, which is insuff…

2022

Spatio-Temporal Recurrent Networks for Event-Based Optical Flow Estimation

AAAI 2022technical

Event camera has offered promising alternative for visual perception, especially in high speed and high dynamic range scenes. Recently, many deep learning methods have shown great success in providing model-free solutions to many event-based problems, such as optical flow estimation. However, existi…

2022

Spike Transformer: Monocular Depth Estimation for Spiking Camera

ECCV 2022poster

"Spiking camera is a bio-inspired vision sensor that mimics the sampling mechanism of the primate fovea, which has shown great potential for capturing high-speed dynamic scenes with a sampling rate of 40,000 Hz. Unlike conventional digital cameras, the spiking camera continuously captures photons an…

2019

Memory-Attended Recurrent Network for Video Captioning

CVPR 2019poster

Typical techniques for video captioning follow the encoder-decoder framework, which can only focus on one source video being processed. A potential disadvantage of such design is that it cannot capture the multiple visual context information of a word appearing in more than one relevant videos in tr…

Cited by 294PDFScholar