← Search

Fang Zhang

14 accepted papers

2026

Equi-RO: A 4D mmWave Radar Odometry via Equivariant Networks

RA-L 2026

Autonomous vehicles and robots rely on accurate odometry estimation in GPS-denied environments. While LiDARs and cameras struggle under extreme weather, 4D mmWave radar emerges as a robust alternative with all-weather operability and velocity measurement. In this paper, we introduce Equi-RO, an equi

Cited by 3SourceScholar
2026

FastHybrid: Accelerating Hybrid Autoregressive Image Generation with Lookahead and Guided Decoding

CVPR 2026

Autoregressive (AR) models have achieved remarkable success in natural language processing, yet their application to image generation faces significant challenges. When implementing VQ-based decoders for autoregressive image generation, the generated images typically preserve semantic information bu

Cited by 0SourceScholar
2026

Trajectory-Consistent Denoising Diffusion Codebook Models for Zero-Shot High-Fidelity Image Compression at Ultra-Low Bitrates

IJCAI 2026

Denoising Diffusion Codebook Models (DDCM) have emerged as a promising framework for zero-shot image compression by replacing stochastic sampling with discrete selection from a reproducible Gaussian codebook. By greedily picking noise vectors that best match the target image, DDCM encodes the genera

Cited by 0Scholar
2026

Unveiling And Addressing Dimensional Collapse In Vector Quantization Models Via Codebook Regularization

ICML 2026poster

While recent advancements in Vector Quantization (VQ) models have successfully achieved complete codebook utilization, a critical bottleneck remains largely unexplored: the effective dimensionality of the codebook embedding space. We observe that discrete codebook representations tend to degenerate …

Cited by 0SourceScholar
2025

Dynamic Prefix as Instructor for Incremental Named Entity Recognition: A Unified Seq2Seq Generation Framework

ACL 2025finding

The Incremental Named Entity Recognition (INER) task aims to update a model to extract entities from an expanding set of entity type candidates due to concerns related to data privacy and scarcity. However, conventional sequence labeling approaches to INER often suffer from the catastrophic forgetti…

2025

PlaneRAS: Learning Planar Primitives for 3D Plane Recovery

ICCV 2025poster

3D plane recovery from monocular images constitutes a fundamental task in indoor scene understanding. Recent methods formulate this problem as 2D pixel-level segmentation through convolutional networks or query-based architectures, which purely rely on 2D pixel features while neglecting the inherent…

Cited by 0SourcePDFScholar
2025

Spherical Scissor-Like Reconfigurable Palm Design in Robotic Hands: Insights from Human Hand Functionality

IROS 2025

The human palm demonstrates spatial reconfigurability during the gripping process and forms a spherical grasping envelope. Based on these observations, this study designs a reconfigurable spherical palm that incorporates a spatial scissor mechanism, which only requires a single actuator to reshape t

Cited by 0SourceScholar
2025

S²MILE: Semantic-and-Structure-Aware Music-Driven Lyric Generation

AAAI 2025technical

The task of music-to-lyric generation aims to create lyrics that can be sung in harmony with the music while capturing the music’s intrinsic meaning. Previous efforts in this area have struggled to effectively handle both the structural and semantic alignments of music and lyrics, often relying on r…

Cited by 0SourcePDFScholar
2024

Empowering Diffusion Models on the Embedding Space for Text Generation

NAACL 2024long

Diffusion models have achieved state-of-the-art synthesis quality on both visual and audio tasks, and recent works further adapt them to textual data by diffusing on the embedding space. In this paper, we conduct systematic studies of the optimization challenges encountered with both the embedding s…

2024

Salpot: A Jet Propulsion Swimmer With Scissor Structure and Bilateral Apertures

RA-L 2024

In recent years, researchers have increasingly turned to marine organisms for inspiration in designing underwater robots. While most robots rely on jet propulsion, akin to squid or jellyfish, using a single posterior aperture for water intake and expulsion, there are few incorporating an additional

Cited by 6SourceScholar
2024

Visual Hallucination Elevates Speech Recognition

AAAI 2024technical

Due to the detrimental impact of noise on the conventional audio speech recognition (ASR) task, audio-visual speech recognition~(AVSR) has been proposed by incorporating both audio and visual video signals. Although existing methods have demonstrated that the aligned visual input of lip movements ca…

Cited by 5SourcePDFScholar
2022

Anderson Acceleration for on-Manifold Iterated Error State Kalman Filters

RA-L 2022

Iterated Extended Kalman Filter is a promising and widely-used estimator for real-time localization applications. It iterates the observation equation to find a better linearization point and, simultaneously, only maintains the state estimation in a single time to save the computation resources. Ins

Cited by 10SourceScholar
2022

Faster-LIO: Lightweight Tightly Coupled Lidar-Inertial Odometry Using Parallel Sparse Incremental Voxels

RA-L 2022

This letter presents an incremental voxel-based lidar-inertial odometry (LIO) method for fast-tracking spinning and solid-state lidar scans. To achieve the high tracking speed, we neither use complicated tree-based structures to divide the spatial point cloud nor the strict k nearest neighbor (k-NN)

Cited by 338SourceScholar
2015

Deep convolutional activation features for large scale Brain Tumor histopathology image classification and segmentation

ICASSP 2015accepted

We propose a simple, efficient and effective method using deep convolutional activation features (CNNs) to achieve stat- of-the-art classification and segmentation for the MICCAI 2014 Brain Tumor Digital Pathology Challenge. Common traits of such medical image challenges are characterized by large i…

Cited by 0SourceScholar