← Search

Zeyu Chen

22 accepted papers

2026

Beyond Detection: A Structure-Aware Framework for Scene Text Tracking

ICML 2026poster

Modern visual object trackers show impressive results on general targets, yet their performance drops substantially when dealing with scene text. Although currently underexplored, tracking text in videos is essential for dynamic text manipulations such as segmentation, removal, and editing. To fill …

Cited by 0SourceScholar
2026

Prototype-Based Pseudo-Label Denoising for Source-Free Domain Adaptation in Remote Sensing Semantic Segmentation

ICASSP 2026oral

Source-Free Domain Adaptation (SFDA) enables domain adaptation for semantic segmentation of Remote Sensing Images (RSIs) using only a well-trained source model and unlabeled target domain data. However, the lack of ground-truth labels in the target domain often leads to the generation of noisy pseud…

Cited by 0SourcePDFScholar
2026

Sortblock: Similarity-Aware Feature Reuse for Diffusion Model

AAAI 2026technical

Diffusion Transformers (DiTs) have demonstrated remarkable generative capabilities, particularly benefiting from Transformer architectures that enhance visual and artistic fidelity. However, their inherently sequential denoising process results in high inference latency, limiting their deployment in

Cited by 0SourcePDFScholar
2025

3D Gaussian Splatting for Fine-Detailed Surface Reconstruction in Large-Scale Scene

IROS 2025

Recent developments in 3D Gaussian Splatting have made significant advances in surface reconstruction. However, scaling these methods to large-scale scenes remains challenging due to high computational demands and the complex dynamic appearances typical of outdoor environments. These challenges hind

Cited by 5SourceScholar
2025

FlashMask: Efficient and Rich Mask Extension of FlashAttention

ICLR 2025poster

The computational and memory demands of vanilla attention scale quadratically with the sequence length $N$, posing significant challenges for processing long sequences in Transformer models. FlashAttention alleviates these challenges by eliminating the $\mathcal{O}(N^2)$ memory dependency and reduci…

2025

KDMOS:Knowledge Distillation for Motion Segmentation

IROS 2025

Motion Object Segmentation (MOS) is crucial for autonomous driving, as it enhances localization, path planning, map construction, scene flow estimation, and future state prediction. While existing methods achieve strong performance, balancing accuracy and real-time inference remains a challenge. To

Cited by 0SourcecodeScholar
2025

Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation

NeurIPS 2025poster

Reinforcement learning (RL) has garnered increasing attention in text-to-image (T2I) generation. However, most existing RL approaches are tailored to either diffusion models or autoregressive models, overlooking an important alternative: masked generative models. In this work, we propose Mask-GRPO,…

Cited by 0SourceScholar
2025

Watermarking Large Language Models: An Unbiased and Low-risk Method

ACL 2025long

Recent advancements in large language models (LLMs) have highlighted the risk of misusing them, raising the need for accurate detection of LLM-generated content. In response, a viable solution is to inject imperceptible identifiers into LLMs, known as watermarks. Our research extends the existing wa…

2024

A Point-Line Features Fusion Method for Fast and Robust Monocular Visual-Inertial Initialization

IROS 2024poster

Fast and robust initialization is essential for highly accurate monocular visual-inertial odometer (VIO), but at present majority of initialization methods rely only on point features, unstable in low texture and blurring situations. Therefore, we propose a novel point-line features fusion method fo…

Cited by 0SourceScholar
2024

FAFA: Frequency-Aware Flow-Aided Self-Supervision for Underwater Object Pose Estimation

ECCV 2024poster

"Although methods for estimating the pose of objects in indoor scenes have achieved great success, the pose estimation of underwater objects remains challenging due to difficulties brought by the complex underwater environment, such as degraded illumination, blurring, and the substantial cost of obt…

2024

G–LIME: Statistical Learning for Local Interpretations of Deep Neural Networks Using Global Priors (Abstract Reprint)

AAAI 2024technical

To explain the prediction result of a Deep Neural Network (DNN) model based on a given sample, LIME [1] and its derivatives have been proposed to approximate the local behavior of the DNN model around the data point via linear surrogates. Though these algorithms interpret the DNN by finding the key…

Cited by 1SourcePDFScholar
2024

ROV6D: 6D Pose Estimation Benchmark Dataset for Underwater Remotely Operated Vehicles

RA-L 2024

Accurately localization between multi-robots is crucial for many underwater applications, such as tracking, convoying and subsea intervention tasks. 6D pose estimation is a fundamental task that enables precise object localization in 3D space with full six degrees of freedom. However, one critical c

Cited by 12SourceScholar
2024

UW-SDF: Exploiting Hybrid Geometric Priors for Neural SDF Reconstruction from Underwater Multi-view Monocular Images

IROS 2024

Due to the unique characteristics of underwater environments, accurate 3D reconstruction of underwater objects poses a challenging problem in tasks such as underwater exploration and mapping. Traditional methods that rely on multiple sensor data for 3D reconstruction are time-consuming and face chal

Cited by 2SourceScholar
2023

LIMI-VC: A Light Weight Voice Conversion Model with Mutual Information Disentanglement

ICASSP 2023accepted

Voice conversion(VC) model aims to convert the source timbre to the target one. Recently, many VC models utilize pre-trained models to enhance the performance and achieve good results. However, pre-trained models could not somehow disentangle the timbre and linguistic information, thus resulting in…

Cited by 0SourceScholar
2022

PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit

NAACL 2022system demonstrations

PaddleSpeech is an open-source all-in-one speech toolkit. It aims at facilitating the development and research of speech processing technologies by providing an easy-to-use command-line interface and a simple code structure. This paper describes the design philosophy and core architecture of PaddleS…

2022

Parameter-Free Style Projection for Arbitrary Image Style Transfer

ICASSP 2022accepted

Arbitrary image style transfer is a challenging task which aims to stylize a content image conditioned on arbitrary style images. In this task the feature-level content-style transformation plays a vital role for proper fusion of features. Existing feature transformation algorithms often suffer from…

Cited by 0SourceScholar
2022

RGL: A Simple yet Effective Relation Graph Augmented Prompt-based Tuning Approach for Few-Shot Learning

NAACL 2022findings

Pre-trained language models (PLMs) can provide a good starting point for downstream applications. However, it is difficult to generalize PLMs to new tasks given a few labeled samples. In this work, we show that Relation Graph augmented Learning (RGL) can improve the performance of few-shot natural l…

2022

Simple and Effective Relation-based Embedding Propagation for Knowledge Representation Learning

IJCAI 2022poster

Relational graph neural networks have garnered particular attention to encode graph context in knowledge graphs (KGs). Although they achieved competitive performance on small KGs, how to efficiently and effectively utilize graph context for large KGs remains an open problem. To this end, we propose…

2020

ReDA:Reinforced Differentiable Attribute for 3D Face Reconstruction

CVPR 2020oral

The key challenge for 3D face shape reconstruction is to build the correct dense face correspondence between the deformable mesh and the single input image. Given the ill-posed nature, previous works heavily rely on prior knowledge (such as 3DMM [2]) to reduce depth ambiguity. Although impressive re…

Cited by 47PDFScholar