← Search

Zeyu Ma

25 accepted papers

2026

DWTSG: Parameter-Efficient Fine-Tuning of Large Pre-trained Models via Discrete Wavelet Transform and Subband Guidance

AAAI 2026technical

Fully fine-tuning large pre-trained models for each downstream task is impractical due to prohibitive memory, computation, and storage costs. Although parameter-efficient fine-tuning (PEFT) methods address this issue, leading methods like LoRA still exhibit linear scaling of trainable parameters wit

Cited by 0SourcePDFScholar
2026

Domain Adaptive Object Detection via Dynamic Causal Refinement

ICML 2026poster

Domain Adaptive Object Detection (DAOD) addresses the challenge of transferring object detectors from labeled source domains to unlabeled target domains. Existing domain adaptation methods primarily rely on feature distribution alignment, which enhances domain-invariant features (statistical invaria…

Cited by 0SourceScholar
2026

FairGC: Fostering Individual and Group Fairness for Deep Graph Clustering

AAAI 2026technical

The widespread adoption of graph neural networks (GNNs) has brought increased attention to fairness issues related to sensitive attributes, such as gender and race, in practical scenarios. However, this concern remains largely unexplored in the context of graph clustering. Conventional fair graph cl

Cited by 0SourcePDFScholar
2026

MaskGuide: Efficient Distillation for Deployable Lightweight Segmentation in Marine Environments

RA-L 2026

The growing demand for efficient image segmentation in marine ecological studies is currently constrained by two key factors: the high computational requirements of models such as the Segment Anything Model (SAM) and the degraded accuracy of lightweight models in underwater environments. To overcome

Cited by 0SourceScholar
2026

OSNIP: Breaking the Privacy-Utility-Efficiency Trilemma in LLM Inference via Obfuscated Semantic Null Space

ICML 2026poster

We propose Obfuscated Semantic Null space Injection for Privacy (OSNIP), a lightweight client-side encryption framework for privacy-preserving LLM inference. Generalizing the geometric intuition of linear kernels to the high-dimensional latent space of LLMs, we formally define the ``Obfuscated Seman…

Cited by 0SourceScholar
2026

TP-Spikformer: Token Pruned Spiking Transformer

ICLR 2026poster

Spiking neural networks (SNNs) offer an energy-efficient alternative to traditional neural networks due to their event-driven computing paradigm. However, recent advancements in spiking transformers have focused on improving accuracy with large-scale architectures, which require significant computat…

Cited by 0SourceScholar
2026

UFO-4D: Unposed Feedforward 4D reconstruction from Two Images

ICLR 2026poster

Dense 4D reconstruction from unposed images remains a critical challenge, with current methods relying on slow test-time optimization or fragmented, task-specific feedforward models. We introduce UFO-4D, a unified feedforward framework to reconstruct a dense, explicit 4D representation from just a p…

Cited by 0SourceScholar
2026

ViTPrompt: Training-Free Prompt Refinement with Visual Tokens for Open-Vocabulary Detection

CVPR 2026

Test-Time Adaptive Object Detection (TTAOD) aims to maintain detection performance under distribution shifts without retraining. While recent vision-language models enable open-vocabulary detection, existing TTAOD methods--whether closed-set or open-vocabulary--focus exclusively on improving classif

Cited by 0SourceScholar
2025

Bipolar Self-attention for Spiking Transformers

NeurIPS 2025spotlight

Harnessing the event-driven characteristic, Spiking Neural Networks (SNNs) present a promising avenue toward energy-efficient Transformer architectures. However, existing Spiking Transformers still suffer significant performance gaps compared to their Artificial Neural Network counterparts. Through…

Cited by 0SourceScholar
2025

Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling

NeurIPS 2025poster

The explosive growth in sequence length has intensified the demand for effective and efficient long sequence modeling. Benefiting from intrinsic oscillatory membrane dynamics, Resonate-and-Fire (RF) neurons can efficiently extract frequency components from input signals and encode them into spatiote…

Cited by 0SourceScholar
2025

Evaluating Robustness of Monocular Depth Estimation with Procedural Scene Perturbations

NeurIPS 2025poster

Recent years have witnessed substantial progress on monocular depth estimation, particularly as measured by the success of large models on standard benchmarks. However, performance on standard benchmarks does not offer a complete assessment, because most evaluate accuracy but not robustness. In thi…

Cited by 0SourcecodeScholar
2025

OMNI-DC: Highly Robust Depth Completion with Multiresolution Depth Integration

ICCV 2025poster

Depth completion (DC) aims to predict a dense depth map from an RGB image and a sparse depth map. Existing DC methods generalize poorly to new datasets or unseen sparse depth patterns, limiting their real-world applications. We propose OMNI-DC, a highly robust DC model that generalizes well zero-sho…

2025

SyncGaussian: Stable 3D Gaussian-Based Talking Head Generation with Enhanced Lip Sync via Discriminative Speech Features

IJCAI 2025

Generating high-fidelity talking heads that maintain stable head poses and achieve robust lip sync remains a significant challenge. Although methods based on 3D Gaussian Splatting (3DGS) offer a promising solution via point-based deformation, they suffer from inconsistent head dynamics and mismatche

Cited by 0SourcePDFScholar
2024

Infinigen Indoors: Photorealistic Indoor Scenes using Procedural Generation

CVPR 2024poster

We introduce Infinigen Indoors a Blender-based procedural generator of photorealistic indoor scenes. It builds upon the existing Infinigen system which focuses on natural scenes but expands its coverage to indoor scenes by introducing a diverse library of procedural indoor assets including furniture…

Cited by 24SourcePDFScholar
2023

CoCo: A Coupled Contrastive Framework for Unsupervised Domain Adaptive Graph Classification

ICML 2023poster

Although graph neural networks (GNNs) have achieved impressive achievements in graph classification, they often need abundant task-specific labels, which could be extensively costly to acquire. A credible solution is to explore additional labeled graphs to enhance unsupervised learning on the target…

Cited by 33SourcePDFScholar
2023

Infinite Photorealistic Worlds Using Procedural Generation

CVPR 2023poster

We introduce Infinigen, a procedural generator of photorealistic 3D scenes of the natural world. Infinigen is entirely procedural: every asset, from shape to texture, is generated from scratch via randomized mathematical rules, using no external source and allowing infinite variation and composition…

2022

Asymmetric Relation Consistency Reasoning for Video Relation Grounding

ECCV 2022poster

"Video relation grounding has attracted growing attention in the fields of video understanding and multimodal learning. While the past years have witnessed remarkable progress in this issue, the difficulties of multi-instance and complex temporal reasoning make it still a challenging task. In this p…

Cited by 5SourcePDFScholar
2022

DHWP: Learning High-Quality Short Hash Codes Via Weight Pruning

ICASSP 2022accepted

Hashing is widely used in large-scale image retrieval because of its efficiency in storage and computation. Although longer hash codes can lead to higher search accuracy, the retrieval cost increases linearly with the increase of the number of hash bits. Most deep hashing methods suffer from the pro…

Cited by 0SourceScholar
2022

Improved Deep Unsupervised Hashing with Fine-grained Semantic Similarity Mining for Multi-Label Image Retrieval

IJCAI 2022poster

In this paper, we study deep unsupervised hashing, a critical problem for approximate nearest neighbor research. Most recent methods solve this problem by semantic similarity reconstruction for guiding hashing network learning or contrastive learning of hash codes. However, in multi-label scenarios,…

Cited by 16SourcePDFScholar
2022

Learning Modal-Invariant and Temporal-Memory for Video-Based Visible-Infrared Person Re-Identification

CVPR 2022poster

Thanks for the cross-modal retrieval techniques, visible-infrared (RGB-IR) person re-identification (Re-ID) is achieved by projecting them into a common space, allowing person Re-ID in 24-hour surveillance systems. However, with respect to the "probe-to-gallery", almost all existing RGB-IR based cro…

Cited by 64PDFcodeScholar
2021

TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation

EMNLP 2021finding

Recent progress in generative language models has enabled machines to generate astonishingly realistic texts. While there are many legitimate applications of such models, there is also a rising need to distinguish machine-generated texts from human-written ones (e.g., fake news detection). However,…