← Search

lu Yang

21 accepted papers

2026

3D MeanFlow: One-Step Point Cloud Completion and Generation via Average-Velocity Transport

ICML 2026poster

Point cloud completion and generation are important across many 3D tasks, where both fidelity and sampling efficiency matter. Prevailing high-fidelity approaches rely on long sampling schedules, which incur substantial inference latency. Few-step alternatives typically use rectification or distillat…

Cited by 0SourceScholar
2026

Exploring Position Encoding Mechanism in Diffusion U-Net for Training-free High-resolution Image Generation

AAAI 2026technical

Denoising higher-resolution latents using a pre-trained U-Net often results in repetitive and disordered image patterns. In this work, we are motivated to reveal the intrinsic cause of such pattern disruption in high-resolution image generation. Through theoretical analysis and empirical studies, we

Cited by 0SourcePDFScholar
2025

A Compact Implicit Neural Representation for Efficient Storage of Massive 4D Functional Magnetic Resonance Imaging

AAAI 2025technical

Functional Magnetic Resonance Imaging (fMRI) data is a widely used kind of four-dimensional biomedical data, which requires effective compression. However, fMRI compressing poses unique challenges due to its intricate temporal dynamics, low signal-to-noise ratio, and complicated underlying redundanc…

Cited by 0SourcePDFScholar
2025

How Far Are We from Optimal Reasoning Efficiency?

NeurIPS 2025poster

Large Reasoning Models (LRMs) demonstrate remarkable problem-solving capabilities through extended Chain-of-Thought (CoT) reasoning but often produce excessively verbose and redundant reasoning traces. This inefficiency incurs high inference costs and limits practical deployment. While existing fine…

Cited by 0SourcecodeScholar
2025

Image is All You Need to Empower Large-scale Diffusion Models for In-Domain Generation

CVPR 2025poster

In-domain generation aims to perform a variety of tasks within a specific domain, such as unconditional generation, text-to-image, image editing, 3D generation, and more. Early research typically required training specialized generators for each unique task and domain, often relying on fully-labeled…

Cited by 0SourcePDFScholar
2025

LLGS: Unsupervised Gaussian Splatting for Image Enhancement and Reconstruction in Pure Dark Environment

ICRA 2025

D Gaussian Splatting has shown remarkable capabilities in novel view rendering tasks and exhibits significant potential for multi-view optimization. However, the original 3D Gaussian Splatting lacks color representation for inputs in lowlight environments. Simply using enhanced images as inputs woul

Cited by 3SourceScholar
2025

Label Drop for Multi-Aspect Relation Modeling in Universal Information Extraction

NAACL 2025long

Universal Information Extraction (UIE) has garnered significant attention due to its ability to address model explosion problems effectively. Extractive UIE can achieve strong performance using a relatively small model, making it widely adopted. Extractive UIEs generally rely on task instructions fo…

2025

NOTA: Multimodal Music Notation Understanding for Visual Large Language Model

NAACL 2025findings

Symbolic music is represented in two distinct forms: two-dimensional, visually intuitive score images, and one-dimensional, standardized text annotation sequences. While large language models have shown extraordinary potential in music, current research has primarily focused on unimodal symbol seque…

Cited by 0SourcePDFScholar
2024

The Music Maestro or The Musically Challenged, A Massive Music Evaluation Benchmark for Large Language Models

ACL 2024findings

Benchmark plays a pivotal role in assessing the advancements of large language models (LLMs). While numerous benchmarks have been proposed to evaluate LLMs’ capabilities, there is a notable absence of a dedicated benchmark for assessing their musical abilities. To address this gap, we present ZIQI-E…

2023

Large-Scale Person Detection and Localization Using Overhead Fisheye Cameras

ICCV 2023oral

Location determination finds wide applications in daily life. Instead of existing efforts devoted to localizing tourist photos captured by perspective cameras, in this article, we focus on developing person positioning solutions using overhead fisheye cameras. Such solutions are advantageous in larg…

Cited by 25PDFScholar
2022

Dynamically Transformed Instance Normalization Network for Generalizable Person Re-identification

ECCV 2022poster

"Existing person re-identification methods often suffer significant performance degradation on unseen domains, which fuels interest in domain generalizable person re-identification (DG-PReID). As an effective technology to alleviate domain variance, the Instance Normalization (IN) has been widely em…

Cited by 51SourcePDFScholar
2022

Joint Dual-Domain Matrix Factorization for ECG Biometric Recognition

ICASSP 2022accepted

Electrocardiogram (ECG) biometrics has aroused extensive attention in the research field of biometric recognition. How-ever, most existing methods either only consider a single do-main (time domain or frequency domain) to extract features or extract multi-features while ignoring the specific proper-…

Cited by 0SourceScholar
2022

Locality-Aware Inter- and Intra-Video Reconstruction for Self-Supervised Correspondence Learning

CVPR 2022poster

Our target is to learn visual correspondence from unlabeled videos. We develop LIIR, a locality-aware inter-and intra-video reconstruction framework that fills in three missing pieces, i.e., instance discrimination, location awareness, and spatial compactness, of self-supervised correspondence learn…

Cited by 54PDFcodeScholar
2020

Efficient Scene Text Detection with Textual Attention Tower

ICASSP 2020accepted

Scene text detection has received attention for years and achieved an impressive performance across various benchmarks. In this work, we propose an efficient and accurate approach to detect multi-oriented text in scene images. The proposed feature fusion mechanism allows us to use a shallower networ…

Cited by 0SourceScholar
2020

Renovating Parsing R-CNN for Accurate Multiple Human Parsing

ECCV 2020poster

Multiple human parsing aims to segment various human parts and associate each part with the corresponding instance simultaneously. This is a very challenging task due to the diverse human appearance, semantic ambiguity of different body parts and clothing, and complex background. Through analysis of…

2019

Vehicle Re-Identification in Aerial Imagery: Dataset and Approach

ICCV 2019poster

In this work, we construct a large-scale dataset for vehicle re-identification (ReID), which contains 137k images of 13k vehicle instances captured by UAV-mounted cameras. To our knowledge, it is the largest UAV-based vehicle ReID dataset. To increase intra-class variation, each vehicle is captured…

Cited by 78PDFScholar