← Search

Jiayu Yang

14 accepted papers

2026

ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recall

ICLR 2026poster

LLMs require efficient knowledge editing (KE) to update factual information, yet existing methods exhibit significant performance decay in multi-hop factual recall. This failure is particularly acute when edits involve intermediate implicit subjects within reasoning chains. Through causal analysis,…

Cited by 0SourcecodeScholar
2026

ClipGStream: Clip-Stream Gaussian Splatting for Any Length and Any Motion Multi-View Dynamic Scene Reconstruction

CVPR 2026

Dynamic 3D scene reconstruction is essential for immersive media such as VR, MR, and XR, yet remains challenging for long multi-view sequences with large-scale motion. Existing dynamic Gaussian approaches are either Frame-Stream, offering scalability but poor temporal stability, or Clip, achieving l

Cited by 0SourceScholar
2026

Joint Shadow Generation and Relighting via Light-Geometry Interaction Maps

ICLR 2026poster

We propose Light–Geometry Interaction (LGI) maps, a novel representation that encodes light-aware occlusion from monocular depth. Unlike ray tracing, which requires full 3D reconstruction, LGI captures essential light–shadow interactions reliably and accurately, computed from off-the-shelf 2.5D dept…

Cited by 0SourceScholar
2026

SSCL: Adversarially Guided Image Compression via Semantic and Spectral Consistency Learning

AAAI 2026technical

Perceptual image compression has recently gained increasing attention, as it aims to reconstruct visually realistic images using generative models. Most existing methods adopt patch-based generative adversarial networks (PatchGAN) for one-step image generation, where adversarial training helps the d

Cited by 0SourcePDFScholar
2026

Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priors

CVPR 2026

We present a novel method for generating geometrically realistic and consistent orbital videos from a single image of an object. Existing video generation works mostly rely on pixel-wise attention to enforce view consistency across frames. However, such mechanism does not impose sufficient constrain

Cited by 0SourceScholar
2025

Compressing Streamable Free-Viewpoint Videos to 0.1 MB per Frame

AAAI 2025technical

The success of 3D Gaussian Splatting (3DGS) in static scenes has inspired numerous attempts to construct Free-Viewpoint Videos (FVVs) of dynamic scenes from multi-view videos. Despite advancements in current techniques, simultaneously achieving photo-realistic view synthesis results, fast on-the-fly…

2025

Instant Gaussian Stream: Fast and Generalizable Streaming of Dynamic Scene Reconstruction via Gaussian Splatting

CVPR 2025highlight

Building Free-Viewpoint Videos in a streaming manner offers the advantage of rapid responsiveness compared to offline training methods, greatly enhancing user experience. However, current streaming approaches face challenges of high per-frame reconstruction time (10s+) and error accumulation, limiti…

2025

LocalDyGS: Multi-view Global Dynamic Scene Modeling via Adaptive Local Implicit Feature Decoupling

ICCV 2025poster

Due to the complex and highly dynamic motions in the real world, synthesizing dynamic videos from multi-view inputs for arbitrary viewpoints is challenging. Previous works based on neural radiance field or 3D Gaussian splatting are limited to modeling fine-scale motion, greatly restricting their app…

Cited by 0SourcePDFScholar
2024

ConsistNet: Enforcing 3D Consistency for Multi-view Images Diffusion

CVPR 2024poster

Given a single image of a 3D object this paper proposes a novel method (named ConsistNet) that can generate multiple images of the same object as if they are captured from different viewpoints while the 3D (multi-view) consistencies among those multiple generated images are effectively exploited. Ce…

2024

Improving Learned Video Compression by Exploring Spatial Redundancy

ICASSP 2024accepted

Learned video compression has developed rapidly and shown promising rate-distortion performance recently. Existing works have made great progress on removing temporal redundancy between inter-frames, while neglecting spatial redundancy within a frame. In this paper, we propose to explore spatial red…

Cited by 0SourceScholar
2023

Parametric Depth Based Feature Representation Learning for Object Detection and Segmentation in Bird's-Eye View

ICCV 2023poster

Recent vision-only perception models for autonomous driving achieved promising results by encoding multi-view image features into Bird's-Eye-View (BEV) space. A critical step and the main bottleneck of these methods is transforming image features into the BEV coordinate frame. This paper focuses on…

Cited by 10PDFcodeScholar
2022

Non-Parametric Depth Distribution Modelling Based Depth Inference for Multi-View Stereo

CVPR 2022poster

Recent cost volume pyramid based deep neural networks have unlocked the potential of efficiently leveraging high-resolution images for depth inference from multi-view stereo. In general, those approaches assume that the depth of each pixel follows a unimodal distribution. Boundary pixels usually fol…

Cited by 45PDFcodeScholar