← Search

Jifei Song

26 accepted papers

2026

Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensing

CVPR 2026

Always-on sensing is essential for next-generation edge/wearable AI systems, yet continuous high-fidelity RGB video capture remains prohibitively expensive for resource-constrained mobile and edge platforms. We present a new paradigm for efficient streaming video understanding: grayscale-always, col

Cited by 0SourcecodeScholar
2026

Diffusion-Based Makeup Transfer with Facial Region-Aware Makeup Features

CVPR 2026

Current diffusion-based makeup transfer methods commonly use the makeup information encoded by off-the-shelf foundation models (e.g., CLIP) as condition to preserve the makeup style of reference image in the generation. Although effective, these works mainly have two limitations: (1) foundation mode

Cited by 0SourcecodeScholar
2026

FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation

CVPR 2026

The development of generalizable Novel View Synthesis (NVS) models is critically limited by the scarcity of large-scale training data featuring diverse and precise camera trajectories. While real-world captures are photorealistic, they are typically sparse and discrete. Conversely, synthetic data sc

Cited by 0SourcecodeScholar
2026

LiteVSR: Enabling Cross-Domain Fine-Grained Detail Generation in Light-Weight Transformers for Video Super-Resolution

ICML 2026poster

Large-scale pre-trained video generators offer powerful priors for Video Super-Resolution (VSR), yet adapting them remains computationally prohibitive. Full fine-tuning demands extensive resources, and ControlNet-style adapters lose their efficiency advantage under modern Diffusion Transformers (DiT…

Cited by 0SourceScholar
2026

Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation

AAAI 2026technical

The performance of egocentric AI agents is fundamentally limited by multimodal intent ambiguity. This challenge arises from a combination of underspecified language, imperfect visual data, and deictic gestures, which frequently leads to task failure. Existing monolithic Vision-Language Models (VLMs)

Cited by 0SourcePDFScholar
2026

ViMo: A Generative Visual GUI World Model for App Agents

ICLR 2026poster

App agents, which autonomously operate mobile Apps through GUIs, have gained significant interest in real-world applications. Yet, they often struggle with long-horizon planning, failing to find the optimal actions for complex tasks with longer steps. To address this, world models are used to predic…

Cited by 0SourceScholar
2025

CaricatureBooth: Data-Free Interactive Caricature Generation in a Photo Booth

CVPR 2025poster

We present CaricatureBooth, a system that transforms caricature creation into a simple interactive experience -- as easy as using a photo booth! A key challenge in caricature generation is two-fold: the scarcity of high-quality caricature data and the difficulty in enabling precise creative control…

2025

Deep Gaussian from Motion: Exploring 3D Geometric Foundation Models for Gaussian Splatting

NeurIPS 2025poster

Neural radiance fields (NeRF) and 3D Gaussian Splatting (3DGS) are popular techniques to reconstruct and render photorealistic images. However, the prerequisite of running Structure-from-Motion (SfM) to get camera poses limits their completeness. Although previous methods can reconstruct a few unpos…

Cited by 0SourceScholar
2025

Frequency-Guided Diffusion for Training-Free Text-Driven Image Translation

ICCV 2025poster

Current training-free text-driven image translation primarily uses diffusion features (convolution and attention) of pre-trained model as guidance to preserve the style/structure of source image in translated image. However, the coarse guidance at feature level struggles with style (e.g., visual pat…

Cited by 0SourcePDFScholar
2025

Learning Precise Affordances from Egocentric Videos for Robotic Manipulation

ICCV 2025poster

Affordance, defined as the potential actions that an object offers, is crucial for embodied AI agents. For example, such knowledge directs an agent to grasp a knife by the handle for cutting or by the blade for safe handover. While existing approaches have made notable progress, affordance research…

2025

Single-view Image to Novel-view Generation for Hand-Object Interactions

AAAI 2025technical

Hand-object interaction modeling from a single RGB image is a significantly challenging task. Previous works typically reconstruct hand-object interactions as texture-less meshes, ignoring photo-realistic image generation. In this work, we introduce the HO123, a novel method to synthesize novel-view…

Cited by 0SourcePDFScholar
2025

UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting

ICLR 2025poster

Recent advancements in multi-modal 3D pre-training methods have shown promising efficacy in learning joint representations of text, images, and point clouds. However, adopting point clouds as 3D representation fails to fully capture the intricacies of the 3D world and exhibits a noticeable gap betwe…

Cited by 0SourcePDFScholar
2025

Unlocking the Potential of Diffusion Priors in Blind Face Restoration

ICCV 2025poster

Although diffusion prior is rising as a powerful solution for blind face restoration (BFR), the inherent gap between the vanilla diffusion model and BFR settings hinders its seamless adaptation. The gap mainly stems from the discrepancy between 1) high-quality (HQ) and low-quality (LQ) images and 2)…

Cited by 0SourcePDFScholar
2024

HeadGaS: Real-Time Animatable Head Avatars via 3D Gaussian Splatting

ECCV 2024poster

"3D head animation has seen major quality and runtime improvements over the last few years, particularly empowered by the advances in differentiable rendering and neural radiance fields. Real-time rendering is a highly desirable goal for real-world applications. We propose HeadGaS, a model that uses…

Cited by 35SourcePDFScholar
2024

Human Gaussian Splatting: Real-time Rendering of Animatable Avatars

CVPR 2024poster

This work addresses the problem of real-time rendering of photorealistic human body avatars learned from multi-view videos. While the classical approaches to model and render virtual humans generally use a textured mesh recent research has developed neural body representations that achieve impressiv…

2024

SAGS: Structure-Aware 3D Gaussian Splatting

ECCV 2024poster

"Following the advent of NeRFs, 3D Gaussian Splatting (3D-GS) has paved the way to real-time neural rendering overcoming the computational burden of volumetric methods. Several extensions of 3D-GS have been proposed to achieve compressible and high-fidelity performance. However, by employing a geome…

Cited by 7SourcePDFScholar
2024

SCRREAM : SCan, Register, REnder And Map: A Framework for Annotating Accurate and Dense 3D Indoor Scenes with a Benchmark

NeurIPS 2024poster

Traditionally, 3d indoor datasets have generally prioritized scale over ground-truth accuracy in order to obtain improved generalization. However, using these datasets to evaluate dense geometry tasks, such as depth rendering, can be problematic as the meshes of the dataset are often incomplete and…

2024

SWinGS: Sliding Windows for Dynamic 3D Gaussian Splatting

ECCV 2024poster

"Novel view synthesis has shown rapid progress recently, with methods capable of producing increasingly photorealistic results. 3D Gaussian Splatting has emerged as a promising method, producing high-quality renderings of scenes and enabling interactive viewing at real-time frame rates. However, it…

Cited by 11SourcePDFScholar
2023

On the Importance of Accurate Geometry Data for Dense 3D Vision Tasks

CVPR 2023poster

Learning-based methods to solve dense 3D vision problems typically train on 3D sensor data. The respectively used principle of measuring distances provides advantages and drawbacks. These are typically not compared nor discussed in the literature due to a lack of multi-modal datasets. Texture-less r…

2022

CroMo: Cross-Modal Learning for Monocular Depth Estimation

CVPR 2022poster

Learning-based depth estimation has witnessed recent progress in multiple directions; from self-supervision using monocular video to supervised methods offering highest accuracy. Complementary to supervision, further boosts to performance and robustness are gained by combining information from multi…

Cited by 19PDFScholar
2019

Generalizable Person Re-Identification by Domain-Invariant Mapping Network

CVPR 2019poster

We aim to learn a domain generalizable person re-identification (ReID) model. When such a model is trained on a set of source domains (ReID datasets collected from different camera networks), it can be directly applied to any new unseen dataset for effective ReID without any model updating. Despite…

Cited by 301PDFScholar
2018

Learning to Sketch With Shortcut Cycle Consistency

CVPR 2018poster

To see is to sketch -- free-hand sketching naturally builds ties between human and machine vision. In this paper, we present a novel approach for translating an object photo to a sketch, mimicking the human sketching process. This is an extremely challenging task because the photo and sketch domains…

Cited by 139SourcePDFScholar
2018

Universal Sketch Perceptual Grouping

ECCV 2018poster

In this work we aim to develop a universal sketch grouper. That is, a grouper that can be applied to sketches of any category in any domain to group constituent strokes/segments into semantically meaningful object parts. The first obstacle to this goal is the lack of large-scale datasets with groupi…

Cited by 58SourcePDFScholar
2017

Deep Spatial-Semantic Attention for Fine-Grained Sketch-Based Image Retrieval

ICCV 2017poster

Human sketches are unique in being able to capture both the spatial topology of a visual object, as well as its subtle appearance details. Fine-grained sketch-based image retrieval (FG-SBIR) importantly leverages on such fine-grained characteristics of sketches to conduct instance-level retrieval of…

Cited by 318PDFScholar