← Search

Yijin Li

18 accepted papers

2026

D-Prism: Differentiable Primitives for Structured Dynamic Modeling

CVPR 2026

Capturing both geometry and rigid motion for structured dynamic objects, like multi-part assemblies or jointed mechanisms, remains a key challenge. Existing dynamic methods, such as deformable meshes or 3DGS, rely on unstructured representations and fail to jointly model suitable geometry and articu

Cited by 0SourcecodeScholar
2026

Neural Field-Based 3D Surface Reconstruction of Microstructures from Multi-Detector Signals in Scanning Electron Microscopy

CVPR 2026

The 3D characterization of microstructures is crucial for understanding and designing functional materials. However, the scanning electron microscope (SEM), widely used in scientific research, captures only 2D electron intensity distributions. Existing SEM 3D reconstruction methods struggle with tex

Cited by 0SourcecodeScholar
2025

BlinkTrack: Feature Tracking over 80 FPS via Events and Images

ICCV 2025poster

Event cameras, known for their high temporal resolution and ability to capture asynchronous changes, have gained significant attention for their potential in feature tracking, especially in challenging conditions. However, event cameras lack the fine-grained texture information that conventional cam…

2025

ETO+: Revisit the Refinement Stage in Efficient Feature Matching

IROS 2025

Recent feature matching approaches like ETO have focused on developing lightweight matching algorithms for real-time applications. However, their lack of cross-image feature interaction and sufficient refinement often lead to a decline in matching accuracy. To address these challenges, we propose ET

Cited by 0SourceScholar
2025

GS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point Tracking

CVPR 2025poster

4D video control is essential in video generation as it enables the use of sophisticated lens techniques, such as multi-camera shooting and dolly zoom, which are currently unsupported by existing methods. Training a video Diffusion Transformer (DiT) directly to control 4D content requires expensive…

2024

"BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation using RGB Frames and Events"

ECCV 2024poster

"Recent advances in event-based vision suggest that they complement traditional cameras by providing continuous observation without frame rate limitations and high dynamic range which are well-suited for correspondence tasks such as optical flow and point tracking. However, so far there is still a l…

Cited by 4SourcePDFScholar
2024

A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding

NeurIPS 2024poster

In this paper, we propose a novel multi-view stereo (MVS) framework that gets rid of the depth range prior. Unlike recent prior-free MVS methods that work in a pair-wise manner, our method simultaneously considers all the source images. Specifically, we introduce a Multi-view Disparity Attention (MD…

Cited by 0SourcePDFScholar
2024

DiffInDScene: Diffusion-based High-Quality 3D Indoor Scene Generation

CVPR 2024poster

We present DiffInDScene a novel framework for tackling the problem of high-quality 3D indoor scene generation which is challenging due to the complexity and diversity of the indoor scene geometry. Although diffusion-based generative models have previously demonstrated impressive performance in image…

2024

ETO:Efficient Transformer-based Local Feature Matching by Organizing Multiple Homography Hypotheses

NeurIPS 2024poster

We tackle the efficiency problem of learning local feature matching.Recent advancements have given rise to purely CNN-based and transformer-based approaches, each augmented with deep learning techniques. While CNN-based methods often excel in matching speed, transformer-based methods tend to provide…

Cited by 3SourcePDFScholar
2024

ZoLA: Zero-Shot Creative Long Animation Generation with Short Video Model

ECCV 2024oral

"Although video generation has made great progress in capacity and controllability and is gaining increasing attention, currently available video generation models still make minimal progress in the video length they can generate. Due to the lack of well-annotated long video data, high training/infe…

2023

BlinkFlow: A Dataset to Push the Limits of Event-Based Optical Flow Estimation

IROS 2023poster

Event cameras provide high temporal precision, low data rates, and high dynamic range visual perception, which are well-suited for optical flow estimation. While data-driven optical flow estimation has obtained great success in RGB cameras, its generalization performance is seriously hindered in eve…

Cited by 38SourcecodeScholar
2023

Context-PIPs: Persistent Independent Particles Demands Spatial Context Features

NeurIPS 2023spotlight

We tackle the problem of Persistent Independent Particles (PIPs), also called Tracking Any Point (TAP), in videos, which specifically aims at estimating persistent long-term trajectories of query points in videos. Previous methods attempted to estimate these trajectories independently to incorporate…

2023

Multi-Modal Neural Radiance Field for Monocular Dense SLAM with a Light-Weight ToF Sensor

ICCV 2023poster

Light-weight time-of-flight (ToF) depth sensors are compact and cost-efficient, and thus widely used on mobile devices for tasks such as autofocus and obstacle detection. However, due to the sparse and noisy depth measurements, these sensors have rarely been considered for dense geometry reconstruct…

Cited by 32PDFcodeScholar
2023

PATS: Patch Area Transportation With Subdivision for Local Feature Matching

CVPR 2023poster

Local feature matching aims at establishing sparse correspondences between a pair of images. Recently, detector-free methods present generally better performance but are not satisfactory in image pairs with large scale differences. In this paper, we propose Patch Area Transportation with Subdivision…

Cited by 42SourcePDFScholar
2022

DELTAR: Depth Estimation from a Light-Weight ToF Sensor and RGB Image

ECCV 2022poster

"Light-weight time-of-flight (ToF) depth sensors are small, cheap, low-energy and have been massively deployed on mobile devices for the purposes like autofocus, obstacle detection, etc. However, due to their specific measurements (depth distribution in a region instead of the depth value at a certa…

2021

Graph-Based Asynchronous Event Processing for Rapid Object Recognition

ICCV 2021poster

Different from traditional video cameras, event cameras capture asynchronous events stream in which each event encodes pixel location, trigger time, and the polarity of the brightness changes. In this paper, we introduce a novel graph-based framework for event cameras, namely SlideGCN. Unlike some r…

Cited by 110PDFScholar
2021

Learning Object-Compositional Neural Radiance Field for Editable Scene Rendering

ICCV 2021poster

Implicit neural rendering techniques have shown promising results for novel view synthesis. However, existing methods usually encode the entire scene as a whole, which is generally not aware of the object identity and limits the ability to the high-level editing tasks such as moving or adding furnit…

Cited by 356PDFScholar
2021

VS-Net: Voting With Segmentation for Visual Localization

CVPR 2021poster

Visual localization is of great importance in robotics and computer vision. Recently, scene coordinate regression based methods have shown good performance in visual localization in small static scenes. However, it still estimates camera poses from many inferior scene coordinates. To address this pr…

Cited by 57PDFcodeScholar