← Search

Di Lin

37 accepted papers

2026

BA-GS: Bayesian Adaptive Gaussian Splatting for SFM-Free 3D Reconstruction

CVPR 2026

3D Gaussian Splatting (3DGS) has demonstrated exceptional performance in reconstruction and novel view synthesis tasks. However, its reliance on Structure-from-Motion preprocessing may lead to degraded performance under sparse-view scenarios. Recent works attempt to address this limitation by levera

Cited by 0SourceScholar
2026

CorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous Driving

AAAI 2026technical

End-to-end planning methods are the de-facto standard of the current autonomous driving system, while the robustness of the data-driven approaches suffers due to the notorious long-tail problem (i.e., rare but safety-critical failure cases). In this work, we explore whether recent diffusion-based vi

Cited by 0SourcePDFScholar
2026

DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning

CVPR 2026

Although multi-modal large language models (MLLMs) have shown strong capabilities across diverse domains, their application in generating fine-grained 3D perception and prediction outputs in autonomous driving remains underexplored. In this paper, we propose DrivePI, a novel spatial-aware 4D MLLM th

Cited by 0SourcecodeScholar
2026

Exploiting Geometric Structures for Modeling Multi-Agent Behaviors: A New Thinking

AAAI 2026technical

In this paper, we rethink model agent behaviors from a geometric structure perspective in multi-agent reinforcement learning. Modeling agent behaviors is essential for understanding how agents interact and facilitating effective decisions. The key lies in capturing the dependencies and sequential re

Cited by 0SourcePDFScholar
2026

Firing Bits Where It Matters: Spiking-Guided Just Recognizable Distortion Modeling for Machine-Centric Video Coding

AAAI 2026technical

Just recognizable distortion (JRD) has emerged as a promising paradigm for machine-centric video coding. However, existing JRD-guided coding methods are limited by coarse annotation granularity and high computational cost, which hinder their deployment. In this paper, we first investigate the impact

Cited by 0SourcePDFScholar
2026

FreeMem: Enhancing Consistency in Long Video Generation via Tuning-Free Memory

AAAI 2026technical

Text-to-Video (T2V) generation has advanced greatly, yet maintaining consistency remains challenging, especially for tuning-free long video generation. We attribute the consistency problem to cumulative deviations for long video generation at three levels: the random noise lacking correlation resu

Cited by 0SourcePDFScholar
2026

HOPS: Hierarchical Open-vocabulary Part Segmentation with Attention-Aware Filtering and Affinity-Guided Enhancement

CVPR 2026

Open-vocabulary part segmentation (OVPS) aims to segment objects into fine-grained parts while generalizing to unseen categories. Existing VLM-based methods face two challenges: (1) object over-segmentation, caused by overly broad semantic activations, and (2) part under-segmentation, resulting from

Cited by 0SourcecodeScholar
2026

Multi-modal Frequency Decomposition Network for Semantic Scene Completion

CVPR 2026

Based on an RGB-D image pair, semantic scene completion (SSC) provides a description for 3D scene understanding by predicting 3D semantic occupancy map. Recent methods extract RGB-D multi-modal features and fuse them in spatial domain, which disregards the misalignment caused by the imperfect raw mu

Cited by 0SourceScholar
2026

OccDriver: Future Occupancy Guided Dual-branch Trajectory Planner in Autonomous Driving

ICLR 2026poster

Trajectory planning for autonomous driving is challenging due to agents' behavioral uncertainty and intricate multi-agent interaction modeling. Most existing studies generate trajectories without explicitly exploiting possible scene evolution, while world models predict consequences from ego behavio…

Cited by 0SourceScholar
2026

RecEdit-Drive: 3D Reconstruction-Guided Spatiotemporal Video Editing for Autonomous Driving Scenes

CVPR 2026

High-quality video editing and processing are crucial in domains such as filmmaking and autonomous driving, where accurate visual refinement and data preparation are essential. However, it is challenging to achieve precise control over dynamic objects while maintaining spatiotemporal consistency. Cu

Cited by 0SourcecodeScholar
2026

Video Generation with Stable Transparency via Shiftable RGB-A Distribution Learner

CVPR 2026

Generating RGB-A videos, which include alpha channels for transparency, has wide applications. However, current methods often suffer from low quality due to confusion between RGB and alpha. In this paper, we address this problem by learning shiftable RGB-A distributions. We adjust both the latent sp

Cited by 0SourcecodeScholar
2025

CVLN-Think: Causal Inference with Counterfactual Style Adaptation for Continuous Vision-and-Language Navigation

IROS 2025

Vision-and-Language Navigation in Continuous Environments (VLN-CE) presents challenges due to environmental variations and domain shifts, making it difficult for agents to generalize beyond seen environments. Most existing methods rely on learning correlations between observations and actions from t

Cited by 0SourceScholar
2025

Generative Hard Example Augmentation for Semantic Point Cloud Segmentation

CVPR 2025poster

The recent progress in semantic point cloud segmentation is attributed to deep networks, which require a large amount of point cloud data for training. However, how to collect substantial point-wise annotations of the point clouds at affordable cost for the end-to-end network training still needs to…

Cited by 0SourcePDFScholar
2025

NoiseController: Towards Consistent Multi-view Video Generation via Noise Decomposition and Collaboration

ICCV 2025poster

High-quality video generation is crucial for many fields, including the film industry and autonomous driving. However, generating videos with spatiotemporal consistencies remains challenging. Current methods typically utilize attention mechanisms or modify noise to achieve consistent videos, neglect…

2025

Open-Vocabulary Part Segmentation via Progressive and Boundary-Aware Strategy

NeurIPS 2025poster

Open-vocabulary part segmentation (OVPS) struggles with structurally connected boundaries due to the inherent conflict between continuous image features and discrete classification mechanism. To address this, we propose PBAPS, a novel training-free framework specifically designed for OVPS. PBAPS lev…

Cited by 0SourcecodeScholar
2025

PhysGCN-DL: Physics-Informed Graph Convolutional Networks with Diversity-Aware Loss Optimization for Multimodal Pedestrian Trajectory Prediction

IROS 2025

Pedestrian trajectory prediction ensures safe navigation in autonomous driving and intelligent robots. Existing methods have shown promising results but still face challenges in handling dynamic environments, social interactions, and high-dimensional data. In this paper, we propose a novel PhysGCN-D

Cited by 0SourceScholar
2025

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

CVPR 2025poster

Large vision-language models (LVLMs) have shown remarkable capabilities in interpreting visual content. While existing works demonstrate these models' vulnerability to deliberately placed adversarial texts, such texts are often easily identifiable as anomalous. In this paper, we present the first ap…

2025

Trajectory-LLM: A Language-based Data Generator for Trajectory Prediction in Autonomous Driving

ICLR 2025poster

Vehicle trajectory prediction is a crucial aspect of autonomous driving, which requires extensive trajectory data to train prediction models to understand the complex, varied, and unpredictable patterns of vehicular interactions. However, acquiring real-world data is expensive, so we advocate using…

2025

mmFAS: Multimodal Face Anti-Spoofing Using Multi-Level Alignment and Switch-Attention Fusion

AAAI 2025technical

The increasing number of presentation attacks on reliable face matching has raised concerns and garnered attention towards face anti-spoofing (FAS). However, existing methods for FAS modeling commonly fuse multiple visual modalities (e.g., RGB, Depth, and Infrared) in a straightforward manner, disre…

Cited by 0SourcePDFScholar
2024

LRR: Language-Driven Resamplable Continuous Representation against Adversarial Tracking Attacks

ICLR 2024poster

Visual object tracking plays a critical role in visual-based autonomous systems, as it aims to estimate the position and size of the object of interest within a live video. Despite significant progress made in this field, state-of-the-art (SOTA) trackers often fail when faced with adversarial pertur…

2024

Sim2Real-Fire: A Multi-modal Simulation Dataset for Forecast and Backtracking of Real-world Forest Fire

NeurIPS 2024poster

The latest research on wildfire forecast and backtracking has adopted AI models, which require a large amount of data from wildfire scenarios to capture fire spread patterns. This paper explores using cost-effective simulated wildfire scenarios to train AI models and apply them to the analysis of re…

Cited by 1SourcePDFScholar
2024

Voxel Proposal Network via Multi-Frame Knowledge Distillation for Semantic Scene Completion

NeurIPS 2024poster

Semantic scene completion is a difficult task that involves completing the geometry and semantics of a scene from point clouds in a large-scale environment. Many current methods use 3D/2D convolutions or attention mechanisms, but these have limitations in directly constructing geometry and accuratel…

Cited by 1SourcePDFScholar
2024

msLPCC: A Multimodal-Driven Scalable Framework for Deep LiDAR Point Cloud Compression

AAAI 2024technical

LiDAR sensors are widely used in autonomous driving, and the growing storage and transmission demands have made LiDAR point cloud compression (LPCC) a hot research topic. To address the challenges posed by the large-scale and uneven-distribution (spatial and categorical) of LiDAR point data, this pa…

Cited by 4SourcePDFScholar
2023

CVSformer: Cross-View Synthesis Transformer for Semantic Scene Completion

ICCV 2023poster

Semantic scene completion (SSC) requires an accurate understanding of the geometric and semantic relationships between the objects in the 3D scene for reasoning the occluded objects. The popular SSC methods voxelize the 3D objects, allowing the deep 3D convolutional network (3D CNN) to learn the obj…

Cited by 9PDFcodeScholar
2023

Leveraging Inpainting for Single-Image Shadow Removal

ICCV 2023poster

Fully-supervised shadow removal methods achieve the best restoration qualities on public datasets but still generate some shadow remnants. One of the reasons is the lack of large-scale shadow & shadow-free image pairs. Unsupervised methods can alleviate the issue but their restoration qualities are…

Cited by 28PDFcodeScholar
2023

Open Compound Domain Adaptation with Object Style Compensation for Semantic Segmentation

NeurIPS 2023poster

Many methods of semantic image segmentation have borrowed the success of open compound domain adaptation. They minimize the style gap between the images of source and target domains, more easily predicting the accurate pseudo annotations for target domain's images that train segmentation network. Th…

Cited by 8SourcePDFScholar
2022

Generative Status Estimation and Information Decoupling for Image Rain Removal

NeurIPS 2022accept

Image rain removal requires the accurate separation between the pixels of the rain streaks and object textures. But the confusing appearances of rains and objects lead to the misunderstanding of pixels, thus remaining the rain streaks or missing the object details in the result. In this paper, we pr…

Cited by 9SourcePDFScholar
2022

MISF: Multi-Level Interactive Siamese Filtering for High-Fidelity Image Inpainting

CVPR 2022poster

Although achieving significant progress, existing deep generative inpainting methods still show low generalization across different scenes. As a result, the generated images usually contain artifacts or the filled pixels differ greatly from the ground truth, making them far from real-world applicati…

Cited by 111PDFcodeScholar
2020

RANet: Region Attention Network for Semantic Segmentation

NeurIPS 2020poster

Recent semantic segmentation methods model the relationship between pixels to construct the contextual representations. In this paper, we introduce the \emph{Region Attention Network} (RANet), a novel attention network for modeling the relationship between object regions. RANet divides the image int…

2019

ZigZagNet: Fusing Top-Down and Bottom-Up Context for Object Segmentation

CVPR 2019poster

Multi-scale context information has proven to be essential for object segmentation tasks. Recent works construct the multi-scale context by aggregating convolutional feature maps extracted by different levels of a deep neural network. This is typically done by propagating and fusing features in a on…

Cited by 86PDFcodeScholar
2018

Multi-Scale Context Intertwining for Semantic Segmentation

ECCV 2018poster

Accurate semantic image segmentation requires the joint consideration of local appearance, semantic information, and global scene context. In today’s age of pre-trained deep networks and their powerful convolutional features, state-of-the-art semantic segmentation approaches differ mostly in how the…

Cited by 211SourcePDFScholar
2017

Cascaded Feature Network for Semantic Segmentation of RGB-D Images

ICCV 2017poster

Fully convolutional network (FCN) has been successfully applied in semantic segmentation of scenes represented with RGB images. Images augmented with depth channel provide more understanding of the geometric information of the scene in the image. The question is how to best exploit this additional i…

Cited by 175PDFScholar
2017

Learning to Aggregate Ordinal Labels by Maximizing Separating Width

ICML 2017poster

While crowdsourcing has been a cost and time efficient method to label massive samples, one critical issue is quality control, for which the key challenge is to infer the ground truth from noisy or even adversarial data by various users. A large class of crowdsourcing problems, such as those involvi…

Cited by 9SourcePDFScholar
2016

ScribbleSup: Scribble-Supervised Convolutional Networks for Semantic Segmentation

CVPR 2016oral

Large-scale data are of crucial importance for learning semantic segmentation models, but annotating per-pixel masks is a tedious and inefficient procedure. We note that for the topic of interactive image segmentation, scribbles are very widely used in academic research and commercial software, and…

Cited by 1333PDFScholar
2015

Deep LAC: Deep Localization, Alignment and Classification for Fine-Grained Recognition

CVPR 2015poster

We propose a fine-grained recognition system that incorporates part localization, alignment, and classification in one deep neural network. This is a nontrivial process, as the input to the classification module should be functions that enable back-propagation in constructing the solver. Our major c…

Cited by 439SourcePDFScholar