← Search

Ke Lin

7 accepted papers

2026

Flexible Trajectory Planning for Autonomous Vehicles Via Environmental Assessment in Extreme Scenarios

ICRA 2026poster

Trajectory planning is a core task in autonomous driving. However, in diverse extreme scenarios characterized by unstructured obstacles, there is a lack of solutions that provide efficient computation, safety, and scene generalization capabilities. To address this issue, we propose a two-stage spati…

Cited by 0Scholar
2025

Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

CVPR 2025poster

With the continuous progress of visual generation technologies, the scale of video datasets has grown exponentially. The quality of these datasets plays a pivotal role in the performance of video generation models. We assert that temporal splitting, detailed captions, and video quality filtering are…

2025

Lost in Overlap: Exploring Logit-based Watermark Collision in LLMs

NAACL 2025findings

The proliferation of large language models (LLMs) in generating content raises concerns about text copyright. Watermarking methods, particularly logit-based approaches, embed imperceptible identifiers into text to address these challenges. However, the widespread usage of watermarking across diverse…

Cited by 0SourcePDFScholar
2025

SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

EMNLP 2025

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual inputs, yet modality alignment remains one of the most challenging aspects. Current MLLMs typically rely on simple adapter architectures and pretraining approaches to bridge vision en

2025

Towards Precise Scaling Laws for Video Diffusion Transformers

CVPR 2025poster

Achieving optimal performance of video diffusion transformers within given data and compute budget is crucial due to their high training costs. This necessitates precisely determining the optimal model size and training hyperparameters before large-scale training. While scaling laws are employed in…

Cited by 3SourcePDFScholar
2021

Augmented Partial Mutual Learning with Frame Masking for Video Captioning

AAAI 2021technical

Recent video captioning work improves greatly due to the invention of various elaborate model architectures. If multiple captioning models are combined into a unified framework not only by simple more ensemble, and each model can benefit from each other, the final captioning might be boosted further…