← Search

Yuxiang Yan

7 accepted papers

2026

GIFT: Global Irreplaceability Frame Targeting for Efficient Video Understanding

CVPR 2026

Video Large Language Models (VLMs) have achieved remarkable success in video understanding, but the significant computational cost from processing dense frames severely limits their practical application. Existing methods alleviate this by selecting keyframes, but their greedy decision-making, combi

Cited by 0SourceScholar
2026

MacTok: Robust Continuous Tokenization for Image Generation

CVPR 2026

Continuous image tokenizers enable efficient visual generation, and those based on variational frameworks can learn smooth, structured latent representations through KL regularization. Yet this often leads to posterior collapse when using fewer tokens, where the encoder fails to encode informative f

Cited by 0SourceScholar
2026

Talk2Code: A Multi-Turn Interaction Benchmark with Dual-Track Evaluation for Code Generation

AAAI 2026technical

While large language models (LLMs) have demonstrated strong capabilities in code generation, current benchmarks primarily focus on single-turn scenarios, neglecting the complexity of multi-turn interactions and user diversity. To address this gap, we introduce Talk2Code, the first benchmark for user

Cited by 0SourcePDFScholar
2025

Dark-ISP: Enhancing RAW Image Processing for Low-Light Object Detection

ICCV 2025poster

Low-light Object detection is crucial for many real-world applications but remains challenging due to degraded image quality. While recent studies have shown that RAW images offer superior potential over RGB images, existing approaches either use RAW-RGB images with information loss or employ comple…

Cited by 0SourcePDFScholar
2025

Learning Spatial-Aware Manipulation Ordering

NeurIPS 2025poster

Manipulation in cluttered environments is challenging due to spatial dependencies among objects, where an improper manipulation order can cause collisions or blocked access. Existing approaches often overlook these spatial relationships, limiting their flexibility and scalability. To address these l…

Cited by 0SourceScholar
2024

PointSSC: A Cooperative Vehicle-Infrastructure Point Cloud Benchmark for Semantic Scene Completion

ICRA 2024poster

Semantic Scene Completion (SSC) aims to jointly generate space occupancies and semantic labels for complex 3D scenes. Most existing SSC models focus on volumetric representations, which are memory-inefficient for large outdoor spaces. Point clouds provide a lightweight alternative but existing bench…

Cited by 5SourcecodeScholar