← Search

Xudong Tan

4 accepted papers

2026

FRISM: Fine-Grained Reasoning Injection via Subspace-Level Model Merging for Vision–Language Models

ICML 2026poster

Efficiently enhancing the reasoning capabilities of Vision-Language Models (VLMs) by merging them with Large Reasoning Models (LRMs) has emerged as a promising direction. However, existing methods typically operate at a coarse-grained layer level, which often leads to a trade-off between injecting r…

Cited by 0SourceScholar
2026

Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach

CVPR 2026

Multimodal large language models suffer from substantial inference overhead since multimodal KV Cache grows proportionally with the visual input length. Existing multimodal KV cache compression methods mostly rely on attention score to reduce cache size, which makes them are incompatible with establ

Cited by 0SourceScholar
2025

Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning

CVPR 2025poster

Model quantization reduces the bit-width of weights and activations, improving memory efficiency and inference speed in diffusion models. However, achieving 4-bit quantization remains challenging. Existing methods, primarily based on integer quantization and post-training quantization fine-tuning, s…

Cited by 0SourcePDFScholar
2023

Unobtrusive Respiratory Monitoring System for Intensive Care

ICASSP 2023accepted

The video-based non-contact respiration detection technology can be used in many application scenarios to unobtrusively and ubiquitously monitor the physical state of living beings, and various researchers are currently working on this technology. The optical flow method in tandem with crossover poi…

Cited by 0SourceScholar