← Search

Kaishen Yuan

4 accepted papers

2026

CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation

ICLR 2026poster

Emotional Image Content Generation (EICG) aims to generate semantically clear and emotionally faithful images based on given emotion categories, with broad application prospects. While recent text-to-image diffusion models excel at generating concrete concepts, they struggle with the complexity of a…

Cited by 0SourcecodeScholar
2026

MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis

CVPR 2026

Accurate and interpretable multi-disease diagnosis remains a critical challenge in medical research, particularly when leveraging heterogeneous multimodal medical data. Current approaches often rely on single-modal data, limiting their ability to comprehensively understand complex diseases. To addre

Cited by 0SourcecodeScholar
2025

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model

CVPR 2025poster

Periodic or quasi-periodic phenomena reveal intrinsic characteristics in various natural processes, such as weather patterns, movement behaviors, traffic flows, and biological signals. Given that these phenomena span multiple modalities, the capabilities of Multimodal Large Language Models (MLLMs) o…

2024

AUFormer: Vision Transformers are Parameter-Efficient Facial Action Unit Detectors

ECCV 2024poster

"Facial Action Units (AU) is a vital concept in the realm of affective computing, and AU detection has always been a hot research topic. Existing methods suffer from overfitting issues due to the utilization of a large number of learnable parameters on scarce AU-annotated datasets or heavy reliance…