← Search

Siyao Cheng

4 accepted papers

2026

Stop Mixing Things Up! BISCUIT Teaches Vision-Language Models to Learn New Concepts from Images on the Spot

AAAI 2026technical

Vision-Language Models (VLMs) have achieved impressive performance across various tasks, but often struggle to apply newly introduced visual concepts during inference. A common failure pattern is what we call Mixing Things Up: VLMs frequently confuse concept names, resulting in vague descriptions an

Cited by 0SourcePDFScholar
2026

Twin-T & TwintVQA: A Reliable Structure-Detail Separating VLM and a Comprehensive Benchmark for Chart and Table Tasks

CVPR 2026

With the rapid development of Vision-Language Models (VLMs), there is a growing demand for automatic analysis of structured visual data. Charts and tables carry quantitative information through regular layouts, explicit numbers, and chart-specific reading patterns, yet current VLMs still underuse th

Cited by 0SourcecodeScholar
2025

Adapting Single-Channel Pre-trained Transformer Models for Multi-Channel Sound Event Localization and Detection

ICASSP 2025accepted

In recent years, the significance of pre-trained transformer audio models has been increasingly recognized. However, existing pre-trained transformer audio models are based on single-channel audio. They cannot be directly applied to multi-channel audio for Sound Event Localization and Detection (SEL…

Cited by 0SourceScholar
2021

Typingwristband: A Human Slight Motion Sensing System Based on Vibration Detection

ICASSP 2021accepted

With the widespread of Human-Cyber-Physical Systems (HCPS), the fine-grained human movement detection becomes more and more important. Especially for the slightly motions of human’s hands, they are not only bring abundant information, but also provide a new way for the interaction between users and…

Cited by 0SourceScholar