← Search

Sihang Cai

5 accepted papers

2026

Scene-Aware Spatiotemporal Generalization: Towards Robust Temporal Action Detection Across Domains

AAAI 2026technical

Temporal Action Detection (TAD) aims to identify specific actions in long, untrimmed videos by determining their start, end times and categories, yet existing models suffer from performance degradation under out-of-distribution scenarios due to unrealistic i.i.d. assumptions. While domain generaliza

Cited by 0SourcePDFScholar
2025

AnomalyCoT: A Multi-Scenario Chain-of-Thought Dataset for Multimodal Large Language Models

NeurIPS 2025poster

Industrial Anomaly Detection (IAD) is an indispensable quality control technology in modern production processes. Recently, on account of the outstanding visual comprehension and cross-domain knowledge transfer capabilities of multimodal large language models (MLLMs), existing studies have explored…

Cited by 0SourcecodeScholar
2025

Chat-Driven Text Generation and Interaction for Person Retrieval

EMNLP 2025

Text-based person search (TBPS) enables the retrieval of person images from large-scale databases using natural language descriptions, offering critical value in surveillance applications. However, a major challenge lies in the labor-intensive process of obtaining high-quality textual annotations, w

Cited by 0SourcePDFScholar
2025

Omni-Chart-600K: A Comprehensive Dataset of Chart Types for Chart Understanding

NAACL 2025findings

To address the deficiencies in chart types and the limited scope of chart tasks in existing datasets, we conducted a comprehensive review of current data collection methodologies. By integrating manual annotation with data generation leveraging GPT-4, we developed a dataset that includes 21 diverse…

Cited by 0SourcePDFScholar
2025

Towards Transformer-Based Aligned Generation with Self-Coherence Guidance

CVPR 2025poster

We introduce a novel, training-free approach for enhancing alignment in Transformer-based Text-Guided Diffusion Models (TGDMs). Existing TGDMs often struggle to generate semantically aligned images, particularly when dealing with complex text prompts or multi-concept attribute binding challenges. Pr…