← Search

Shihao Zou

10 accepted papers

2026

Appearance Discrepancy-guided Sequence Hybrid Masking for Robust Scene Text Recognition

AAAI 2026technical

Masked Image Modeling (MIM) has been widely recognized as a powerful self-supervised paradigm for learning general-purpose visual representations. However, standard MIM based on random masking tends to underperform in domain-specific tasks like Scene Text Recognition (STR), due to challenges such as

Cited by 0SourcePDFScholar
2026

TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition

CVPR 2026

Understanding complex surgical scenes requires recognizing multiple interdependent entities--such as instruments, actions, and targets--and maintaining their relational consistency across time. Existing surgical triplet recognition methods struggle to jointly model intra-frame label dependencies and

Cited by 0SourceScholar
2025

Modal Feature Optimization Network with Prompt for Multimodal Sentiment Analysis

COLING 2025main

Multimodal sentiment analysis(MSA) is mostly used to understand human emotional states through multimodal. However, due to the fact that the effective information carried by multimodal is not balanced, the modality containing less effective information cannot fully play the complementary role betwee…

2025

Multi-Condition Guided Diffusion Network for Multimodal Emotion Recognition in Conversation

NAACL 2025findings

Emotion recognition in conversation (ERC) involves identifying emotional labels associated with utterances within a conversation, a task that is essential for developing empathetic robots. Current research emphasizes contextual factors, the speaker’s influence, and extracting complementary informati…

Cited by 0SourcePDFScholar
2025

SpikeVideoFormer: An Efficient Spike-Driven Video Transformer with Hamming Attention and $\mathcal{O}(T)$ Complexity

ICML 2025poster

Spiking Neural Networks (SNNs) have shown competitive performance to Artificial Neural Networks (ANNs) in various vision tasks, while offering superior energy efficiency. However, existing SNN-based Transformers primarily focus on single-image tasks, emphasizing spatial features while not effectivel…

2024

Tri-Modal Motion Retrieval by Learning a Joint Embedding Space

CVPR 2024highlight

Text-to-motion tasks have been the focus of recent advancements in the human motion domain. However the performance of text-to-motion tasks have not reached its potential primarily due to the lack of motion datasets and the pronounced gap between the text and motion modalities. To mitigate this chal…

Cited by 5SourcePDFScholar
2022

Generating Diverse and Natural 3D Human Motions From Text

CVPR 2022poster

Automated generation of 3D human motions from text is a challenging problem. The generated motions are expected to be sufficiently diverse to explore the text-grounded motion space, and more importantly, accurately depicting the content in prescribed text descriptions. Here we tackle this problem wi…

Cited by 615PDFcodeScholar
2021

EventHPE: Event-Based 3D Human Pose and Shape Estimation

ICCV 2021poster

Event camera is an emerging imaging sensor for capturing dynamics of moving objects as events, which motivates our work in estimating 3D human pose and shape from the event signals. Events, on the other hand, have their unique challenges: rather than capturing static body postures, the event signals…

Cited by 58PDFcodeScholar
2020

3D Human Shape Reconstruction from a Polarization Image

ECCV 2020poster

This paper tackles the problem of estimating 3D body shape of clothed humans from single polarized 2D images, i.e. polarization images. Polarization images are known to be able to capture polarized reflected lights that preserve rich geometric cues of an object, which has motivated its recent applic…

Cited by 56SourcePDFScholar