← Search

Zhiwen Shao

9 accepted papers

2026

DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal Modeling

AAAI 2026technical

Video shadow detection confronts two entwined difficulties: distinguishing shadows from complex backgrounds and modeling dynamic shadow deformations under varying illumination. To address shadow-background ambiguity, we leverage linguistic priors through the proposed Vision-language Match Module (VM

Cited by 0SourcePDFScholar
2025

Fast and Slow Streams for Online Time Series Forecasting Without Information Leakage

ICLR 2025poster

Current research in online time series forecasting (OTSF) faces two significant issues. The first is information leakage, where models make predictions and are then evaluated on historical time steps that have already been used in backpropagation for parameter updates. The second is practicality: wh…

2025

G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o

AAAI 2025technical

Evaluation metric of visual captioning is important yet not thoroughly explored. Traditional metrics like BLEU, METEOR, CIDEr, and ROUGE often miss semantic depth, while trained metrics such as CLIP-Score, PAC-S, and Polos are limited in zero-shot scenarios. Advanced Language Model-based metrics als…

2025

Modality-Guided Dynamic Graph Fusion and Temporal Diffusion for Self-Supervised RGB-T Tracking

IJCAI 2025

To reduce the reliance on large-scale annotations, self-supervised RGB-T tracking approaches have garnered significant attention. However, the omission of the object region by erroneous pseudo-label or the introduction of background noise affects the efficiency of modality fusion, while pseudo-label

2024

LRANet: Towards Accurate and Efficient Scene Text Detection with Low-Rank Approximation Network

AAAI 2024technical

Recently, regression-based methods, which predict parameterized text shapes for text localization, have gained popularity in scene text detection. However, the existing parameterized text shape methods still have limitations in modeling arbitrary-shaped texts due to ignoring the utilization of text-…

2023

IterativePFN: True Iterative Point Cloud Filtering

CVPR 2023poster

The quality of point clouds is often limited by noise introduced during their capture process. Consequently, a fundamental 3D vision task is the removal of noise, known as point cloud filtering or denoising. State-of-the-art learning based methods focus on training neural networks to infer filtered…

2022

Show, Deconfound and Tell: Image Captioning With Causal Inference

CVPR 2022poster

The transformer-based encoder-decoder framework has shown remarkable performance in image captioning. However, most transformer-based captioning methods ever overlook two kinds of elusive confounders: the visual confounder and the linguistic confounder, which generally lead to harmful bias, induce t…

Cited by 67PDFcodeScholar
2018

Deep Adaptive Attention for Joint Facial Action Unit Detection and Face Alignment

ECCV 2018poster

Facial action unit (AU) detection and face alignment are two highly correlated tasks since facial landmarks can provide precise AU locations to facilitate the extraction of meaningful local features for AU detection. Most existing AU detection works often treat face alignment as a preprocessing and…

Cited by 223SourcePDFScholar
2016

Face alignment by deep convolutional network with adaptive learning rate

ICASSP 2016accepted

Deep convolutional network has been widely used in face recognition while not often used in face alignment. One of the most important reasons of this is the lack of training images annotated with landmarks due to fussy and time-consuming annotation work. To overcome this problem, we propose a novel…

Cited by 0SourceScholar