← Search

Donglai Wei

19 accepted papers

2025

See the World, Discover Knowledge: A Chinese Factuality Evaluation for Large Vision Language Models

ACL 2025finding

The evaluation of factual accuracy in large vision language models (LVLMs) has lagged behind their rapid development, making it challenging to fully reflect these models’ knowledge capacity and reliability. In this paper, we introduce the first factuality-based visual question-answering benchmark in…

Cited by 0SourcePDFScholar
2024

LLCP: Learning Latent Causal Processes for Reasoning-based Video Question Answer

ICLR 2024poster

Current approaches to Video Question Answering (VideoQA) primarily focus on cross-modality matching, which is limited by the requirement for extensive data annotations and the insufficient capacity for causal reasoning (e.g. attributing accidents). To address these challenges, we introduce a causal…

Cited by 2SourcePDFScholar
2024

Learning Causal Domain-Invariant Temporal Dynamics for Few-Shot Action Recognition

ICML 2024poster

Few-shot action recognition aims at quickly adapting a pre-trained model to the novel data with a distribution shift using only a limited number of samples. Key challenges include how to identify and leverage the transferable knowledge learned by the pre-trained model. We therefore propose CDTD, or…

Cited by 2SourcePDFScholar
2024

R^2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding

ECCV 2024poster

"Video temporal grounding (VTG) is a fine-grained video understanding problem that aims to ground relevant clips in untrimmed videos given natural language queries. Most existing VTG models are built upon frame-wise final-layer CLIP features, aided by additional temporal backbones (, SlowFast) with…

2024

SocialGPT: Prompting LLMs for Social Relation Reasoning via Greedy Segment Optimization

NeurIPS 2024poster

Social relation reasoning aims to identify relation categories such as friends, spouses, and colleagues from images. While current methods adopt the paradigm of training a dedicated network end-to-end using labeled image data, they are limited in terms of generalizability and interpretability. To ad…

2023

CLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine Translation

ICCV 2023poster

There has been a growing interest in developing multimodal machine translation (MMT) systems that enhance neural machine translation (NMT) with visual knowledge. This problem setup involves using images as auxiliary information during training, and more recently, eliminating their use during inferen…

Cited by 11PDFcodeScholar
2023

QuantArt: Quantizing Image Style Transfer Towards High Visual Fidelity

CVPR 2023poster

The mechanism of existing style transfer algorithms is by minimizing a hybrid loss function to push the generated image toward high similarities in both content and style. However, this type of approach cannot guarantee visual fidelity, i.e., the generated artworks should be indistinguishable from r…

2023

Why Is the Winner the Best?

CVPR 2023poster

International benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and…

Cited by 29SourcePDFScholar
2022

Learning Task-Specific Representation for Video Anomaly Detection with Spatial-Temporal Attention

ICASSP 2022accepted

The automatic detection of abnormal events in surveillance videos with weak supervision has been formulated as a multiple instance learning task, which aims to localize the clips containing abnormal events temporally with the video-level labels. However, most existing methods rely on the features ex…

Cited by 0SourceScholar
2022

Look, Listen and Pay More Attention: Fusing Multi-Modal Information for Video Violence Detection

ICASSP 2022accepted

Violence detection is an essential and challenging problem in the computer vision community. Most existing works focus on single modal data analysis, which is not effective when multi-modality is available. Therefore, we propose a two-stage multi-modal information fusion method for violence detectio…

Cited by 0SourceScholar
2022

Texture-Based Error Analysis for Image Super-Resolution

CVPR 2022poster

Evaluation practices for image super-resolution (SR) use a single-value metric, the PSNR or SSIM, to determine model performance. This provides little insight into the source of errors and model behavior. Therefore, it is beneficial to move beyond the conventional approach and reconceptualize evalua…

Cited by 19PDFScholar
2022

YouMVOS: An Actor-Centric Multi-Shot Video Object Segmentation Dataset

CVPR 2022poster

Many video understanding tasks require analyzing multi-shot videos, but existing datasets for video object segmentation (VOS) only consider single-shot videos. To address this challenge, we collected a new dataset---YouMVOS---of 200 popular YouTube videos spanning ten genres, where each video is on…

Cited by 2PDFcodeScholar
2021

Context Reasoning Attention Network for Image Super-Resolution

ICCV 2021poster

Deep convolutional neural networks (CNNs) are achieving great successes for image super-resolution (SR), where global context is crucial for accurate restoration. However, the basic convolutional layer in CNNs is designed to extract local patterns, lacking the ability to model global context. Many e…

Cited by 89PDFScholar
2021

Dynamic High-Pass Filtering and Multi-Spectral Attention for Image Super-Resolution

ICCV 2021poster

Deep convolutional neural networks (CNNs) have pushed forward the frontier of super-resolution (SR) research. However, current CNN models exhibit a major flaw: they are biased towards learning low-frequency signals. This bias becomes more problematic for the image SR task which targets reconstructin…

Cited by 109PDFScholar
2021

Learning to Generate Realistic Noisy Images via Pixel-level Noise-aware Adversarial Training

NeurIPS 2021poster

Existing deep learning real denoising methods require a large amount of noisy-clean image pairs for supervision. Nonetheless, capturing a real noisy-clean dataset is an unacceptable expensive and cumbersome procedure. To alleviate this problem, this work investigates how to generate realistic noisy…

Cited by 79SourcePDFScholar
2020

Super-BPD: Super Boundary-to-Pixel Direction for Fast Image Segmentation

CVPR 2020poster

Image segmentation is a fundamental vision task and still remains a crucial step for many applications. In this paper, we propose a fast image segmentation method based on a novel super boundary-to-pixel direction (super-BPD) and a customized segmentation algorithm with super-BPD. Precisely, we defi…

Cited by 31PDFcodeScholar
2020

Two Stream Active Query Suggestion for Active Learning in Connectomics

ECCV 2020poster

For large-scale vision tasks in biomedical images, the labeled data is often limited to train effective deep models. Active learning is a common solution, where a query suggestion method selects representative unlabeled samples for annotation, and the new labels are used to improve the base model. H…

2019

Biologically-Constrained Graphs for Global Connectomics Reconstruction

CVPR 2019poster

Most current state-of-the-art connectome reconstruction pipelines have two major steps: initial pixel-based segmentation with affinity prediction and watershed transform, and refined segmentation by merging over-segmented regions. These methods rely only on local context and are typically agnostic t…

Cited by 28PDFScholar