← Search

Wenmin Wang

13 accepted papers

2026

MCGS: Markov Chain Gaussian Splatting for Dynamic Scenes Reconstruction

AAAI 2026technical

We present MCGS (Markov Chain Gaussian Splatting), a novel approach for high-fidelity dynamic scene reconstruction via combining Markov chain and 3D Gaussian splatting. Our method addresses the critical challenge of artifact-free temporal consistency in dynamic neural rendering. By integrating a Mar

Cited by 0SourcePDFScholar
2025

QSpell 250K: A Large-Scale, Practical Dataset for Chinese Search Query Spell Correction

NAACL 2025industry

Chinese Search Query Spell Correction is a task designed to autonomously identify and correct typographical errors within queries in the search engine. Despite the availability of comprehensive datasets like Microsoft Speller and Webis, their monolingual nature and limited scope pose significant cha…

2024

Local Information Guided Global Integration for Infrared Small Target Detection

ICASSP 2024accepted

Infrared small targets often exhibit small scale and weak semantic features, which makes it a great challenge to their detection. To address this situation, we propose a novel network for infrared small target detection that combines local details information and global contextual information. To pr…

Cited by 0SourceScholar
2023

Shadow Removal of Text Document Images Using Background Estimation and Adaptive Text Enhancement

ICASSP 2023accepted

This paper proposes a simple yet effective method to re-move shadows from text document images. It mainly includes several parts. Firstly, we propose a text elimination-based background extraction strategy to estimate shadow map. It indicates the shadow regions accurately and helps to predict global…

Cited by 0SourceScholar
2020

Exploring Entity-Level Spatial Relationships for Image-Text Matching

ICASSP 2020accepted

Exploring the entity-level (i.e., objects in an image, words in a text) spatial relationship contributes to understanding multimedia content precisely. The ignorance of spatial information in previous works probably leads to misunderstandings of image contents. For instance, sentences `Boats are on…

Cited by 0SourceScholar
2019

Adaptively Aligned Image Captioning via Adaptive Attention Time

NeurIPS 2019poster

Recent neural models for image captioning usually employ an encoder-decoder framework with an attention mechanism. However, the attention mechanism in such a framework aligns one single (attended) image feature vector to one caption word, assuming one-to-one mapping from source image regions and tar…

2019

Multi-step Self-attention Network for Cross-modal Retrieval Based on a Limited Text Space

ICASSP 2019accepted

Cross-modal retrieval has been recently proposed to find an appropriate subspace where the similarity among different modalities, such as image and text, can be directly measured. In this paper, we propose Multi-step Self-Attention Network (MSAN) to perform cross-modal retrieval in a limited text sp…

Cited by 0SourceScholar
2017

Cross-modality matching based on Fisher Vector with neural word embeddings and deep image features

ICASSP 2017accepted

Cross-modal retrieval, which aims to solve the problem that the query and the retrieved results are from different modality, becomes more and more essential with the development of the Internet. In this paper, we mainly focus on the exploration of high-level semantic representation of image and text…

Cited by 0SourceScholar
2016

Deep Alternative Neural Network: Exploring Contexts as Early as Possible for Action Recognition

NeurIPS 2016poster

Contexts are crucial for action recognition in video. Current methods often mine contexts after extracting hierarchical local features and focus on their high-order encodings. This paper instead explores contexts as early as possible and leverages their evolutions for action recognition. In particul…

Cited by 27SourcePDFScholar