← Search

Weiming Dong

25 accepted papers

2026

HLD: Approximate Hierarchical Linguistic Distribution Modeling for LLM-Generated Text Detection

ICLR 2026poster

The widespread deployment of large language models (LLMs) has made the reliable detection of AI-generated text a crucial task. However, existing zero-shot detectors typically rely on proxy models to approximate probability distributions of unknown source models at a single token level. Such approach…

Cited by 0SourcecodeScholar
2026

Image Guides Images: Consistent Video Amodal Completion with Rectified In-Context Exemplar Guidance

CVPR 2026

Video amodal completion (VAC) aims to mimic the human brain's ability to implicitly perceive the complete appearance of partially occluded objects, thereby facilitating recognition and understanding. Existing VAC methods finetune video generation models on custom datasets, yet these datasets often h

Cited by 0SourcecodeScholar
2026

SongEcho: Towards Cover Song Generation via Instance-Adaptive Element-wise Linear Modulation

ICLR 2026poster

Cover songs constitute a vital aspect of musical culture, preserving the core melody of an original composition while reinterpreting it to infuse novel emotional depth and thematic emphasis. Although prior research has explored the reinterpretation of instrumental music through melody-conditioned te…

Cited by 0SourcecodeScholar
2025

Bridging Class Imbalance and Partial Labeling via Spectral-Balanced Energy Propagation for Skeleton-based Action Recognition

ICCV 2025poster

Skeleton-based action recognition faces class imbalance and insufficient labeling problems in real-world applications. Existing methods typically address these issues separately, lacking a unified framework that can effectively handle both issues simultaneously while considering their inherent relat…

Cited by 0SourcePDFScholar
2025

LiveStar: Live Streaming Assistant for Real-World Online Video Understanding

NeurIPS 2025poster

Despite significant progress in Video Large Language Models (Video-LLMs) for offline video understanding, existing online Video-LLMs typically struggle to simultaneously process continuous frame-by-frame inputs and determine optimal response timing, often compromising real-time responsiveness and na…

Cited by 0SourcecodeScholar
2025

Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation

NeurIPS 2025poster

In recent years, Multimodal Large Language Models (MLLMs) have been extensively utilized for multimodal reasoning tasks, including Graphical User Interface (GUI) automation. Unlike general offline multimodal tasks, GUI automation is executed in online interactive environments, necessitating step-by-…

Cited by 0SourcecodeScholar
2025

SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding

ICLR 2025spotlight

Despite the significant advancements of Large Vision-Language Models (LVLMs) on established benchmarks, there remains a notable gap in suitable evaluation regarding their applicability in the emerging domain of long-context streaming video understanding. Current benchmarks for video understanding ty…

2024

Lighting Image/Video Style Transfer Methods by Iterative Channel Pruning

ICASSP 2024accepted

Deploying style transfer methods on resource-constrained devices is challenging, which limits their real-world applicability. To tackle this issue, we propose using pruning techniques to accelerate various visual style transfer methods. We argue that typical pruning methods may not be well-suited fo…

Cited by 0SourceScholar
2024

Music Style Transfer with Time-Varying Inversion of Diffusion Models

AAAI 2024technical

With the development of diffusion models, text-guided image style transfer has demonstrated great controllable and high-quality results. However, the utilization of text for diverse music style transfer poses significant challenges, primarily due to the limited availability of matched audio-text dat…

2024

Revealing the Two Sides of Data Augmentation: An Asymmetric Distillation-based Win-Win Solution for Open-Set Recognition

IJCAI 2024poster

In this paper, we reveal the two sides of data augmentation: enhancements in closed-set recognition correlate with a significant decrease in open-set recognition. Through empirical investigation, we find that multi-sample-based augmentations would contribute to reducing feature discrimination, there…

Cited by 1SourcePDFScholar
2024

Three Heads Are Better than One: Complementary Experts for Long-Tailed Semi-supervised Learning

AAAI 2024technical

We address the challenging problem of Long-Tailed Semi-Supervised Learning (LTSSL) where labeled data exhibit imbalanced class distribution and unlabeled data follow an unknown distribution. Unlike in balanced SSL, the generated pseudo-labels are skewed towards head classes, intensifying the trainin…

2024

Z*: Zero-shot Style Transfer via Attention Reweighting

CVPR 2024poster

Despite the remarkable progress in image style transfer formulating style in the context of art is inherently subjective and challenging. In contrast to existing methods this study shows that vanilla diffusion models can directly extract style information and seamlessly integrate the generative prio…

2023

Inversion-Based Style Transfer With Diffusion Models

CVPR 2023poster

The artistic style within a painting is the means of expression, which includes not only the painting material, colors, and brushstrokes, but also the high-level attributes, including semantic elements and object shapes. Previous arbitrary example-guided artistic image generation methods often fail…

2022

Evo-ViT: Slow-Fast Token Evolution for Dynamic Vision Transformer

AAAI 2022technical

Vision transformers (ViTs) have recently received explosive popularity, but the huge computational cost is still a severe issue. Since the computation complexity of ViT is quadratic with respect to the input sequence length, a mainstream paradigm for computation reduction is to reduce the number of…

2022

StyTr2: Image Style Transfer With Transformers

CVPR 2022poster

The goal of image style transfer is to render an image with artistic features guided by a style reference while maintaining the original content. Owing to the locality in convolutional neural networks (CNNs), extracting and maintaining the global information of input images is difficult. Therefore,…

Cited by 379PDFcodeScholar
2021

Arbitrary Video Style Transfer via Multi-Channel Correlation

AAAI 2021technical

Video style transfer is attracting increasing attention from the artificial intelligence community because of its numerous applications, such as augmented reality and animation production. Relative to traditional image style transfer, video style transfer presents new challenges, including how to ef…

2021

Unveiling the Potential of Structure Preserving for Weakly Supervised Object Localization

CVPR 2021poster

Weakly supervised object localization (WSOL) remains an open problem due to the deficiency of finding object extent information using a classification network. While prior works struggle to localize objects by various spatial regularization strategies, we argue that how to extract object structural…

Cited by 110PDFcodeScholar
2020

Dynamic Refinement Network for Oriented and Densely Packed Object Detection

CVPR 2020oral

Object detection has achieved remarkable progress in the past decade. However, the detection of oriented and densely packed objects remains challenging because of following inherent reasons: (1) receptive fields of neurons are all axis-aligned and of the same shape, whereas objects are usually of di…

Cited by 411PDFcodeScholar
2019

Joint Representation and Estimator Learning for Facial Action Unit Intensity Estimation

CVPR 2019poster

Facial action unit (AU) intensity is an index to characterize human expressions. Accurate AU intensity estimation depends on three major elements: image representation, intensity estimator, and supervisory information. Most existing methods learn intensity estimator with fixed image representation,…

Cited by 41PDFScholar
2019

LGM-Net: Learning to Generate Matching Networks for Few-Shot Learning

ICML 2019oral

In this work, we propose a novel meta-learning approach for few-shot classification, which learns transferable prior knowledge across tasks and directly produces network parameters for similar unseen tasks with training samples. Our approach, called LGM-Net, includes two key modules, namely, TargetN…

2018

Bilateral Ordinal Relevance Multi-Instance Regression for Facial Action Unit Intensity Estimation

CVPR 2018poster

Automatic intensity estimation of facial action units (AUs) is challenging in two aspects. First, capturing subtle changes of facial appearance is quiet difficult. Second, the annotation of AU intensity is scarce and expensive. Intensity annotation requires strong domain knowledge thus only experts…

Cited by 53SourcePDFScholar
2018

Classifier Learning With Prior Probabilities for Facial Action Unit Recognition

CVPR 2018poster

Facial action units (AUs) play an important role in human emotion understanding. One big challenge for data-driven AU recognition approaches is the lack of enough AU annotations, since AU annotation requires strong domain expertise. To alleviate this issue, we propose a knowledge-driven method for j…

Cited by 63SourcePDFScholar
2018

Weakly-Supervised Deep Convolutional Neural Network Learning for Facial Action Unit Intensity Estimation

CVPR 2018poster

Facial action unit (AU) intensity estimation plays an important role in affective computing and human-computer interaction. Recent works have introduced deep neural networks for AU intensity estimation, but they require a large amount of intensity annotations. AU annotation needs strong domain exper…

Cited by 64SourcePDFScholar
2015

UniHIST: A Unified Framework for Image Restoration With Marginal Histogram Constraints

CVPR 2015poster

Marginal histograms provide valuable information for various computer vision problems. However, current image restoration methods do not fully exploit the potential of marginal histograms, in particular, their role as ensemble constraints on the marginal statistics of the restored image. In this pap…

Cited by 12SourcePDFScholar