← Search

Wei Ke

30 accepted papers

2026

CDMIQA: A Cross-Domain Perceptual Method and Benchmark Dataset for Medical Image Quality Assessment

IJCAI 2026

Medical image quality assessment (IQA) serves as a critical safeguard for precise clinical diagnosis and treatment. However, existing methods still face challenges arising from data scarcity and heterogeneity across imaging domains, which confine solutions to domain-specific designs and limit their

Cited by 0Scholar
2026

EfficientFlow: Efficient Equivariant Flow Policy Learning for Embodied AI

AAAI 2026technical

Generative modeling has recently shown remarkable promise for visuomotor policy learning, enabling flexible and expressive control across diverse embodied AI tasks. However, existing generative policies often struggle with data inefficiency, requiring large-scale demonstrations, and sampling ineffic

Cited by 0SourcePDFScholar
2026

FlexiVideo: Variation-Aware Temporal Dynamics Modeling for Efficient Video Understanding

CVPR 2026

Natural videos exhibit heterogeneous temporal dynamics, with certain segments undergoing high-dynamic scene transitions and others dominated by low-dynamic visual changes. However, treating all frames identically, a common practice in most MLLMs, leads to redundant visual encoding, which results in

Cited by 0SourcecodeScholar
2026

UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register

CVPR 2026

Representation learning with Vision Transformers (ViTs) has advanced rapidly, yet the utility of large-scale models in spatially sensitive tasks is hindered by spurious tokens. Prior efforts to mitigate this have been limited, often defining these artifacts narrowly, for example, as simple high-norm

Cited by 0SourceScholar
2026

Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable AI-Generated Image Detection

ICLR 2026poster

Current AI-Generated Image (AIGI) detection approaches predominantly rely on binary classification to distinguish real from synthetic images, often lacking interpretable or convincing evidence to substantiate their decisions. This limitation stems from existing AIGI detection benchmarks, which, desp…

Cited by 0SourcecodeScholar
2025

Are High-Quality AI-Generated Images More Difficult for Models to Detect?

ICML 2025poster

The remarkable evolution of generative models has enabled the generation of high-quality, visually attractive images, often perceptually indistinguishable from real photographs to human eyes. This has spurred significant attention on AI-generated image (AIGI) detection. Intuitively, higher image qua…

2025

Enhancing Zero-shot Object Counting via Text-guided Local Ranking and Number-evoked Global Attention

ICCV 2025poster

Text-guided zero-shot object counting leverages vision-language models (VLMs) to count objects of an arbitrary class given by a text prompt. Existing approaches for this challenging task only utilize local patch-level features to fuse with text feature, ignoring the important influence of the global…

2025

Generating Multimodal Driving Scenes via Next-Scene Prediction

CVPR 2025poster

Generative models in Autonomous Driving (AD) enable diverse scenario creation, yet existing methods fall short by only capturing a limited range of modalities, restricting the capability of generating controllable scenes for comprehensive evaluation of AD systems. In this paper, we introduce a multi…

2025

Refining CLIP's Spatial Awareness: A Visual-Centric Perspective

ICLR 2025poster

Contrastive Language-Image Pre-training (CLIP) excels in global alignment with language but exhibits limited sensitivity to spatial information, leading to strong performance in zero-shot classification tasks but underperformance in tasks requiring precise spatial understanding. Recent approaches ha…

Cited by 0SourcePDFScholar
2024

A Parameterized Generative Adversarial Network Using Cyclic Projection for Explainable Medical Image Classifications

ICASSP 2024accepted

Although current data augmentation methods are successful to alleviate the data insufficiency, conventional augmentation are primarily intra-domain while advanced generative adversarial networks (GANs) generate images remaining uncertain, particularly in small-scale datasets. In this paper, we propo…

Cited by 0SourceScholar
2024

Mind Your Augmentation: The Key to Decoupling Dense Self-Supervised Learning

ICLR 2024poster

Dense Self-Supervised Learning (SSL) creates positive pairs by building positive paired regions or points, thereby aiming to preserve local features, for example of individual objects. However, existing approaches tend to couple objects by leaking information from the neighboring contextual regions…

Cited by 2SourcePDFScholar
2024

Mitigating Object Dependencies: Improving Point Cloud Self-Supervised Learning through Object Exchange

CVPR 2024poster

In the realm of point cloud scene understanding particularly in indoor scenes objects are arranged following human habits resulting in objects of certain semantics being closely positioned and displaying notable inter-object correlations. This can create a tendency for neural networks to exploit the…

2024

Multi-Attention Enhanced Discriminator for GAN-Based Anomalous Sound Detection

ICASSP 2024accepted

Generative adversarial networks (GAN) have been regarded as promising for anomalous sound detection (ASD) by training an unsupervised one-class classifier to pick out the anomalous sample. Existing GAN-based anomaly detection models usually focus on the generator to reduce the reconstruction error.…

Cited by 0SourceScholar
2023

Spatiotemporal Self-Supervised Learning for Point Clouds in the Wild

CVPR 2023poster

Self-supervised learning (SSL) has the potential to benefit many applications, particularly those where manually annotating data is cumbersome. One such situation is the semantic segmentation of point clouds. In this context, existing methods employ contrastive learning strategies and define positiv…

2022

CoupAlign: Coupling Word-Pixel with Sentence-Mask Alignments for Referring Image Segmentation

NeurIPS 2022accept

Referring image segmentation aims at localizing all pixels of the visual objects described by a natural language sentence. Previous works learn to straightforwardly align the sentence embedding and pixel-level embedding for highlighting the referred objects, but ignore the semantic consistency of pi…

Cited by 33SourcePDFScholar
2022

Leverage Your Local and Global Representations: A New Self-Supervised Learning Strategy

CVPR 2022poster

Self-supervised learning (SSL) methods aim to learn view-invariant representations by maximizing the similarity between the features extracted from different crops of the same image regardless of cropping size and content. In essence, this strategy ignores the fact that two crops may truly contain d…

Cited by 40PDFcodeScholar
2021

Error-Aware Density Isomorphism Reconstruction for Unsupervised Cross-Domain Crowd Counting

AAAI 2021technical

This paper focuses on the unsupervised domain adaptation problem for video-based crowd counting, in which we use labeled data as source domain and unlabelled video data as target domain. It is challenging as there is a huge gap between the source and the target domain and no annotations of samples a…

2021

Kohonen Self-Organizing Map based Route Planning: A Revisit

IROS 2021poster

In this paper, we revisit the long-standing Traveling Salesman Problem (TSP) and focus on the challenging, yet practical route planning problem with limited computational resources. We make contributions to TSP, one of the most famous NP-hard problems by providing a new improved approximate solution…

Cited by 15SourceScholar
2020

Multiple Anchor Learning for Visual Object Detection

CVPR 2020poster

Classification and localization are two pillars of visual object detectors. However, in CNN-based detectors, these two modules are usually optimized under a fixed set of candidate (or anchor) bounding boxes. This configuration significantly limits the possibility to jointly optimize classification a…

Cited by 125PDFcodeScholar
2020

Weakly-Supervised Action Localization with Expectation-Maximization Multi-Instance Learning

ECCV 2020poster

Weakly-supervised action localization requires training a model to localize the action segments in the video given only video level action label. It can be solved under the Multiple Instance Learning (MIL) framework, where a bag (video) contains multiple instances (action segments). Since only the b…

2019

C-MIL: Continuation Multiple Instance Learning for Weakly Supervised Object Detection

CVPR 2019oral

Weakly supervised object detection (WSOD) is a challenging task when provided with image category supervision but required to simultaneously learn object locations and object detectors. Many WSOD approaches adopt multiple instance learning (MIL) and have non-convex loss functions which are prone to…

Cited by 299PDFcodeScholar
2019

Orthogonal Decomposition Network for Pixel-Wise Binary Classification

CVPR 2019poster

The weight sharing scheme and spatial pooling operations in Convolutional Neural Networks (CNNs) introduce semantic correlation to neighboring pixels on feature maps and therefore deteriorate their pixel-wise classification performance. In this paper, we implement an Orthogonal Decomposition Unit (O…

Cited by 10PDFScholar
2017

SRN: Side-output Residual Network for Object Symmetry Detection in the Wild

CVPR 2017oral

In this paper, we establish a baseline for object symmetry detection in complex backgrounds by presenting a new benchmark and an end-to-end deep learning approach, opening up a promising direction for symmetry detection in the wild. The new benchmark, named Sym-PASCAL, spans challenges including obj…

Cited by 120PDFcodeScholar
2017

Self-Learning Scene-Specific Pedestrian Detectors Using a Progressive Latent Model

CVPR 2017poster

In this paper, a self-learning approach is proposed towards solving scene-specific pedestrian detection problem without any human' annotation involved. The self-learning approach is deployed as progressive steps of object discovery, object enforcement, and label propagation. In the learning procedur…

Cited by 41PDFScholar
2015

Pedestrian detection via PCA filters based convolutional channel features

ICASSP 2015accepted

In this paper, we propose a kind of image representation, named PCA filters based convolutional channel features (PCA-CCF) for pedestrian detection. The motivation is to use the convolutional network architecture with orthogonal PCA filters to enhance the state-of-the-art aggregate channel features…

Cited by 0SourceScholar