← Search

Fei Shen

39 accepted papers

2026

Benchmarking Dense and Indiscernible Object Counting with Blueberries

ICML 2026poster

Real-world agricultural counting often operates in the extreme regime of \textbf{Dense and Indiscernible Object Counting (DIOC)}, where targets are tiny, clustered, and highly camouflaged. To facilitate research in this domain, we introduce \textbf{DIOCblueberry}, a large-scale benchmark that pushes…

Cited by 0SourceScholar
2026

Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models

AAAI 2026technical

Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-lab

Cited by 0SourcePDFScholar
2026

DNA: Uncovering Universal Latent Forgery Knowledge

ICML 2026poster

As generative AI achieves hyper-realism, superficial artifact detection has become obsolete. While prevailing methods rely on resource-intensive fine-tuning of black-box backbones, we propose that forgery detection capability is already encoded within pre-trained models rather than requiring end-to-…

Cited by 0SourceScholar
2026

DiT-Distill: Open-Set Fine-Grained Retrieval via Generative Curriculum Knowledge

CVPR 2026

Open-set fine-grained retrieval (OSFR) is a challenging task where models must generalize to unseen subcategories. Existing methods often fail this, as they embed category-specific semantics from closed-set training labels. Recently, diffusion transformers (DiT) have shown promise by encoding attrib

Cited by 0SourceScholar
2026

Explainable Forensics of Manipulated Segments in Untrimmed Long Videos

ICML 2026poster

The rapid advancement of AI-driven video generation has transformed content creation, while simultaneously increasing the risk of misinformation through localized manipulations in long-form videos. Existing video forensic methods predominantly operate on short, independent clips, and thus fail to ca…

Cited by 0SourceScholar
2026

HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment

AAAI 2026technical

Contrastive vision-language models like CLIP have achieved impressive results in image-text retrieval by aligning image and text representations in a shared embedding space. However, these models often treat text as flat sequences, limiting their ability to handle complex, compositional, and long-fo

Cited by 0SourcePDFScholar
2026

IMAGGarment+: Efficient Attribute-Wise Diffusion for Garment Generation

AAAI 2026technical

Diffusion models have advanced fine-grained garment generation, yet balancing controllability, efficiency, and texture fidelity remains challenging. Adapter-based methods often yield incoherent details, while full fine-tuning is computationally expensive and prone to overwriting pretrained priors. T

Cited by 0SourcePDFScholar
2026

Jointly Conditioned Diffusion Model for Multi-View Pose-Guided Person Image Synthesis

ICASSP 2026poster

Pose-guided human image generation is limited by incomplete textures from single reference views and the absence of explicit cross-view interaction. We present jointly conditioned diffusion model (JCDM), a jointly conditioned diffusion framework that exploits multi-view priors. The appearance prior…

Cited by 0SourcePDFScholar
2026

Reasoning-VLA: An Efficient and Spatial-Guided General Vision-Language-Action Reasoning Model for Autonomous Driving

ICML 2026poster

Vision-Language-Action (VLA) models have recently shown strong decision-making capabilities in autonomous driving. However, existing VLAs often struggle with achieving efficient inference and generalizing to novel autonomous vehicle configurations and driving scenarios. In this paper, we propose Rea…

Cited by 0SourceScholar
2026

SGMHand: Structure-Guided Modulation for Structure-Aware Hand Inpainting

AAAI 2026technical

Diffusion-based generative models have demonstrated remarkable capabilities in image synthesis, yet realistic hand generation remains a persistent challenge due to complex articulations, self-occlusion, and the lack of explicit structural guidance. To address these issues, we present SGMHand, a nov

Cited by 0SourcePDFScholar
2026

Seeing through the Conflict: Transparent Knowledge Conflict Handling in Retrieval-Augmented Generation

AAAI 2026technical

Large language models (LLMs) equipped with retrieval—the Retrieval-Augmented Generation (RAG) paradigm—should combine their parametric knowledge with external evidence, yet in practice they often hallucinate, over-trust noisy snippets, or ignore vital context. We introduce TCR (Transparent Conflict

Cited by 0SourcePDFScholar
2026

StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback

AAAI 2026technical

The advancement of intelligent agents has revolutionized problem-solving across diverse domains, yet solutions for personalized fashion styling remain underexplored, which holds immense promise for promoting shopping experiences. In this work, we present StyleTailor, the first collaborative agent fr

Cited by 0SourcePDFScholar
2026

Threshold-Guided Optimization for Visual Generative Models

ICML 2026poster

Aligning large visual generative models with human feedback is often performed through pairwise preference optimization. While such approaches are conceptually simple, they fundamentally rely on annotated pairs, limiting scalability in settings where feedback is collected as independent scalar ratin…

Cited by 0SourceScholar
2026

TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention

ICML 2026poster

Despite their capabilities, large foundation models (LFMs) remain susceptible to adversarial manipulation. Current defenses predominantly rely on the ``locality hypothesis", suppressing isolated neurons or features. However, harmful semantics act as distributed, cross-layer circuits, rendering such …

Cited by 0SourceScholar
2026

Transport and Merge: Cross-Architecture Merging for Large Language Models

ICML 2026poster

Large language models (LLMs) achieve strong capabilities by scaling model capacity and training data, yet many real-world deployments rely on smaller models trained or adapted from low-resource data. This gap motivates the need for mechanisms to transfer knowledge from large, high-resource models to…

Cited by 0SourceScholar
2026

Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation

CVPR 2026

Multilingual text-to-image (T2I) models have advanced rapidly in terms of visual realism and semantic alignment, and are now widely utilised. Yet outputs vary across cultural contexts: because language carries cultural connotations, images synthesized from multilingual prompts should preserve cross-

Cited by 0SourceScholar
2026

Who Transfers Safety? Identifying and Targeting Cross-Lingual Shared Safety Neurons

ICML 2026poster

Multilingual safety remains significantly imbalanced, leaving non-high-resource (NHR) languages vulnerable compared to robust high-resource (HR) ones. Moreover, the neural mechanisms driving safety alignment remain unclear despite observed cross-lingual representation transfer.In this paper, we find…

Cited by 0SourceScholar
2026

WildActor: Unconstrained Identity-Preserving Video Generation

ICML 2026poster

Production-ready human video generation requires digital actors to maintain strictly consistent full-body identities across dynamic shots, viewpoints and motions, a setting that remains challenging for existing methods. Prior methods often suffer from face-centric behavior that neglects body-level c…

Cited by 0SourceScholar
2025

Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition Tokenization

NeurIPS 2025spotlight

Sign Language Video Generation (SLVG) seeks to generate identity-preserving sign language videos from spoken language texts. Existing methods primarily rely on the single coarse condition (e.g., skeleton sequences) as the intermediary to bridge the translation model and the video generation model, w…

Cited by 0SourcecodeScholar
2025

Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models

AAAI 2025technical

Recent research showcases the considerable potential of conditional diffusion models for generating consistent stories. However, current methods, which primarily generate stories in a caption-dependent manner, often overlook the importance of contextual consistency and the relevance of frames durin…

2025

CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action Model

NeurIPS 2025poster

Autonomous driving represents a prominent application of artificial intelligence. Recent approaches have shifted from focusing solely on common scenarios to addressing complex, long-tail situations such as subtle human behaviors, traffic accidents, and non-compliant driving patterns. Given the demon…

Cited by 0SourceScholar
2025

DVP-MVS: Synergize Depth-Edge and Visibility Prior for Multi-View Stereo

AAAI 2025technical

Patch deformation-based methods have recently exhibited substantial effectiveness in multi-view stereo, due to the incorporation of deformable and expandable perception to reconstruct textureless areas. However, such approaches typically focus on exploring correlative reliable pixels to alleviate m…

Cited by 4SourcePDFScholar
2025

DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary Lookup

ICCV 2025poster

Recent vision-language models (e.g., CLIP) have demonstrated remarkable class-generalizable ability to unseen classes in few-shot anomaly segmentation (FSAS), leveraging supervised prompt learning or fine-tuning on seen classes. However, their cross-category generalization largely depends on prior k…

2025

Ensembling Diffusion Models via Adaptive Feature Aggregation

ICLR 2025poster

The success of the text-guided diffusion model has inspired the development and release of numerous powerful diffusion models within the open-source community. These models are typically fine-tuned on various expert datasets, showcasing diverse denoising capabilities. Leveraging multiple high-qualit…

2025

Exploring the Potential of Large Vision-Language Models for Unsupervised Text-Based Person Retrieval

AAAI 2025technical

The aim of text-based person retrieval is to identify pedestrians using natural language descriptions within a large-scale image gallery. Traditional methods rely heavily on manually annotated image-text pairs, which are resource-intensive to obtain. With the emergence of Large Vision-Language Model…

Cited by 0SourcePDFScholar
2025

FaceShot: Bring Any Character into Life

ICLR 2025poster

In this paper, we present ***FaceShot***, a novel training-free portrait animation framework designed to bring any character into life from any driven video without fine-tuning or retraining. We achieve this by offering precise and robust reposed landmark sequences from an appearance-guided landmark…

2025

IMAGDressing-v1: Customizable Virtual Dressing

AAAI 2025technical

Existing virtual try-on (VTON) methods provide only limited user control over garment attributes and generally overlook essential factors such as face, pose, and scene context. To address these limitations, we introduce the virtual dressing (VD) task, which aims to synthesize freely editable human i…

2025

Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model

ICML 2025poster

Recent advances in conditional diffusion models have shown promise for generating realistic TalkingFace videos, yet challenges persist in achieving consistent head movement, synchronized facial expressions, and accurate lip synchronization over extended generations. To address these, we introduce th…

Cited by 16SourcePDFScholar
2025

MSP-MVS: Multi-Granularity Segmentation Prior Guided Multi-View Stereo

AAAI 2025technical

Recently, patch deformation-based methods have demonstrated significant strength in multi-view stereo by adaptively expanding the reception field of patches to help reconstruct textureless areas. However, such methods mainly concentrate on searching for pixels without matching ambiguity (i.e., reli…

Cited by 6SourcePDFScholar
2025

PrimeK-Net: Multi-scale Spectral Learning via Group Prime-Kernel Convolutional Neural Networks for Single Channel Speech Enhancement

ICASSP 2025accepted

Single-channel speech enhancement is a challenging ill-posed problem focused on estimating clean speech from degraded signals. Existing studies have demonstrated the competitive performance of combining convolutional neural networks (CNNs) with Transformers in speech enhancement tasks. However, exis…

Cited by 16SourceScholar
2025

SPEAK: Speech-Driven Pose and Emotion-Adjustable Talking Head Generation

ICASSP 2025accepted

Most earlier researches on talking face generation have focused on the synchronization of lip motion and speech content. However, head pose and facial emotions are equally important characteristics of natural faces. While audio-driven talking face generation has seen notable advancements, existing m…

Cited by 0SourceScholar
2025

SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency

NeurIPS 2025poster

Recent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus on character consistency, they often overlook the crucial role of scenes in storytelling, which restricts their creativit…

Cited by 0SourceScholar
2025

UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation

ICML 2025poster

Although significant advancements have been achieved in the progress of keypoint-guided Text-to-Image diffusion models, existing mainstream keypoint-guided models encounter challenges in controlling the generation of more general non-rigid objects beyond humans (e.g., animals). Moreover, it is diffi…

Cited by 0SourcePDFScholar
2024

Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models

ICLR 2024poster

Recent work has showcased the significant potential of diffusion models in pose-guided person image synthesis. However, owing to the inconsistency in pose between the source and target images, synthesizing an image with a distinct pose, relying exclusively on the source image and target pose informa…

2024

VCP-CLIP: A visual context prompting model for zero-shot anomaly segmentation

ECCV 2024poster

"Recently, large-scale vision-language models such as CLIP have demonstrated immense potential in zero-shot anomaly segmentation (ZSAS) task, utilizing a unified model to directly detect anomalies on any unseen product with painstakingly crafted text prompts. However, existing methods often assume t…

2016

An energy-aware auction for hybrid access in heterogeneous networks under QoS requirements

ICASSP 2016accepted

We consider a heterogeneous network (HetNet) in which multiple small cell base stations (SBSs) aim to offload a quantity of macro cell user equipments (MUEs) to reduce the energy consumption of the network while guaranteeing the QoS requirements of all UEs. We design an ascending-bid auction mechani…

Cited by 0SourceScholar