← Search

Linlin Shen

69 accepted papers

2026

Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection

CVPR 2026

Face forgery detection faces a critical challenge: a persistent gap between offline benchmarks and real-world efficacy, which we attribute to the ecological invalidity of training data. This work introduces Agent4FaceForgery to address two fundamental problems: (1) how to capture the diverse intents

Cited by 0SourceScholar
2026

EchoVDiff: Cardiac-Cycle Echocardiography Video Generation from Arbitrary Single Frame

CVPR 2026

Reconstructing a physiologically plausible cardiac video from a single image remains a fundamental challenge in generative modeling, owing to the complex and nonlinear periodic dynamics of echocardiography. Previous image-to-video (I2V) approaches primarily focus on temporal continuity, yet often st

Cited by 0SourcecodeScholar
2026

FineXtrol: Controllable Motion Generation via Fine-Grained Text

AAAI 2026technical

Recent works have sought to enhance the controllability and precision of text-driven motion generation. Some approaches leverage large language models (LLMs) to produce more detailed texts, while others incorporate global 3D coordinate sequences as additional control signals. However, the former oft

Cited by 0SourcePDFScholar
2026

Gamba: Mamba-based graph convolutional network with dynamic graph topology learning for action recognition

CVPR 2026

Existing graph models predominantly utilize self-attention mechanisms to model feature correlations between the joints of each sample, which not only neglects dynamic relation dependencies in temporal dimension but also leads to redundant computation and difficulty in establishing a unified framewor

Cited by 0SourcecodeScholar
2026

H2-Surv: Hierarchical Hyperbolic Multimodal Representation Learning for Survival Prediction

CVPR 2026

Cancer survival prediction through multimodal learning that combines histopathology images with genomic data represents a promising research direction. However, current approaches still suffer from two key limitations. First, most methods operate in a Euclidean feature space, which makes it difficul

Cited by 0SourceScholar
2026

Incomplete Multi-view Diabetic Retinopathy Grading via Self-Supervised Inter- and Intra-View Restoration

AAAI 2026technical

Multi-view diabetic retinopathy (DR) grading has achieved remarkable performance by capturing more comprehensive pathological features than single-view methods. However, complete multi-view fundus images are often difficult to obtain in clinical practice, and the performance degrades significantly w

Cited by 0SourcePDFScholar
2026

Maximizing Incremental Information Entropy for Contrastive Learning

ICLR 2026poster

Contrastive learning has achieved remarkable success in self-supervised representation learning, often guided by information-theoretic objectives such as mutual information maximization. Motivated by the limitations of static augmentations and rigid invariance constraints, we propose IE-CL (Incremen…

Cited by 0SourceScholar
2026

More than the Sum: Panorama-Language Models for Adverse Omni-Scenes

CVPR 2026

Existing vision-language models (VLMs) are tailored for pinhole imagery, stitching multiple narrow field-of-view inputs to piece together a complete omni-scene understanding. Yet, such multi-view perception overlooks the holistic spatial and contextual relationships that a single panorama inherently

Cited by 0SourcecodeScholar
2026

OralGPT-Omni: A Versatile Dental Multimodal Large Language Model

CVPR 2026

Multimodal Large Language Models (MLLMs) have exhibited immense potential across numerous medical specialties, yet dentistry remains underexplored, in part due to limited domain-specific data, scarce dental expert annotations, insufficient modality-specific modeling, and challenges in reliability. I

Cited by 0SourceScholar
2026

PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing

ICLR 2026poster

Remote photoplethysmography (rPPG) enables non-contact physiological measurement but remains highly susceptible to illumination changes, motion artifacts, and limited temporal modeling. Large Language Models (LLMs) excel at capturing long-range dependencies, offering a potential solution but struggl…

Cited by 0SourceScholar
2026

ProConMV: Provenance-Enabled Conceptual Framework for Interpretable Multi-View Diabetic Retinopathy Diagnosis

ICML 2026poster

Existing deep learning models have demonstrated potential in Diabetic retinopathy (DR) diagnosis, but they still suffer from three key challenges: reliance on single-source inputs, opaque and untraceable reasoning processes, and the absence of a mechanism for result verification. Thus, we propose a …

Cited by 0SourceScholar
2026

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark

CVPR 2026

Fine-grained spatiotemporal reasoning on surgical videos is critical, yet the capabilities of Multi-modal Large Language Models (MLLMs) in this domain remain largely unexplored. To bridge this gap, we introduce SurgCoT, a unified benchmark for evaluating chain-of-thought (CoT) reasoning in MLLMs acr

Cited by 0SourcecodeScholar
2026

VISION KAN: TOWARDS AN ATTENTION-FREE BACKBONE FOR VISION WITH KOLMOGOROV-ARNOLD NETWORKS

ICASSP 2026poster

Attention mechanisms have become a key module in modern vision backbones due to their ability to model long-range dependencies. However, their quadratic complexity in sequence length and the difficulty of interpreting attention weights limit both scalability and clarity. Recent attention-free archit…

Cited by 0SourcePDFScholar
2026

X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis

CVPR 2026

Despite significant progress in Multi-modal Large Language Models (MLLMs), their clinical reasoning capacity for multi-modal diagnosis remains largely unexamined. Current benchmarks, mostly single-modality data, can't evaluate progressive reasoning and cross-modal integration essential for clinical

Cited by 0SourcecodeScholar
2025

Asclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models

ACL 2025long

The significant breakthroughs of Medical Multi-Modal Large Language Models (Med-MLLMs) renovate modern healthcare with robust information synthesis and medical decision support. However, these models are often evaluated on benchmarks that are unsuitable for the Med-MLLMs due to the intricate nature…

2025

Big-Moe: Bypassing Isolated Gating For Generalized Multimodal Face Anti-Spoofing

ICASSP 2025accepted

In the domain of facial recognition security, multimodal Face Anti-Spoofing (FAS) is essential for countering presentation attacks. However, existing technologies encounter challenges due to modality biases and imbalances, as well as domain shifts. Our research introduces a Mixture of Experts (MoE)…

Cited by 0SourceScholar
2025

CA-Edit: Causality-Aware Condition Adapter for High-Fidelity Local Facial Attribute Editing

AAAI 2025technical

For efficient and high-fidelity local facial attribute editing, most existing editing methods either require additional fine-tuning for different editing effects or tend to affect beyond the editing regions. Alternatively, inpainting methods can edit the target image region while preserving external…

2025

DAMPER: A Dual-Stage Medical Report Generation Framework with Coarse-Grained MeSH Alignment and Fine-Grained Hypergraph Matching

AAAI 2025technical

Medical report generation is crucial for clinical diagnosis and patient management, summarizing diagnoses and recommendations based on medical imaging. However, existing work often overlook the clinical pipeline involved in report writing, where physicians typically conduct an initial quick review f…

Cited by 0SourcePDFScholar
2025

DAP-MAE: Domain-Adaptive Point Cloud Masked Autoencoder for Effective Cross-Domain Learning

ICCV 2025poster

Compared to 2D data, the scale of point cloud data in different domains available for training, is quite limited. Researchers have been trying to combine these data of different domains for masked autoencoder (MAE) pre-training to leverage such a data scarcity issue. However, the prior knowledge lea…

2025

DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face Synthesis

ICASSP 2025accepted

Accurately synthesizing talking face videos and capturing fine facial features for individuals with long hair presents a significant challenge. To tackle these challenges in existing methods, we propose a decomposed per-embedding Gaussian fields (DEGSTalk), a 3D Gaussian Splatting (3DGS)-based talki…

Cited by 0SourceScholar
2025

DeeperForward: Enhanced Forward-Forward Training for Deeper and Better Performance

ICLR 2025poster

While backpropagation effectively trains models, it presents challenges related to bio-plausibility, resulting in high memory demands and limited parallelism. Recently, Hinton (2022) proposed the Forward-Forward (FF) algorithm for high-parallel local updates. FF leverages squared sums as the local u…

Cited by 0SourcePDFScholar
2025

EAGLE: Expert-Guided Self-Enhancement for Preference Alignment in Pathology Large Vision-Language Model

ACL 2025long

Recent advancements in Large Vision Language Models (LVLMs) show promise for pathological diagnosis, yet their application in clinical settings faces critical challenges of multimodal hallucination and biased responses. While preference alignment methods have proven effective in general domains, acq…

2025

FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMs

CVPR 2025poster

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in various tasks. However, effectively evaluating these MLLMs on face perception remains largely unexplored. To address this gap, we introduce FaceBench, a dataset featuring hierarchical multi-view and multi-level att…

2025

FineMotion: A Dataset and Benchmark with both Spatial and Temporal Annotation for Fine-grained Motion Generation and Editing

ICCV 2025poster

Generating realistic human motions from textual descriptions has undergone significant advancements. However, existing methods often overlook specific body part movements and their timing. In this paper, we address this issue by enriching the textual description with more details. Specifically, we p…

Cited by 0SourcePDFScholar
2025

High-Fidelity Editable Portrait Synthesis with 3D GAN Inversion

ICASSP 2025accepted

The 3D generative adversarial network (GAN) inversion converts an image into 3D representation to attain high-fidelity reconstruction and facilitate realistic image manipulation within the 3D latent space. However, previous approaches face challenges regarding the trade-off between the reconstructio…

Cited by 0SourceScholar
2025

Like an Ophthalmologist: Dynamic Selection Driven Multi-View Learning for Diabetic Retinopathy Grading

AAAI 2025technical

Diabetic retinopathy (DR), with its large patient population, has become a formidable threat to human visual health. In the clinical diagnosis of DR, multi-view fundus images are considered to be more suitable for DR diagnosis because of the wide coverage of the field of view. Therefore, different f…

2025

MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities

CVPR 2025poster

Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the overall semantics of an entire motion sequence in just a few…

2025

MedChain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive Sequence

NeurIPS 2025spotlight

Clinical decision making (CDM) is a complex, dynamic process crucial to healthcare delivery, yet it remains a significant challenge for artificial intelligence systems. While Large Language Model (LLM)-based agents have been tested on general medical knowledge using licensing exams and knowledge que…

Cited by 0SourceScholar
2025

PerReactor: Offline Personalised Multiple Appropriate Facial Reaction Generation

AAAI 2025technical

In dyadic human-human interactions, individuals may express multiple different facial reactions in response to the same/similar behaviours expressed by their conversational partners depending on their personalised behaviour patterns. As a result, frequently-employed reconstruction loss-based strateg…

2025

SynFER: Towards Boosting Facial Expression Recognition with Synthetic Data

ICCV 2025poster

Facial expression datasets remain limited in scale due to privacy concerns, the subjectivity of annotations, and the labor-intensive nature of data collection. This limitation poses a significant challenge for developing modern deep learning-based facial expression analysis models, particularly foun…

Cited by 0SourcePDFScholar
2025

S³-Mamba: Small-Size-Sensitive Mamba for Lesion Segmentation

AAAI 2025technical

Small lesions play a critical role in early disease diagnosis and intervention of severe infections. Popular models often face challenges in segmenting small lesions, as it occupies only a minor portion of an image, while down-sampling operations may inevitably lose focus on local features of small…

Cited by 1SourcePDFScholar
2025

WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image

ICCV 2025poster

Recent advances in computational pathology have introduced whole slide image (WSI)-level multimodal large language models (MLLMs) for automated pathological analysis. However, current WSI-level MLLMs face two critical challenges: limited explainability in their decision-making process and insufficie…

Cited by 0SourcePDFScholar
2024

APSeg: Auto-Prompt Network for Cross-Domain Few-Shot Semantic Segmentation

CVPR 2024poster

Few-shot semantic segmentation (FSS) endeavors to segment unseen classes with only a few labeled samples. Current FSS methods are commonly built on the assumption that their training and application scenarios share similar domains and their performances degrade significantly while applied to a disti…

Cited by 16SourcePDFScholar
2024

Boosting Adversarial Transferability across Model Genus by Deformation-Constrained Warping

AAAI 2024technical

Adversarial examples generated by a surrogate model typically exhibit limited transferability to unknown target systems. To address this problem, many transferability enhancement approaches (e.g., input transformation and model augmentation) have been proposed. However, they show poor performances i…

2024

Circular Decomposition and Cross-Modal Recombination for Multimodal Sentiment Analysis

ICASSP 2024accepted

Multimodal Sentiment Analysis is a burgeoning research area, leveraging various modalities to predict the sentiment score. Nevertheless, previous studies have disregarded the impact of noise interference on specific modal sentiments during video recording, thereby compromising the accuracy of sentim…

Cited by 0SourceScholar
2024

Dynamic Data Sampler for Cross-Language Transfer Learning in Large Language Models

ICASSP 2024accepted

Large Language Models (LLMs) have gained significant attention in the field of natural language processing (NLP) due to their wide range of applications. However, training LLMs for languages other than English poses significant challenges, due to the difficulty in acquiring large-scale corpus and th…

Cited by 0SourceScholar
2024

Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation

ACL 2024long

Fine-grained vision-language models (VLM) have been widely used for inter-modality local alignment between the predefined fixed patches and textual words. However, in medical analysis, lesions exhibit varying sizes and positions, and using fixed patches may cause incomplete representations of lesion…

Cited by 13SourcePDFScholar
2024

HairDiffusion: Vivid Multi-Colored Hair Editing via Latent Diffusion

NeurIPS 2024poster

Hair editing is a critical image synthesis task that aims to edit hair color and hairstyle using text descriptions or reference images, while preserving irrelevant attributes (e.g., identity, background, cloth). Many existing methods are based on StyleGAN to address this task. However, due to the li…

Cited by 0SourcePDFScholar
2024

MERG: Multi-Dimensional Edge Representation Generation Layer for Graph Neural Networks

ICASSP 2024accepted

Edges are essential in describing relationships among nodes. While existing graphs frequently use a single-value edge to describe association between each pair of node vectors, crucial relationships may be disregarded if they are not linearly correlated, which may limit graph analysis performance. A…

Cited by 0SourceScholar
2024

MTaDCS: Moving Trace and Feature Density-based Confidence Sample Selection under Label Noise

ECCV 2024poster

"Learning from noisy labels is a challenging task, as noisy labels can compromise decision boundaries and result in suboptimal generalization performance. Most previous approaches for dealing noisy labels are based on sample selection, which utilized the small loss criterion to reduce the adverse ef…

2024

Scale-Free And Task-Generic Attack: Generating Photo-Realistic Adversarial Patterns With Patch Quilting Generator

ICASSP 2024accepted

Recent CNN generator-based attack approaches can synthe-size unrestricted and semantically meaningful entities to the image, which are able to improve the transferability and robustness. However, such methods attack images by either synthesizing local adversarial entities, which are only suitable fo…

Cited by 0SourceScholar
2024

Towards Combating Frequency Simplicity-biased Learning for Domain Generalization

NeurIPS 2024poster

Domain generalization methods aim to learn transferable knowledge from source domains that can generalize well to unseen target domains. Recent studies show that neural networks frequently suffer from a simplicity-biased learning behavior which leads to over-reliance on specific frequency sets, nam…

2024

Tune-An-Ellipse: CLIP Has Potential to Find What You Want

CVPR 2024highlight

Visual prompting of large vision language models such as CLIP exhibits intriguing zero-shot capabilities. A manually drawn red circle commonly used for highlighting can guide CLIP's attention to the surrounding region to identify specific objects within an image. Without precise object proposals how…

2023

Learning Visual Prior via Generative Pre-Training

NeurIPS 2023poster

Various stuff and things in visual data possess specific traits, which can be learned by deep neural networks and are implicitly represented as the visual prior, e.g., object location and shape, in the model. Such prior potentially impacts many vision tasks. For example, in conditional image synthes…

2023

Shift from Texture-bias to Shape-bias: Edge Deformation-based Augmentation for Robust Object Recognition

ICCV 2023poster

Recent studies have shown the vulnerability of CNNs under perturbation noises, which is partially caused by the reason that the well-trained CNNs are too biased toward the object texture, i.e., they make predictions mainly based on texture cues. To reduce this texture-bias, current studies resort to…

Cited by 7PDFcodeScholar
2023

StyleGene: Crossover and Mutation of Region-Level Facial Genes for Kinship Face Synthesis

CVPR 2023highlight

High-fidelity kinship face synthesis has many potential applications, such as kinship verification, missing child identification, and social media analysis. However, it is challenging to synthesize high-quality descendant faces with genetic relations due to the lack of large-scale, high-quality anno…

2023

UniFace: Unified Cross-Entropy Loss for Deep Face Recognition

ICCV 2023poster

As a widely used loss function in deep face recognition, the softmax loss cannot guarantee that the minimum positive sample-to-class similarity is larger than the maximum negative sample-to-class similarity. As a result, no unified threshold is available to separate positive sample-to-class pairs fr…

Cited by 29PDFcodeScholar
2023

UniTSFace: Unified Threshold Integrated Sample-to-Sample Loss for Face Recognition

NeurIPS 2023poster

Sample-to-class-based face recognition models can not fully explore the cross-sample relationship among large amounts of facial images, while sample-to-sample-based models require sophisticated pairing processes for training. Furthermore, neither method satisfies the requirements of real-world face…

2022

C2AM: Contrastive Learning of Class-Agnostic Activation Map for Weakly Supervised Object Localization and Semantic Segmentation

CVPR 2022poster

While class activation map (CAM) generated by image classification network has been widely used for weakly supervised object localization (WSOL) and semantic segmentation (WSSS), such classifiers usually focus on discriminative object regions. In this paper, we propose Contrastive learning for Class…

Cited by 142PDFcodeScholar
2022

CLIMS: Cross Language Image Matching for Weakly Supervised Semantic Segmentation

CVPR 2022poster

It has been widely known that CAM (Class Activation Map) usually only activates discriminative object regions and falsely includes lots of object-related backgrounds. As only a fixed set of image-level object labels are available to the WSSS (weakly supervised semantic segmentation) model, it could…

Cited by 183PDFcodeScholar
2022

CSL: A Large-scale Chinese Scientific Literature Dataset

COLING 2022main

Scientific literature serves as a high-quality corpus, supporting a lot of Natural Language Processing (NLP) research. However, existing datasets are centered around the English language, which restricts the development of Chinese scientific NLP. In this work, we present CSL, a large-scale Chinese S…

2022

Frequency-Driven Imperceptible Adversarial Attack on Semantic Similarity

CVPR 2022poster

Current adversarial attack research reveals the vulnerability of learning-based classifiers against carefully crafted perturbations. However, most existing attack methods have inherent limitations in cross-dataset generalization as they rely on a classification layer with a closed set of categories.…

Cited by 137PDFcodeScholar
2022

Learning Multi-dimensional Edge Feature-based AU Relation Graph for Facial Action Unit Recognition

IJCAI 2022poster

The activations of Facial Action Units (AUs) mutually influence one another. While the relationship between a pair of AUs can be complex and unique, existing approaches fail to specifically and explicitly represent such cues for each pair of AUs in each facial display. This paper proposes an AU rela…

2022

RamGAN: Region Attentive Morphing GAN for Region-Level Makeup Transfer

ECCV 2022poster

"In this paper, we propose a region adaptive makeup transfer GAN, called RamGAN, for precise region-level makeup transfer. Compared to face-level transfer methods, our RamGAN uses spatial-aware Region Attentive Morphing Module (RAMM) to encode Region Attentive Matrices (RAMs) for local regions like…

2022

Scene Consistency Representation Learning for Video Scene Segmentation

CVPR 2022poster

A long-term video, such as a movie or TV show, is composed of various scenes, each of which represents a series of shots sharing the same semantic story. Spotting the correct scene boundary from the long-term video is a challenging task, since a model must understand the storyline of the video to fi…

Cited by 20PDFcodeScholar
2021

Adversarial Defence by Diversified Simultaneous Training of Deep Ensembles

AAAI 2021technical

Learning-based classifiers are susceptible to adversarial examples. Existing defence methods are mostly devised on individual classifiers. Recent studies showed that it is viable to increase adversarial robustness by promoting diversity over an ensemble of models. In this paper, we propose adversari…

2021

FS-Net: Fast Shape-Based Network for Category-Level 6D Object Pose Estimation With Decoupled Rotation Mechanism

CVPR 2021poster

In this paper, we focus on category-level 6D pose and size estimation from a monocular RGB-D image. Previous methods suffer from inefficient category-level pose feature extraction, which leads to low accuracy and inference speed. To tackle this problem, we propose a fast shape-based network (FS-Net)…

Cited by 197PDFcodeScholar
2021

Group-Wise Inhibition Based Feature Regularization for Robust Classification

ICCV 2021poster

The convolutional neural network (CNN) is vulnerable to degraded images with even very small variations (e.g. corrupted and adversarial samples). One of the possible reasons is that CNN pays more attention to the most discriminative regions, but ignores the auxiliary features when learning, leading…

Cited by 18PDFcodeScholar
2021

Online Refinement of Low-Level Feature Based Activation Map for Weakly Supervised Object Localization

ICCV 2021poster

We present a two-stage learning framework for weakly supervised object localization (WSOL). While most previous efforts rely on high-level feature based CAMs (Class Activation Maps), this paper proposes to localize objects using the low-level feature based activation maps. In the first stage, an act…

Cited by 78PDFcodeScholar
2021

Translate the Facial Regions You Like Using Self-Adaptive Region Translation

AAAI 2021technical

With the progression of Generative Adversarial Networks (GANs), image translation methods has achieved increasingly remarkable performance. However, most available methods can only achieve image level translation, which is unable to precisely control the regions to be translated. In this paper, we p…

Cited by 9SourcePDFScholar
2020

Geometry Constrained Weakly Supervised Object Localization

ECCV 2020poster

We propose a geometry constrained network, termed GCNet, for weakly supervised object localization (WSOL). GC-Net consists of three modules: a detector, a generator and a classifier. The detector predicts the object location defined by a set of coefficients describing a geometric shape (i.e. ellipse or…

2020

Self-Supervised CycleGAN for Object-Preserving Image-to-Image Domain Adaptation

ECCV 2020poster

Recent generative adversarial network (GAN) based methods (e.g., CycleGAN) are prone to fail at preserving image-objects in image-to-image translation, which reduces their practicality on tasks such as domain adaptation. Some frameworks have been proposed to adopt a segmentation network as the auxil…

Cited by 34SourcePDFScholar
2017

Multi-Way Multi-Level Kernel Modeling for Neuroimaging Classification

CVPR 2017poster

Owing to prominence as a diagnostic tool for probing the neural correlates of cognition, neuroimaging tensor data has been the focus of intense investigation. Although many supervised tensor learning approaches have been proposed, they either cannot capture the nonlinear relationships of tensor data…

Cited by 31PDFScholar