← Search

Weisheng Li

26 accepted papers

2026

ClinAlign: Clinical Workflow Aligned Memory Retrieval for Radiology Report Generation

IJCAI 2026

Automated radiology report generation aims to create clear and clinically correct diagnostic reports from medical images. Existing retrieval enhancement methods primarily focus on reusing textual knowledge, neglecting the crucial role of local visual pattern memory in clinical diagnosis. Furthermore

Cited by 0Scholar
2025

A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly Understanding

NeurIPS 2025poster

While unmanned aerial vehicles (UAVs) offer wide-area, high-altitude coverage for anomaly detection, they face challenges such as dynamic viewpoints, scale variations, and complex scenes. Existing datasets and methods, mainly designed for fixed ground-level views, struggle to adapt to these conditio…

Cited by 0SourcecodeScholar
2025

Bidirectional Reference Image Quality Assessment via Content-Quality Correlation Modeling

ICASSP 2025accepted

The emphasis on no-reference image quality assessment has often overshadowed the significance of Full-Reference Image Quality Assessment (FR-IQA), which generally better reflects human contrastive perception mechanism. However, FRIQA presents challenges in obtaining content-aligned reference images.…

Cited by 0SourceScholar
2025

CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training

AAAI 2025technical

Benefiting from large-scale pre-training of text-video pairs, current text-to-video (T2V) diffusion models can generate high-quality videos from the text description. Besides, given some reference images or videos, the parameter-efficient fine-tuning method, i.e. LoRA, can generate high-quality cust…

2025

Incremental Transformer: Efficient Encoder for Incremented Text Over MRC and Conversation Tasks

COLING 2025main

Some encoder inputs such as conversation histories are frequently extended with short additional inputs like new responses. However, to obtain the real-time encoding of the extended input, existing Transformer-based encoders like BERT have to encode the whole extended input again without utilizing t…

Cited by 0SourcePDFScholar
2025

Learning Preconditioners in Gates-controlled Deep Unfolding Networks based on Quasi-Newton Methods For Accelerated MRI Reconstruction

ICASSP 2025accepted

Deep unfolding networks (DUNs) have made significant progress in MRI reconstruction, successfully tackling the problem of prolonged imaging time. However, the ill-conditioned nature of MRI reconstruction often causes slow convergence in iterative optimization, potentially compromising the performanc…

Cited by 0SourceScholar
2025

Self-Geometry-Guided Direct Pose Regression Based on Dual Perspective Fusion for 2D-3D Cross Dimensional Spinal Surgery Navigation

ICASSP 2025accepted

2D-3D cross-dimensional registration for spinal surgery navigation, which aims to achieve real-time visual navigation of preoperative 3D vertebrae based on intraoperative 2D fluoroscopy images, faces significant challenges due to semantic and dimensional gaps. Traditional 2D-3D registration methods…

Cited by 0SourceScholar
2025

Stacking U-Nets in U-shape: Redesigning the Information Flow in Model-based Networks for MRI Reconstruction

ICASSP 2025accepted

Model-based networks have shown convincing performance in MRI reconstruction. However, the unrolled cascades within the networks are constrained to solely obtain information from the preceding counterpart, resulting in potential error accumulation. Moreover, the linear structure fails to address the…

Cited by 0SourceScholar
2025

Subsampling Decomposition based k-Space Refinement for Accelerated MRI Reconstruction

ICASSP 2025accepted

In accelerated MRI reconstruction problem, directly recovering all the missing k-space data from undersampled measurements is highly ill-posed and often leads to suboptimal performance. To address the problem, we propose a novel deep unfolding network (DUN) with subsampling decomposition (SD) based…

Cited by 0SourceScholar
2025

Towards Universal AI-Generated Image Detection by Variational Information Bottleneck Network

CVPR 2025poster

The rapid advancement of generative models has significantly improved the quality of generated images. Meanwhile, it challenges information authenticity and credibility. Current generated image detection methods based on large-scale pre-trained multimodal models have achieved impressive results. Alt…

2024

CT and MRI Fusion with Anisotropic Guided Filtering

ICASSP 2024accepted

The combination of CT and MRI can provide more accurate images of lesions, yielding a significantly higher diagnostic value compared to single-modality pathological images. However, in CT-MRI fusion, preserving the gray-scale distribution of the source image while avoiding ‘detail halos’ poses a cha…

Cited by 0SourceScholar
2024

Facial Aesthetic Enhancement Network for Asian Faces Based on Differential Facial Aesthetic Activations

ICASSP 2024accepted

In this paper, we addressed facial aesthetic enhancement (FAE). Although existing methods have made great progress, the beautified images generated by them are highly prone to poor beautification, which limits their application to real-world scenes. To tackle this problem, we proposed a new method c…

Cited by 0SourceScholar
2024

Fast Intra Mode Prediction Algorithms for SCBS in VVC SCC

ICASSP 2024accepted

Versatile Video Coding (VVC) now supports Screen Content Coding (SCC) by integrating two efficient coding modes: Intra Block Copy (IBC) and Palette (PLT). However, the numerous modes and the Quad-Tree Plus Multi-Type Tree (QTMT) structure inherent to VVC contribute to a very high coding complexity.…

Cited by 0SourceScholar
2024

Using My Artistic Style? You Must Obtain My Authorization

ECCV 2024poster

"Artistic images typically contain the unique creative styles of artists. However, it is easy to transfer an artist’s style to arbitrary target images using style transfer techniques. To protect styles, some researchers use adversarial attacks to safeguard artists’ artistic style images. Prior metho…

2024

Window-Based Convolutional Sparse Coding: Towards A Unified Framework

ICASSP 2024accepted

Sparse Coding (SC) and Convolution Sparse Coding (CSC) are two widely studied sparse methods in computer vision and signal processing. SC encodes the image patches independently, however fails to utilize the correlation among them. CSC adopts a convolution operator to connect the overlapping patches…

Cited by 0SourceScholar
2023

A Novel Mode Selection-Based Fast Intra Prediction Algorithm for Spatial SHVC

ICASSP 2023accepted

Due to multi-layer encoding and Inter-layer prediction, Spatial Scalable High-Efficiency Video Coding (SSHVC) has extremely high coding complexity. It is very crucial to improve its coding speed so as to promote widespread and cost-effective SSHVC applications. In this paper, we have proposed a nove…

Cited by 0SourceScholar
2023

DLBD: A Self-Supervised Direct-Learned Binary Descriptor

CVPR 2023poster

For learning-based binary descriptors, the binarization process has not been well addressed. The reason is that the binarization blocks gradient back-propagation. Existing learning-based binary descriptors learn real-valued output, and then it is converted to binary descriptors by their proposed bin…

2023

MCF: Mutual Correction Framework for Semi-Supervised Medical Image Segmentation

CVPR 2023poster

Semi-supervised learning is a promising method for medical image segmentation under limited annotation. However, the model cognitive bias impairs the segmentation performance, especially for edge regions. Furthermore, current mainstream semi-supervised medical image segmentation (SSMIS) methods lack…

2023

Self-Supervised Image Local Forgery Detection by JPEG Compression Trace

AAAI 2023technical

For image local forgery detection, the existing methods require a large amount of labeled data for training, and most of them cannot detect multiple types of forgery simultaneously. In this paper, we firstly analyzed the JPEG compression traces which are mainly caused by different JPEG compression c…

Cited by 7SourcePDFScholar
2021

Poolingformer: Long Document Modeling with Pooling Attention

ICML 2021spotlight

In this paper, we introduce a two-level attention schema, Poolingformer, for long document modeling. Its first level uses a smaller sliding window pattern to aggregate information from neighbors. Its second level employs a larger window to increase receptive fields with pooling attention to reduce b…

2021

Say ‘YES’ to Positivity: Detecting Toxic Language in Workplace Communications

EMNLP 2021finding

Workplace communication (e.g. email, chat, etc.) is a central part of enterprise productivity. Healthy conversations are crucial for creating an inclusive environment and maintaining harmony in an organization. Toxic communications at the workplace can negatively impact overall job satisfaction and…

Cited by 26SourcePDFScholar