← Search

Jiang liu

43 accepted papers

2026

DeLightMono: Enhancing Self-Supervised Monocular Depth Estimation in Endoscopy by Decoupling Uneven Illumination

AAAI 2026technical

Self-supervised monocular depth estimation serves as a key task in the development of endoscopic navigation systems. However, performance degradation persists due to uneven illumination inherent in endoscopic images, particularly in low-intensity regions. Existing low-light enhancement techniques fa

Cited by 0SourcePDFScholar
2026

ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning

ICLR 2026poster

The rapid advancement of text-to-image (T2I) models has increased the need for reliable human preference modeling, a demand further amplified by recent progress in reinforcement learning for preference alignment. However, existing approaches typically quantify the quality of a generated image using…

Cited by 0SourceScholar
2026

MOVi: Training-free Text-conditioned Multi-Object Video Generation

ICASSP 2026oral

Recent advances in diffusion-based text-to-video (T2V) models have demonstrated remarkable progress, but these models still face challenges in generating videos with multiple objects. Most models struggle with accurately capturing complex object interactions, often treating some objects as static ba…

Cited by 0SourcePDFScholar
2026

OmniCT: Towards a Unified Slice-Volume LVLM for Comprehensive CT Analysis

ICLR 2026poster

Computed Tomography (CT) is one of the most widely used and diagnostically information-dense imaging modalities, covering critical organs such as the heart, lungs, liver, and colon. Clinical interpretation relies on both \textbf{slice-driven} local features (e.g., sub-centimeter nodules, lesion boun…

Cited by 0SourcecodeScholar
2026

PERCEPTUAL QUALITY OPTIMIZATION OF IMAGE SUPER-RESOLUTION

ICASSP 2026poster

Single-image super-resolution (SR) has achieved remarkable progress with deep learning, yet most approaches rely on distortion-oriented losses or heuristic perceptual priors, which often lead to a trade-off between fidelity and visual quality. To address this issue, we propose an \textit{Efficient P…

Cited by 0SourcePDFScholar
2026

Real2Sim2Real: RetinalDepth-64K for Depth Estimation in Posterior Segment Ophthalmic Surgery

CVPR 2026

Accurate depth estimation is crucial for 3D reconstruction and precise navigation in posterior segment ophthalmic surgery. However, acquiring annotated data remains challenging due to the impracticality of depth sensors under surgical microscopes. To overcome this limitation, we introduce RetinalDep

Cited by 0SourceScholar
2026

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis

ICML 2026poster

The advancement of Medical Vision-Language Models (VLMs) for 3D Computed Tomography (CT) analysis is hindered by a misalignment between optimization objectives and clinical rigor. Current Reinforcement Learning (RL) paradigms rely on lexical proxy signals that induce ``\textbf{evaluation hallucinati…

Cited by 0SourceScholar
2026

RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse

ICML 2026poster

The increasing complexity of AI tasks has shifted the paradigm from monolithic models toward multi-agent large language model (LLM) systems. However, these collaborative architectures introduce a critical bottleneck: redundant prefill computation for shared content generated by previous agents, whic…

Cited by 0SourceScholar
2026

TumorChain: Interleaved Multimodal Chain-of-Thought Reasoning for Traceable Clinical Tumor Analysis

ICLR 2026poster

Accurate tumor analysis is central to clinical radiology and precision oncology, where early detection, reliable lesion characterization, and pathology-level risk assessment directly guide diagnosis, staging, and treatment planning. Chain-of-Thought (CoT) reasoning is particularly critical in this s…

Cited by 0SourcecodeScholar
2026

VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking

CVPR 2026

Video agentic models have advanced challenging video-language tasks. However, most agentic approaches still heavily rely on greedy parsing over densely sampled video frames, resulting in high computational cost. We present VideoSeek, a long-horizon video agent that leverages video logic flow to acti

Cited by 3SourcecodeScholar
2026

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models

ICLR 2026poster

Omni-modal large language models (OLLMs) aim to unify audio, vision, and text understanding within a single framework. While existing benchmarks have advanced multimodal evaluation, it remains unclear whether OLLMs achieve modality-invariant reasoning or inherit modality-specific biases. We introduc…

Cited by 0SourceScholar
2026

Yours or Mine? Overwriting Attacks Against Neural Audio Watermarking

AAAI 2026technical

As generative audio models are rapidly evolving, AI-generated audios increasingly raise concerns about copyright infringement and misinformation spread. Audio watermarking, as a proactive defense, can embed secret messages into audio for copyright protection and source verification. However, current

Cited by 0SourcePDFScholar
2025

AIF-SFDA: Autonomous Information Filter Driven Source-Free Domain Adaptation for Medical Image Segmentation

AAAI 2025technical

Decoupling domain-variant information (DVI) from domain-invariant information (DII) serves as a prominent strategy for mitigating domain shifts in the practical implementation of deep learning algorithms. However, in medical settings, concerns surrounding data collection and privacy often restrict a…

2025

Agent Laboratory: Using LLM Agents as Research Assistants

EMNLP 2025

Historically, scientific discovery has been a lengthy and costly process, demanding substantial time and resources from initial conception to final results. To accelerate scientific discovery, reduce research costs, and improve research quality, we introduce Agent Laboratory, an autonomous LLM-based

Cited by 0SourcePDFScholar
2025

Align2LLaVA: Cascaded Human and Large Language Model Preference Alignment for Multi-modal Instruction Curation

ACL 2025finding

Recent advances in Multi-modal Large Language Models (MLLMs), such as LLaVA-series models, are driven by massive machine-generated instruction-following data tuning. Such automatic instruction collection pipelines, however, inadvertently introduce significant variability in data quality. This paper…

2025

Exploring Temporal Constraints for Unsupervised Iris Motion Tracking in AS-OCT Videos

ICASSP 2025accepted

Iris motion tracking is critical for discriminating the iris stiffness and developmental stage of primary angle-closure disease (PACD). Anterior segment optical coherence tomography (AS-OCT) video is a highly efficient approach to observe the morphological determinant in iris motion. However, the ir…

Cited by 0SourceScholar
2025

NCRE: A Benchmark for Document-level Nominal Compound Relation Extraction

COLING 2025main

Entity and relation extraction is a conventional task in the field of information extraction. Existing work primarily focuses on detecting specific relations between entities, often constrained to particular fields and lacking general applicability. In response, we propose a novel task: nominal comp…

2025

RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection

ACL 2025long

Large language models (LLMs) have demonstrated remarkable capabilities in various domains, including radiology report generation. Previous approaches have attempted to utilize multimodal LLMs for this task, enhancing their performance through the integration of domain-specific knowledge retrieval. H…

2025

Self-Taught Agentic Long Context Understanding

ACL 2025long

Answering complex, long-context questions remains a major challenge for large language models (LLMs) as it requires effective question clarifications and context retrieval. We propose Agentic Long-Context Understanding (AgenticLU), a framework designed to enhance an LLM’s understanding of such queri…

2025

SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer

CVPR 2025poster

Efficient image tokenization with high compression ratios remains a critical challenge for training generative models.We present SoftVQ-VAE, a continuous image tokenizer that leverages soft categorical posteriors to aggregate multiple codewords into each latent token, substantially increasing the re…

2025

TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games

EMNLP 2025

Large reasoning models (LRMs) have demonstrated impressive reasoning capabilities across a broad range of tasks including Olympiad-level mathematical problems, indicating evidence of their complex reasoning abilities. While many reasoning benchmarks focus on the STEM domain, the ability of LRMs to r

Cited by 0SourcePDFScholar
2025

TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition

ACL 2025long

While Parameter-Efficient Fine-Tuning (PEFT) methods like Low-Rank Adaptation (LoRA) effectively address resource constraints during fine-tuning, their performance often falls short, especially in multidimensional task scenarios. To address this issue, one straightforward solution is to introduce ta…

2025

Unleashing Hour-Scale Video Training for Long Video-Language Understanding

NeurIPS 2025spotlight

Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has left the training of hour-long Video-LMMs underexplored. To close this gap, we present VideoMarathon, a large-scale hou…

Cited by 0SourceScholar
2024

Accelerating Non-Maximum Suppression: A Graph Theory Perspective

NeurIPS 2024poster

Non-maximum suppression (NMS) is an indispensable post-processing step in object detection. With the continuous optimization of network models, NMS has become the ``last mile'' to enhance the efficiency of object detection. This paper systematically analyzes NMS from a graph theory perspective for t…

2024

Flattening Singular Values of Factorized Convolution for Medical Images

ICASSP 2024accepted

Convolutional neural networks (CNNs) have long been the paradigm of choice for robust medical image processing (MIP). Therefore, it is crucial to effectively and efficiently deploy CNNs on devices with different computing capabilities to support computer-aided diagnosis. Many methods employ factoriz…

Cited by 0SourceScholar
2024

ICON: Improving Inter-Report Consistency in Radiology Report Generation via Lesion-aware Mixup Augmentation

EMNLP 2024finding

Previous research on radiology report generation has made significant progress in terms of increasing the clinical accuracy of generated reports. In this paper, we emphasize another crucial quality that it should possess, i.e., inter-report consistency, which refers to the capability of generating c…

2024

Scale Optimization Using Evolutionary Reinforcement Learning for Object Detection on Drone Imagery

AAAI 2024technical

Object detection in aerial imagery presents a significant challenge due to large scale variations among objects. This paper proposes an evolutionary reinforcement learning agent, integrated within a coarse-to-fine object detection framework, to optimize the scale for more effective detection of obje…

2023

ORGAN: Observation-Guided Radiology Report Generation via Tree Reasoning

ACL 2023long

This paper explores the task of radiology report generation, which aims at generating free-text descriptions for a set of radiographs. One significant challenge of this task is how to correctly maintain the consistency between the images and the lengthy report. Previous research explored solving thi…

2023

Oct Image Blind Despeckling Based on Gradient Guided Filter with Speckle Statistical Prior

ICASSP 2023accepted

Optical coherence tomography (OCT) imaging technique has been widely used for ocular disease diagnosis. However, speckles occur in OCT images due to the property of coherent imaging, inevitably affecting the visual quality and clinical analysis. To alleviate this problem, we propose a novel gradient…

Cited by 0SourceScholar
2023

PolyFormer: Referring Image Segmentation As Sequential Polygon Generation

CVPR 2023poster

In this work, instead of directly predicting the pixel-level segmentation masks, the problem of referring image segmentation is formulated as sequential polygon generation, and the predicted polygons can be later converted into segmentation masks. This is enabled by a new sequence-to-sequence framew…

2023

RECAP: Towards Precise Radiology Report Generation via Dynamic Disease Progression Reasoning

EMNLP 2023long findings

Automating radiology report generation can significantly alleviate radiologists’ workloads. Previous research has primarily focused on realizing highly concise observations while neglecting the precise attributes that determine the severity of diseases (e.g., small pleural effusion). Since incorrect…

Cited by 37SourcecodeScholar
2022

SS3D: Sparsely-Supervised 3D Object Detection From Point Cloud

CVPR 2022poster

Conventional deep learning based methods for 3D object detection require a large amount of 3D bounding box annotations for training, which is expensive to obtain in practice. Sparsely annotated object detection, which can largely reduce the annotations, is very challenging since the missingannotated…

Cited by 28PDFcodeScholar
2022

Segment and Complete: Defending Object Detectors Against Adversarial Patch Attacks With Robust Patch Detection

CVPR 2022poster

Object detection plays a key role in many security-critical systems. Adversarial patch attacks, which are easy to implement in the physical world, pose a serious threat to state-of-the-art object detectors. Developing reliable defenses for object detectors against patch attacks is critical but sever…

Cited by 111PDFcodeScholar
2022

Spatial-Context-Aware Deep Neural Network for Multi-Class Image Classification

ICASSP 2022accepted

Multi-label image classification is a fundamental but challenging task in computer vision. Over the past few decades, solutions exploring relationships between semantic labels have made great progress. However, the underlying spatial-contextual information of labels is under-exploited. To tackle thi…

Cited by 0SourceScholar
2022

Unified Named Entity Recognition as Word-Word Relation Classification

AAAI 2022technical

So far, named entity recognition (NER) has been involved with three major types, including flat, overlapped (aka. nested), and discontinuous NER, which have mostly been studied individually. Recently, a growing interest has been built for unified NER, tackling the above three jobs concurrently with…

2021

Interference Analysis in Reconfigurable Intelligent Surface-Assisted Multiple-Input Multiple-Output Systems

ICASSP 2021accepted

Reconfigurable intelligent surfaces (RISs) are regarded as an emerging technology for the next generation of wireless communications. In this paper, we consider a multiple-input multiple-output network where each base station serves a user equipment with the aid of an RIS equipped with N reconfigura…

Cited by 0SourceScholar
2020

Encoding Structure-Texture Relation with P-Net for Anomaly Detection in Retinal Images

ECCV 2020poster

Anomaly detection in retinal image refers to the identification of abnormality caused by various retinal diseases/lesions, by only leveraging normal images in training phase. Normal images from healthy subjects often have regular structures (e.g., the structured blood vessels in the fundus image, or…

2019

Topology Reconstruction of Tree-Like Structure in Images via Structural Similarity Measure and Dominant Set Clustering

CVPR 2019poster

The reconstruction and analysis of tree-like topological structures in the biomedical images is crucial for biologists and surgeons to understand biomedical conditions and plan surgical procedures. The underlying tree-structure topology reveals how different curvilinear components are anatomically…

Cited by 13PDFScholar
2018

DecideNet: Counting Varying Density Crowds Through Attention Guided Detection and Density Estimation

CVPR 2018poster

In real-world crowd counting applications, the crowd densities vary greatly in spatial and temporal domains. A detection based counting method will estimate crowds accurately in low density scenes, while its reliability in congested areas is downgraded. A regression based approach, on the other hand…

Cited by 448SourcePDFScholar
2015

A Low-Dimensional Step Pattern Analysis Algorithm With Application to Multimodal Retinal Image Registration

CVPR 2015poster

Existing feature descriptor-based methods on retinal image registration are mainly based on scale-invariant feature transform (SIFT) or partial intensity invariant feature descriptor (PIIFD). While these descriptors are often being exploited, they do not work very well upon unhealthy multimodal imag…

Cited by 45SourcePDFScholar