← Search

Kang Li

23 accepted papers

2026

Evolution Strategies at the Hyperscale

ICML 2026poster

Evolution Strategies (ES) is a class of powerful black-box optimisation methods that are highly parallelisable and can handle non-differentiable and noisy objectives. However, naïve ES becomes prohibitively expensive at scale on GPUs due to the low arithmetic intensity of batched matrix multiplicati…

Cited by 0SourceScholar
2026

GHPT: Real-Time Relightable Gaussian Splatting using Hybrid Path Tracing

CVPR 2026

3D Gaussian splatting (3DGS) has emerged as a promising approach for high-fidelity 3D scene representation. However, relighting and composition of Gaussian splatting remain challenging because path tracing is not directly applicable. Existing relighting methods for Gaussian splatting typically adopt

Cited by 0SourcecodeScholar
2026

StoryTailor:A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual Narratives

CVPR 2026

Generating multi-frame, action-rich visual narratives without fine-tuning faces a threefold tension: action text faithfulness, subject identity fidelity, and cross frame background continuity. We propose StoryTailor, a zero-shot pipeline that runs on a single RTX 4090 (24 GB) and produces temporally

Cited by 0SourceScholar
2026

Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt Framework

CVPR 2026

To continuously enhance model adaptability in surgical video scene parsing, recent studies incrementally update it to progressively learn to segment an increasing number of surgical instruments over time. However, prior works constantly overlooked the potential of positive forward knowledge transfer

Cited by 0SourceScholar
2026

ViMo: A Generative Visual GUI World Model for App Agents

ICLR 2026poster

App agents, which autonomously operate mobile Apps through GUIs, have gained significant interest in real-world applications. Yet, they often struggle with long-horizon planning, failing to find the optimal actions for complex tasks with longer steps. To address this, world models are used to predic…

Cited by 0SourceScholar
2025

Guiding Medical Vision-Language Models with Diverse Visual Prompts: Framework Design and Comprehensive Exploration of Prompt Variations

NAACL 2025long

While mainstream vision-language models (VLMs) have advanced rapidly in understanding image-level information, they still lack the ability to focus on specific areas designated by humans. Rather, they typically rely on large volumes of high-quality image-text paired data to learn and generate poster…

Cited by 0SourcePDFScholar
2025

LOB-Bench: Benchmarking Generative AI for Finance - an Application to Limit Order Book Data

ICML 2025poster

While financial data presents one of the most challenging and interesting sequence modelling tasks due to high noise, heavy tails, and strategic interactions, progress in this area has been hindered by the lack of consensus on quantitative evaluation paradigms. To address this, we present **LOB-Ben…

2025

iDPA: Instance Decoupled Prompt Attention for Incremental Medical Object Detection

ICML 2025poster

Existing prompt-based approaches have demonstrated impressive performance in continual learning, leveraging pre-trained large-scale models for classification tasks; however, the tight coupling between foreground-background information and the coupled attention between prompts and image-text tokens p…

Cited by 0SourcePDFScholar
2024

CReStyler: Text-Guided Single Image Style Transfer Method Based on CNN and Restormer

ICASSP 2024accepted

Text-guided image style transfer methods have gradually become a research hotspot. However, existing text-guided style transfer method suffers from content information missing and artifacts in the generated stylized images. Therefore, we propose CReStyler, a text-guided image style method based on t…

Cited by 0SourceScholar
2024

Memory-Efficient Prompt Tuning for Incremental Histopathology Classification

AAAI 2024technical

Recent studies have made remarkable progress in histopathology classification. Based on current successes, contemporary works proposed to further upgrade the model towards a more generalizable and robust direction through incrementally learning from the sequentially delivered domains. Unlike previou…

Cited by 1SourcePDFScholar
2024

One-to-Normal: Anomaly Personalization for Few-shot Anomaly Detection

NeurIPS 2024poster

Traditional Anomaly Detection (AD) methods have predominantly relied on unsupervised learning from extensive normal data. Recent AD methods have evolved with the advent of large pre-trained vision-language models, enhancing few-shot anomaly detection capabilities. However, these latest AD methods st…

Cited by 0SourcePDFScholar
2023

AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer

ICASSP 2023accepted

In this paper, we propose an effective sound event detection (SED) method based on the audio spectrogram transformer (AST) model, pretrained on the large-scale AudioSet for audio tagging (AT) task, termed AST-SED. Pretrained AST models have recently shown promise on DCASE2022 challenge task4 where t…

Cited by 0SourceScholar
2023

Gender-Cartoon: Image Cartoonization Method Based on Gender Classification

ICASSP 2023accepted

Qin Opera art is one of China’s intangible cultural heritage, and its influence is gradually declining. The cartoonization of Qin Opera is one of the feasible methods. However, current cartoonization methods suffer from the inability to classify and accurately cartoonize Qinqiang portraits by gender…

Cited by 0SourceScholar
2023

MEDICAL IMAGE UNDERSTANDING WITH PRETRAINED VISION LANGUAGE MODELS: A COMPREHENSIVE STUDY

ICLR 2023poster

The large-scale pre-trained vision language models (VLM) have shown remarkable domain transfer capability on natural images. However, it remains unknown whether this capability can also apply to the medical image domain. This paper thoroughly studies the knowledge transferability of pre-trained VLMs…

2022

Depth Estimation Matters Most: Improving Per-Object Depth Estimation for Monocular 3D Detection and Tracking

ICRA 2022poster

Monocular image-based 3D perception has become an active research area in recent years owing to its applications in autonomous driving. Approaches to monocular 3D perception including detection and tracking, however, often yield inferior performance when compared to LiDAR-based techniques. Through s…

Cited by 24SourceScholar
2022

MMT: Multi-way Multi-modal Transformer for Multimodal Learning

IJCAI 2022poster

The heart of multimodal learning research lies the challenge of effectively exploiting fusion representations among multiple modalities.However, existing two-way cross-modality unidirectional attention could only exploit the intermodal interactions from one source to one target modality. This indeed…

Cited by 21SourcePDFScholar
2021

CTFN: Hierarchical Learning for Multimodal Sentiment Analysis Using Coupled-Translation Fusion Network

ACL 2021long

Multimodal sentiment analysis is the challenging research area that attends to the fusion of multiple heterogeneous modalities. The main challenge is the occurrence of some missing modalities during the multimodal fusion procedure. However, the existing techniques require all modalities as input, th…

2021

On Explainability of Graph Neural Networks via Subgraph Explorations

ICML 2021spotlight

We consider the problem of explaining the predictions of graph neural networks (GNNs), which otherwise are considered as black boxes. Existing methods invariably focus on explaining the importance of graph nodes or edges but ignore the substructures of graphs, which are more intuitive and human-inte…

2017

Motion evaluation of a modified multi-link robotic rat

IROS 2017poster

The interaction test between a robotic rat and living rat is considered as a possible way to quantitatively characterize the rat sociality. In such robot-rat interactions, the robot should be designed to fully replicate a real rat in terms of morphological and behavioral characteristics. To address…

Cited by 6SourceScholar