← Search

Kai He

19 accepted papers

2026

ChronoEdit: Towards Temporal Reasoning for In-Context Image Editing and World Simulation

ICLR 2026poster

Recent advances in large generative models have significantly advanced image editing and in-context image generation, yet a critical gap remains in ensuring physical consistency, where edited objects must remain coherent. This capability is especially vital for world simulation related tasks. In thi…

Cited by 0SourcecodeScholar
2026

DPsurv: Dual-Prototype Evidential Fusion for Uncertainty-Aware and Interpretable Whole Slide Image Survival Prediction

ICML 2026poster

Whole-slide images (WSIs) are widely used for cancer survival analysis because of their comprehensive histopathological information at both cellular and tissue levels, enabling quantitative, large-scale, and prognostically rich tumor feature analysis. However, most existing WSI survival analysis met…

Cited by 0SourceScholar
2026

MARS: Multi-Agent Adaptive Reasoning with Socratic Guidance for Automated Prompt Optimization

AAAI 2026technical

Large language models (LLMs) typically operate in a question-answering paradigm, where the quality of the input prompt critically affects the response. Automated Prompt Optimization (APO) aims to overcome the cognitive biases of manually crafted prompts and explore a broader prompt design space. How

Cited by 0SourcePDFScholar
2026

Recovering Coherent Affective Patterns: Addressing Modality Missing in Multimodal Sentiment Analysis

AAAI 2026technical

Multimodal sentiment analysis (MSA) seeks to decode human emotions by integrating heterogeneous modalities. However, real-world scenarios often involve missing or misaligned data due to sensor failures or transmission errors, leading to disrupted temporal dynamics and degraded cross-modal correlatio

Cited by 0SourcePDFScholar
2025

CTRL-D: Controllable Dynamic 3D Scene Editing with Personalized 2D Diffusion

CVPR 2025poster

Achieving controllable and consistent editing in dynamic 3D scenes remains a significant challenge. Previous work is largely constrained by its editing backbones, resulting in inconsistent edits and limited controllability. We propose to address this challenge using personalized diffusion models. In…

Cited by 0SourcePDFScholar
2025

Crab: A Novel Configurable Role-Playing LLM with Assessing Benchmark

ACL 2025long

This study introduces Crab, a novel Configurable Role-Playing (RP) LLM with Assessing Benchmark, which consists of Role-Centric Dataset Curation, Persona-Embodying LLM Construction, and Comprehensive Benchmark Creation for RP dialogue generation. Distinct from traditional RP models that employ only…

Cited by 0SourcePDFScholar
2025

DivScore: Zero-Shot Detection of LLM-Generated Text in Specialized Domains

EMNLP 2025

Detecting LLM-generated text in specialized and high-stakes domains like medicine and law is crucial for combating misinformation and ensuring authenticity. However, current zero-shot detectors, while effective on general text, often fail when applied to specialized content due to domain shift. We p

2025

GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images

NeurIPS 2025poster

While recent multimodal large language models (MLLMs) have advanced automated ECG interpretation, they still face two key limitations: (1) insufficient multimodal synergy between ECG time series and ECG images, and (2) limited explainability in linking diagnoses to granular waveform evidence. We int…

Cited by 0SourcecodeScholar
2025

LuxDiT: Lighting Estimation with Video Diffusion Transformer

NeurIPS 2025poster

Estimating scene lighting from a single image or video remains a longstanding challenge in computer vision and graphics. Learning-based approaches are constrained by the scarcity of ground-truth HDR environment maps, which are expensive to capture and limited in diversity. While recent generative mo…

Cited by 0SourceScholar
2025

Self-supervised Quantized Representation for Seamlessly Integrating Knowledge Graphs with Large Language Models

ACL 2025long

Due to the presence of the natural gap between Knowledge Graph (KG) structures and the natural language, the effective integration of holistic structural information of KGs with Large Language Models (LLMs) has emerged as a significant question. To this end, we propose a two-stage framework to learn…

Cited by 0SourcePDFScholar
2025

UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting

NeurIPS 2025spotlight

We address the challenge of relighting a single image or video, a task that demands precise scene intrinsic understanding and high-quality light transport synthesis. Existing end-to-end relighting models are often limited by the scarcity of paired multi-illumination data, restricting their ability t…

Cited by 0SourceScholar
2024

MetaPro 2.0: Computational Metaphor Processing on the Effectiveness of Anomalous Language Modeling

ACL 2024findings

Metaphor interpretation is a difficult task in natural language understanding. The development of relevant techniques in this domain is slow, mostly because of the lack of large annotated datasets and effective pre-trained language models (PLMs) for metaphor learning. Thus, we propose a large annota…

Cited by 21SourcePDFScholar
2023

A Novel Omnidirectional Swimming Robot With Articulated-Compliant Legs

RA-L 2023

Stability, adaptability, and maneuverability are critical performance indexes for underwater biomimetic robots, especially in narrow spaces. However, these aspects can sometimes be contradictory. This letter presents an omnidirectional swimming robot inspired by the whirligig beetle and designed to

Cited by 3SourceScholar
2023

COCA: COllaborative CAusal Regularization for Audio-Visual Question Answering

AAAI 2023technical

Audio-Visual Question Answering (AVQA) is a sophisticated QA task, which aims at answering textual questions over given video-audio pairs with comprehensive multimodal reasoning. Through detailed causal-graph analyses and careful inspections of their learning processes, we reveal that AVQA models ar…

Cited by 21SourcePDFScholar
2023

Neuro-Symbolic Sentiment Analysis with Dynamic Word Sense Disambiguation

EMNLP 2023long findings

Sentiment analysis is a task that highly depends on the understanding of word senses. Traditional neural network models are black boxes that represent word senses as vectors that are uninterpretable for humans. On the other hand, the application of Word Sense Disambiguation (WSD) systems in downstre…

Cited by 0SourceScholar
2023

Relightable Neural Human Assets From Multi-View Gradient Illuminations

CVPR 2023poster

Human modeling and relighting are two fundamental problems in computer vision and graphics, where high-quality datasets can largely facilitate related research. However, most existing human datasets only provide multi-view human images captured under the same illumination. Although valuable for mode…

2022

COPNER: Contrastive Learning with Prompt Guiding for Few-shot Named Entity Recognition

COLING 2022main

Distance metric learning has become a popular solution for few-shot Named Entity Recognition (NER). The typical setup aims to learn a similarity metric for measuring the semantic similarity between test samples and referents, where each referent represents an entity class. The effect of this setup m…

2022

Netrca: An Effective Network Fault Cause Localization Algorithm

ICASSP 2022accepted

Localizing the root cause of network faults is crucial to network operation and maintenance. However, due to the complicated network architectures and wireless environments, as well as limited labeled data, accurately localizing the true root cause is challenging. In this paper, we propose a novel a…

Cited by 0SourceScholar
2021

Learning Interpretable Decision Rule Sets: A Submodular Optimization Approach

NeurIPS 2021spotlight

Rule sets are highly interpretable logical models in which the predicates for decision are expressed in disjunctive normal form (DNF, OR-of-ANDs), or, equivalently, the overall model comprises an unordered collection of if-then decision rules. In this paper, we consider a submodular optimization bas…

Cited by 35SourcePDFScholar