← Search

Kai Ye

21 accepted papers

2026

A Difference-in-Difference Approach to Detecting AI-Generated Images

CVPR 2026

Diffusion models are able to produce AI-generated images that are almost indistinguishable from real ones, raising concerns about their potential misuse and posing substantial challenges for detecting them. Many existing detectors rely on reconstruction error -- the difference between the input imag

Cited by 0SourcecodeScholar
2026

Learn-to-Distance: Distance Learning for Detecting LLM-Generated Text

ICLR 2026poster

Modern large language models (LLMs) such as GPT, Claude, and Gemini have transformed the way we learn, work, and communicate. Yet, their ability to produce highly human-like text raises serious concerns about misinformation and academic integrity, making it an urgent need for reliable algorithms to…

Cited by 0SourcecodeScholar
2026

REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot Alignment

ICML 2026poster

Knowledge-intensive Visual Question Answering (KI-VQA) frequently suffers from severe knowledge conflicts caused by the inherent limitations of open-domain retrieval. However, existing paradigms face critical limitations, including the lack of generalizable conflict detection and intra-model constra…

Cited by 0SourceScholar
2026

RIS-LAD: A Benchmark and Model for Referring Image Segmentation in Low-Altitude Drone Imagery

AAAI 2026technical

Referring Image Segmentation (RIS), which aims to segment specific objects based on natural language descriptions, plays an essential role in vision-language understanding. Despite its progress in remote sensing applications, RIS under Low-Altitude Drone (LAD) scenarios remains underexplored, as exi

Cited by 0SourcePDFScholar
2026

Scaling Agents via Continual Pre-training

ICLR 2026poster

Large language models (LLMs) have evolved into agentic systems capable of autonomous tool use and multi-step reasoning for complex problem-solving. However, post-training approaches building upon general-purpose foundation models consistently underperform in agentic tasks, particularly in open-sourc…

Cited by 0SourcecodeScholar
2026

S²Teacher: Step-by-step Teacher for Sparsely Annotated Oriented Object Detection

AAAI 2026technical

Although fully-supervised oriented object detection has made significant progress in remote sensing image understanding, it comes at the cost of labor-intensive annotation. Recent studies have explored weakly and semi-supervised learning to alleviate this burden. However, these methods overlook the

Cited by 0SourcePDFScholar
2026

The Less You Depend, The More You Learn: Synthesizing Novel Views from Sparse, Unposed Images without Any 3D Knowledge

ICLR 2026poster

Recent advances in feed-forward Novel View Synthesis (NVS) have led to a divergence between two design philosophies: bias-driven methods, which rely on explicit 3D knowledge, such as handcrafted 3D representations (e.g., NeRF and 3DGS) and camera poses annotated by Structure-from-Motion algorithms,…

Cited by 0SourceScholar
2025

AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees

NeurIPS 2025poster

We study the problem of determining whether a piece of text has been authored by a human or by a large language model (LLM). Existing state of the art logits-based detectors make use of statistics derived from the log-probability of the observed text evaluated using the distribution function of a gi…

Cited by 0SourcecodeScholar
2025

CAKE: Category Aware Knowledge Extraction for Open-Vocabulary Object Detection

AAAI 2025technical

Open vocabulary object detection (OVOD) task aims to detect objects of novel categories beyond the base categories in the training set. To this end, the detector needs to access image-text pairs containing rich semantic information or the visual language pre-trained model (VLM) learned on them. Rece…

Cited by 0SourcePDFScholar
2025

Doubly Robust Alignment for Large Language Models

NeurIPS 2025poster

This paper studies reinforcement learning from human feedback (RLHF) for aligning large language models with human preferences. While RLHF has demonstrated promising results, many algorithms are highly sensitive to misspecifications in the underlying preference model (e.g., the Bradley-Terry model),…

Cited by 0SourcecodeScholar
2025

E2E-VGuard: Adversarial Prevention for Production LLM-based End-To-End Speech Synthesis

NeurIPS 2025poster

Recent advancements in speech synthesis technology have enriched our daily lives, with high-quality and human-like audio widely adopted across real-world applications. However, malicious exploitation like voice-cloning fraud poses severe security risks. Existing defense techniques struggle to addres…

Cited by 0SourcecodeScholar
2025

GeoSplatting: Towards Geometry Guided Gaussian Splatting for Physically-based Inverse Rendering

ICCV 2025poster

Recent 3D Gaussian Splatting (3DGS) representations have demonstrated remarkable performance in novel view synthesis; further, material-lighting disentanglement on 3DGS warrants relighting capabilities and its adaptability to broader applications. While the general approach to the latter operation l…

Cited by 0SourcePDFScholar
2025

STINR: Deciphering Spatial Transcriptomics via Implicit Neural Representation

CVPR 2025poster

Spatial transcriptomics (ST) are emerging technologies that reveal spatial distributions of gene expressions within tissues, serving as important ways to uncover biological insights. However, the irregular spatial profiles and variability of genes make it challenging to integrate spatial information…

2024

An Interpretable Approach to the Solutions of High-Dimensional Partial Differential Equations

AAAI 2024technical

In recent years, machine learning algorithms, especially deep learning, have shown promising prospects in solving Partial Differential Equations (PDEs). However, as the dimension increases, the relationship and interaction between variables become more complex, and existing methods are difficult to…

2023

Learning Visual Prior via Generative Pre-Training

NeurIPS 2023poster

Various stuff and things in visual data possess specific traits, which can be learned by deep neural networks and are implicitly represented as the visual prior, e.g., object location and shape, in the model. Such prior potentially impacts many vision tasks. For example, in conditional image synthes…

2022

CLIMS: Cross Language Image Matching for Weakly Supervised Semantic Segmentation

CVPR 2022poster

It has been widely known that CAM (Class Activation Map) usually only activates discriminative object regions and falsely includes lots of object-related backgrounds. As only a fixed set of image-level object labels are available to the WSSS (weakly supervised semantic segmentation) model, it could…

Cited by 183PDFcodeScholar
2022

Multi-Robot Active Mapping via Neural Bipartite Graph Matching

CVPR 2022poster

We study the problem of multi-robot active mapping, which aims for complete scene map construction in minimum time steps. The key to this problem lies in the goal position estimation to enable more efficient robot movements. Previous approaches either choose the frontier as the goal position via a m…

Cited by 34PDFScholar