← Search

Kai Guo

22 accepted papers

2026

Consistent Text-to-Image Generation via Scene De-Contextualization

ICLR 2026poster

Consistent text-to-image (T2I) generation seeks to produce identity-preserving images of the same subject across diverse scenes, yet it often fails due to a phenomenon called identity (ID) shift. Previous methods have tackled this issue, but typically rely on the unrealistic assumption of knowing al…

Cited by 0SourcecodeScholar
2026

Fix Before Search: Benchmarking Agentic Visual Query Pre-processing in Multimodal Retrieval-augmented Generation

ICML 2026poster

Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a key paradigm for grounding MLLMs with external knowledge. While query pre-processing (e.g., rewriting) is standard in text-based RAG, existing MRAG pipelines predominantly treat visual inputs as static and immutable, implicitly assumi…

Cited by 0SourceScholar
2026

Revealing the Invisible: Latent Structure Modeling for Semantically Consistent Cloud Removal

AAAI 2026technical

Cloud removal (CR) in remote sensing imagery is a critical yet challenging task due to complex cloud patterns and diverse underlying ground structures. Despite recent progress in generative models such as diffusion models, CR remains limited by their inadequate capability to perceive and reconstruct

Cited by 0SourcePDFScholar
2026

TerAdapt: Proprioceptive Terrain-Adaptive Locomotion via Codebook Aligned Representation Learning

RA-L 2026

Humanoid robots aim to achieve human-like locomotion in unstructured environments. However, designing a controller for such robots is highly challenging due to their inherent instability and the requirement to adapt to diverse terrains. To address this problem, we present TerAdapt, a proprioceptive

Cited by 0SourceScholar
2026

VPIES: Variational Privileged Information Encoder as Scaffold for Legged Locomotion Learning

RA-L 2026

Legged robots face significant challenges in complex terrains due to partial observability. While teacher-student frameworks address this through imitation, they often cause representation mismatch and covariate shift, limiting deployment robustness. To address these limitations, we propose the Vari

Cited by 0SourceScholar
2026

When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression

ICML 2026poster

Reasoning hallucinations in large language models (LLMs) often appear as fluent yet unsupported conclusions that violate either the given context or underlying factual knowledge. Although such failures are widely observed, the mechanisms by which decoder-only Transformers produce them remain poorly …

Cited by 0SourceScholar
2025

A Practical Gated Recurrent Transformer Network Incorporating Multiple Fusions for Video Denoising

ICASSP 2025accepted

State-of-the-art (SOTA) video denoising methods employ multi-frame simultaneous denoising mechanisms, resulting in significant delays (e.g., 16 frames), making them impractical for real-time cameras. To overcome this limitation, we propose a multi-fusion gated recurrent Transformer network (GRTN) th…

Cited by 0SourceScholar
2025

Disentangling Multi-view Representations via Curriculum Learning with Learnable Prior

IJCAI 2025

Multi-view representation learning methods typically follow a consistent-and-specific pipeline that aims at extracting latent representations for an entity from its multiple observable views to facilitate downstream tasks. However, most of them overlook the complex underlying correlation between dif

2025

Empowering GraphRAG with Knowledge Filtering and Integration

EMNLP 2025

In recent years, large language models (LLMs) have revolutionized the field of natural language processing. However, they often suffer from knowledge gaps and hallucinations. Graph retrieval-augmented generation (GraphRAG) enhances LLM reasoning by integrating structured knowledge from external grap

Cited by 0SourcePDFScholar
2025

From Sequence to Structure: Uncovering Substructure Reasoning in Transformers

NeurIPS 2025poster

Recent studies suggest that large language models (LLMs) possess the capability to solve graph reasoning tasks. Notably, even when graph structures are embedded within textual descriptions, LLMs can still effectively answer related questions. This raises a fundamental question: How can a decoder-onl…

Cited by 0SourceScholar
2025

Improving Contextual ASR with Enhanced Phrase-Level Representation Based on MCTC Loss

ICASSP 2025accepted

Contextual biasing is essential for addressing scenario-specific challenges in End-to-End (E2E) Automatic Speech Recognition (ASR) systems. Prior contextual E2E ASR methods, such as the contextual bias with CPP Network, have utilized bias CTC loss for explicit supervision of bias tasks, However, the…

Cited by 0SourceScholar
2025

Learning Robust Multi-view Representation Using Dual-masked VAEs

IJCAI 2025

Most existing multi-view representation learning methods assume view-completeness and noise-free data. However, such assumptions are strong in real-world applications. Despite advances in methods tailored to view-missing or noise problems individually, a one-size-fits-all approach that concurrently

2025

Token-Level Contextual Network with Ladder-Shaped Attention for End-to-End ASR

ICASSP 2025accepted

Contextual automatic speech recognition (ASR) plays an increasingly important role in addressing the long-tail issues of general ASR. In the past, contextual ASR mainly focused on phrase-level discussions, providing a convenient way to handle biasing phrases. This paper introduces a new contextual n…

Cited by 0SourceScholar
2025

Towards Context-Robust LLMs: A Gated Representation Fine-tuning Approach

ACL 2025long

Large Language Models (LLMs) enhanced with external contexts, such as through retrieval-augmented generation (RAG), often face challenges in handling imperfect evidence. They tend to over-rely on external knowledge, making them vulnerable to misleading and unhelpful contexts. To address this, we pro…

Cited by 0SourcePDFScholar
2024

Analysis and Verification on Backstepping Control of Antagonistic Variable Stiffness Actuators

RA-L 2024

Variable stiffness actuators (VSAs) are an important class of compliant actuators that can benefit the robustness and task adaptability of robotic systems. However, controlling VSAs is still challenging as VSAs are difficult to model exactly and have highly nonlinear characteristics. This study appl

Cited by 3SourceScholar
2024

Biomimetic Robotic Remora With Hitchhiking Ability: Design, Control and Experiment

RA-L 2024

Remora, which is well known for it's ’hitchhiking’ behavior, can attach to diverse marine animals and travel with them for a long distance with low energy consumption due to its special disc. In this letter, inspired by the unique ’hitchhiking’ behavior, a new prototype of a robotic remora with good

Cited by 2SourceScholar
2023

Design and Development of a Rapidly Deployable Low-Cost Tensegrity In-Pipe Robot

IROS 2023poster

Existing in-pipe robots have insufficient adaptability when dealing with accidents in unfamiliar pipe environments. Developing a pipe robot that can be designed and manufactured quickly is one solution. The tensegrity structure is a self-stressing spatial structure formed by the interaction of rigid…

Cited by 0SourceScholar
2023

Multi-Head Uncertainty Inference for Adversarial Attack Detection

ICASSP 2023accepted

Deep neural networks (DNNs) are sensitive and susceptible to tiny perturbations by adversarial attacks which cause erroneous predictions. Various methods, including adversarial defense and uncertainty inference (UI), have been developed to overcome adversarial attacks in recent years. In this paper,…

Cited by 0SourceScholar
2023

Out-of-Distribution Generalization of Federated Learning via Implicit Invariant Relationships

ICML 2023poster

Out-of-distribution generalization is challenging for non-participating clients of federated learning under distribution shifts. A proven strategy is to explore those invariant relationships between input and target variables, working equally well for non-participating clients. However, learning inv…

Cited by 35SourcePDFScholar
2020

Non-Uniform Video Time-Lapse Method Based on Motion Scenario and Stabilization Constraint

ICASSP 2020accepted

Time-lapse of user captured video becomes popular in many applications recently, non-uniform sampling and digital video stabilization (VS) are usually two independent steps to keep meaningful contents and provide stabilized output. However, non-uniform sampling may produce large sampling interval an…

Cited by 0SourceScholar
2020

Robust Full-Fov Depth Estimation in Tele-Wide Camera System

ICASSP 2020accepted

Tele-wide camera system with different Field of View (FoV) lenses becomes very popular in recent mobile devices. Usually it is difficult to obtain full-FoV depth based on traditional stereo-matching methods. Pure Deep Neural Network (DNN) based depth estimation methods can obtain full-FoV depth, but…

Cited by 0SourceScholar