← Search

Quyu Kong

12 accepted papers

2026

De-Linearizing Agent Traces: Bayesian Inference of Latent Partial Orders for Efficient Execution

ICML 2026poster

AI agents increasingly execute procedural workflows as sequential action traces, which obscures latent concurrency and induces repeated step-by-step reasoning. We introduce BPOP, a Bayesian framework that infers a latent dependency partial order from noisy linearized traces. BPOP models traces as st…

Cited by 0SourceScholar
2026

Fair Bayesian Data Selection via Generalized Discrepancy Measures

AAAI 2026technical

Fairness concerns are increasingly critical as machine learning models are deployed in high-stakes applications. While existing fairness-aware methods typically intervene at the model level, they often suffer from high computational costs, limited scalability, and poor generalization. To address the

Cited by 0SourcePDFScholar
2026

Long-range Modeling and Processing of Multimodal Event Sequences

ICLR 2026poster

Temporal point processes (TPPs) have emerged as powerful tools for modeling asynchronous event sequences. While recent advances have extended TPPs to handle textual information, existing approaches are limited in their ability to generate rich, multimodal content and reason about event dynamics. A k…

Cited by 0SourcecodeScholar
2026

Negative Binomial Variational Autoencoders for Overdispersed Latent Modeling

CVPR 2026

Although artificial neural networks are often described as brain-inspired, their representations typically rely on continuous activations, such as the continuous latent variables in variational autoencoders (VAEs), which limits their biological plausibility compared to the discrete spike-based signa

Cited by 0SourceScholar
2026

UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction as Reasoning

ICLR 2026poster

GUI grounding, which maps natural-language instructions to actionable UI elements, is a core capability of GUI agents. Prior work largely treats instructions as a static proxy for user intent, overlooking the impact of instruction diversity on grounding performance. Through a careful investigation o…

Cited by 0SourcecodeScholar
2025

DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding

NeurIPS 2025poster

We introduce DanmakuTPPBench, a comprehensive benchmark designed to advance multi-modal Temporal Point Process (TPP) modeling in the era of Large Language Models (LLMs). While TPPs have been widely studied for modeling temporal event sequences, existing datasets are predominantly unimodal, hinderin…

Cited by 0SourcecodeScholar
2025

Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions

CVPR 2025poster

Grounding 3D object affordance is a task that locates objects in 3D space where they can be manipulated, which links perception and action for embodied intelligence. For example, for an intelligent robot, it is necessary to accurately ground the affordance of an object and grasp it according to huma…

Cited by 0SourcePDFScholar
2025

TPP-SD: Accelerating Transformer Point Process Sampling with Speculative Decoding

NeurIPS 2025poster

We propose TPP-SD, a novel approach that accelerates Transformer temporal point process (TPP) sampling by adapting speculative decoding (SD) techniques from language models. By identifying the structural similarities between thinning algorithms for TPPs and speculative decoding for language models,…

Cited by 0SourceScholar
2024

Enhancing Closed-Loop Performance in Learning-Based Vehicle Motion Planning by Integrating Rule-Based Insights

RA-L 2024

This letter introduces an innovative vehicle motion planning method that leverages the integration of rule-based insights to significantly improve closed-loop performance within a learning-based framework. We first employ rule-based methods to heuristically search and generate a diverse set of traje

Cited by 2SourceScholar
2024

OTVIC: A Dataset with Online Transmission for Vehicle-to-Infrastructure Cooperative 3D Object Detection

IROS 2024poster

Vehicle-to-infrastructure cooperative 3D object detection (VIC3D) is a task that leverages both vehicle and roadside sensors to jointly perceive the surrounding environment. However, considering the high speed of vehicles, the real-time requirements, and the limitations of communication bandwidth, r…

Cited by 1SourceScholar
2023

Integration-free Training for Spatio-temporal Multimodal Covariate Deep Kernel Point Processes

NeurIPS 2023poster

In this study, we propose a novel deep spatio-temporal point process model, Deep Kernel Mixture Point Processes (DKMPP), that incorporates multimodal covariate information. DKMPP is an enhanced version of Deep Mixture Point Processes (DMPP), which uses a more flexible deep kernel to model complex re…

Cited by 9SourcePDFScholar