← Search

Tian Xia

26 accepted papers

2026

BuildArena: A Physics‑Aligned Interactive Benchmark of LLMs for Engineering Construction

ICML 2026poster

Engineering construction automation aims to transform natural language specifications into physically viable structures, requiring complex integrated reasoning under strict physical constraints. While modern LLMs possess broad knowledge and strong reasoning capabilities that make them promising cand…

Cited by 0SourceScholar
2026

Factored Classifier-Free Guidance

ICML 2026poster

Counterfactual generation aims to simulate realistic hypothetical outcomes under causal interventions. Diffusion models have emerged as a powerful tool for this task, combining DDIM inversion with conditional generation and classifier-free guidance (CFG). In this work, we identify a key limitation o…

Cited by 0SourceScholar
2026

FrontierCS: Evolving Challenges for Evolving Intelligence

ICML 2026poster

We introduce FrontierCS, a benchmark of 240 open-ended problems across diverse areas of computer science, designed and reviewed by experts, including CS PhDs and top-tier competitive programming participants and problem setters. Unlike existing benchmarks that focus on tasks with known optimal solut…

Cited by 0SourceScholar
2026

Phys-Liquid: A Physics-Informed Dataset for Estimating 3D Geometry and Volume of Transparent Deformable Liquids

AAAI 2026technical

Estimating the geometric and volumetric properties of transparent deformable liquids is challenging due to optical complexities and dynamic surface deformations induced by container movements. Autonomous robots performing precise liquid manipulation tasks—such as dispensing, aspiration, and mixing—m

Cited by 0SourcePDFScholar
2025

Diffusion Counterfactual Generation with Semantic Abduction

ICML 2025poster

Counterfactual image generation presents significant challenges, including preserving identity, maintaining perceptual quality, and ensuring faithfulness to an underlying causal model. While existing auto-encoding frameworks admit semantic latent spaces which can be manipulated for causal control, t…

2025

Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoning

EMNLP 2025

Recent advances in Large Language Models (LLMs) have enabled them to process increasingly longer sequences, ranging from 2K to 2M tokens and even beyond. However, simply extending the input sequence length does not necessarily lead to effective long-context understanding. In this study, we integrate

2025

GeGS-PCR: Fast and Robust Color 3D Point Cloud Registration with Two-Stage Geometric-3DGS Fusion

NeurIPS 2025poster

We address the challenge of point cloud registration using color information, where traditional methods relying solely on geometric features often struggle in low-overlap and incomplete scenarios. To overcome these limitations, we propose GeGS-PCR, a novel two-stage method that combines geometric, c…

Cited by 0SourceScholar
2025

IDE: A Multi-Agent-Driven Iterative Framework for Dynamic Evaluation of LLMs

ICASSP 2025accepted

With the widespread use of large language models (LLMs) in natural language processing, traditional evaluation methods based on static datasets have become inadequate to fully capture their performance and generalization capabilities. To address this challenge, we propose an Iterative Dynamic Evalua…

Cited by 0SourceScholar
2025

MiniMal: Hard-Label Adversarial Attack Against Static Malware Detection with Minimal Perturbation

IJCAI 2025

Static malware detectors based on machine learning are integral to contemporary antivirus systems, but they are vulnerable to adversarial attacks. While existing research has demonstrated success with adversarial attacks in black-box hard-label scenarios, challenges such as high perturbation rates a

2025

PlanGenLLMs: A Modern Survey of LLM Planning Capabilities

ACL 2025long

LLMs have immense potential for generating plans, transforming an initial world state into a desired goal state. A large body of research has explored the use of LLMs for various planning tasks, from web navigation to travel planning and database querying. However, many of these systems are tailored…

Cited by 0SourcePDFScholar
2024

Measuring Bargaining Abilities of LLMs: A Benchmark and A Buyer-Enhancement Method

ACL 2024findings

Bargaining is an important and unique part of negotiation between humans. As LLM-driven agents learn to negotiate and act like real humans, how to evaluate agents’ bargaining abilities remains an open problem.For the first time, we formally described the Bargaining task as an asymmetric incomplete i…

2024

R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

EMNLP 2024finding

Large language models (LLMs) have exhibited great potential in autonomously completing tasks across real-world applications. Despite this, these LLM agents introduce unexpected safety risks when operating in interactive environments. Instead of centering on the harmlessness of LLM-generated content…

2023

Constrained Market Share Maximization by Signal-Guided Optimization

AAAI 2023technical

With the rapid development of the airline industry, maximizing the market share with a constrained budget is an urgent econometric problem for an airline. We investigate the problem by adjusting flight frequencies on different flight routes. Owing to the large search space of solutions and the diffi…

2023

High Fidelity Image Counterfactuals with Probabilistic Causal Models

ICML 2023poster

We present a general causal generative modelling framework for accurate estimation of high fidelity image counterfactuals with deep structural causal models. Estimation of interventional and counterfactual queries for high-dimensional structured variables, such as images, remains a challenging task.…

2021

A Neural Transition-based Joint Model for Disease Named Entity Recognition and Normalization

ACL 2021long

Disease is one of the fundamental entities in biomedical research. Recognizing such entities from biomedical text and then normalizing them to a standardized disease vocabulary offer a tremendous opportunity for many downstream applications. Previous studies have demonstrated that joint modeling of…

Cited by 19SourcePDFScholar
2020

Aligntts: Efficient Feed-Forward Text-to-Speech System Without Explicit Alignment

ICASSP 2020accepted

Targeting at both high efficiency and performance, we propose AlignTTS to predict the mel-spectrum in parallel. AlignTTS is based on a Feed-Forward Transformer which generates mel-spectrum from a sequence of characters, and the duration of each character is determined by a duration predictor. Instea…

Cited by 0SourceScholar
2020

Learning Multi-Robot Decentralized Macro-Action-Based Policies via a Centralized Q-Net

ICRA 2020poster

In many real-world multi-robot tasks, high-quality solutions often require a team of robots to perform asynchronous actions under decentralized control. Decentralized multi-agent reinforcement learning methods have difficulty learning decentralized policies because of the environment appearing to be…

Cited by 39SourceScholar
2019

Adversarial Discrete Sequence Generation without Explicit NeuralNetworks as Discriminators

AISTATS 2019poster

This paper presents a novel approach to train GANs for discrete sequence generation without resorting to an explicit neural network as the discriminator. We show that when an alternative mini-max optimization procedure is performed for the value function where a closed form solution for the discrimi…

2015

Learning From Massive Noisy Labeled Data for Image Classification

CVPR 2015poster

Large-scale supervised datasets are crucial to train convolutional neural networks (CNNs) for various computer vision problems. However, obtaining a massive amount of well-labeled data is usually very expensive and time consuming. In this paper, we introduce a general framework to train CNNs with on…

Cited by 1493SourcePDFScholar