← Search

YAO ZHANG

61 accepted papers

2026

On the Evaluation of Capability Estimation Methods for Large Language Models

AAAI 2026technical

The emergence of large language models (LLMs) marks a transformative era in artificial intelligence~(AI). However, systematically evaluating the capability of LLMs is challenging due to the necessity of a large number of labeled test data. To tackle this problem, in the conventional AI field, AutoEv

Cited by 0SourcePDFScholar
2026

ProCURE: Addressing the Programming Concept Understanding Gap for Code Generation in LLMs via Concept-Aware Consistency Learning

IJCAI 2026

Although Large Language Models (LLMs) excel at code generation, recent research reveals that they exhibit an insufficient grasp of core programming concepts, such as data flow and control flow. This limitation undermines their robustness when encountering variations in these concepts in practice; ho

Cited by 0Scholar
2026

ProactiveLLM: Learning Active Interaction for Streaming Large Language Models

ICML 2026poster

Standard Large Language Models (LLMs) operate on a ''read-then-generate'' paradigm, incurring avoidable latency and computational redundancy. Recently, streaming LLMs have attempted to overcome these bottlenecks by allowing input and output to unfold synchronously. However, this introduces a critica…

Cited by 0SourceScholar
2026

T-SKM-Net: Trainable Neural Network Framework for Linear Constraint Satisfaction via Sampling Kaczmarz-Motzkin Method

AAAI 2026technical

Neural network constraint satisfaction is crucial for safety-critical applications such as power system optimization, robotic path planning, and autonomous driving. However, existing constraint satisfaction methods face efficiency-applicability trade-offs, with hard constraint methods suffering from

Cited by 0SourcePDFScholar
2026

Task-Specific Distance Correlation Matching for Few-Shot Action Recognition

AAAI 2026technical

Few-shot action recognition (FSAR) has recently made notable progress through set matching and efficient adaptation of large-scale pre-trained models. However, two key limitations persist. First, existing set matching metrics typically rely on cosine similarity to measure inter-frame linear dependen

Cited by 0SourcePDFScholar
2026

VIL2C: Value-of-Information Aware Low-Latency Communication for Multi-Agent Reinforcement Learning

AAAI 2026technical

Inter-agent communication serves as an effective mechanism for enhancing performance in collaborative multi-agent reinforcement learning (MARL) systems. However, the inherent communication latency in practical systems induces both action decision delays and outdated information sharing, impeding MAR

Cited by 0SourcePDFScholar
2026

WebArbiter: A Generative Reasoning Process Reward Model for Web Agents

ICLR 2026poster

Web agents hold great potential for automating complex computer tasks, yet their interactions involve long horizons, multi-step decisions, and actions that can be irreversible. In such settings, outcome-based supervision is sparse and delayed, often rewarding incorrect trajectories and failing to su…

Cited by 0SourceScholar
2025

DoGA: Enhancing Grounded Object Detection via Grouped Pre-Training with Attributes

AAAI 2025technical

Recent advances in vision-language pre-training have significantly enhanced the model capabilities on grounded object detection. However, these studies often pre-train with coarse-grained text prompts, such as plain category names and brief grounded phrases. This limitation curtails the model's capa…

2025

FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models

CVPR 2025poster

One-Shot Federated Learning (OSFL), a special decentralized machine learning paradigm, has recently gained significant attention. OSFL requires only a single round of client data or model upload, which reduces communication costs and mitigates privacy threats compared to traditional FL. Despite thes…

2025

Generating Questions, Answers, and Distractors for Videos: Exploring Semantic Uncertainty of Object Motions

ACL 2025finding

Video Question-Answer-Distractors (QADs) show promising values for assessing the performance of systems in perceiving and comprehending multimedia content. Given the significant cost and labor demands of manual annotation, existing large-scale Video QADs benchmarks are typically generated automatica…

Cited by 0SourcePDFScholar
2025

Incentives for Early Arrival in Cooperative Games (Extended Abstract)

IJCAI 2025

We study cooperative games where players join sequentially, and the value generated by those who have joined at any point must be irrevocably divided among these players. We introduce two desiderata for the value division mechanism: that the players should have incentives to join as early as possibl

Cited by 0SourcePDFScholar
2025

Listening to Patients: Detecting and Mitigating Patient Misreport in Medical Dialogue System

ACL 2025finding

Medical Dialogue Systems (MDSs) have emerged as promising tools for automated healthcare support through patient-agent interactions. Previous efforts typically relied on an idealized assumption — patients can accurately report symptoms aligned with their actual health conditions. However, in reality…

Cited by 0SourcePDFScholar
2025

Portcullis: A Scalable and Verifiable Privacy Gateway for Third-Party LLM Inference

AAAI 2025technical

Businesses using third-party LLMs face privacy risks from exposed prompts. This paper presents Portcullis, a privacy-preserving gateway that safeguards sensitive data while supporting efficient and accurate LLM responses. Portcullis functions as a mediator, anonymizing sensitive data in prompts thro…

Cited by 0SourcePDFScholar
2025

SCOP: Evaluating the Comprehension Process of Large Language Models from a Cognitive View

ACL 2025long

Despite the great potential of large language models (LLMs) in machine comprehension, it is still disturbing to fully count on them in real-world scenarios. This is probably because there is no rational explanation for whether the comprehension process of LLMs is aligned with that of experts. In thi…

2025

SwarmAgentic: Towards Fully Automated Agentic System Generation via Swarm Intelligence

EMNLP 2025

The rapid progress of Large Language Models has advanced agentic systems in decision-making, coordination, and task execution. Yet, existing agentic system generation frameworks lack full autonomy, missing from-scratch agent generation, self-optimizing agent functionality, and collaboration, limitin

Cited by 0SourcePDFScholar
2025

WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration

AAAI 2025technical

LLM-based autonomous agents often fail to execute complex web tasks that require dynamic interaction, largely due to the inherent uncertainty and complexity of these environments. Existing LLM-based web agents typically rely on rigid, expert-designed policies specific to certain states and actions,…

Cited by 21SourcePDFScholar
2024

A Lightweight Powered Knee Prosthesis Replicating Early-Stance Knee Flexion During Level Walking

RA-L 2024

Powered knee prostheses promise to improve the mobility of transfemoral amputees by imitating the biomechanics of the missing knee joint. Unfortunately, the heavy weight and short battery life severely limit the application of powered prostheses. Here, we present a lightweight powered knee prosthesi

Cited by 2SourceScholar
2024

Can We Learn Question, Answer, and Distractors All from an Image? A New Task for Multiple-choice Visual Question Answering

COLING 2024main

Multiple-choice visual question answering (MC VQA) requires an answer picked from a list of distractors, based on a question and an image. This research has attracted wide interest from the fields of visual question answering, visual question generation, and visual distractor generation. However, th…

Cited by 4SourcePDFScholar
2024

DESectBot: Design and Validation of a Novel Two-Segment Decoupled Continuum Robotic System for Endoscopic Submucosal Dissection

IROS 2024poster

Endoscopic Submucosal Dissection (ESD) is a minimally invasive procedure designed to remove precancerous and cancerous lesions from the gastrointestinal (GI) tract. Given the GI tract’s tortuous and narrow shape, along with the need for varied movements during dissection, this requires highly flexib…

Cited by 0SourceScholar
2024

Exploring Union and Intersection of Visual Regions for Generating Questions, Answers, and Distractors

EMNLP 2024main

Multiple-choice visual question answering (VQA) is to automatically choose a correct answer from a set of choices after reading an image. Existing efforts have been devoted to a separate generation of an image-related question, a correct answer, or challenge distractors. By contrast, we turn to a ho…

2024

FedDAT: An Approach for Foundation Model Finetuning in Multi-Modal Heterogeneous Federated Learning

AAAI 2024technical

Recently, foundation models have exhibited remarkable advancements in multi-modal learning. These models, equipped with millions (or billions) of parameters, typically require a substantial amount of data for finetuning. However, collecting and centralizing training data from diverse sectors becomes…

2024

GroupCover: A Secure, Efficient and Scalable Inference Framework for On-device Model Protection based on TEEs

ICML 2024poster

Due to the high cost of training DNN models, how to protect the intellectual property of DNN models, especially when the models are deployed to users' devices, is becoming an important topic. One practical solution is to use Trusted Execution Environments (TEEs) and researchers have proposed various…

Cited by 2SourcePDFScholar
2023

Adaptive Structure Induction for Aspect-based Sentiment Analysis with Spectral Perspective

EMNLP 2023long findings

Recently, incorporating structure information (e.g. dependency syntactic tree) can enhance the performance of aspect-based sentiment analysis (ABSA). However, this structure information is obtained from off-the-shelf parsers, which is often sub-optimal and cumbersome. Thus, automatically learning ad…

Cited by 0SourceScholar
2023

ECOLA: Enhancing Temporal Knowledge Embeddings with Contextualized Language Representations

ACL 2023findings

Since conventional knowledge embedding models cannot take full advantage of the abundant textual information, there have been extensive research efforts in enhancing knowledge embedding using texts. However, existing enhancement approaches cannot apply to temporal knowledge graphs (tKGs), which cont…

2023

Explaining Temporal Graph Models through an Explorer-Navigator Framework

ICLR 2023poster

While GNN explanation has recently received significant attention, existing works are consistently designed for static graphs. Due to the prevalence of temporal graphs, many temporal graph models have been proposed, but explaining their predictions remains to be explored. To bridge the gap, in this…

Cited by 20SourcePDFScholar
2023

KeFVP: Knowledge-enhanced Financial Volatility Prediction

EMNLP 2023long findings

Financial volatility prediction is vital for indicating a company's risk profile. Transcripts of companies' earnings calls are important unstructured data sources to be utilized to access companies' performance and risk profiles. However, current works ignore the role of financial metrics knowledge…

Cited by 0SourceScholar
2023

Learning How to Learn Domain-Invariant Parameters for Domain Generalization

ICASSP 2023accepted

Due to domain shift, deep neural networks (DNNs) usually fail to generalize well on unknown test data in practice. Domain generalization (DG) aims to overcome this issue by capturing domain-invariant representations from source domains. Motivated by the insight that only partial parameters of DNNs a…

Cited by 0SourceScholar
2023

SAP-DETR: Bridging the Gap Between Salient Points and Queries-Based Transformer Detector for Fast Model Convergency

CVPR 2023poster

Recently, the dominant DETR-based approaches apply central-concept spatial prior to accelerating Transformer detector convergency. These methods gradually refine the reference points to the center of target objects and imbue object queries with the updated central reference information for spatially…

2023

Task Allocation on Networks with Execution Uncertainty (Extended Abstract)∗

IJCAI 2023poster

We study a single task allocation problem where each worker connects to some other workers to form a network and the task requester only connects to some of the workers. The goal is to design an allocation mechanism such that each worker is incentivized to invite her neighbours to join the allocatio…

Cited by 0SourcePDFScholar
2023

Well Begun is Half Done: Generator-agnostic Knowledge Pre-Selection for Knowledge-Grounded Dialogue

EMNLP 2023long main

Accurate knowledge selection is critical in knowledge-grounded dialogue systems. Towards a closer look at it, we offer a novel perspective to organize existing literature, i.e., knowledge selection coupled with, after, and before generation. We focus on the third under-explored category of study,…

Cited by 0SourcecodeScholar
2022

Composition-based Heterogeneous Graph Multi-channel Attention Network for Multi-aspect Multi-sentiment Classification

COLING 2022main

Aspect-based sentiment analysis (ABSA) has drawn more and more attention because of its extensive applications. However, towards the sentence carried with more than one aspect, most existing works generate an aspect-specific sentence representation for each aspect term to predict sentiment polarity,…

2022

Deep-Learning-Based Compliant Motion Control of a Pneumatically-Driven Robotic Catheter

RA-L 2022

In cardiovascular interventions, when steering catheters and especially robotic catheters, great care should be paid to prevent applying too large forces on the vessel walls as this could dislodge calcifications, induce scars or even cause perforation. To address this challenge, this paper presents

Cited by 39SourceScholar
2022

Design and Validation of a Polycentric Hybrid Knee Prosthesis With Electromagnet-Controlled Mode Transition

RA-L 2022

A hybrid knee prosthesis is proposed in this letter, which consists of a polycentric structure in passive mode for low-torque activities and a single-axis structure in active mode for high-torque activities. A novel mode transition mechanism controls self-holding electromagnets for switching modes b

Cited by 5SourceScholar
2022

Diffusion Incentives in Cooperative Games

IJCAI 2022poster

We study a cooperative game setting where we want to gather more players through their social connections. Social connections can be modeled as a graph, and initially, only a subset of the players are in the game. We want to introduce diffusion incentives in such a cooperative game, i.e., incentiviz…

Cited by 0SourcePDFScholar
2022

Fact-Tree Reasoning for N-ary Question Answering over Knowledge Graphs

ACL 2022findings

Current Question Answering over Knowledge Graphs (KGQA) task mainly focuses on performing answer reasoning upon KGs with binary facts. However, it neglects the n-ary facts, which contain more than two entities. In this work, we highlight a more challenging but under-explored task: n-ary KGQA, i.e.,…

Cited by 8SourcePDFScholar
2022

Identifiable Energy-based Representations: An Application to Estimating Heterogeneous Causal Effects

AISTATS 2022poster

Conditional average treatment effects (CATEs) allow us to understand the effect heterogeneity across a large population of individuals. However, typical CATE learners assume all confounding variables are measured in order for the CATE to be identifiable. This requirement can be satisfied by collecti…

2022

Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine Comprehension

ACL 2022long

Procedural Multimodal Documents (PMDs) organize textual instructions and corresponding images step by step. Comprehending PMDs and inducing their representations for the downstream reasoning tasks is designated as Procedural MultiModal Machine Comprehension (M3C). In this study, we approach Procedur…

2021

Argument Mining Driven Analysis of Peer-Reviews

AAAI 2021technical

Peer reviewing is a central process in modern research and essential for ensuring high quality and reliability of published work. At the same time, it is a time-consuming process and increasing interest in emerging fields often results in a high review workload, especially for senior researchers in…

2021

BAMBOO: A Multi-instance Multi-label Approach Towards VDI User Logon Behavior Modeling

IJCAI 2021poster

Different to traditional on-premise VDI , the virtual desktops in DaaS (Desktop as a Service) are hosted in public cloud where virtual machines are charged based on usage. Accordingly, an adaptive power management system which can turn off spare virtual machines without sacrificing end user experien…

Cited by 2SourcePDFScholar
2021

GMH: A General Multi-hop Reasoning Model for KG Completion

EMNLP 2021main

Knowledge graphs are essential for numerous downstream natural language processing applications, but are typically incomplete with many facts missing. This results in research efforts on multi-hop reasoning task, which can be formulated as a search process and current models typically perform short…

Cited by 17SourcePDFScholar
2021

Generalized Relation Learning with Semantic Correlation Awareness for Link Prediction

AAAI 2021technical

Developing link prediction models to automatically complete knowledge graphs has recently been the focus of significant research interest. The current methods for the link prediction task have two natural problems: 1) the relation distributions in KGs are usually unbalanced, and 2) there are many un…

Cited by 18SourcePDFScholar
2021

Hysteresis Modeling of Robotic Catheters Based on Long Short-Term Memory Network for Improved Environment Reconstruction

RA-L 2021

Catheters are increasingly being used to tackle problems in the cardiovascular system. However, positioning precision of the catheter tip is negatively affected by hysteresis. To ensure tissue damage due to imprecise positioning is avoided, hysteresis is to be understood and compensated for. This wo

Cited by 52SourceScholar
2021

MIRACLE: Causally-Aware Imputation via Learning Missing Data Mechanisms

NeurIPS 2021poster

Missing data is an important problem in machine learning practice. Starting from the premise that imputation methods should preserve the causal structure of the data, we develop a regularization scheme that encourages any baseline imputation method to be causally consistent with the underlying data…

2021

Reinforcement Learning Enhanced Explainer for Graph Neural Networks

NeurIPS 2021poster

Graph neural networks (GNNs) have recently emerged as revolutionary technologies for machine learning tasks on graphs. In GNNs, the graph structure is generally incorporated with node representation via the message passing scheme, making the explanation much more challenging. Given a trained GNN mod…

Cited by 81SourcePDFScholar
2021

SyncTwin: Treatment Effect Estimation with Longitudinal Outcomes

NeurIPS 2021poster

Most of the medical observational studies estimate the causal treatment effects using electronic health records (EHR), where a patient's covariates and outcomes are both observed longitudinally. However, previous methods focus only on adjusting for the covariates while neglecting the temporal struct…

2020

CASTLE: Regularization via Auxiliary Causal Graph Discovery

NeurIPS 2020poster

Regularization improves generalization of supervised models to out-of-sample data. Prior works have shown that prediction in the causal direction (effect from cause) results in lower testing error than the anti-causal direction. However, existing regularization methods are agnostic of causality. We…

2020

Learning Overlapping Representations for the Estimation of Individualized Treatment Effects

AISTATS 2020poster

The choice of making an intervention depends on its potential benefit or harm in comparison to alternatives. Estimating the likely outcome of alternatives from observational data is a challenging problem as all outcomes are never observed, and selection bias precludes the direct comparison of differ…

2020

Learning outside the Black-Box: The pursuit of interpretable models

NeurIPS 2020poster

Machine learning has proved its ability to produce accurate models -- but the deployment of these models outside the machine learning community has been hindered by the difficulties of interpreting these models. This paper proposes an algorithm that produces a continuous global interpretation of any…

2020

Robust Recursive Partitioning for Heterogeneous Treatment Effects with Uncertainty Quantification

NeurIPS 2020poster

Subgroup analysis of treatment effects plays an important role in applications from medicine to public policy to recommender systems. It allows physicians (for example) to identify groups of patients for whom a given drug or treatment is likely to be effective and groups of patients for which it is…

2020

Stepwise Model Selection for Sequence Prediction via Deep Kernel Learning

AISTATS 2020poster

An essential problem in automated machine learning (AutoML) is that of model selection. A unique challenge in the sequential setting is the fact that the optimal model itself may vary over time, depending on the distribution of features and labels available up to each point in time. In this paper, w…

2020

VIME: Extending the Success of Self- and Semi-supervised Learning to Tabular Domain

NeurIPS 2020poster

Self- and semi-supervised learning frameworks have made significant progress in training machine learning models with limited labeled data in image and language domains. These methods heavily rely on the unique structure in the domain datasets (such as spatial relationships in images or semantic rel…

2015

Pedestrian detection via PCA filters based convolutional channel features

ICASSP 2015accepted

In this paper, we propose a kind of image representation, named PCA filters based convolutional channel features (PCA-CCF) for pedestrian detection. The motivation is to use the convolutional network architecture with orthogonal PCA filters to enhance the state-of-the-art aggregate channel features…

Cited by 0SourceScholar