← Search

Siqi Liu

30 accepted papers

2026

A Tri-Axial FBG-Based Force Sensor at the Tool Tip of a Continuum Manipulator for Single-Port Access Surgery

ICRA 2026poster

Abstract— The absence of force feedback remains a major bottleneck in the development of robotic laparoendoscopic single-site (R-LESS) surgery, reducing the control precision of surgical instruments and increasing the risk of tissue damage. To address this challenge, we propose a miniature triaxial …

Cited by 0Scholar
2026

Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models

ICML 2026poster

1-bit LLM quantization offers significant advantages in reducing storage and computational costs. However, existing methods typically train 1-bit LLMs from scratch, failing to fully leverage pre-trained models. This results in high training costs and notable accuracy degradation. We identify that th…

Cited by 0SourceScholar
2026

Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language Models

CVPR 2026

As large language models (LLMs) continue to advance, there is increasing interest in their ability to infer human mental states and demonstrate a human-like Theory of Mind (ToM). Most existing ToM evaluations, however, are centered on text-based inputs, while scenarios relying solely on visual infor

Cited by 0SourceScholar
2025

Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning

ICML 2025poster

Behavioral diversity, expert imitation, fairness, safety goals and others give rise to preferences in sequential decision making domains that do not decompose additively across time. We introduce the class of convex Markov games that allow general convex preferences over occupancy measures. Despite…

Cited by 0SourcePDFScholar
2025

Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement

ICASSP 2025accepted

Deep learning-based speech enhancement (SE) models have recently outperformed traditional techniques, yet their deployment on resource-constrained devices remains challenging due to high computational and memory demands. This paper introduces a novel dynamic frequency-adaptive knowledge distillation…

Cited by 0SourceScholar
2025

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models

ICML 2025poster

As large language models evolve, there is growing anticipation that they will emulate human-like Theory of Mind (ToM) to assist with routine tasks. However, existing methods for evaluating machine ToM focus primarily on unimodal models and largely treat these models as black boxes, lacking an interp…

Cited by 0SourcePDFScholar
2025

Integrating Co-Training with Edge Discrimination to Enhance Graph Neural Networks Under Heterophily

AAAI 2025technical

Graph Neural Networks (GNNs) have recently achieved significant success in several graph-related tasks. However, traditional GNNs and their variants are constantly limited by the implicit homophily, assuming neighboring nodes belong to the same class. This results in weak performance on heterophilic…

Cited by 0SourcePDFScholar
2025

Re-evaluating Open-ended Evaluation of Large Language Models

ICLR 2025poster

Evaluation has traditionally focused on ranking candidates for a specific skill. Modern generalist models, such as Large Language Models (LLMs), decidedly outpace this paradigm. Open-ended evaluation systems, where candidate models are compared on user-submitted prompts, have emerged as a popular so…

Cited by 1SourcePDFScholar
2024

A Facial Expression Transfer Method Based on 3DMM and Diffusion Models

ICASSP 2024accepted

Due to the complex geometric features of facial expressions, achieving realistic facial expression transfer is a challenging task. This paper proposes a two-stage facial expression transfer method based on 3D Morphable Model (3DMM) and diffusion models. In the first stage, 3DMM-based DECA model is e…

Cited by 0SourceScholar
2024

Build Your Own Robot Friend: An Open-Source Learning Module for Accessible and Engaging AI Education

AAAI 2024technical

As artificial intelligence (AI) is playing an increasingly important role in our society and global economy, AI education and literacy have become necessary components in college and K-12 education to prepare students for an AI-powered society. However, current AI curricula have not yet been made ac…

Cited by 9SourcePDFScholar
2024

NfgTransformer: Equivariant Representation Learning for Normal-form Games

ICLR 2024poster

Normal-form games (NFGs) are the fundamental model of *strategic interaction*. We study their representation using neural networks. We describe the inherent equivariance of NFGs --- any permutation of strategies describes an equivalent game --- as well as the challenges this poses for representation…

2024

Primitive-Based 3D Human-Object Interaction Modelling and Programming

AAAI 2024technical

Embedding Human and Articulated Object Interaction (HAOI) in 3D is an important direction for a deeper human activity understanding. Different from previous works that use parametric and CAD models to represent humans and objects, in this work, we propose a novel 3D geometric primitive-based languag…

Cited by 3SourcePDFScholar
2024

Prompting Vision Foundation Models for Pathology Image Analysis

CVPR 2024poster

The rapid increase in cases of non-alcoholic fatty liver disease (NAFLD) in recent years has raised significant public concern. Accurately identifying tissue alteration regions is crucial for the diagnosis of NAFLD but this task presents challenges in pathology image analysis particularly with small…

2024

Upper-body Hierarchical Graph for Skeleton Based Emotion Recognition in Assistive Driving

ECCV 2024poster

"Emotion recognition plays a crucial role in enhancing the safety and enjoyment of assistive driving experiences. By enabling intelligent systems to perceive and understand human emotions, we can significantly improve human-machine interactions. Current research in emotion recognition emphasizes fac…

2024

XFibrosis: Explicit Vessel-Fiber Modeling for Fibrosis Staging from Liver Pathology Images

CVPR 2024poster

The increasing prevalence of non-alcoholic fatty liver disease (NAFLD) has caused public concern in recent years. The high prevalence and risk of severe complications make monitoring NAFLD progression a public health priority. Fibrosis staging from liver biopsy images plays a key role in demonstrati…

Cited by 2SourcePDFScholar
2023

Beyond Object Recognition: A New Benchmark towards Object Concept Learning

ICCV 2023poster

Understanding objects is a central building block of AI, especially for embodied AI. Even though object recognition excels with deep learning, current machines struggle to learn higher-level knowledge, e.g., what attributes an object has, and what we can do with it. Here, we propose a challenging Ob…

Cited by 9PDFScholar
2023

EDIS: Entity-Driven Image Search over Multimodal Web Content

EMNLP 2023long main

Making image retrieval methods practical for real-world search applications requires significant progress in dataset scales, entity comprehension, and multimodal information fusion. In this work, we introduce Entity-Driven Image Search (EDIS), a challenging dataset for cross-modal image search in th…

Cited by 0SourcecodeScholar
2022

Boundary-Aware Bias Loss for Transformer-Based Aerial Image Segmentation Model

ICASSP 2022accepted

Inspired by the tremendous success of the transformer-based model in natural language processing (NLP), many efforts introduce the transformer-based model into the image processing tasks. However, naive transformer models have to down-sample the image resolution to satisfy computational restrictions…

Cited by 0SourceScholar
2022

Learning Sparse Interpretable Features For NAS Scoring From Liver Biopsy Images

IJCAI 2022poster

Liver biopsy images play a key role in the diagnosis of global non-alcoholic fatty liver disease (NAFLD). The NAFLD activity score (NAS) on liver biopsy images grades the amount of histological findings that reflect the progression of NAFLD. However, liver biopsy image analysis remains a challenging…

Cited by 0SourcePDFScholar
2022

NeuPL: Neural Population Learning

ICLR 2022poster

Learning in strategy games (e.g. StarCraft, poker) requires the discovery of diverse policies. This is often achieved by iteratively training new policies against existing ones, growing a policy population that is robust to exploit. This iterative approach suffers from two issues in real-world games…

Cited by 24SourcePDFScholar
2022

Simplex Neural Population Learning: Any-Mixture Bayes-Optimality in Symmetric Zero-sum Games

ICML 2022spotlight

Learning to play optimally against any mixture over a diverse set of strategies is of important practical interests in competitive games. In this paper, we propose simplex-NeuPL that satisfies two desiderata simultaneously: i) learning a population of strategically diverse basis policies, represente…

Cited by 18SourcePDFScholar
2022

Turbocharging Solution Concepts: Solving NEs, CEs and CCEs with Neural Equilibrium Solvers

NeurIPS 2022accept

Solution concepts such as Nash Equilibria, Correlated Equilibria, and Coarse Correlated Equilibria are useful components for many multiagent machine learning algorithms. Unfortunately, solving a normal-form game could take prohibitive or non-deterministic time to converge, and could fail. We introdu…

Cited by 21SourcePDFScholar
2020

A Generalized Training Approach for Multiagent Learning

ICLR 2020talk

This paper investigates a population-based training regime based on game-theoretic principles called Policy-Spaced Response Oracles (PSRO). PSRO is general in the sense that it (1) encompasses well-known algorithms such as fictitious play and double oracle as special cases, and (2) in principle appl…

Cited by 127SourcecodeScholar
2020

V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control

ICLR 2020poster

Some of the most successful applications of deep reinforcement learning to challenging domains in discrete and continuous control have used policy gradient methods in the on-policy setting. However, policy gradients can suffer from large variance that may limit performance, and in practice require c…

Cited by 136SourceScholar
2019

Emergent Coordination Through Competition

ICLR 2019poster

We study the emergence of cooperative behaviors in reinforcement learning agents by introducing a challenging competitive multi-agent soccer environment with continuous simulated physics. We demonstrate that decentralized, population-based training with co-play can lead to a progression in agents' b…

Cited by 186SourcePDFScholar
2019

Hierarchical Visuomotor Control of Humanoids

ICLR 2019poster

We aim to build complex humanoid agents that integrate perception, motor control, and memory. In this work, we partly factor this problem into low-level motor control from proprioception and high-level coordination of the low-level skills informed by vision. We develop an architecture capable of sur…

Cited by 124SourcePDFScholar
2019

Nonparametric Regressive Point Processes Based on Conditional Gaussian Processes

NeurIPS 2019poster

Real-world event sequences consist of complex mixtures of different types of events occurring in time. An event may depend on past events of the same type, as well as, the other types. Point processes define a general class of models for event sequences. ``Regressive point processes'' refer to point…

2017

Improved Image Captioning via Policy Gradient Optimization of SPIDEr

ICCV 2017spotlight

Current image captioning methods are usually trained via maximum likelihood estimation. However, the log-likelihood score of a caption does not correlate well with human assessments of quality. Standard syntactic evaluation metrics, such as BLEU, METEOR and ROUGE, are also not well correlated. The n…

Cited by 577PDFcodeScholar