← Search

Jiaqi Liu

42 accepted papers

2026

ARCHE: A Novel Task to Evaluate LLMs on Latent Reasoning Chain Extraction

AAAI 2026technical

Large language models (LLMs) are increasingly used in scientific domains. While they can produce reasoning-like content via methods such as chain-of-thought prompting, these outputs are typically unstructured and informal, obscuring whether models truly understand the fundamental reasoning paradigms

Cited by 0SourcePDFScholar
2026

Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning

ICML 2026oral

Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal reasoning tasks; however, their learning remains constrained by the limitations of human-annotated supervision. Recent self-rewarding approaches attempt to overcome this constraint by allowing models to act as their…

Cited by 27SourceScholar
2026

AgentGym-RL: An Open-Source Framework to Train LLM Agents for Long-Horizon Decision Making via Multi-Turn RL

ICLR 2026oral

Training LLM agents for complex multi-turn decision-making tasks requires extensive exploration within their environment, with reinforcement learning (RL) as a natural way. However, the open-source community currently lacks a unified RL framework capable of training agents from scratch across divers…

Cited by 0SourcecodeScholar
2026

Does Reinforcement Fine-Tuning Improve Generalization of LLM Agents? An Empirical Study

ICML 2026poster

Reinforcement fine-tuning (RFT) has shown promise for training LLM agents to perform multi-turn decision-making based on environment feedback. However, most existing evaluations remain largely in-domain—training and testing are conducted in the same environment or even on the same tasks. In real-wor…

Cited by 0SourceScholar
2026

Dynamic Cognitive Planning for Cognitive-Functional Dialogue: A Case Study in Emotional Support Conversation

AAAI 2026technical

Cognitive-functional dialogues, such as those for persuasion, consultation, and question-answering, are prevalent throughout human social interaction. The core difference between these dialogues and casual chat lies in their objective: to guide a person

Cited by 0SourcePDFScholar
2026

FedSDR: Federated Graph Learning with Structural Noise Detection and Reconstruction

CVPR 2026

Federated Graph Learning (FGL) has emerged as a principled framework for decentralized training of Graph Neural Networks (GNNs) while preserving data privacy. In subgraph-FL scenarios, however, structural noise arising from data collection and storage can damage the GNN message-passing scheme of cli

Cited by 0SourcecodeScholar
2026

GeoDiT: A Diffusion-based Vision-Language Model for Geospatial Understanding

CVPR 2026

Autoregressive models are structurally misaligned with the inherently parallel nature of geospatial understanding, forcing a rigid sequential narrative onto scenes and fundamentally hindering the generation of structured and coherent outputs. We challenge this paradigm by reframing geospatial genera

Cited by 0SourceScholar
2026

MLM: Learning Multi-Task Loco-Manipulation Whole-Body Control for Quadruped Robot With Arm

RA-L 2026

Whole-body loco-manipulation for quadruped robots with arms remains a challenging problem, particularly in achieving multi-task control. To address this, we propose MLM, a reinforcement learning framework driven by both real-world and simulation data. It enables a six-DoF robotic arm–equipped quadru

Cited by 4SourceScholar
2026

Paper2Figure: A Multi-Agent Collaborative System for Figure Generation Towards Academic Research Paper

CVPR 2026

Automatically generating clear and accurate figures for research papers remains challenging, as it requires semantic understanding, precise structure, and visual aesthetics. Existing approaches struggle to balance fidelity and quality: large language model (LLM) code-based methods (e.g., SVG, Mermai

Cited by 0SourceScholar
2026

QueryMe: Query-Driven Open-Vocabulary 3D Object Affordances Grounding from Multimodal Evidence

CVPR 2026

Open-vocabulary 3D object affordance grounding aims to identify functional regions of objects given arbitrary semantic descriptions. However, existing methods often rely on fixed training categories and geometric priors, lacking geometric invariance and analogical reasoning capabilities. Since there

Cited by 0SourceScholar
2026

SimpleMem: Efficient Lifelong Memory for LLM Agents

ICML 2026poster

To support long-term interaction in complex environments, LLM agents require memory systems that manage historical experiences. Existing approaches either retain full interaction histories via passive context extension, leading to substantial redundancy, or rely on iterative reasoning to filter nois…

Cited by 0SourceScholar
2026

SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts

AAAI 2026technical

The emergence of large vision-language models (VLMs) has significantly enhanced the efficiency and flexibility of geospatial interpretation. However, general-purpose VLMs remain suboptimal for remote sensing (RS) tasks. Existing geospatial VLMs typically adopt a unified modeling strategy and struggl

Cited by 0SourcePDFScholar
2026

Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping

ICLR 2026poster

Reinforcement learning (RL) has recently become the core paradigm for aligning and strengthening large language models (LLMs). Yet, applying RL in off-policy settings—where stale data from past policies are used for training—improves sample efficiency, but remains challenging: policy entropy decline…

Cited by 0SourcecodeScholar
2026

Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models

ICLR 2026poster

Vision-Language Models (VLMs) in remote sensing often fail at complex analytical tasks, a limitation stemming from their end-to-end training paradigm that bypasses crucial reasoning steps and leads to unverifiable outputs. To address this limitation, we introduce the Perceptually-Grounded Geospatial…

Cited by 0SourcecodeScholar
2026

Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck Perspective

CVPR 2026

The quest for incremental unified multimodal anomaly detection seeks to empower a single model with the ability to systematically detect anomalies across all categories and support incremental learning to accommodate emerging objects/categories. Central to this pursuit is resolving the catastrophic

Cited by 0SourcecodeScholar
2025

ActiveHAI: Active Collection Based Human-AI Diagnosis with Limited Expert Predictions

IJCAI 2025

Recent studies indicate that human-AI collaboration performs better than either alone, particularly in medical diagnosis. Beyond collaboration methods that focus on assigning tasks to humans or AI, like deferral, combining human and AI decisions with their confidence scores is emerging as a promisin

2025

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

NeurIPS 2025poster

In this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 100k university-level questions drawn from 300 UNESCO-defined subjects, spanning diverse formats—multip…

Cited by 0SourceScholar
2025

Language-Driven Policy Distillation for Cooperative Driving in Multi-Agent Reinforcement Learning

RA-L 2025

The cooperative driving technology of Connected and Autonomous Vehicles (CAVs) is crucial for improving the efficiency and safety of transportation systems. Learning-based methods, such as Multi-Agent Reinforcement Learning (MARL), have demonstrated strong capabilities in cooperative decision-making

Cited by 22SourceScholar
2025

Multi-focal Conditioned Latent Diffusion for Person Image Synthesis

CVPR 2025poster

The Latent Diffusion Model (LDM) has demonstrated strong capabilities in high-resolution image generation and has been widely employed for Pose-Guided Person Image Synthesis (PGPIS), yielding promising results. However, the compression process of LDM often results in the deterioration of details, pa…

2025

Revisiting Multimodal Fusion for 3D Anomaly Detection from an Architectural Perspective

AAAI 2025technical

Existing efforts to boost multimodal fusion of 3D anomaly detection (3D-AD) primarily concentrate on devising more effective multimodal fusion strategies. However, little attention was devoted to analyzing the role of multimodal fusion architecture (topology) design in contributing to 3D-AD. In this…

2025

Towards Generalizable Neural Simulators: Addressing Distribution Shifts Induced by Environmental and Temporal Variations

IJCAI 2025

With advancements in deep learning, neural simulators have become increasingly important for improving the efficiency and effectiveness of simulating complex dynamical systems in various scientific and technological fields. This paper presents a novel neural simulator called Context-informed Polymor

2025

Trade-offs in Image Generation: How Do Different Dimensions Interact?

ICCV 2025poster

Model performance in text-to-image (T2I) and image-to-image (I2I) generation often depends on multiple aspects, including quality, alignment, diversity, and robustness. However, models' complex trade-offs among these dimensions have been rarely explored due to (1) the lack of datasets that allow fin…

2024

Tendon Driven Bistable Origami Flexible Gripper for High-Speed Adaptive Grasping

RA-L 2024

This paper introduces a novel bistable origami flexible gripper, which is based on a single-vertex and multi-crease (SVMC) origami structure that has sxeveral advantages, including a simple structure, low cost, and strong deformation capacity. This design addresses the drawbacks of slow response spe

Cited by 39SourceScholar
2024

Unsupervised Continual Anomaly Detection with Contrastively-Learned Prompt

AAAI 2024technical

Unsupervised Anomaly Detection (UAD) with incremental training is crucial in industrial manufacturing, as unpredictable defects make obtaining sufficient labeled data infeasible. However, continual learning methods primarily rely on supervised annotations, while the application in UAD is limited due…

2023

Certified Minimax Unlearning with Generalization Rates and Deletion Capacity

NeurIPS 2023poster

We study the problem of $(\epsilon,\delta)$-certified machine unlearning for minimax models. Most of the existing works focus on unlearning from standard statistical learning models that have a single variable and their unlearning steps hinge on the direct Hessian-based conventional Newton update. W…

Cited by 23SourcePDFScholar
2023

Pushing the Limits of Fewshot Anomaly Detection in Industry Vision: Graphcore

ICLR 2023poster

In the area of few-shot anomaly detection (FSAD), efficient visual feature plays an essential role in the memory bank $\mathcal{M}$-based methods. However, these methods do not account for the relationship between the visual feature and its rotated visual feature, drastically limiting the anomaly de…

Cited by 79SourcePDFScholar
2023

Real3D-AD: A Dataset of Point Cloud Anomaly Detection

NeurIPS 2023poster

High-precision point cloud anomaly detection is the gold standard for identifying the defects of advancing machining and precision manufacturing. Despite some methodological advances in this area, the scarcity of datasets and the lack of a systematic benchmark hinder its development. We introduce Re…

2023

Soft Robotic Arm With Extensible Stiffening Layer

RA-L 2023

When talking about soft robots, softness is considered the most important feature, which brings dexterity and safety in interactive tasks with humans and environments. Such softness sometimes limits the real application of soft robots because load capability and rigidity are widely needed on many oc

Cited by 17SourceScholar
2023

Visual-Kinematics Graph Learning for Procedure-Agnostic Instrument Tip Segmentation in Robotic Surgeries

IROS 2023poster

Accurate segmentation of surgical instrument tip is an important task for enabling downstream applications in robotic surgery, such as surgical skill assessment, tool-tissue interaction and deformation modeling, as well as surgical autonomy. However, this task is very challenging due to the small si…

Cited by 2SourceScholar
2022

Egocentric Human Trajectory Forecasting With a Wearable Camera and Multi-Modal Fusion

RA-L 2022

In this letter, we address the problem of forecasting the trajectory of an egocentric camera wearer (ego-person) in crowded spaces. The trajectory forecasting ability learned from the data of different camera wearers walking around in the real world can be transferred to assist visually impaired peo

Cited by 24SourcecodeScholar
2021

Discriminative Asymmetric Learning for Efficient Surgical Instrument Parsing

ICRA 2021poster

Semantic segmentation of surgical instruments provides essential priors for autonomous surgery. This task is however challenging since the fine-structure of surgical instruments requires the accurate segmentation of detailed regions in images. As the visual guidance for autonomous surgery, the algor…

Cited by 0SourceScholar
2020

A 1 mm-Thick Miniatured Mobile Soft Robot With Mechanosensation and Multimodal Locomotion

RA-L 2020

The miniature soft robots have many promising applications, including micro-manipulations, endoscopy, and microsurgery, etc. Nevertheless, it remains challenging to fabricate a miniatured robot device that is thin, flexible, and can perform multimodal locomotor mobility with sensory capacity. In thi

Cited by 20SourceScholar
2017

Optimization of compound regularization parameters based on Stein's unbiased risk estimate

ICASSP 2017accepted

Recently, the type of compound regularizers has become a popular choice for signal reconstruction. The estimation quality is generally sensitive to the values of multiple regularization parameters. In this work, based on BDF algorithm, we develop a data-driven optimization scheme based on minimizati…

Cited by 0SourceScholar