← Search

JIAN XU

52 accepted papers

2026

A Human Finger-Inspired Rigid-Soft Hybrid Gripper for Damage-Free and Fast Grasping

ICRA 2026poster

Rigid-soft hybrid grippers show good protection and high-payload capacity for fragile and heavy objects. However, because of inadequate actuation speed, it is still challenging for hybrid grippers to grasp moving objects in unstructured environments. To solve the limitation, this article presents a …

Cited by 0SourceScholar
2026

ATPO: ADAPTIVE TREE POLICY OPTIMIZATION FOR MULTI-TURN MEDICAL DIALOGUE

ICLR 2026poster

Effective information seeking in multi-turn medical dialogues is critical for accurate diagnosis, especially when dealing with incomplete information. Aligning Large Language Models (LLMs) for these interactive scenarios is challenging due to the uncertainty inherent in user-agent interactions, whic…

Cited by 0SourceScholar
2026

An Open-Ended Benchmark and Formal Framework for Adjuvant Research with MLLM

ICLR 2026poster

Adjuvants play a critical role in modulating immune responses and are central to the development of vaccines and immunotherapies. Yet progress in this field is constrained by data scarcity and incomplete understanding of mechanisms of action, which limit the transition from experience-based design t…

Cited by 0SourceScholar
2026

Can LLMs Move Beyond Short Exchanges to Realistic Therapy Conversations?

ICLR 2026poster

Recent incidents have revealed that large language models (LLMs) deployed in mental health contexts can generate unsafe guidance, including reports of chatbots encouraging self-harm. Such risks highlight the urgent need for rigorous, clinically valid evaluation before integration into care. However,…

Cited by 0SourceScholar
2026

CatalystBench: A Comprehensive Multi-Task Benchmark for Advancing Language Models in Catalysis Science

ICLR 2026poster

The discovery of novel catalytic materials is a cornerstone of chemical engineering and sustainable energy, yet it remains a complex, knowledge-intensive process. While Large Language Models (LLMs) have demonstrated remarkable potential in various scientific domains, their application to catalysis i…

Cited by 0SourceScholar
2026

E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs

ICML 2026poster

E-commerce short videos represent a high-revenue segment of the online video industry characterized by a goal-driven format and dense multi-modal signals. Current models often struggle with these videos because existing benchmarks focus primarily on general-purpose tasks and neglect the reasoning of…

Cited by 1SourceScholar
2026

Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search

ICLR 2026oral

Auto-bidding serves as a critical tool for advertisers to improve their advertising performance. Recent progress has demonstrated that AI-Generated Bidding (AIGB), which learns a conditional generative planner from offline data, achieves superior performance compared to typical offline reinforcement…

Cited by 0SourceScholar
2026

HVR-Met: A Hypothesis-Verification-Replaning Agentic System for Extreme Weather Diagnosis

ICML 2026poster

While deep learning-based weather forecasting paradigms have made significant strides, addressing extreme weather diagnostics remains a formidable challenge. This gap exists primarily because the diagnostic process demands sophisticated multi-step logical reasoning, dynamic tool invocation, and expe…

Cited by 0SourceScholar
2026

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling

ICML 2026poster

Multimodal Large Language Models (MLLMs) have made substantial progress on visual understanding tasks, yet they still perform poorly on high-resolution images. Prior work often attributes this limitation to perceptual constraints, arguing that MLLMs fail to recognize small objects and therefore rely…

Cited by 0SourceScholar
2026

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

CVPR 2026

Recent Multimodal Large Language Models (MLLMs) have significantly advanced e-commerce product understanding. However, they still face three challenges: (i) the modality imbalance induced by modality mixed training; (ii) underutilization of the intrinsic alignment relationships among visual and text

Cited by 0SourceScholar
2026

MeteorPred: A Meteorological Multimodal Large Model and Dataset for Severe Weather Event Prediction

CVPR 2026

Timely and accurate forecasts of severe weather events are essential for early warning and for constraining downstream analysis and decision-making. Since severe weather events prediction still depends on subjective, time-consuming expert interpretation, end-to-end "AI weather station" systems are e

Cited by 0SourcecodeScholar
2026

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding

CVPR 2026

While Multimodal Large Language Models (MLLMs) excel at single-image understanding, they exhibit significantly degraded performance in multi-image reasoning scenarios. Multi-image reasoning presents fundamental challenges including complex inter-relationships between images and scattered critical in

Cited by 0SourceScholar
2025

A Human Finger-Inspired Rigid-Soft Hybrid Gripper for Damage-Free and Fast Grasping

RA-L 2025

Rigid-soft hybrid grippers show good protection and high-payload capacity for fragile and heavy objects. However, because of inadequate actuation speed, it is still challenging for hybrid grippers to grasp moving objects in unstructured environments. To address this limitation, this article presents

Cited by 1SourceScholar
2025

Bayesian Gaussian Process ODEs via Double Normalizing Flows

AISTATS 2025poster

Gaussian processes have been used to model the vector field of continuous dynamical systems, which are characterized by a probabilistic ordinary differential equation (GP-ODE). Bayesian inference for these models has been extensively studied and applied in tasks such as time series prediction. Howev…

Cited by 0SourceScholar
2025

CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models

COLING 2025main

With the rapid advancements in multimodal large language models, evaluating their multimodal mathematical capabilities continues to receive wide attention. Although datasets such as MathVista have been introduced for evaluating mathematical capabilities in multimodal scenarios, there remains a lack…

2025

Federated Continual Instruction Tuning

ICCV 2025poster

A vast amount of instruction tuning data is crucial for the impressive performance of Large Multimodal Models (LMMs), but the associated computational costs and data collection demands during supervised fine-tuning make it impractical for most researchers. Federated learning (FL) has the potential t…

2025

HiE-VL: A Large Vision-Language Model with Hierarchical Adapter for Handwritten Mathematical Expression Recognition

ICASSP 2025accepted

Large Vision-Language Models (LVLMs) have shown impressive capabilities across various domains, but existing LVLMs have limited performance in dense perception and structured learning problems, such as Handwritten Mathematical Expression Recognition (HMER). The primary challenges stem from the compl…

Cited by 0SourceScholar
2025

LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating

ACL 2025long

Large vision language models (LVLMs) have improved the document understanding capabilities remarkably, enabling the handling of complex document elements, longer contexts, and a wider range of tasks. However, existing document understanding benchmarks have been limited to handling only a small numbe…

2025

Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text Information

AAAI 2025technical

With the advancement of large-scale language modeling techniques, large multimodal models combining visual encoders with large language models have demonstrated exceptional performance in various visual tasks. Most of the current large multimodal models achieve this by mapping visual features obtain…

2025

Think before Recommendation: Autonomous Reasoning-enhanced Recommender

NeurIPS 2025poster

The core task of recommender systems is to learn user preferences from historical user-item interactions. With the rapid development of large language models (LLMs), recent research has explored leveraging the reasoning capabilities of LLMs to enhance rating prediction tasks. However, existing disti…

Cited by 0SourceScholar
2025

Variational Learning of Gaussian Process Latent Variable Models through Stochastic Gradient Annealed Importance Sampling

UAI 2025

Gaussian Process Latent Variable Models (GPLVMs) have become increasingly popular for unsupervised tasks such as dimensionality reduction and missing data recovery due to their flexibility and non-linear nature. An importance-weighted version of the Bayesian GPLVMs has been proposed to obtain a tigh

Cited by 0SourcePDFScholar
2025

pFedGPA: Diffusion-based Generative Parameter Aggregation for Personalized Federated Learning

AAAI 2025technical

Federated Learning (FL) offers a decentralized approach to model training, where data remains local and only model parameters are shared between the clients and the central server. Traditional methods, such as Federated Averaging (FedAvg), linearly aggregate these parameters which are usually traine…

Cited by 0SourcePDFScholar
2024

A Manta Ray-Inspired Fast-Swimming Soft Electrohydraulic Robotic Fish

RA-L 2024

Underwater soft robots inspired by marine life have shown great potential in ocean exploration, monitoring, scientific research, etc., due to their excellent safety, compatibility and adaptability when interacting with underwater environments. However, most of their soft actuators suffer performance

Cited by 11SourceScholar
2024

AuctionNet: A Novel Benchmark for Decision-Making in Large-Scale Games

NeurIPS 2024spotlight

Decision-making in large-scale games is an essential research area in artificial intelligence (AI) with significant real-world impact. However, the limited access to realistic large-scale game environments has hindered research progress in this area. In this paper, we present AuctionNet, a benchmark…

Cited by 3SourcecodeScholar
2024

COALA: A Practical and Vision-Centric Federated Learning Platform

ICML 2024poster

We present COALA, a vision-centric Federated Learning (FL) platform, and a suite of benchmarks for practical FL scenarios, which we categorize as task, data, and model levels. At the task level, COALA extends support from simple classification to 15 computer vision tasks, including object detection,…

2024

FAVOR: Full-Body AR-Driven Virtual Object Rearrangement Guided by Instruction Text

AAAI 2024technical

Rearrangement operations form the crux of interactions between humans and their environment. The ability to generate natural, fluid sequences of this operation is of essential value in AR/VR and CG. Bridging a gap in the field, our study introduces FAVOR: a novel dataset for Full-body AR-driven Virt…

2024

NAC: Mitigating Noisy Correspondence in Cross-Modal Matching Via Neighbor Auxiliary Corrector

ICASSP 2024accepted

The presence of noisy correspondence within cross-modal matching has significantly undermined the performance of existing matching methods. In this paper, we introduce a robust framework named Neighbor Auxiliary Corrector (NAC) for alleviating noise by utilizing the neighbors, which are indicative o…

Cited by 0SourceScholar
2024

SIMMKD: Simple Mask-Flow Keypoint Detection for Both Typhoon Detection and Typhoon Eye Location

ICASSP 2024accepted

Recently, deep learning-based methods has gained increasing attention in typhoon tasks. Due to different optimization targets, existing works apply multi-object detection to typhoon detection and keypoint detection to typhoon eye location. However, such two-stage methods ignored the internal connect…

Cited by 0SourceScholar
2024

Sparse Inducing Points in Deep Gaussian Processes: Enhancing Modeling with Denoising Diffusion Variational Inference

ICML 2024oral

Deep Gaussian processes (DGPs) provide a robust paradigm in Bayesian deep learning. In DGPs, a set of sparse integration locations called inducing points are selected to approximate the posterior distribution of the model. This is done to reduce computational complexity and improve model efficiency.…

Cited by 3SourcePDFScholar
2024

Type-Aware Decoding Via Explicitly Aggregating Event Information for Document-Level Event Extraction

ICASSP 2024accepted

Document-level event extraction (DEE) faces two main challenges: arguments-scattering and multi-event. Although previous methods attempt to address these challenges, they overlook the interference of event-unrelated sentences during event detection and neglect the mutual interference of different ev…

Cited by 0SourceScholar
2024

Unleashing Region Understanding in Intermediate Layers for MLLM-based Referring Expression Generation

NeurIPS 2024poster

The Multi-modal Large Language Model (MLLM) based Referring Expression Generation (REG) task has gained increasing popularity, which aims to generate an unambiguous text description that applies to exactly one object or region in the image by leveraging foundation models. We empirically found that t…

2023

CHORD: Category-level Hand-held Object Reconstruction via Shape Deformation

ICCV 2023poster

In daily life, humans utilize hands to manipulate objects. Modeling the shape of objects that are manipulated by the hand is essential for AI to comprehend daily tasks and to learn manipulation skills. However, previous approaches have encountered difficulties in reconstructing the precise shapes of…

Cited by 15PDFcodeScholar
2023

Curriculum Multi-Level Learning for Imbalanced Live-Stream Recommendation

IJCAI 2023poster

In large-scale e-commerce live-stream recommendation, streamers are classified into different levels based on their popularity and other metrics for marketing. Several top streamers at the head level occupy a considerable amount of exposure, resulting in an unbalanced data distribution. A unified mo…

Cited by 1SourcePDFScholar
2023

POEM: Reconstructing Hand in a Point Embedded Multi-View Stereo

CVPR 2023poster

Enable neural networks to capture 3D geometrical-aware features is essential in multi-view based vision tasks. Previous methods usually encode the 3D information of multi-view stereo into the 2D features. In contrast, we present a novel method, named POEM, that directly operates on the 3D POints Emb…

2023

Personalized Federated Learning with Feature Alignment and Classifier Collaboration

ICLR 2023top-5%

Data heterogeneity is one of the most challenging issues in federated learning, which motivates a variety of approaches to learn personalized models for participating clients. One such approach in deep neural networks based tasks is employing a shared feature representation and learning a customized…

2023

SL-MoE: A Two-Stage Mixture-of-Experts Sequence Learning Framework for Forecasting Rapid Intensification of Tropical Cyclone

ICASSP 2023accepted

Forecasting rapid intensification (RI) of tropical cyclones (TC) is an important and challenging task. However, existing RI forecast methods pay little attention to the imbalanced distribution of RI with dynamic statistical models or ma-chine learning methods. Actually, RI prediction is a class-imba…

Cited by 0SourceScholar
2023

Truthful Auctions for Automated Bidding in Online Advertising

IJCAI 2023poster

Automated bidding, an emerging intelligent decision-making paradigm powered by machine learning, has become popular in online advertising. Advertisers in automated bidding evaluate the cumulative utilities and have private financial constraints over multiple ad auctions in a long-term period. Based…

Cited by 11SourcePDFScholar
2022

APG: Adaptive Parameter Generation Network for Click-Through Rate Prediction

NeurIPS 2022accept

In many web applications, deep learning-based CTR prediction models (deep CTR models for short) are widely adopted. Traditional deep CTR models learn patterns in a static manner, i.e., the network parameters are the same across all the instances. However, such a manner can hardly characterize each…

Cited by 37SourcePDFScholar
2022

CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

ACL 2022long

Artificial Intelligence (AI), along with the recent progress in biomedical language understanding, is gradually offering great promise for medical practice. With the development of biomedical language understanding benchmarks, AI applications are widely used in the medical field. However, most bench…

2022

GBA: A Tuning-free Approach to Switch between Synchronous and Asynchronous Training for Recommendation Models

NeurIPS 2022accept

High-concurrency asynchronous training upon parameter server (PS) architecture and high-performance synchronous training upon all-reduce (AR) architecture are the most commonly deployed distributed training modes for recommendation models. Although synchronous AR training is designed to have higher…

Cited by 3SourcePDFScholar
2022

Object Level Depth Reconstruction for Category Level 6D Object Pose Estimation from Monocular RGB Image

ECCV 2022poster

"Recently, RGBD-based category-level 6D object pose estimation has achieved promising improvement in performance, however, the requirement of depth information prohibits broader applications. In order to relieve this problem, this paper proposes a novel approach named Object Level Depth reconstructi…

Cited by 34SourcePDFScholar
2022

Sustainable Online Reinforcement Learning for Auto-bidding

NeurIPS 2022accept

Recently, auto-bidding technique has become an essential tool to increase the revenue of advertisers. Facing the complex and ever-changing bidding environments in the real-world advertising system (RAS), state-of-the-art auto-bidding policies usually leverage reinforcement learning (RL) algorithms t…

2022

TO-FLOW: Efficient Continuous Normalizing Flows With Temporal Optimization Adjoint With Moving Speed

CVPR 2022poster

Continuous normalizing flows (CNFs) construct invertible mappings between an arbitrary complex distribution and an isotropic Gaussian distribution using Neural Ordinary Differential Equations (neural ODEs). It has not been tractable on large datasets due to the incremental complexity of the neural O…

Cited by 5PDFcodeScholar
2020

Dynamic Knapsack Optimization Towards Efficient Multi-Channel Sequential Advertising

ICML 2020poster

In E-commerce, advertising is essential for merchants to reach their target users. The typical objective is to maximize the advertiser’s cumulative revenue over a period of time under a budget constraint. In real applications, an advertisement (ad) usually needs to be exposed to the same user multip…

Cited by 29SourcePDFScholar
2020

Learning to Accelerate Heuristic Searching for Large-Scale Maximum Weighted b-Matching Problems in Online Advertising

IJCAI 2020poster

Bipartite b-matching is fundamental in algorithm design, and has been widely applied into diverse applications, such as economic markets, labor markets, etc. These practical problems usually exhibit two distinct features: large-scale and dynamic, which requires the matching algorithm to be repeatedl…

Cited by 0SourcePDFScholar
2019

Joint Optimization of Tree-based Index and Deep Model for Recommender Systems

NeurIPS 2019poster

Large-scale industrial recommender systems are usually confronted with computational problems due to the enormous corpus size. To retrieve and recommend the most relevant items to users under response time limits, resorting to an efficient index structure is an effective and practical solution. Th…

2018

Joint & Progressive Learning from High-Dimensional Data for Multi-Label Classification

ECCV 2018poster

Despite the fact that nonlinear subspace learning techniques (e.g. manifold learning) have successfully applied to data representation, there is still room for improvement in explainability (explicit mapping), generalization (out-of-samples), and cost-effectiveness (linearization). To this end, a no…

Cited by 39SourcePDFScholar