← Search

Mingkun Xu

17 accepted papers

2026

$\textit{S}$-SPPO: Semantic-Calibrated Self-Play Preference Optimization

ICML 2026poster

Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO). However, the standard Bradley-Terry instantiation of DPO is limited in modeling common departures from transitivity in human preferences. To address this, recent work has introd…

Cited by 0SourceScholar
2026

Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks

ICLR 2026poster

Recent advances in vision-language models (VLMs) have achieved remarkable performance on standard medical benchmarks, yet their true clinical reasoning ability remains unclear. Existing datasets predominantly emphasize classification accuracy, creating an evaluation illusion in which models appear p…

Cited by 0SourceScholar
2026

Biologically-Inspired Evolutionary Domain Symbiosis for Few-shot and Zero-shot Point Cloud Semantic Segmentation

AAAI 2026technical

Few-shot and zero-shot point cloud semantic segmentation aim to accurately segment novel categories using limited or no labeled samples, respectively. However, existing methods face significant challenges including domain shifts between support and query sets and the inability to handle both few-sho

Cited by 0SourcePDFScholar
2026

Boosting Knowledge Transfer and Retention with Brain-inspired Multi-View Incremental Learning

IJCAI 2026

Traditional multi-view learning models are primarily designed for static datasets with fixed views. However, in dynamic incremental view environments, this approach inevitably leads to view forgetting, where the introduction of new views weakens previously acquired knowledge. In contrast, the human

Cited by 0Scholar
2026

MMPG: MoE-based Adaptive Multi-Perspective Graph Fusion for Protein Representation Learning

AAAI 2026technical

Graph Neural Networks (GNNs) have been widely adopted for Protein Representation Learning (PRL), as residue interaction networks can be naturally represented as graphs. Current GNN-based PRL methods typically rely on single-perspective graph construction strategies, which capture partial properties

Cited by 0SourcePDFScholar
2026

Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language Models

ICML 2026poster

Pre-trained vision-language models (VLMs) such as CLIP have demonstrated strong zero-shot capabilities across diverse domains, yet remain highly vulnerable to adversarial perturbations that disrupt image-text alignment and compromise reliability. Existing defenses typically rely on adversarial fine-…

Cited by 0SourceScholar
2026

Streaming Video Crime Anticipation with Spatio-Temporal Causal Reasoning

CVPR 2026

Crime anticipation enables proactive public safety interventions, yet existing video security systems remain largely reactive, unable to detect precursors of crime. While current visual language models (VLM)-based video understanding methods show promise in high-level reasoning, they are not designe

Cited by 0SourceScholar
2026

TopAdapter: Topology-Aware Prompt Tuning for Efficient Point Cloud Understanding

ICML 2026poster

Point cloud data, with its inherent geometric and topological structures, plays a critical role in 3D vision tasks. However, existing parameter-efficient fine-tuning (PEFT) methods predominantly focus on input token prompting, overlooking the intrinsic geometric information. To address this limitati…

Cited by 0SourceScholar
2025

BIG-FUSION: Brain-Inspired Global-Local Context Fusion Framework for Multimodal Emotion Recognition in Conversations

AAAI 2025technical

Considering the importance of capturing both global conversational topics and local speaker dependencies for multimodal emotion recognition in conversations, current approaches first utilize sequence models like Transformer to extract global context information, then apply Graph Neural Networks to m…

Cited by 0SourcePDFScholar
2025

ClingTP: Curriculum Learning based Multi-style Title Prefix Generation

ICASSP 2025accepted

An informative, creative title prefix is memorable, capable of capturing the attention of readers, and significantly enhances the potential for increased citations. In this work, we pioneer the exploration of the significance of title prefixes in academic papers and propose a controllable title pref…

Cited by 0SourceScholar
2025

Efficient ANN-SNN Conversion with Error Compensation Learning

ICML 2025poster

Artificial neural networks (ANNs) have demonstrated outstanding performance in numerous tasks, but deployment in resource-constrained environments remains a challenge due to their high computational and memory requirements. Spiking neural networks (SNNs) operate through discrete spike events and off…

Cited by 0SourcePDFScholar
2025

Enhancing Graph Contrastive Learning for Protein Graphs from Perspective of Invariance

ICML 2025poster

Graph Contrastive Learning (GCL) improves Graph Neural Network (GNN)-based protein representation learning by enhancing its generalization and robustness. Existing GCL approaches for protein representation learning rely on 2D topology, where graph augmentation is solely based on topological features…

Cited by 0SourcePDFScholar
2025

G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object Manipulation

CVPR 2025poster

Recent advances in imitation learning for 3D robotic manipulation have shown promising results with diffusion-based policies. However, achieving human-level dexterity requires seamless integration of geometric precision and semantic understanding. We present G3Flow, a novel framework that constructs…

Cited by 9SourcePDFScholar
2025

Multi-View Incremental Learning with Structured Hebbian Plasticity for Enhanced Fusion Efficiency

AAAI 2025technical

The rapid evolution of multimedia technology has revolutionized human perception, paving the way for multi-view learning. However, traditional multi-view learning approaches are tailored for scenarios with fixed data views, falling short of emulating the intricate cognitive procedures of the human b…

Cited by 1SourcePDFScholar
2025

RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins

CVPR 2025highlight

In the rapidly advancing field of robotics, dual-arm coordination and complex object manipulation are essential capabilities for developing advanced autonomous systems. However, the scarcity of diverse, high-quality demonstration data and real-world-aligned evaluation benchmarks severely limits such…

Cited by 4SourcePDFScholar
2024

LAMBDA: Large Language Model-Based Data Augmentation for Multi-Modal Machine Translation

EMNLP 2024finding

Multi-modal machine translation (MMT) can reduce ambiguity and semantic distortion compared with traditional machine translation (MT) by utilizing auxiliary information such as images. However, current MMT methods face two primary challenges. The first is their underperformance compared to MT method…

2021

Exploiting Spiking Dynamics with Spatial-temporal Feature Normalization in Graph Learning

IJCAI 2021poster

Biological spiking neurons with intrinsic dynamics underlie the powerful representation and learning capabilities of the brain for processing multimodal information in complex environments. Despite recent tremendous progress in spiking neural networks (SNNs) for handling Euclidean-space tasks, it st…

Cited by 31SourcePDFScholar