← Search

Kai Ma

29 accepted papers

2026

Agentic Reinforced Policy Optimization

ICLR 2026poster

Large-scale reinforcement learning with verifiable rewards (RLVR) has proven effective in harnessing the potential of large language models (LLMs) for single-turn reasoning tasks. In realistic reasoning scenarios, LLMs often rely on external tools to assist in task-solving processes. However, curren…

Cited by 0SourcecodeScholar
2026

DemoFunGrasp: Universal Dexterous Functional Grasping via Demonstration-Editing Reinforcement Learning

CVPR 2026

Reinforcement learning (RL) has achieved great success in dexterous grasping, significantly improving grasp performance and generalization from simulation to the real world. However, fine-grained functional grasping, which is essential for downstream manipulation tasks, remains underexplored and fac

Cited by 0SourcecodeScholar
2026

PositionIC: Unified Position and Identity Consistency for Image Customization

CVPR 2026

Recent subject-driven image customization excels in fidelity, yet fine-grained instance-level spatial control remains an elusive challenge, hindering real-world applications. This limitation stems from two factors: a scarcity of scalable, position-annotated datasets, and the entanglement of identity

Cited by 0SourcecodeScholar
2026

RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment Generalization

ICML 2026poster

Vision-Language-Action (VLA) models hold promise for generalist robotics but currently struggle with data scarcity, architectural inefficiencies, and the inability to generalize across different hardware platforms. We introduce RDT2, a robotic foundation model built upon a 7B parameter VLM designed …

Cited by 0SourceScholar
2025

IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation

ICCV 2025poster

While Text-to-Image (T2I) diffusion models excel at generating visually appealing images of individual instances, they struggle to accurately position and control the features generation of multiple instances. The Layout-to-Image (L2I) task was introduced to address the positioning challenges by inc…

Cited by 0SourcePDFScholar
2025

Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement

ACL 2025finding

Large Language Models (LLMs) have become essential for offensive language detection, yet their ability to handle annotation disagreement remains underexplored. Disagreement samples, which arise from subjective interpretations, pose a unique challenge due to their ambiguous nature. Understanding how…

2025

Make-It-Animatable: An Efficient Framework for Authoring Animation-Ready 3D Characters

CVPR 2025highlight

3D characters are essential to modern creative industries, but making them animatable often demands extensive manual work in tasks like rigging and skinning. Existing automatic rigging tools face several limitations, including the necessity for manual annotations, rigid skeleton topologies, and limi…

2025

Sarcasm-R1: Enhancing Sarcasm Detection through Focused Reasoning

EMNLP 2025

Sarcasm detection is a crucial yet challenging task in natural language processing. Existing methods primarily rely on supervised learning or prompt engineering, which often struggle to capture the complex reasoning process required for effective sarcasm detection. This paper proposes a novel approa

2025

Towards Patronizing and Condescending Language in Chinese Videos: A Multimodal Dataset and Detector

ICASSP 2025accepted

Patronizing and Condescending Language (PCL) is a form of discriminatory toxic speech targeting vulnerable groups, threatening both online and offline safety. While toxic speech research has mainly focused on overt toxicity, such as hate speech, microaggressions in the form of PCL remain underexplor…

Cited by 0SourceScholar
2025

VRoPE: Rotary Position Embedding for Video Large Language Models

EMNLP 2025

Rotary Position Embedding (RoPE) has shown strong performance in text-based Large Language Models (LLMs), but extending it to video remains a challenge due to the intricate spatiotemporal structure of video frames. Existing adaptations, such as RoPE-3D, attempt to encode spatial and temporal dimensi

2024

Fourier Controller Networks for Real-Time Decision-Making in Embodied Learning

ICML 2024poster

Transformer has shown promise in reinforcement learning to model time-varying features for obtaining generalized low-level robot policies on diverse robotics datasets in embodied learning. However, it still suffers from the issues of low data efficiency and high inference latency. In this paper, we…

2024

StableDrag: Stable Dragging for Point-based Image Editing

ECCV 2024poster

"Point-based image editing has attracted remarkable attention since the emergence of DragGAN. Recently, DragDiffusion further pushes forward the generative quality via adapting this dragging technique to diffusion models. Despite these great success, this dragging scheme exhibits two major drawbacks…

Cited by 12SourcePDFScholar
2024

VFIMamba: Video Frame Interpolation with State Space Models

NeurIPS 2024poster

Inter-frame modeling is pivotal in generating intermediate frames for video frame interpolation (VFI). Current approaches predominantly rely on convolution or attention-based models, which often either lack sufficient receptive fields or entail significant computational overheads. Recently, Selectiv…

2023

Approximation Error Back-Propagation for Q-Function in Scalable Reinforcement Learning with Tree Dependence Structure

ICASSP 2023accepted

This paper applies the exponential decay property of scalable RL theory to a specific scenario where the network structure is a tree, and use KL (Kullback-Leibler) divergence to analyze the propagation of approximation error along the structure over time, in order to quantify its backtracking result…

Cited by 0SourceScholar
2022

Boost Supervised Pretraining for Visual Transfer Learning: Implications of Self-Supervised Contrastive Representation Learning

AAAI 2022technical

Unsupervised pretraining based on contrastive learning has made significant progress recently and showed comparable or even superior transfer learning performance to traditional supervised pretraining on various tasks. In this work, we first empirically investigate when and why unsupervised pretrain…

2022

Dense Cross-Query-and-Support Attention Weighted Mask Aggregation for Few-Shot Segmentation

ECCV 2022poster

"Research into Few-shot Semantic Segmentation (FSS) has attracted great attention, with the goal to segment target objects in a query image given only a few annotated support images of the target class. A key to this challenging task is to fully utilize the information in the support images by explo…

2021

Alternative Baselines for Low-Shot 3D Medical Image Segmentation—An Atlas Perspective

AAAI 2021technical

Low-shot (one/few-shot) segmentation has attracted increasing attention as it works well with limited annotation. State-of-the-art low-shot segmentation methods on natural images usually focus on implicit representation learning for each novel class, such as learning prototypes, deriving guidance fe…

Cited by 5SourcePDFScholar
2021

Learning Calibrated Medical Image Segmentation via Multi-Rater Agreement Modeling

CVPR 2021poster

In medical image analysis, it is typical to collect multiple annotations, each from a different clinical expert or rater, in the expectation that possible diagnostic errors could be mitigated. Meanwhile, from the computer vision practitioner viewpoint, it has been a common practice to adopt the grou…

Cited by 188PDFcodeScholar
2021

Multi-Anchor Active Domain Adaptation for Semantic Segmentation

ICCV 2021poster

Unsupervised domain adaption has proven to be an effective approach for alleviating the intensive workload of manual annotation by aligning the synthetic source-domain data and the real-world target-domain samples. Unfortunately, mapping the target-domain distribution to the source-domain unconditio…

Cited by 60PDFcodeScholar
2020

Cross-denoising Network against Corrupted Labels in Medical Image Segmentation with Domain Shift

IJCAI 2020poster

Deep convolutional neural networks (DCNNs) have contributed many breakthroughs in segmentation tasks, especially in the field of medical imaging. However, domain shift and corrupted annotations, which are two common problems in medical imaging, dramatically degrade the performance of DCNNs in practi…

Cited by 0SourcePDFScholar
2020

Deep Image Clustering with Category-Style Representation

ECCV 2020poster

Deep clustering which adopts deep neural networks to obtain optimal representations for clustering has been widely studied recently. In this paper, we propose a novel deep image clustering framework to learn a category-style latent representation in which the category information is disentangled fro…

2020

Dual Adversarial Network for Deep Active Learning

ECCV 2020poster

Active learning, reducing the cost and workload of annotations, attracts increasing attentions from the community. Current active learning approaches commonly adopted uncertainty-based acquisition functions for the data selection due to their effectiveness. However, data selection based on uncertain…

Cited by 40SourcePDFScholar
2020

LT-Net: Label Transfer by Learning Reversible Voxel-Wise Correspondence for One-Shot Medical Image Segmentation

CVPR 2020poster

We introduce a one-shot segmentation method to alleviate the burden of manual annotation for medical images. The main idea is to treat one-shot segmentation as a classical atlas-based segmentation problem, where voxel-wise correspondence from the atlas to the unlabelled data is learned. Subsequently…

Cited by 96PDFScholar
2020

Self-Supervised CycleGAN for Object-Preserving Image-to-Image Domain Adaptation

ECCV 2020poster

Recent generative adversarial network (GAN) based methods (e.g., CycleGAN) are prone to fail at preserving image-objects in image-to-image translation, which reduces their practicality on tasks such as domain adaptation. Some frameworks have been proposed to adopt a segmentation network as the auxil…

Cited by 34SourcePDFScholar
2019

X2CT-GAN: Reconstructing CT From Biplanar X-Rays With Generative Adversarial Networks

CVPR 2019poster

Computed tomography (CT) can provide a 3D view of the patient's internal organs, facilitating disease diagnosis, but it incurs more radiation dose to a patient and a CT scanner is much more cost prohibitive than an X-ray machine too. Traditional CT reconstruction methods require hundreds of X-ray pr…

Cited by 297PDFcodeScholar
2018

Generating Synthetic X-Ray Images of a Person From the Surface Geometry

CVPR 2018poster

We present a novel framework that learns to predict human anatomy from body surface. Specifically, our approach generates a synthetic X-ray image of a person only from the person's surface geometry. Furthermore, the synthetic X-ray image is parametrized and can be manipulated by adjusting a set of b…

2015

BodyPrint: Pose Invariant 3D Shape Matching of Human Bodies

ICCV 2015poster

3D human body shape matching has large potential on many real world applications, especially with the recent advances in the 3D range sensing technology. We address this problem by proposing a novel holistic human body shape descriptor called BodyPrint. To compute the bodyprint for a given body scan…

Cited by 3PDFScholar