← Search

Ke Ma

33 accepted papers

2026

A Wearable Isokinetic Training Robot for Enhanced Bedside Knee Rehabilitation

ICRA 2026poster

Knee pain is prevalent in over 20% of the population, limiting the mobility of those affected. In turn, isokinetic dynamometers and robots have been used to facilitate rehabilitation for those still capable of ambulation. However, there are at most only a few wearable robots capable of delivering is…

Cited by 0SourceScholar
2026

ACD-CLIP: DECOUPLING REPRESENTATION AND DYNAMIC FUSION FOR ZERO-SHOT ANOMALY DETECTION

ICASSP 2026poster

Pre-trained Vision-Language Models (VLMs) struggle with Zero-Shot Anomaly Detection (ZSAD) due to a critical adaptation gap: they lack the local inductive biases required for dense prediction and employ inflexible feature fusion paradigms. We address these limitations through an Architectural Co-Des…

Cited by 0SourcePDFScholar
2026

Hidden Dangers of Compositional Generation: Diagnosing Semantic Safety Failures in Text-to-Image Models

CVPR 2026

Text-to-Image (T2I) models have achieved significant progress in generating high-quality images, with compositional visual generation emerging as an important capability that enables them to synthesize coherent, natural scenes from multiple discrete concepts. However, this powerful compositionality,

Cited by 0SourceScholar
2026

IntuFly: Intuitive Continuous Hand–Gaze Control for UAVs

ICRA 2026poster

Operating Unmanned Aerial Vehicles (UAVs) remains challenging for non-experts because single-modality interfaces distort intent: gesture-only systems depend on discrete vocabularies and mode switches that break continuity and raise cognitive load, while gaze-only control offers limited dimensionalit…

Cited by 0codeScholar
2026

LidarPainter: One-Step Away from Any Lidar View to Novel Guidance

AAAI 2026technical

Dynamic driving scene reconstruction is of great importance in fields like digital twin system and autonomous driving simulation. However, unacceptable degradation occurs when the view deviates from the input trajectory, leading to corrupted background and vehicle models. To improve reconstruction q

Cited by 0SourcePDFScholar
2026

Localize and Neutralize: Gradient-Guided Token Suppression Against Visual Prompt Injection Attack

ICML 2026poster

Adversarial images pose a severe security threat to multimodal large language models through prompt injection. Existing defenses largely lack a principled understanding of the underlying mechanisms and struggle to balance efficiency and fidelity. In this work, we show that successful adversarial att…

Cited by 0SourceScholar
2026

Phys-Liquid: A Physics-Informed Dataset for Estimating 3D Geometry and Volume of Transparent Deformable Liquids

AAAI 2026technical

Estimating the geometric and volumetric properties of transparent deformable liquids is challenging due to optical complexities and dynamic surface deformations induced by container movements. Autonomous robots performing precise liquid manipulation tasks—such as dispensing, aspiration, and mixing—m

Cited by 0SourcePDFScholar
2026

Rhythm: Learning Interactive Whole-Body Control for Dual Humanoids

RSS 2026poster

Realizing interactive whole-body control for multi-humanoid systems is critical for unlocking complex collaborative capabilities in shared environments. Although recent advancements have significantly enhanced the agility of individual robots, bridging the gap to physically coupled multi-humanoid in…

Cited by 1SourceScholar
2026

SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation

CVPR 2026

Diffusion-based video motion customization facilitates the acquisition of human motion representations from a few video samples, while achieving arbitrary subjects transfer through precise textual conditioning. Existing approaches often rely on semantic-level alignment, expecting the model to learn

Cited by 0SourceScholar
2026

TaskIT: Memory-Efficient Fine-Tuning of Multi-LoRA LLMs via Cross-Task Importance Transfer

CVPR 2026

On-device AI systems increasingly adopt a single foundation model equipped with task-specific Low-Rank Adaptation (LoRA) modules, forming a multi-LoRA LLM that supports multiple tasks.We study how to adapt such a model to a new task on memory-constrainted devices.Although LoRA reduces trainable para

Cited by 0SourcecodeScholar
2025

A Visual Servo System for Robotic on-Orbit Servicing Based on 3D Perception of Non-Cooperative Satellite

ICRA 2025

The 3D perception of satellites, including both their shape and pose, is a key foundation for robotic on-orbit servicing. However, the demanding space environment-such as intense and dim illumination-presents significant challenges. Previous non-cooperative methods focus on specific geometric featur

Cited by 0SourceScholar
2025

Adapt Once, Thrive with Updates: Transferable Parameter-Efficient Fine-Tuning on Evolving Base Models

ACL 2025long

Parameter-efficient fine-tuning (PEFT) has become a common method for fine-tuning large language models, where a base model can serve multiple users through PEFT module switching. To enhance user experience, base models require periodic updates. However, once updated, PEFT modules fine-tuned on prev…

Cited by 0SourcePDFScholar
2025

BrainChat: Interactive Semantic Information Decoding from fMRI Using Large-Scale Vision-Language Pretrained Models

ICASSP 2025accepted

Semantic information is crucial for human awareness. The ability to extract such information interactively from brain activity using non-invasive technologies like functional Magnetic Resonance Imaging (fMRI) is valuable for medical assistive technologies. However, research in this domain remains re…

Cited by 0SourceScholar
2025

Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs

ICML 2025poster

Despite the remarkable performance of Large Language Models (\textbf{LLMs}), they remain vulnerable to jailbreak attacks, which can compromise their safety mechanisms. Existing studies often rely on brute-force optimization or manual design, failing to uncover potential risks in real-world scenarios…

Cited by 0SourcePDFScholar
2025

Diffusion-based Adversarial Purification from the Perspective of the Frequency Domain

ICML 2025spotlight

The diffusion-based adversarial purification methods attempt to drown adversarial perturbations into a part of isotropic noise through the forward process, and then recover the clean images through the reverse process. Due to the lack of distribution information about adversarial perturbations in th…

Cited by 0SourcePDFScholar
2025

Divide and Conquer: Heterogeneous Noise Integration for Diffusion-based Adversarial Purification

CVPR 2025poster

Existing diffusion-based purification methods aim to disrupt adversarial perturbations by introducing a certain amount of noise through a forward diffusion process, followed by a reverse process to recover clean examples. However, this approach is fundamentally flawed: the uniform operation of the f…

Cited by 2SourcePDFScholar
2025

Exploring Query Efficient Data Generation Towards Data-Free Model Stealing in Hard Label Setting

AAAI 2025technical

Data-free model stealing involves replicating the functionality of a target model into a substitute model without accessing the target model's structure, parameters, or training data. Instead, the adversary can only access the target model's predictions for generated samples. Once the substitute mod…

Cited by 1SourcePDFScholar
2025

Real-World Automated Vehicle Longitudinal Stability Analysis: Controller Design and Field Test

ICRA 2025

Although extensive research has been conducted on modeling the stable longitudinal controller of automated vehicles (AVs) to dampen traffic oscillations, the real-world performance of these controllers in actual vehicles remains uncertain. In the operation of real-world AVs, the delay between actual

Cited by 0SourceScholar
2025

SURGEON: Memory-Adaptive Fully Test-Time Adaptation via Dynamic Activation Sparsity

CVPR 2025highlight

Despite the growing integration of deep models into mobile terminals, the accuracy of these models declines significantly due to various deployment interferences. Test-time adaptation (TTA) has emerged to improve the performance of deep models by adapting them to unlabeled target data online. Yet, t…

2025

SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device

CVPR 2025poster

We have witnessed the unprecedented success of diffusion-based video generation over the past year. Recently proposed models from the community have wielded the power to generate cinematic and high-resolution videos with smooth motions from arbitrary input prompts. However, as a supertask of image g…

Cited by 2SourcePDFScholar
2024

"Towards Dual Transparent Liquid Level Estimation in Biomedical Lab: Dataset, Methods and Practice"

ECCV 2024poster

"“Dual Transparent Liquid” refers to a liquid and its container, both being transparent. Accurately estimating the levels of such a liquid from arbitrary viewpoints is fundamental and crucial, especially in AI-guided autonomous biomedical laboratories for tasks like liquid dispensing, aspiration, an…

2024

HAWK: Learning to Understand Open-World Video Anomalies

NeurIPS 2024poster

Video Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs. However, current VAD systems are often limited by their superficial semantic understanding of scenes and minimal user interaction. Additionally, the prev…

2024

Slicing Vision Transformer for Flexible Inference

NeurIPS 2024poster

Vision Transformers (ViT) is known for its scalability. In this work, we target to scale down a ViT to fit in an environment with dynamic-changing resource constraints. We observe that smaller ViTs are intrinsically the sub-networks of a larger ViT with different widths. Thus, we propose a general f…

2023

Inspired by Physical Intelligence of an Elephant Trunk: Biomimetic Soft Robot With Pre-Programmable Localized Stiffness

RA-L 2023

Soft robots exhibit promising dexterity and adaptability for manipulation because of their high compliance. However, the existing soft robots with invariant stiffness hardly interact with cluttered environments with varying curvatures. In this study, inspired by the maneuverability of an elephant tr

Cited by 35SourceScholar
2022

LAS-AT: Adversarial Training With Learnable Attack Strategy

CVPR 2022oral

Adversarial training (AT) is always formulated as a minimax problem, of which the performance depends on the inner optimization that involves the generation of adversarial examples (AEs). Most previous methods adopt Projected Gradient Decent (PGD) with manually specifying attack parameters for AE ge…

Cited by 196PDFcodeScholar
2022

Learning an Isometric Surface Parameterization for Texture Unwrapping

ECCV 2022poster

"In this paper, we present a novel approach to learn texture mapping for an isometrically deformed 3D surface and apply it for texture unwrapping of documents or other objects. Recent work on differentiable rendering techniques for implicit surfaces has shown high-quality 3D scene reconstruction and…

2022

Prior-Guided Adversarial Initialization for Fast Adversarial Training

ECCV 2022poster

"Fast adversarial training (FAT) effectively improves the efficiency of standard adversarial training (SAT). However, initial FAT encounters catastrophic overfitting, i.e., the robust accuracy against adversarial attacks suddenly decreases to 0% during training. Though several FAT variants spare no…

2021

What to Select: Pursuing Consistent Motion Segmentation from Multiple Geometric Models

AAAI 2021technical

Motion segmentation aims at separating motions of different moving objects in a video sequence. Facing the complicated real-world scenes, recent studies reveal that combining multiple geometric models would be a more effective way than just employing a single one. This motivates a new wave of model-…

2019

DewarpNet: Single-Image Document Unwarping With Stacked 3D and 2D Regression Networks

ICCV 2019poster

Capturing document images with hand-held devices in unstructured environments is a common practice nowadays. However, "casual" photos of documents are usually unsuitable for automatic information extraction, mainly due to physical distortion of the document paper, as well as various camera positions…

Cited by 92PDFScholar
2018

Finding Global Optima in Nonconvex Stochastic Semidefinite Optimization with Variance Reduction

AISTATS 2018poster

There is a recent surge of interest in nonconvex reformulations via low-rank factorization for stochastic convex semidefinite optimization problem in the purpose of efficiency and scalability. Compared with the original convex formulations, the nonconvex ones typically involve much fewer variables,…

Cited by 0SourcePDFScholar
2017

Robot mapping and localisation in metal water pipes using hydrophone induced vibration and map alignment by dynamic time warping

ICRA 2017poster

Water is a highly valuable resource so asset management of associated infrastructure is of critical importance. Water distribution pipe networks are usually buried, and so are difficult to access. Robots are therefore appealing for performing inspection and detecting damage to target repairs. Howeve…

Cited by 0SourceScholar