← Search

Ming Ma

18 accepted papers

2026

A Tale of Two Graphs: Separating Knowledge Exploration from Outline Structure for Open-Ended Deep Research

ICML 2026poster

Open-Ended Deep Research (OEDR) pushes LLM agents beyond short-form QA toward long-horizon workflows that iteratively search, connect, and synthesize evidence into structured reports. However, existing OEDR agents largely follow either linear "search-then-generate" accumulation or outline-centric pl…

Cited by 1SourceScholar
2026

CG-Reasoner: Centroid-Guided Positional Reasoning Segmentation for Medical Imaging with a Robust Visual-Text Consistency Metric

CVPR 2026

Accurate and interpretable medical image segmentation remains a major challenge, as existing deep learning models primarily optimize pixel-level accuracy while overlooking positional reasoning--an essential component for automated report generation and clinical interpretability. We introduce CG-Reas

Cited by 0SourcecodeScholar
2026

Causal Discovery for Irregularly Time Series with Consistency Guarantees

ICML 2026poster

This paper studies causal discovery in irregularly sampled time series—a key challenge in risk-sensitive domains like finance, healthcare, and climate science, where missing data and inconsistent sampling frequencies distort causal mechanisms. The main challenge comes from the interdependence betwee…

Cited by 0SourceScholar
2026

DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems

ICLR 2026poster

Large language model (LLM)–based multi-agent systems are challenging to debug because failures often arise from long, branching interaction traces. The prevailing practice is to leverage LLMs for log-based failure localization, attributing errors to a specific agent and step. However, this paradigm…

Cited by 0SourceScholar
2025

DSDN-Net: An Effective Network for Semantic Segmentation in Open-Pit Coal Mining Areas for Land Cover Recognition

ICASSP 2025accepted

Currently, existing methods in land cover recognition in open-pit coal mining areas face the issue of insufficient accuracy due to multiscale and blurred boundaries when processing remote sensing images. This paper introduces a remote sensing image semantic segmentation network, DSDN-Net, to tackle…

Cited by 0SourceScholar
2025

From End-to-end to Step-by-step: Learning to Abstract via Abductive Reinforcement Learning

IJCAI 2025

Abstraction is a critical technique in general problem-solving, allowing complex tasks to be decomposed into smaller, manageable sub-tasks. While traditional symbolic planning relies on predefined primitive symbols to construct structured abstractions, its reliance on formal representations limits a

2025

Generalizing Causal Effects from Randomized Controlled Trials to Target Populations across Diverse Environments

ICML 2025poster

Generalizing causal effects from Randomized Controlled Trials (RCTs) to target populations across diverse environments is of significant practical importance, as RCTs are often costly and logistically complex to conduct. A key challenge is environmental shift, defined as changes in the distribution…

Cited by 0SourcePDFScholar
2025

Global Static Pruning via Adaptive Sample Complexity Awareness

ICASSP 2025accepted

Dynamic pruning leverage the feature information of each input sample to dynamically adjust the network structure, generating multiple subnetworks suitable for different sample complexity. However, it inevitably introduces higher computational complexity and increased memory consumption. In addition…

Cited by 0SourceScholar
2025

Interactive and Balanced Multimodal Learning via Cross Attention and Gradient Modulation for Compressed Video Action Recognition

ICASSP 2025accepted

Compressed video action recognition is a crucial task in video processing. Compared with traditional methods, it directly processes RGB (I-frames) and motion (motion vectors and residuals) modalities, which effectively alleviates computational burdens. However, this task suffers from dynamic noise a…

Cited by 0SourceScholar
2025

JointSwinUNETR: an Efficient Feature-enhanced Architecture for Small Intestine Cine MRI Segmentation

ICASSP 2025accepted

The Cine MRI of the small intestine is a dynamic magnetic resonance imaging technique used to observe and evaluate small intestine motility. It captures sequential images of the organ in motion over time through rapid imaging. The Transformer architecture is highly effective at capturing long-range…

Cited by 0SourceScholar
2025

Robust and High-Fidelity 3D Gaussian Splatting: Fusing Pose Priors and Geometry Constraints for Texture-Deficient Outdoor Scenes

IROS 2025

3D Gaussian Splatting (3DGS) has emerged as a key rendering pipeline for digital asset creation due to its balance between efficiency and visual quality. To address the issues of unstable pose estimation and scene representation distortion caused by geometric texture inconsistency in large outdoor s

Cited by 2SourcecodeScholar
2025

SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning

NeurIPS 2025poster

Despite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for integrating spatial cues, such as point clouds or depth, either require specialized sensors or fail to effectively exploit d…

Cited by 0SourcecodeScholar
2025

Semi-distributed Cross-modal Air-Ground Relative Localization

IROS 2025

Efficient, accurate, and flexible relative localization is crucial in air-ground collaborative tasks. However, current approaches for robot relative localization are primarily realized in the form of distributed multi-robot SLAM systems with the same sensor configuration, which are tightly coupled w

Cited by 0SourcecodeScholar
2024

MoGU: A Framework for Enhancing Safety of LLMs While Preserving Their Usability

NeurIPS 2024poster

Large Language Models (LLMs) are increasingly deployed in various applications. As their usage grows, concerns regarding their safety are rising, especially in maintaining harmless responses when faced with malicious instructions. Many defense strategies have been developed to enhance the safety of…

Cited by 4SourcePDFScholar
2023

MTFD: Multi-Teacher Fusion Distillation for Compressed Video Action Recognition

ICASSP 2023accepted

As an important work in computer vision, some recent representative works such as Two-stream networks, 3D ConvNets, and Transformer-based networks have achieved outstanding performance. However, due to the high computational cost, the explosion of computation time and parameters, they cannot meet th…

Cited by 0SourceScholar