← Search

Zhiyuan Ma

41 accepted papers

2026

AlignCVC: Aligning Cross-View Consistency for Single-Image-to-3D Generation

AAAI 2026technical

Single-image-to-3D models typically follow a sequential generation and reconstruction workflow. However, intermediate multi-view images synthesized by pre-trained generation models often lack cross-view consistency (CVC), significantly degrading 3D reconstruction performance. While recent methods at

Cited by 0SourcePDFScholar
2026

CCAHCL: Multi-Level Hypergraph Contrastive Learning for Connected Component Awareness

AAAI 2026technical

Hypergraph contrastive learning has emerged as a powerful unsupervised paradigm for hypergraph representation learning. Traditional hypergraph contrastive learning methods typically leverage neighbor aggregation strategy to obtain entity (node and hyperedge) representations within each connected com

Cited by 0SourcePDFScholar
2026

CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability

ICML 2026oral

Evaluating and improving the security capabilities of code agents requires high-quality, executable vulnerability tasks. However, existing works rely on costly, unscalable manual reproduction and suffer from outdated data distributions. To address these, we present CVE-Factory, the first multi-agent…

Cited by 0SourceScholar
2026

MARTI: A Framework for Multi-Agent LLM Systems Reinforced Training and Inference

ICLR 2026poster

We present MARTI (Multi-Agent Reinforced Training and Inference), an open-source framework designed to facilitate scalable and efficient learning of multi-agent LLM systems. MARTI supports centralized multi-agent interactions and distributed policy training, with the added capability of multi-turn a…

Cited by 0SourcecodeScholar
2026

OPERA: A Reinforcement Learning--Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop Retrieval

AAAI 2026technical

Recent advances in large language models (LLMs) and dense retrievers have driven significant progress in retrieval-augmented generation (RAG). However, existing approaches face significant challenges in complex reasoning-oriented multi-hop retrieval tasks: 1) Ineffective reasoning-oriented planning:

Cited by 0SourcePDFScholar
2026

One2Scene: Geometric Consistent Explorable 3D Scene Generation from a Single Image

ICLR 2026poster

Generating explorable 3D scenes from a single image is a highly challenging problem in 3D vision. Existing methods struggle to support free exploration, often producing severe geometric distortions and noisy artifacts when the viewpoint moves far from the original perspective. We introduce One2Scene…

Cited by 0SourcecodeScholar
2026

Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement

CVPR 2026

Although recent 3D-native generators have made great progress in synthesizing reliable geometry, they still fall short in achieving realistic appearances. A key obstacle lies in the lack of diverse and high-quality real-world 3D assets with rich surface details, since capturing such data is intrinsi

Cited by 0SourcecodeScholar
2025

Automated Creation of Reusable and Diverse Toolsets for Enhancing LLM Reasoning

AAAI 2025technical

Augmenting large language models (LLMs) with tools significantly enhances their problem-solving potential across multifaceted tasks. However, current tools automatically created by LLMs often serve as a mere summary of specific problems or solutions, which face two main issues: 1) Low reusability:…

2025

CADGrasp: Learning Contact and Collision Aware General Dexterous Grasping in Cluttered Scenes

NeurIPS 2025poster

Dexterous grasping in cluttered environments presents substantial challenges due to the high degrees of freedom of dexterous hands, occlusion, and potential collisions arising from diverse object geometries and complex layouts. To address these challenges, we propose CADGrasp, a two-stage algorithm…

Cited by 0SourceScholar
2025

Gumbel Reranking: Differentiable End-to-End Reranker Optimization

ACL 2025long

RAG systems rely on rerankers to identify relevant documents. However, fine-tuning these models remains challenging due to the scarcity of annotated query-document pairs. Existing distillation-based approaches suffer from training-inference misalignment and fail to capture interdependencies among ca…

Cited by 0SourcePDFScholar
2025

MVBoost: Boost 3D Reconstruction with Multi-View Refinement

CVPR 2025poster

Recent advancements in 3D object reconstruction have been remarkable, yet most current 3D models rely heavily on existing 3D datasets. The scarcity of diverse 3D datasets results in limited generalization capabilities of 3D reconstruction models. In this paper, we propose a novel framework for boost…

2025

Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach

CVPR 2025poster

Diffusion prior-based methods have shown impressive results in real-world image super-resolution (SR). However, most existing methods entangle pixel-level and semantic-level SR objectives in the training process, struggling to balance pixel-wise fidelity and perceptual quality. Meanwhile, users have…

2025

Progressive Rendering Distillation: Adapting Stable Diffusion for Instant Text-to-Mesh Generation without 3D Data

CVPR 2025poster

It is highly desirable to obtain a model that can generate high-quality 3D meshes from text prompts in just seconds. While recent attempts have adapted pre-trained text-to-image diffusion models, such as Stable Diffusion (SD), into generators of 3D representations (e.g., Triplane), they often suffer…

2025

Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines

AAAI 2025technical

Retrieval-augmented generation (RAG) has emerged to address the knowledge-intensive visual question answering (VQA) task. Current methods mainly employ separate retrieval and generation modules to acquire external knowledge and generate answers, respectively. We propose ReAuSE, an alternative to the…

2025

UniTransfer: Video Concept Transfer via Progressive Spatio-Temporal Decomposition

NeurIPS 2025poster

Recent advancements in video generation models have enabled the creation of diverse and realistic videos, with promising applications in advertising and film production. However, as one of the essential tasks of video generation models, video concept transfer remains significantly challenging. Exist…

Cited by 0SourceScholar
2025

VideoDirector: Precise Video Editing via Text-to-Video Models

CVPR 2025poster

Despite the typical inversion-then-editing paradigm using text-to-image (T2I) models has demonstrated promising results, directly extending it to text-to-video (T2V) models still suffers severe artifacts such as color flickering and content distortion. Consequently, current video editing methods pri…

Cited by 0SourcePDFScholar
2025

Zero-Shot Blind-Spot Image Denoising via Cross-Scale Non-Local Pixel Refilling

NeurIPS 2025poster

Blind-spot denoising (BSD) method is a powerful paradigm for zero-shot image denoising by training models to predict masked target pixels from their neighbors. However, they struggle with real-world noise exhibiting strong local correlations, where efforts to suppress noise correlation often weaken…

Cited by 0SourceScholar
2025

Zero-Shot Blind-spot Image Denoising via Implicit Neural Sampling

CVPR 2025poster

The blind-spot principle has been a widely used tool in zero-shot image denoising but faces challenges with real-world noise that exhibits strong local correlations. Existing methods focus on reducing noise correlation, which also weaken the pixel correlations needed for accurately estimating missin…

Cited by 0SourcePDFScholar
2024

AdapEdit: Spatio-Temporal Guided Adaptive Editing Algorithm for Text-Based Continuity-Sensitive Image Editing

AAAI 2024technical

With the great success of text-conditioned diffusion models in creative text-to-image generation, various text-driven image editing approaches have attracted the attentions of many researchers. However, previous works mainly focus on discreteness-sensitive instructions such as adding, removing or re…

2024

Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding

ECCV 2024poster

"Recent vision-language pre-training models have exhibited remarkable generalization ability in zero-shot recognition tasks. Previous open-vocabulary 3D scene understanding methods mostly focus on training 3D models using either image or text supervision while neglecting the collective strength of a…

2024

Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models

CVPR 2024poster

With the emergence of pre-trained vision-language models like CLIP how to adapt them to various downstream classification tasks has garnered significant attention in recent research. The adaptation strategies can be typically categorized into three paradigms: zero-shot adaptation few-shot adaptation…

2024

Enhancing Distantly Supervised Named Entity Recognition with Strong Label Guided Lottery Training

COLING 2024main

In low-resource Named Entity Recognition (NER) scenarios, only a limited quantity of strongly labeled data is available, while a vast amount of weakly labeled data can be easily acquired through distant supervision. However, weakly labeled data may fail to improve the model performance or even harm…

Cited by 0SourcePDFScholar
2024

Exploring Adversarial Robustness of Deep State Space Models

NeurIPS 2024poster

Deep State Space Models (SSMs) have proven effective in numerous task scenarios but face significant security challenges due to Adversarial Perturbations (APs) in real-world deployments. Adversarial Training (AT) is a mainstream approach to enhancing Adversarial Robustness (AR) and has been validate…

2024

Generative Multi-Modal Knowledge Retrieval with Large Language Models

AAAI 2024technical

Knowledge retrieval with multi-modal queries plays a crucial role in supporting knowledge-intensive multi-modal applications. However, existing methods face challenges in terms of their effectiveness and training efficiency, especially when it comes to training and integrating multiple retrievers to…

2024

LMD: Faster Image Reconstruction with Latent Masking Diffusion

AAAI 2024technical

As a class of fruitful approaches, diffusion probabilistic models (DPMs) have shown excellent advantages in high-resolution image reconstruction. On the other hand, masked autoencoders (MAEs), as popular self-supervised vision learners, have demonstrated simpler and more effective image reconstructi…

2024

Mirror-Consistency: Harnessing Inconsistency in Majority Voting

EMNLP 2024finding

Self-Consistency, a widely-used decoding strategy, significantly boosts the reasoning capabilities of Large Language Models (LLMs). However, it depends on the plurality voting rule, which focuses on the most frequent answer while overlooking all other minority responses. These inconsistent minority…

Cited by 2SourcePDFScholar
2024

Neural Residual Diffusion Models for Deep Scalable Vision Generation

NeurIPS 2024poster

The most advanced diffusion models have recently adopted increasingly deep stacked networks (e.g., U-Net or Transformer) to promote the generative emergence capabilities of vision generation models similar to large language models (LLMs). However, progressively deeper stacked networks will intuitive…

Cited by 5SourcePDFScholar
2024

One-Step Effective Diffusion Network for Real-World Image Super-Resolution

NeurIPS 2024poster

The pre-trained text-to-image diffusion models have been increasingly employed to tackle the real-world image super-resolution (Real-ISR) problem due to their powerful generative image priors. Most of the existing methods start from random noise to reconstruct the high-quality (HQ) image under the g…

2024

UltraMedical: Building Specialized Generalists in Biomedicine

NeurIPS 2024spotlight

Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains and are moving towards more specialized areas. Recent advanced proprietary models such as GPT-4 and Gemini have achieved significant advancements in biomedicine, which have also raised privacy and security…

2023

HybridPrompt: Bridging Language Models and Human Priors in Prompt Tuning for Visual Question Answering

AAAI 2023technical

Visual Question Answering (VQA) aims to answer the natural language question about a given image by understanding multimodal content. However, the answer quality of most existing visual-language pre-training (VLP) methods is still limited, mainly due to: (1) Incompatibility. Upstream pre-training ta…

2023

Noise-Robust Training with Dynamic Loss and Contrastive Learning for Distantly-Supervised Named Entity Recognition

ACL 2023findings

Distantly-supervised named entity recognition (NER) aims at training networks with distantly-labeled data, which is automatically obtained by matching entity mentions in the raw text with entity types in a knowledge base. Distant supervision may induce incomplete and noisy labels, so recent state-of…

Cited by 6SourcePDFScholar
2023

OTAvatar: One-Shot Talking Face Avatar With Controllable Tri-Plane Rendering

CVPR 2023poster

Controllability, generalizability and efficiency are the major objectives of constructing face avatars represented by neural implicit field. However, existing methods have not managed to accommodate the three requirements simultaneously. They either focus on static portraits, restricting the represe…

2023

Weakly Supervised Referring Expression Grounding via Target-Guided Knowledge Distillation

ICRA 2023poster

Weakly supervised referring expression grounding aims to train a model without the manual labels between image regions and referring expressions during the training phase. Current predominant models often adopt deep structures to reconstruct the region-expression correspondence. A crucial deficiency…

Cited by 4SourcecodeScholar
2022

GLAF: Global-to-Local Aggregation and Fission Network for Semantic Level Fact Verification

COLING 2022main

Accurate fact verification depends on performing fine-grained reasoning over crucial entities by capturing their latent logical relations hidden in multiple evidence clues, which is generally lacking in existing fact verification models. In this work, we propose a novel Global-to-Local Aggregation a…

2022

UniTranSeR: A Unified Transformer Semantic Representation Framework for Multimodal Task-Oriented Dialog System

ACL 2022long

As a more natural and intelligent interaction manner, multimodal task-oriented dialog system recently has received great attention and many remarkable progresses have been achieved. Nevertheless, almost all existing studies follow the pipeline to first learn intra-modal features separately and then…

2021

Intention Reasoning Network for Multi-Domain End-to-end Task-Oriented Dialogue

EMNLP 2021main

Recent years has witnessed the remarkable success in end-to-end task-oriented dialog system, especially when incorporating external knowledge information. However, the quality of most existing models’ generated response is still limited, mainly due to their lack of fine-grained reasoning on determin…

2021

Model Adaptation through Hypothesis Transfer with Gradual Knowledge Distillation

IROS 2021poster

The ability to adapt their perception to changing environments is a core characterization of intelligent robots. At present, Unsupervised Domain Adaptation (UDA) methods are used to address this problem where the adaptation task is formulated as a transfer problem from a well-described scenario (sou…

Cited by 21SourceScholar