← Search

Chenyu Zhou

11 accepted papers

2026

Dehallu3D: Hallucination-Mitigated 3D Generation from a Single Image via Cyclic View Consistency Refinement

CVPR 2026

Large 3D reconstruction models have revolutionized the 3D content generation field, enabling broad applications in virtual reality and gaming. Just like other large models, large 3D reconstruction models suffer from hallucinations as well, introducing structural outliers (e.g., odd holes or protrusi

Cited by 0SourceScholar
2026

MemDecoder: Enhancing Test-Time Compute for LLM Agents via Reinforced Memory Decoding

ICML 2026poster

Agentic memory—conditioning large language and vision–language models on past cases, external knowledge, or meta‑experiences—has become a key mechanism for improving inference‑time reasoning. However, existing approaches largely rely on heuristic retrieval or expensive LLM‑based reranking, and do no…

Cited by 0SourceScholar
2026

Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain Rules

ICLR 2026poster

E-commerce agents contribute greatly to helping users complete their e-commerce needs. To promote further research and application of e-commerce agents, benchmarking frameworks are introduced for evaluating LLM agents in the e-commerce domain. Despite the progress, current benchmarks lack evaluating…

Cited by 0SourceScholar
2026

Resilient UAV Swarm with Fast Connectivity Recovery and Extensive Coverage

AAAI 2026technical

To address partial node failures in unmanned aerial vehicle swarms, self-healing communication techniques are commonly employed to restore backbone connectivity while preserving area coverage. However, existing heuristic methods struggle to scale under large-scale failures and dynamic conditions, wh

Cited by 0SourcePDFScholar
2026

StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language Models

ICLR 2026poster

Large Language Models (LLMs) have shown promising capabilities for solving Operations Research (OR) problems. While reinforcement learning serves as a powerful paradigm for LLM training on OR problems, existing works generally face two key limitations. First, outcome reward suffers from the $\texti…

Cited by 0SourcecodeScholar
2025

Learning Interleaved Image-Text Comprehension in Vision-Language Large Models

ICLR 2025poster

The swift progress of Multi-modal Large Models (MLLMs) has showcased their impressive ability to tackle tasks blending vision and language. Yet, most current models and benchmarks cater to scenarios with a narrow scope of visual and textual contexts. These models often fall short when faced with com…

Cited by 0SourcePDFScholar
2025

MST-HA: Multi-Modal Signal Fusion with Bayesian Optimization for Robust Industrial Robot Joint Health Assessment

ICASSP 2025accepted

This paper presents a novel multi-modal deep learning framework for industrial robot joint health assessment and prediction, leveraging non-invasive signal fusion and Bayesian optimization. The proposed method addresses the challenges of comprehensive joint state monitoring in complex industrial env…

Cited by 0SourceScholar
2025

Meta-Conscious Driven Domain-Aware Federated Learning

ICASSP 2025accepted

Cross-domain collaboration can drive comprehensive knowledge innovation and foster synergistic advancements. Federated learning (FL) enables such collaboration while ensuring data security. However, cross-domain FL often faces challenges due to knowledge interference between domains, which can resul…

Cited by 0SourceScholar
2025

Neuron-based Multifractal Analysis of Neuron Interaction Dynamics in Large Models

ICLR 2025poster

In recent years, there has been increasing attention on the capabilities of large-scale models, particularly in handling complex tasks that small-scale models are unable to perform. Notably, large language models (LLMs) have demonstrated ``intelligent'' abilities such as complex reasoning and abstra…

2025

Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

CVPR 2025highlight

In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus remains on developing their capabilities in static image understanding. The potential of MLLMs to process sequential visual…

Cited by 368SourcePDFScholar
2023

On Strengthening and Defending Graph Reconstruction Attack with Markov Chain Approximation

ICML 2023poster

Although powerful graph neural networks (GNNs) have boosted numerous real-world applications, the potential privacy risk is still underexplored. To close this gap, we perform the first comprehensive study of graph reconstruction attack that aims to reconstruct the adjacency of nodes. We show that a…