← Search

Yucheng Zhou

26 accepted papers

2026

From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation

ICLR 2026poster

Combining Chain-of-Thought (CoT) with Reinforcement Learning (RL) improves text-to-image (T2I) generation, yet the underlying interaction between CoT's exploration and RL's optimization remains unclear. We present a systematic entropy-based analysis that yields three key insights: (1) CoT expands th…

Cited by 0SourcecodeScholar
2026

HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning

CVPR 2026

Recent advances in diffusion models have demonstrated impressive capability in generating high-quality images for simple prompts. However, when confronted with complex prompts involving multiple objects and hierarchical structures, existing models struggle to accurately follow instructions, leading

Cited by 0SourcecodeScholar
2026

Less Is More: Vision Representation Compression for Efficient Video Generation with Large Language Models

AAAI 2026technical

Video generation using Large Language Models (LLMs) has shown promising potential, effectively leveraging the extensive LLM infrastructure to provide a unified framework for multimodal understanding and content generation. However, these methods face critical challenges, i.e., token redundancy and i

Cited by 0SourcePDFScholar
2026

Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity Masks

AAAI 2026technical

Despite significant progress in pixel-level medical image analysis, existing medical image segmentation models rarely explore medical segmentation and diagnosis tasks jointly. However, it is crucial for patients that models can provide explainable diagnoses along with medical segmentation results. I

Cited by 0SourcePDFScholar
2025

DC-ControlNet: Decoupling Inter- and Intra-Element Conditions in Image Generation with Diffusion Models

ICCV 2025poster

In this paper, we introduce DC (Decouple)-ControlNet, a highly flexible and precisely controllable framework for multi-condition image generation. The core idea behind DC-ControlNet is to decouple control conditions, transforming global control into a hierarchical system that integrates distinct ele…

Cited by 0SourcePDFScholar
2025

Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

ACL 2025long

Existing Medical Large Vision-Language Models (Med-LVLMs), encapsulating extensive medical knowledge, demonstrate excellent capabilities in understanding medical images. However, there remain challenges in visual localization in medical images, which is crucial for abnormality detection and interpre…

Cited by 0SourcePDFScholar
2025

InsectMamba: State Space Model with Adaptive Composite Features for Insect Recognition

ICASSP 2025accepted

The recognition of insect pests is a critical task in agricultural technology, vital for ensuring food security and environmental sustainability. However, due to factors like high camouflage and species diversity, the complexity of pest identification poses significant obstacles. Existing methods st…

Cited by 0SourceScholar
2025

MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

ACL 2025finding

Recent advancements in medical Large Language Models (LLMs) have showcased their powerful reasoning and diagnostic capabilities. Despite their success, current unified multimodal medical LLMs face limitations in knowledge update costs, comprehensiveness, and flexibility. To address these challenges,…

2025

Safety Alignment via Constrained Knowledge Unlearning

ACL 2025long

Despite significant progress in safety alignment, large language models (LLMs) remain susceptible to jailbreak attacks. Existing defense mechanisms have not fully deleted harmful knowledge in LLMs, which allows such attacks to bypass safeguards and produce harmful outputs. To address this challenge,…

2025

Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation

ACL 2025finding

Text-to-image models are powerful for producing high-quality images based on given text prompts, but crafting these prompts often requires specialized vocabulary. To address this, existing methods train rewriting models with supervision from large amounts of manually annotated data and trained aesth…

Cited by 0SourcePDFScholar
2025

Semantic Causality-Aware Vision-Based 3D Occupancy Prediction

ICCV 2025poster

Vision-based 3D semantic occupancy prediction is a critical task in 3D vision that integrates volumetric 3D reconstruction with semantic understanding. Existing methods, however, often rely on modular pipelines. These modules are typically optimized independently or use pre-configured inputs, leadin…

2025

Towards Stabilized and Efficient Diffusion Transformers through Long-Skip-Connections with Spectral Constraints

ICCV 2025poster

Diffusion Transformers (DiT) have emerged as a powerful architecture for image and video generation, offering superior quality and scalability. However, their practical application suffers from inherent dynamic feature instability, leading to error amplification during cached inference. Through syst…

2025

Weak to Strong Generalization for Large Language Models with Multi-capabilities

ICLR 2025poster

As large language models (LLMs) grow in sophistication, some of their capabilities surpass human abilities, making it essential to ensure their alignment with human values and intentions, i.e., Superalignment. This superalignment challenge is particularly critical for complex tasks, as annotations p…

Cited by 88SourcePDFScholar
2024

Fine-Grained Distillation for Long Document Retrieval

AAAI 2024technical

Long document retrieval aims to fetch query-relevant documents from a large-scale collection, where knowledge distillation has become de facto to improve a retriever by mimicking a heterogeneous yet powerful cross-encoder. However, in contrast to passages or sentences, retrieval on long documents su…

Cited by 52SourcePDFScholar
2024

MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution

NeurIPS 2024poster

In software development, resolving the emergent issues within GitHub repositories is a complex challenge that involves not only the incorporation of new code but also the maintenance of existing code. Large Language Models (LLMs) have shown promise in code generation but face difficulties in resolvi…

Cited by 39SourcePDFScholar
2024

SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information

EMNLP 2024main

Large Vision-Language Models (LVLMs) have become pivotal at the intersection of computer vision and natural language processing. However, the full potential of LVLMs’ Retrieval-Augmented Generation (RAG) capabilities remains underutilized. Existing works either focus solely on the text modality or a…

2024

Visual In-Context Learning for Large Vision-Language Models

ACL 2024findings

In Large Visual Language Models (LVLMs), the efficacy of In-Context Learning (ICL) remains limited by challenges in cross-modal interactions and representation disparities. To overcome these challenges, we introduce a novel Visual In-Context Learning (VICL) method comprising Visual Demonstration Ret…

Cited by 106SourcePDFScholar
2023

Towards Robust Ranker for Text Retrieval

ACL 2023findings

A neural ranker plays an indispensable role in the de facto ‘retrieval & rerank’ pipeline, but its training still lags behind due to the weak negative mining during contrastive learning. Compared to retrievers boosted by self-adversarial (i.e., in-distribution) negative mining, the ranker’s heavy st…

Cited by 53SourcePDFScholar
2022

ASDOT: Any-Shot Data-to-Text Generation with Pretrained Language Models

EMNLP 2022finding

Data-to-text generation is challenging due to the great variety of the input data in terms of domains (e.g., finance vs sports) or schemata (e.g., diverse predicates). Recent end-to-end neural methods thus require substantial training examples to learn to disambiguate and describe the data. Yet, rea…

2022

ClarET: Pre-training a Correlation-Aware Context-To-Event Transformer for Event-Centric Generation and Classification

ACL 2022long

Generating new events given context with correlated ones plays a crucial role in many event-centric reasoning tasks. Existing works either limit their scope to specific scenarios or overlook event-level correlations. In this paper, we propose to pre-train a general Correlation-aware context-to-Event…

2022

Sketch Storytelling

ICASSP 2022accepted

Sketch storytelling aims to generate a story for a given sketch. Although image captioning based on deep learning has great progress, describing the sketch in a story style is still a challenge. The reason is that there is currently no paired sketch-story data which is expensive to acquire. Therefor…

Cited by 0SourceScholar
2021

Improving Zero-Shot Cross-lingual Transfer for Multilingual Question Answering over Knowledge Graph

NAACL 2021long

Multilingual question answering over knowledge graph (KGQA) aims to derive answers from a knowledge graph (KG) for questions in multiple languages. To be widely applicable, we focus on its zero-shot transfer setting. That is, we can only access training data in a high-resource language, while need t…