← Search

Zhuohang Dang

10 accepted papers

2026

AutoGPS: Automated Geometry Problem Solving via Multimodal Formalization and Deductive Reasoning

ICLR 2026poster

Geometry problem solving presents distinctive challenges in artificial intelligence, requiring exceptional multimodal comprehension and rigorous mathematical reasoning capabilities. Existing approaches typically fall into two categories: neural-based and symbolic-based methods, both of which exhibit…

Cited by 0SourcecodeScholar
2026

Beyond Layer-Wise Merging: Chain-of-Merging for Vision-Language Models

CVPR 2026

While model merging has demonstrated remarkable success for large language models (LLMs), its application to vision-language models (VLMs) remains largely underexplored. Recent methods attempt to enhance VLM reasoning capabilities by integrating specialized LLM parameters through layer-wise merging.

Cited by 0SourceScholar
2026

Correspondence Coverage Matters for Multi-Modal Dataset Distillation

AAAI 2026technical

Multi-modal dataset distillation (DD) condenses large datasets into compact ones that retain task efficacy by capturing correspondence patterns, i.e., shared semantics between paired modalities. However, such patterns rely on cross-modal similarity and cannot be faithfully captured by intra-modal si

Cited by 0SourcePDFScholar
2026

PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling

CVPR 2026

Consistent image generation requires faithfully preserving identities, styles, and logical coherence across multiple images,which is essential for applications such as storytelling and character design.Supervised training approaches struggle with this task due to the lack of large-scale datasets cap

Cited by 0SourcecodeScholar
2025

AgentStore: Scalable Integration of Heterogeneous Agents As Specialized Generalist Computer Assistant

ACL 2025finding

Digital agents capable of automating complex computer tasks have attracted considerable attention. However, existing agent methods exhibit deficiencies in their generalization and specialization capabilities, especially in handling open-ended computer tasks in real-world environments. Inspired by th…

Cited by 0SourcePDFScholar
2025

ChatGen: Automatic Text-to-Image Generation From FreeStyle Chatting

CVPR 2025poster

Despite the significant advancements in text-to-image (T2I) generative models, users often face a trial-and-error challenge in practical scenarios. This challenge arises from the complexity and uncertainty of tedious steps such as crafting suitable prompts, selecting appropriate models, and configur…

Cited by 1SourcePDFScholar
2025

CoFFT: Chain of Foresight-Focus Thought for Visual Language Models

NeurIPS 2025poster

Despite significant advances in Vision Language Models (VLMs), they remain constrained by the complexity and redundancy of visual input. When images contain large amounts of irrelevant information, VLMs are susceptible to interference, thus generating excessive task-irrelevant reasoning processes or…

Cited by 0SourceScholar
2024

Noisy Correspondence Learning with Self-Reinforcing Errors Mitigation

AAAI 2024technical

Cross-modal retrieval relies on well-matched large-scale datasets that are laborious in practice. Recently, to alleviate expensive data collection, co-occurring pairs from the Internet are automatically harvested for training. However, it inevitably includes mismatched pairs, i.e., noisy corresponde…

Cited by 7SourcePDFScholar
2024

SSMG: Spatial-Semantic Map Guided Diffusion Model for Free-Form Layout-to-Image Generation

AAAI 2024technical

Despite significant progress in Text-to-Image (T2I) generative models, even lengthy and complex text descriptions still struggle to convey detailed controls. In contrast, Layout-to-Image (L2I) generation, aiming to generate realistic and complex scene images from user-specified layouts, has risen to…

Cited by 16SourcePDFScholar
2023

Towards Real-Time Person Search with Invariant Feature Learning

ICASSP 2023accepted

Person search aims to locate a query person in a gallery of unconstrained scene images, which has many real-world applications. However, existing methods directly build off of advances in object detection for better performance rather than efficiency. Complex designs in heavy-weight detectors are re…

Cited by 0SourceScholar