← Search

Zhe Hu

26 accepted papers

2025

Debate-to-Write: A Persona-Driven Multi-Agent Framework for Diverse Argument Generation

COLING 2025main

Writing arguments is a challenging task for both humans and machines. It entails incorporating high-level beliefs from various perspectives on the topic, along with deliberate reasoning and planning to construct a coherent narrative. Current language models often generate outputs autoregressively, l…

2025

Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers

EMNLP 2025

While language models (LMs) paired with residual vector quantization (RVQ) tokenizers have shown promise in text-to-audio (T2A) generation, they still lag behind diffusion-based models by a non-trivial margin. We identify a critical dilemma underpinning this gap: incorporating more RVQ layers improv

2025

Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning

NeurIPS 2025poster

Vision Language Models exhibit impressive performance for various tasks, yet they often lack the sophisticated situational reasoning required for complex decision-making. This paper shows that VLMs can achieve surprisingly strong decision-making performance when visual scenes are replaced by textual…

Cited by 0SourceScholar
2025

Synchronized Video-to-Audio Generation via Mel Quantization-Continuum Decomposition

CVPR 2025poster

Video-to-audio generation is essential for synthesizing realistic audio tracks that synchronize effectively with silent videos.Following the perspective of extracting essential signals from videos that can precisely control the mature text-to-audio generative diffusion models, this paper presents ho…

Cited by 0SourcePDFScholar
2024

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions

NeurIPS 2024oral

Recent advancements in large vision language models have demonstrated remarkable proficiency across a wide range of tasks. Yet, these models still struggle with understanding the nuances of human humor through juxtaposition, particularly when it involves nonlinear narratives that underpin many joke…

Cited by 4SourcePDFScholar
2024

Language-Augmented Symbolic Planner for Open-World Task Planning

RSS 2024poster

Enabling robotic agents to perform complex long-horizon tasks has been a long-standing goal in robotics and artificial intelligence (AI). Despite the potential shown by large language models (LLMs), their planning capabilities remain limited to short-horizon tasks and they are unable to replace the…

2022

MOCHA: A Multi-Task Training Approach for Coherent Text Generation from Cognitive Perspective

EMNLP 2022main

Teaching neural models to generate narrative coherent texts is a critical problem. Recent pre-trained language models have achieved promising results, but there is still a gap between human written texts and machine-generated outputs. In this work, we propose a novel multi-task training strategy for…

2022

PLANET: Dynamic Content Planning in Autoregressive Transformers for Long-form Text Generation

ACL 2022long

Despite recent progress of pre-trained language models on generating fluent text, existing methods still suffer from incoherence problems in long-form text generation tasks that require proper content control and planning to form a coherent high-level logical flow. In this work, we propose PLANET, a…

Cited by 42SourcePDFScholar
2021

A Computational Framework for Robot Hand Design via Reinforcement Learning

IROS 2021poster

Robot hand is essential for a fully functional robot and designing a good robot hand is a sophisticated job that challenges the designer’s knowledge and experience. This paper presents a computational framework for automatic optimal robot hand design based on reinforcement learning (RL), which consi…

Cited by 8SourceScholar
2020

An Actor-Critic Approach for Legible Robot Motion Planner

ICRA 2020poster

In human-robot collaboration, it is crucial for the robot to make its intentions clear and predictable to the human partners. Inspired by the mutual learning and adaptation of human partners, we suggest an actor-critic approach for a legible robot motion planner. This approach includes two neural ne…

Cited by 24SourceScholar
2020

Multi-Scale Boosted Dehazing Network With Dense Feature Fusion

CVPR 2020poster

In this paper, we propose a Multi-Scale Boosted Dehazing Network with Dense Feature Fusion based on the U-Net architecture. The proposed method is designed based on two principles, boosting and error feedback, and we show that they are suitable for the dehazing problem. By incorporating the Strength…

Cited by 1034PDFcodeScholar
2020

Non-Local Spatial Propagation Network for Depth Completion

ECCV 2020poster

In this paper, we propose a robust and efficient end-to-end non-local spatial propagation network for depth completion. The proposed network takes RGB and sparse depth images as inputs and estimates non-local neighbors and their affinities of each pixel, as well as an initial depth map with pixel-wi…

2017

Grasp quality evaluation and planning for objects with negative curvature

ICRA 2017poster

We consider the problem of grasping concave objects, i.e., objects whose surface includes regions with negative curvature. When a multifingered hand is used to restrain these objects, these areas can be advantageously used to determine grasps capable of more robustly resisting to external disturbanc…

Cited by 3SourceScholar
2016

A Comparative Study for Single Image Blind Deblurring

CVPR 2016spotlight

Numerous single image blind deblurring algorithms have been proposed to restore latent sharp images under camera motion. However, these algorithms are mainly evaluated using either synthetic datasets or few selected real blurred images. It is thus unclear how these algorithms would perform on images…

Cited by 524PDFScholar