← Search

Zerui Chen

14 accepted papers

2026

From Retrieval to Translation: Translating Query into Graph-level Clues for Retrieval-Augmented Generation

ICML 2026poster

Retrieval-Augmented Generation (RAG) has recently been enhanced with tree or graph structures to match user intent for precise passage retrieval, which facilitates large language models (LLMs) in effectively mitigating hallucinations by leveraging external knowledge. However, we identify that existi…

Cited by 0SourceScholar
2025

HORT: Monocular Hand-held Objects Reconstruction with Transformers

ICCV 2025poster

Reconstructing hand-held objects in 3D from monocular images remains a significant challenge in computer vision. Most existing approaches rely on implicit 3D representations, which produce overly smooth reconstructions and are time-consuming to generate explicit 3D shapes. While more recent methods…

Cited by 0SourcePDFScholar
2025

How do Language Models Reshape Entity Alignment? A Survey of LM-Driven EA Methods: Advances, Benchmarks, and Future

EMNLP 2025

Entity alignment (EA), critical for knowledge graph (KG) integration, identifies equivalent entities across different KGs. Traditional methods often face challenges in semantic understanding and scalability. The rise of language models (LMs), particularly large language models (LLMs), has provided p

Cited by 0SourcePDFScholar
2025

Simulation-Free Hierarchical Latent Policy Planning for Proactive Dialogues

AAAI 2025technical

Recent advancements in proactive dialogues have garnered significant attention, particularly for more complex objectives (e.g. emotion support and persuasion). Unlike traditional task-oriented dialogues, proactive dialogues demand advanced policy planning and adaptability, requiring rich scenarios a…

Cited by 1SourcePDFScholar
2025

ViViDex: Learning Vision-Based Dexterous Manipulation from Human Videos

ICRA 2025

In this work, we aim to learn a unified vision-based policy for multi-fingered robot hands to manipulate a variety of objects in diverse poses. Though prior work has shown benefits of using human videos for policy learning, performance gains have been limited by the noise in estimated trajectories.

Cited by 36SourcecodeScholar
2024

GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension

IJCAI 2024poster

There are substantial instructional videos on the Internet, which provide us tutorials for completing various tasks. Existing instructional video datasets only focus on specific steps at the video level, lacking experiential guidelines at the task level, which can lead to beginners struggling to lea…

Cited by 1SourcePDFScholar
2024

Infrared-LLaVA: Enhancing Understanding of Infrared Images in Multi-Modal Large Language Models

EMNLP 2024finding

Expanding the understanding capabilities of multi-modal large language models (MLLMs) for infrared modality is a challenge due to the single-modality nature and limited amount of training data. Existing methods typically construct a uniform embedding space for cross-modal alignment and leverage abun…

Cited by 1SourcePDFScholar
2024

Learning Explicit Contact for Implicit Reconstruction of Hand-Held Objects from Monocular Images

AAAI 2024technical

Reconstructing hand-held objects from monocular RGB images is an appealing yet challenging task. In this task, contacts between hands and objects provide important cues for recovering the 3D geometry of the hand-held objects. Though recent works have employed implicit functions to achieve impressive…

2024

Planning Like Human: A Dual-process Framework for Dialogue Planning

ACL 2024long

In proactive dialogue, the challenge lies not just in generating responses but in steering conversations toward predetermined goals, a task where Large Language Models (LLMs) typically struggle due to their reactive nature. Traditional approaches to enhance dialogue planning in LLMs, ranging from el…

2024

Pseudo Labels Regularization for Imbalanced Partial-Label Learning

ICASSP 2024accepted

Partial-label learning (PLL) is an important branch of weakly supervised learning where the single ground truth resides in a set of candidate labels, while the research rarely considers the label imbalance. A recent study for imbalanced PLL propose that the combinatorial challenge of partial-label l…

Cited by 0SourceScholar
2023

gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object Reconstruction

CVPR 2023poster

Signed distance functions (SDFs) is an attractive framework that has recently shown promising results for 3D shape reconstruction from images. SDFs seamlessly generalize to different shape resolutions and topologies but lack explicit modelling of the underlying 3D geometry. In this work, we exploit…

2022

AlignSDF: Pose-Aligned Signed Distance Fields for Hand-Object Reconstruction

ECCV 2022poster

"Recent work achieved impressive progress towards joint reconstruction of hands and manipulated objects from monocular color images. Existing methods focus on two alternative representations in terms of either parametric meshes or signed distance fields (SDFs). On one side, parametric models can ben…

2020

Prediction and Recovery for Adaptive Low-Resolution Person Re-Identification

ECCV 2020poster

Low-resolution person re-identification (LR re-id) is a challenging task with low-resolution probes and high-resolution gallery images. To address the resolution mismatch, existing methods typically recover missing details for low-resolution probes by super-resolution. However, they usually pre-spec…

Cited by 32SourcePDFScholar
2020

Towards Part-aware Monocular 3D Human Pose Estimation: An Architecture Search Approach

ECCV 2020poster

Even though most existing monocular 3D pose estimation approaches achieve very competitive results, they ignore the heterogeneity among human body parts by estimating them with the same network architecture. To accurately estimate 3D poses of different body parts, we attempt to build a part-aware 3D…

Cited by 32SourcePDFScholar