← Search

Jiarui Zhang

20 accepted papers

2026

AReaL-DTA: Dynamic Tree Attention for Efficient Reinforcement Learning of Large Language Models

ICML 2026poster

Reinforcement learning (RL) based post-training for large language models (LLMs) is computationally expensive, as it generates many rollout sequences that could frequently share long token prefixes. Existing RL frameworks usually process these sequences independently, repeatedly recomputing identica…

Cited by 0SourceScholar
2026

DiffGraph: An Automated Agent-driven Model Merging Framework for In-the-Wild Text-to-Image Generation

CVPR 2026

The rapid growth of the text-to-image (T2I) community has fostered a thriving online ecosystem of expert models, which are variants of pretrained diffusion models specialized for diverse generative capabilities. Yet, existing model merging methods remain limited in fully leveraging abundant online e

Cited by 0SourceScholar
2026

M4Human: A Large-Scale Multimodal mmWave Radar Benchmark for Human Mesh Reconstruction

CVPR 2026

Human mesh reconstruction (HMR) provides direct insights into body-environment interaction, enabling various immersive applications. However, existing large-scale HMR benchmarks largely rely on line-of-sight RGB sensing, causing HMR systems to inherit the limitations of vision-based systems, includi

Cited by 0SourcecodeScholar
2026

SAM-Veteran: An MLLM-Based Human-like SAM Agent for Reasoning Segmentation

ICLR 2026poster

Significant progress has been made in reasoning segmentation by combining multi-modal large language models (MLLMs) with the Segment Anything Model (SAM): the former excel in reasoning and vision–language alignment, while the latter offers powerful pixel-level understanding. However, current paradig…

Cited by 0SourceScholar
2026

mmPred: Radar-based Human Motion Prediction in the Dark

AAAI 2026technical

Existing Human Motion Prediction (HMP) methods based on RGB(D) cameras are sensitive to lighting conditions and raise privacy concerns, limiting their real-world applications such as firefighting and elderly care. Motivated by the robustness and privacy-preserving nature of millimeter-wave (mmWave)

Cited by 0SourcePDFScholar
2025

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind

NeurIPS 2025poster

Large Multimodal Models (LMMs) has demonstrated capabilities across various domains, but comprehensive benchmarks for agricultural remote sensing (RS) remain scarce. Existing benchmarks designed for agricultural RS scenarios exhibit notable limitations, primarily in terms of insufficient scene diver…

Cited by 0SourcecodeScholar
2025

EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUs

AAAI 2025technical

Egocentric human pose estimation (HPE) using wearable sensors is essential for VR/AR applications. Most methods rely solely on either egocentric-view images or sparse Inertial Measurement Unit (IMU) signals, leading to inaccuracies due to self-occlusion in images or the sparseness and drift of inert…

Cited by 2SourcePDFScholar
2025

MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs

ICLR 2025poster

Multimodal Large Language Models (MLLMs) have experienced rapid progress in visual recognition tasks in recent years. Given their potential integration into many critical applications, it is important to understand the limitations of their visual perception. In this work, we study whether MLLMs can…

2025

RAGRouter: Learning to Route Queries to Multiple Retrieval-Augmented Language Models

NeurIPS 2025poster

Retrieval-Augmented Generation (RAG) significantly improves the performance of Large Language Models (LLMs) on knowledge-intensive tasks. However, varying response quality across LLMs under RAG necessitates intelligent routing mechanisms, which select the most suitable model for each query from mult…

Cited by 0SourcecodeScholar
2024

Human Palm Performance Evaluation and the Palm Design of Humanoid Robotic Hands

RA-L 2024

Losing the ability to deform the palm is unthinkable for a human hand, and the same is true for a humanoid robotic hand. Here, we aim to evaluate the performance of human palms and develop a palm-movable humanoid robotic hand with a few actuators. We quantified the palm morphological characteristics

Cited by 13SourceScholar
2024

MARVEL: Multidimensional Abstraction and Reasoning through Visual Evaluation and Learning

NeurIPS 2024poster

While multi-modal large language models (MLLMs) have shown significant progress across popular visual reasoning benchmarks, whether they possess abstract visual reasoning abilities remains an open question. Similar to the Sudoku puzzles, abstract visual reasoning (AVR) problems require finding high-…

2024

Teaching Large Language Models to Translate on Low-resource Languages with Textbook Prompting

COLING 2024main

Large Language Models (LLMs) have achieved impressive results in Machine Translation by simply following instructions, even without training on parallel data. However, LLMs still face challenges on low-resource languages due to the lack of pre-training data. In real-world situations, humans can beco…

Cited by 17SourcePDFScholar
2024

Trend-Aware Supervision: On Learning Invariance for Semi-supervised Facial Action Unit Intensity Estimation

AAAI 2024technical

With the increasing need for facial behavior analysis, semi-supervised AU intensity estimation using only keyframe annotations has emerged as a practical and effective solution to relieve the burden of annotation. However, the lack of annotations makes the spurious correlation problem caused by AU c…

Cited by 0SourcePDFScholar
2023

Deformable Model-Driven Neural Rendering for High-Fidelity 3D Reconstruction of Human Heads Under Low-View Settings

ICCV 2023poster

Reconstructing 3D human heads in low-view settings presents technical challenges, mainly due to the pronounced risk of overfitting with limited views and high-frequency signals. To address this, we propose geometry decomposition and adopt a two-stage, coarse-to-fine training strategy, allowing for p…

Cited by 9PDFcodeScholar
2023

OmniObject3D: Large-Vocabulary 3D Object Dataset for Realistic Perception, Reconstruction and Generation

CVPR 2023poster

Recent advances in modeling 3D objects mostly rely on synthetic datasets due to the lack of large-scale real-scanned 3D databases. To facilitate the development of 3D perception, reconstruction, and generation in the real world, we propose OmniObject3D, a large vocabulary 3D object dataset with mass…

Cited by 214SourcePDFScholar
2022

Clickbait Detection via Contrastive Variational Modelling of Text and Label

IJCAI 2022poster

Clickbait refers to deliberately created sensational or deceptive text for tricking readers into clicking, which severely hurts the web ecosystem. With a growing number of clickbaits on social media, developing automatic detection methods becomes essential. Nonetheless, the performance of existing n…

Cited by 6SourcePDFScholar
2019

Multiview 2D/3D Rigid Registration via a Point-Of-Interest Network for Tracking and Triangulation

CVPR 2019poster

We propose to tackle the problem of multiview 2D/3D rigid registration for intervention via a Point-Of-Interest Network for Tracking and Triangulation (POINT^2). POINT^2 learns to establish 2D point-to-point correspondences between the pre- and intra-intervention images by tracking a set of random P…

Cited by 65PDFScholar