← Search

Kai Jia

7 accepted papers

2026

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance

ICML 2026poster

Large language models (LLMs) have recently advanced in reasoning when optimized with reinforcement learning (RL) under verifiable rewards. Existing methods primarily rely on outcome-based supervision to strengthen internal LLM reasoning, often leading to inefficient exploration and sparse rewards. T…

Cited by 0SourceScholar
2026

PosterAgent: Agentic Poster Generation via Stage-Aware Reinforcement Learning

ICML 2026poster

Poster generation is a complex task demanding a harmonious integration of visual aesthetics and information hierarchy. While recent text-to-image models have advanced visual synthesis, they remain non-editable and struggle with precise text rendering. Conversely, existing layout-generation methods o…

Cited by 0SourceScholar
2026

Thinking with Programming Vision: Towards a Unified View for Thinking with Images

CVPR 2026

Multimodal large language models (MLLMs) that "think with images" can interactively use tools to reason about visual inputs, but current approaches often rely on a narrow set of tools with limited real-world necessity and scalability. In this work, we first reveal a critical and previously overlooke

Cited by 0SourcecodeScholar
2025

PrimHOI: Compositional Human-Object Interaction via Reusable Primitives

ICCV 2025accepted

Synthesizing realistic Human-Object Interaction (HOI) motions is essential for creating believable digital characters and intelligent robots. Existing approaches rely on data-intensive learning models that struggle with the compositional structure of daily HOI motions, particularly for complex multi…

Cited by 0SourcePDFScholar
2023

Delving Deep into Pixel Alignment Feature for Accurate Multi-View Human Mesh Recovery

AAAI 2023technical

Regression-based methods have shown high efficiency and effectiveness for multi-view human mesh recovery. The key components of a typical regressor lie in the feature extraction of input views and the fusion of multi-view features. In this paper, we present Pixel-aligned Feedback Fusion (PaFF) for a…

2018

MegDet: A Large Mini-Batch Object Detector

CVPR 2018poster

The development of object detection in the era of deep learning, from R-CNN [11], Fast/Faster R-CNN [10, 31] to recent Mask R-CNN [14] and RetinaNet [24], mainly come from novel network, new framework, or loss design. How- ever, mini-batch size, a key factor for the training of deep neural networks,…

Cited by 408SourcePDFScholar