← Search

Jingke Meng

8 accepted papers

2025

ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction Generation

CVPR 2025poster

We propose ChainHOI, a novel approach for text-driven human-object interaction (HOI) generation that explicitly models interactions at both the joint and kinetic chain levels. Unlike existing methods that implicitly model interactions using full-body poses as tokens, we argue that explicitly mode…

Cited by 2SourcePDFScholar
2025

Distilling LLM Prior to Flow Model for Generalizable Agent’s Imagination in Object Goal Navigation

NeurIPS 2025poster

The Object Goal Navigation (ObjectNav) task challenges agents to locate a specified object in an unseen environment by imagining unobserved regions of the scene. Prior approaches rely on deterministic and discriminative models to complete semantic maps, overlooking the inherent uncertainty in indoor…

Cited by 0SourceScholar
2025

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

CVPR 2025highlight

Recent open-vocabulary detectors achieve promising performance with abundant region-level annotated data. In this work, we show that an open-vocabulary detector co-training with a large language model by generating image-level detailed captions for each image can further improve performance. To achi…

2025

Person De-reidentification: A Variation-guided Identity Shift Modeling

CVPR 2025poster

Person re-identification (ReID) is to associate images of individuals from different camera views against cross-view variations. Like other surveillance technologies, Re-ID faces serious privacy challenges, particularly the potential for unauthorized tracking. Although various tasks (e.g., face reco…

Cited by 0SourcePDFScholar
2025

VIPerson: Flexibly Generating Virtual Identity for Person Re-Identification

ICCV 2025poster

Person re-identification (ReID) is to match the person images under different camera views. Training ReID models necessitates a substantial amount of labeled real-world data, leading to high labeling costs and privacy issues. Although several ReID data synthetic methods are proposed to address these…

2025

monoVLN: Bridging the Observation Gap between Monocular and Panoramic Vision and Language Navigation

ICCV 2025poster

Vision and Language Navigation(VLN) requires agents to navigate 3D environments by following natural language instructions. While existing methods predominantly assume access to panoramic observations, many practical robotics are equipped with monocular RGBD cameras, creating a significant configura…

Cited by 0SourcePDFScholar
2023

Event-Guided Procedure Planning from Instructional Videos with Text Supervision

ICCV 2023poster

In this work, we focus on the task of procedure planning from instructional videos with text supervision, where a model aims to predict an action sequence to transform the initial visual state into the goal visual state. A critical challenge of this task is the large semantic gap between observed vi…

Cited by 20PDFScholar