← Search

Junjie Zhang

22 accepted papers

2026

A Simple "Motivation" Can Enhance Reinforcement Finetuning of Large Reasoning Models

ICLR 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful learn-to-reason paradigm for Large Reasoning Models to tackle complex tasks. However, current RLVR paradigm is still not efficient enough, as it works in a trial-and-error manner. To perform better, the model needs to e…

Cited by 0SourcecodeScholar
2026

Beyond Sequences: A Dynamic Hierarchical Heterogeneous Spatio-Temporal Graph for Irregular Multivariate Time Series Forecasting

IJCAI 2026

Irregular Multivariate Time Series (IMTS) analysis is a challenging task as asynchronous irregular sampling disrupts intra-variable temporal consistency and cross-variable alignment. Most existing methods model multivariate correlations at either the variable or observation level in static ways. The

Cited by 0Scholar
2026

Bridging LLMs and SAT Solving: Automated Evolution of High-Performance Heuristics

IJCAI 2026

Despite decades of intensive research and optimization, modern Boolean Satisfiability (SAT) solvers have reached a plateau where significant performance gains are increasingly difficult to achieve. While Large Language Models (LLMs) have demonstrated remarkable capabilities in pattern recognition an

Cited by 0Scholar
2026

DriveMamba: Task-Centric Scalable State Space Model for Efficient End-to-End Autonomous Driving

ICLR 2026poster

Recent advances towards End-to-End Autonomous Driving (E2E-AD) focus on integrating modular designs into a unified framework for joint optimization. Most of these advances follow a sequential paradigm (i.e., perception-prediction-planning) based on separable Transformer decoders and rely on dense BE…

Cited by 0SourceScholar
2026

Frequency-Aware Vision-Language Multimodality Generalization Network for Remote Sensing Image Classification

AAAI 2026technical

The booming remote sensing (RS) technology is giving rise to a novel multimodality generalization task, which requires the model to overcome data heterogeneity while possessing powerful cross-scene generalization ability. Moreover, most vision-language models usually describe surface materials using

Cited by 0SourcePDFScholar
2026

Generalizable Heterogeneity-aware Federated Feature and Basic-matrix Consistency Learning

AAAI 2026technical

As an emerging distributed learning paradigm, Federated Learning (FL) facilitates collaborative training among multiple clients without sharing raw data. However, the classic FL still faces significant challenges due to feature/model heterogeneity and catastrophic forgetting, which seriously hinder

Cited by 0SourcePDFScholar
2025

A Pre-trained Plug-in Mixture-of-LoRAs Model for Transferable Sequential Recommendation

ICASSP 2025accepted

The goal of transferable sequential recommendation (TSR) is to improve the performance of sequential recommenders in multiple target domains leveraging knowledge transferred from source domains. Most existing transferable sequential recommenders rely on item modality information but pay insufficient…

Cited by 0SourceScholar
2025

FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection Transformers

ICCV 2025poster

Detecting 3D objects accurately from multi-view 2D images is a challenging yet essential task in the field of autonomous driving. Current methods resort to integrating depth prediction to recover the spatial information for object query decoding, which necessitates explicit supervision from LiDAR po…

Cited by 0SourcePDFScholar
2025

Multi-modal Multi-platform Person Re-Identification: Benchmark and Method

ICCV 2025poster

Conventional person re-identification (ReID) research is often limited to single-modality sensor data from static cameras, which fails to address the complexities of real-world scenarios where multi-modal signals are increasingly prevalent. For instance, consider an urban ReID system integrating sta…

Cited by 0SourcePDFScholar
2025

NS4S: Neighborhood Search for Scheduling Problems Via Large Language Models

IJCAI 2025

Large Language Models (LLMs) have emerged as a promising technology for solving combinatorial optimization problems. However, their direct application to scheduling problems remains limited due to the inherent complexity of these problems. This paper proposes an LLMs-based neighborhood search method

2025

SimpleDeepSearcher: Deep Information Seeking via Web-Powered Reasoning Trajectory Synthesis

EMNLP 2025

Retrieval-augmented generation (RAG) systems have advanced large language models (LLMs) in complex deep search scenarios requiring multi-step reasoning and iterative information retrieval. However, existing approaches face critical limitations that lack high-quality training trajectories or suffer f

2025

Supervised Optimism Correction: Be Confident When LLMs Are Sure

ACL 2025finding

In this work, we establish a novel theoretical connection between supervised fine-tuning and offline reinforcement learning under the token-level Markov decision process, revealing that large language models indeed learn an implicit Q-function for inference.Through this theoretical lens, we demonstr…

Cited by 0SourcePDFScholar
2025

Towards Robust Sensor-Fusion Ground SLAM: A Comprehensive Benchmark and A Resilient Framework

IROS 2025

Considerable advancements have been achieved in SLAM methods tailored for structured environments, yet their robustness under challenging corner cases remains a critical limitation. Although multi-sensor fusion approaches integrating diverse sensors have shown promising performance improvements, the

Cited by 8SourceScholar
2024

3DBench: A Scalable 3D Benchmark and Instruction-Tuning Dataset

IJCAI 2024poster

Evaluating the performance of Multi-modal Large Language Models (MLLMs), integrating both point cloud and language, presents significant challenges. The lack of a comprehensive assessment hampers determining whether these models truly represent advancements, thereby impeding further progress in the…

2024

AuriSRec: Adversarial User Intention Learning in Sequential Recommendation

EMNLP 2024finding

With recommender systems broadly deployed in various online platforms, many efforts have been devoted to learning user preferences and building effective sequential recommenders. However, existing work mainly focuses on capturing user implicit preferences from historical interactions and simply matc…

Cited by 3SourcePDFScholar
2024

SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation

ICML 2024poster

Acquiring a multi-task imitation policy in 3D manipulation poses challenges in terms of scene understanding and action prediction. Current methods employ both 3D representation and multi-view 2D representation to predict the poses of the robot’s end-effector. However, they still require a considerab…

Cited by 11SourcePDFScholar
2024

Voxel or Pillar: Exploring Efficient Point Cloud Representation for 3D Object Detection

AAAI 2024technical

Efficient representation of point clouds is fundamental for LiDAR-based 3D object detection. While recent grid-based detectors often encode point clouds into either voxels or pillars, the distinctions between these approaches remain underexplored. In this paper, we quantify the differences between t…

Cited by 8SourcePDFScholar
2022

Genre-Conditioned Long-Term 3D Dance Generation Driven by Music

ICASSP 2022accepted

Dancing to music is an artistic behavior of humans, however, letting machines generate dances from music is still challenging. Most existing works have been made progress in tackling the problem of motion prediction conditioned by music, yet they rarely consider the importance of the musical genre.…

Cited by 0SourceScholar
2021

Neural Architecture Search for Joint Human Parsing and Pose Estimation

ICCV 2021poster

Human parsing and pose estimation are crucial for the understanding of human behaviors. Since these tasks are closely related, employing one unified model to perform two tasks simultaneously allows them to benefit from each other. However, since human parsing is a pixel-wise classification process w…

Cited by 26PDFcodeScholar
2021

PTN: A Poisson Transfer Network for Semi-supervised Few-shot Learning

AAAI 2021technical

The predicament in semi-supervised few-shot learning (SSFSL) is to maximize the value of the extra unlabeled data to boost the few-shot learner. In this paper, we propose a Poisson Transfer Network (PTN) to mine the unlabeled information for SSFSL from two aspects. First, the Poisson Merriman–Bence–…

Cited by 31SourcePDFScholar
2019

Mind Your Neighbours: Image Annotation With Metadata Neighbourhood Graph Co-Attention Networks

CVPR 2019poster

As the visual reflections of our daily lives, images are frequently shared on the social network, which generates the abundant 'metadata' that records user interactions with images. Due to the diverse contents and complex styles, some images can be challenging to recognise when neglecting the contex…

Cited by 25PDFScholar
2018

Goal-Oriented Visual Question Generation via Intermediate Rewards

ECCV 2018poster

Despite significant progress in a variety of vision-and-language problems, developing a method capable of asking intelligent, goal-oriented questions about images is proven to be an inscrutable challenge. Towards this end, we propose a Deep Reinforcement Learning framework based on three new interme…

Cited by 47SourcePDFScholar