← Search

Junle Wang

17 accepted papers

2026

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech

ICML 2026poster

While existing text-to-speech (TTS) models exhibit high expressiveness, fine-grained control over composite instructions remains challenging due to the structural mismatch between discrete textual intents and continuous acoustic realizations. Inspired by human cognitive decoupling, we introduce Agen…

Cited by 0SourceScholar
2026

LongHorizonUI: A Unified Framework for Robust long-horizon Task Automation of GUI Agent

ICLR 2026poster

Although agents based on multimodal large language models (MLLMs) demonstrate proficiency in general short-term graphical user interface (GUI) tasks, their robustness remains a significant challenge for handling complex long-horizon tasks in dynamic environments . In response, the LongHorizonUI fram…

Cited by 0SourcecodeScholar
2026

Progressive Subexpression Reuse in Symbolic Regression: Insights from RL-based Search and a Genetic Programming Realization

IJCAI 2026

Symbolic regression (SR) aims to recover compact and interpretable mathematical expressions from data. Genetic programming (GP) directly searches over symbolic structures, but its population dynamics can make it difficult to reliably preserve and accumulate useful subexpressions. In contrast, reinfo

Cited by 0Scholar
2025

Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning

NeurIPS 2025poster

Process reward model (PRM) has been proven effective in test-time scaling of LLM on challenging reasoning tasks. However, the reward hacking induced by PRM hinders its successful applications in reinforcement fine-tuning. We find the primary cause of reward hacking induced by PRM is that: the canoni…

Cited by 0SourcecodeScholar
2024

Attacking Transformers with Feature Diversity Adversarial Perturbation

AAAI 2024technical

Understanding the mechanisms behind Vision Transformer (ViT), particularly its vulnerability to adversarial perturbations, is crucial for addressing challenges in its real-world applications. Existing ViT adversarial attackers rely on labels to calculate the gradient for perturbation, and exhibit lo…

Cited by 5SourcePDFScholar
2024

Dynamic Feature Pruning and Consolidation for Occluded Person Re-identification

AAAI 2024technical

Occluded person re-identification (ReID) is a challenging problem due to contamination from occluders. Existing approaches address the issue with prior knowledge cues, such as human body key points and semantic segmentations, which easily fail in the presence of heavy occlusion and other humans as o…

2024

FreeMan: Towards Benchmarking 3D Human Pose Estimation under Real-World Conditions

CVPR 2024poster

Estimating the 3D structure of the human body from nat- ural scenes is a fundamental aspect of visual perception. 3D human pose estimation is a vital step in advancing fields like AIGC and human-robot interaction serving as a crucial tech- nique for understanding and interacting with human actions i…

2024

Narrative Action Evaluation with Prompt-Guided Multimodal Interaction

CVPR 2024poster

In this paper we investigate a new problem called narrative action evaluation (NAE). NAE aims to generate professional commentary that evaluates the execution of an action. Unlike traditional tasks such as score-based action quality assessment and video captioning involving superficial sentences NAE…

2023

C2F2NeUS: Cascade Cost Frustum Fusion for High Fidelity and Generalizable Neural Surface Reconstruction

ICCV 2023poster

There is an emerging effort to combine the two popular 3D frameworks using Multi-View Stereo (MVS) and Neural Implicit Surfaces (NIS) with a specific focus on the few-shot / sparse view setting. In this paper, we introduce a novel integration scheme that combines the multi-view stereo with neural si…

Cited by 10PDFScholar
2023

NeMF: Inverse Volume Rendering with Neural Microflake Field

ICCV 2023poster

Recovering the physical attributes of an object's appearance from its images captured under an unknown illumination is challenging yet essential for photo-realistic rendering.Recent approaches adopt the emerging implicit scene representations and have shown impressive results.However, they unanimous…

Cited by 26PDFcodeScholar
2023

REC-MV: REconstructing 3D Dynamic Cloth From Monocular Videos

CVPR 2023poster

Reconstructing dynamic 3D garment surfaces with open boundaries from monocular videos is an important problem as it provides a practical and low-cost solution for clothes digitization. Recent neural rendering methods achieve high-quality dynamic clothed human reconstruction results from monocular vi…

2023

Semantic Human Parsing via Scalable Semantic Transfer Over Multiple Label Domains

CVPR 2023poster

This paper presents Scalable Semantic Transfer (SST), a novel training paradigm, to explore how to leverage the mutual benefits of the data from different label domains (i.e. various levels of label granularity) to train a powerful human parsing network. In practice, two common application scenarios…

2022

Considering User Agreement in Learning to Predict the Aesthetic Quality

ICASSP 2022accepted

How to robustly rank the aesthetic quality of given images has been a long-standing ill-posed topic. Such challenge stems mainly from the diverse subjective opinions of different observers about the varied types of content. There is a growing interest in estimating the user agreement by considering…

Cited by 0SourceScholar
2022

High-Resolution Face Swapping via Latent Semantics Disentanglement

CVPR 2022poster

We present a novel high-resolution face swapping method using the inherent prior knowledge of a pre-trained GAN model. Although previous research can leverage generative priors to produce high-resolution results, their quality can suffer from the entangled semantics of the latent space. We explicitl…

Cited by 96PDFcodeScholar
2022

Subjective And Objective Quality Assessment Of Mobile Gaming Video

ICASSP 2022accepted

Nowadays, with the vigorous expansion and development of gaming video streaming techniques and services, the expectation of users, especially the mobile phone users, for higher quality of experience is also growing swiftly. As most of the existing research focuses on traditional video streaming, the…

Cited by 0SourceScholar
2018

Hybrid-MST: A Hybrid Active Sampling Strategy for Pairwise Preference Aggregation

NeurIPS 2018poster

In this paper we present a hybrid active sampling strategy for pairwise preference aggregation, which aims at recovering the underlying rating of the test candidates from sparse and noisy pairwise labeling. Our method employs Bayesian optimization framework and Bradley-Terry model to construct the u…