← Search

Yong Wang

63 accepted papers

2026

AIR: Post-training Data Selection for Reasoning via Attention Head Influence

ICML 2026poster

LLMs achieve remarkable multi-step reasoning capabilities, yet effectively transferring these skills via post-training distillation remains challenging. Existing data selection methods, ranging from manual curation to heuristics based on length, entropy, or overall loss, fail to capture the causal i…

Cited by 0SourceScholar
2026

AdaCuRL: Adaptive Curriculum Reinforcement Learning with Invalid Sample Mitigation and Historical Revisiting

AAAI 2026technical

Reinforcement learning (RL) has demonstrated considerable potential for enhancing reasoning in large language models (LLMs). However, existing methods suffer from Gradient Starvation and Policy Degradation when training directly on samples with mixed difficulty. To mitigate this, prior approaches l

Cited by 18SourcePDFScholar
2026

Anchored Policy Optimization: Mitigating Exploration Collapse via Support-Constrained Rectification

ICML 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) is increasingly viewed as a tree pruning mechanism. However, we identify a systemic pathology termed Recursive Space Contraction (RSC), an irreversible collapse driven by the combined dynamics of positive sharpening and negative squeezing, where …

Cited by 6SourceScholar
2026

BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation

ICLR 2026poster

LLM-as-a-Judge has been widely adopted across various research and practical applications, yet the robustness and reliability of its evaluation remain a critical issue. A core challenge it faces is bias, which has primarily been studied in terms of known biases and their impact on evaluation outcome…

Cited by 0SourcecodeScholar
2026

CaT-Diff: Cascaded Text-enhanced Diffusion Model for Time-Series Imputation

AAAI 2026technical

Most state-of-the-art time series imputation methods can leverage textual information to improve imputation quality, but they often struggle because they fail to effectively filter noisy information from large language model (LLM) derived textual information. Some existing solutions only filter over

Cited by 0SourcePDFScholar
2026

D²Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning

ICML 2026poster

Reinforcement learning (RL) has demonstrated potential for enhancing reasoning in large language models (LLMs). However, effective RL training, which requires medium-difficulty training samples, faces two fundamental challenges: Effective Data Scarcity and Dynamic Difficulty Shifts, where medium-dif…

Cited by 0SourceScholar
2026

Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models

ICLR 2026poster

Text-to-image (T2I) models have achieved remarkable success in generating high-fidelity images, but they often fail in handling complex spatial relationships, e.g., spatial perception, reasoning, or interaction. These critical aspects are largely overlooked by current benchmarks due to their short o…

Cited by 0SourcecodeScholar
2026

FASA: FREQUENCY-AWARE SPARSE ATTENTION

ICLR 2026poster

The deployment of Large Language Models (LLMs) faces a critical bottleneck when handling lengthy inputs: the prohibitive memory footprint of the Key Value (KV) cache. To address this bottleneck, the token pruning paradigm leverages attention sparsity to selectively retain a small, critical subset of…

Cited by 0SourceScholar
2026

GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning

ICLR 2026poster

Reinforcement Learning (RL) can directly enhance the reasoning capabilities of large language models without extensive reliance on Supervised Fine-Tuning (SFT). In this work, we revisit the traditional Policy Gradient (PG) mechanism and propose a minimalist RL approach termed Group Policy Gradient (…

Cited by 0SourcecodeScholar
2026

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation

ICML 2026poster

Unified Multimodal Models (UMMs) integrate both visual understanding and generation within a single framework. Their ultimate aspiration is to create a cycle where understanding and generation mutually reinforce each other. While recent post-training methods have successfully leveraged understanding…

Cited by 0SourceScholar
2026

Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation

ICLR 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) offers a robust mechanism for enhancing mathematical reasoning in large models. However, we identify a systematic lack of emphasis on more challenging questions in existing methods from both algorithmic and data perspectives, despite their import…

Cited by 0SourcecodeScholar
2026

Implicit Action Chunking for Smooth Continuous Control

ICML 2026poster

Reinforcement learning often produces high-frequency oscillatory control signals that undermine the safety and stability required for physical deployment. Explicit action chunking addresses this by predicting fixed-horizon trajectories but increases the policy output dimension to R^hd, leading to op…

Cited by 0SourceScholar
2026

Journey to the Centre of Cluster: Harnessing Interior Nodes for A/B Testing under Network Interference

ICLR 2026poster

A/B testing on platforms often faces challenges from network interference, where a unit's outcome depends not only on its own treatment but also on the treatments of its network neighbors. To address this, cluster-level randomization has become standard, enabling the use of network-aware estimators.…

Cited by 0SourcecodeScholar
2026

PFGNet: A Fully Convolutional Frequency-Guided Peripheral Gating Network for Efficient Spatiotemporal Predictive Learning

CVPR 2026

Spatiotemporal predictive learning (STPL) aims to forecast future frames from past observations and is essential across a wide range of applications. Compared with recurrent or hybrid architectures, pure convolutional models offer superior efficiency and full parallelism, yet their fixed receptive f

Cited by 0SourcecodeScholar
2026

Shedding the Facades, Connecting the Domains: Detecting Shifting Multimodal Hate Video with Test-Time Adaptation

AAAI 2026technical

Hate Video Detection (HVD) is crucial for online ecosystems. Existing methods assume identical distributions between training (source) and inference (target) data. However, hateful content often evolves into irregular and ambiguous forms to evade censorship, resulting in substantial semantic drift a

Cited by 0SourcePDFScholar
2026

Tree Search for LLM Agent Reinforcement Learning

ICLR 2026poster

Recent advances in reinforcement learning (RL) have significantly enhanced the agentic capabilities of large language models (LLMs). In long-term and multi-turn agent tasks, existing approaches driven solely by outcome rewards often suffer from the problem of sparse supervision. To address the chall…

Cited by 0SourcecodeScholar
2026

Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning

AAAI 2026technical

Large Multimodal Models (LMMs) have shown promising in-context learning (ICL) capabilities, but scaling to many-shot settings remains difficult due to limited context length and high inference cost. To address these challenges, task-vector-based methods have been explored by inserting compact repres

Cited by 0SourcePDFScholar
2026

mmWave-Diffusion: A Novel Framework for Respiration Sensing Using Observation-Anchored Conditional Diffusion Model

ICASSP 2026oral

Millimeter-wave (mmWave) radar enables contactless respiratory sensing,yet fine-grained monitoring is often degraded by nonstationary interference from body micromotions.To achieve micromotion interference removal,we propose mmWave-Diffusion,an observation-anchored conditional diffusion framework th…

Cited by 0SourcePDFScholar
2025

Cross-MoE: An Efficient Temporal Prediction Framework Integrating Textual Modality

EMNLP 2025

It has been demonstrated that incorporating external information as textual modality can effectively improve time series forecasting accuracy. However, current multi-modal models ignore the dynamic and different relations between time series patterns and textual features, which leads to poor perform

2025

Deep Support Vein Machine for Lung Parcellation

ICASSP 2025accepted

Pulmonary segments parcellation is essential to thoracoscopic segmentectomy. Surgeons manually outline pulmonary segments from preoperative images before surgery, which is a time-consuming, labor-intensive and mental-stress procedure. This work proposes a novel small learning model of deep support v…

Cited by 0SourceScholar
2025

E-Verify: A Paradigm Shift to Scalable Embedding-based Factuality Verification

EMNLP 2025

Large language models (LLMs) exhibit remarkable text-generation capabilities, yet struggle with factual consistency, motivating growing interest in factuality verification. Existing factuality verification methods typically follow a Decompose-Then-Verify paradigm, which improves granularity but suff

2025

GeMIMO: Searching the Cores of X-formers for Time Series Forecasting

ICASSP 2025accepted

In recent years, Transformer-based models have been widely used in time series forecasting tasks, demonstrating exceptional performance. However, these models lack interpretability, making it difficult to identify which components play a core role in predictions and which are redundant. To address t…

Cited by 0SourceScholar
2025

HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocation

EMNLP 2025

Self-taught reasoners (STaRs) enhance the mathematical reasoning abilities of large language models (LLMs) by leveraging self-generated responses for self-training. Recent studies have incorporated reward models to guide response selection or decoding, aiming to obtain higher-quality data. However,

Cited by 0SourcePDFScholar
2025

Keep Your Friends Close, and Your Enemies Farther: Distance-aware Voxel-wise Contrastive Learning for Semi-supervised Multi-organ Segmentation

ICCV 2025poster

Based on pseudo-labels, voxel-wise contrastive learning (VCL) is a prominent approach designed to learn effective feature representations for semi-supervised medical image segmentation. However, in multi-organ segmentation (MoS), the complex anatomical structures of certain organs often lead to many…

Cited by 0SourcePDFScholar
2025

Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition

CVPR 2025poster

Visual Emotion Recognition (VER) is a critical yet challenging task aimed at inferring emotional states of individuals based on visual cues. However, existing works focus on single domains, e.g., realistic images or stickers, limiting VER models' cross-domain generalizability. To fill this gap, we…

Cited by 0SourcePDFScholar
2025

METDrive: Multimodal End-to-End Autonomous Driving with Temporal Guidance

ICRA 2025

Multimodal end-to-end autonomous driving has shown promising advancements in recent work. By embedding more modalities into end-to-end networks, the system's understanding of both static and dynamic aspects of the driving environment is enhanced, thereby improving the safety of autonomous driving. I

Cited by 4SourceScholar
2025

POSITION BIAS MITIGATES POSITION BIAS: Mitigate Position Bias Through Inter-Position Knowledge Distillation

EMNLP 2025

Positional bias (PB), manifesting as non-uniform sensitivity across different contextual locations, significantly impairs long-context comprehension and processing capabilities. Previous studies have addressed PB either by modifying the underlying architectures or by employing extensive contextual a

2025

SelaFD: Seamless Adaptation of Vision Transformer Fine-tuning for Radar-based Human Activity Recognition

ICASSP 2025accepted

Human Activity Recognition (HAR) such as fall detection has become increasingly critical due to the aging population, necessitating effective monitoring systems to prevent serious injuries and fatalities associated with falls. This study focuses on fine-tuning the Vision Transformer (ViT) model spec…

Cited by 0SourceScholar
2025

UML: A Unified Multimodal Learning Framework for Cataract Postoperative Visual Acuity Prediction with Uncertain Missing Modalities

ICASSP 2025accepted

Cataracts are the leading cause of blindness worldwide, with surgery as the only effective treatment. Accurate prediction of Best Corrected Visual Acuity (BCVA) is crucial for surgical planning. In this paper, we propose a novel Unified Multimodal Learning (UML) framework for BCVA prediction with un…

Cited by 0SourceScholar
2025

USP: Unified Self-Supervised Pretraining for Image Generation and Understanding

ICCV 2025poster

Recent studies have highlighted the interplay between diffusion models and representation learning. Intermediate representations from diffusion models can be leveraged for downstream visual tasks, while self-supervised vision models can enhance the convergence and generation quality of diffusion mod…

2024

Classifier Clustering and Feature Alignment for Federated Learning under Distributed Concept Drift

NeurIPS 2024poster

Data heterogeneity is one of the key challenges in federated learning, and many efforts have been devoted to tackling this problem. However, distributed concept drift with data heterogeneity, where clients may additionally experience different concept drifts, is a largely unexplored area. In this wo…

2024

Diffusion-Based Pose Refinement and Multi-Hypothesis Generation for 3D Human Pose Estimation

ICASSP 2024accepted

Previous probabilistic models for 3D Human Pose Estimation (3DHPE) aimed to enhance pose accuracy by generating multiple hypotheses. However, most of the hypotheses generated deviate substantially from the true pose. Compared to deterministic models, the excessive uncertainty in probabilistic models…

Cited by 0SourceScholar
2024

E-GNN: An Enhanced Method for Multi-Object Tracking with Collective Motion Patterns

RA-L 2024

The long-term consistent visual tracking of large-scale moving swarms of animals or autonomous moving robots (AMR) is extremely challenging when the three factors are involved: 1) similar appearance of animals or AMR, 2) frequent and unpredictable occlusions, and 3) non-linear maneuvers. When facing

Cited by 3SourceScholar
2024

Exploring Self-Explainable Street-Level IP Geolocation with Graph Information Bottleneck

ICASSP 2024accepted

Accurate IP geolocation is crucial for location-aware applications. While recent advances in router-centric IP graph methods have garnered attention, they face two persistent challenges: (1) the sparsity problem of IP graphs in rural areas and (2) the limited explainability of current IP geolocation…

Cited by 0SourceScholar
2024

FaceCom: Towards High-fidelity 3D Facial Shape Completion via Optimization and Inpainting Guidance

CVPR 2024poster

We propose FaceCom a method for 3D facial shape completion which delivers high-fidelity results for incomplete facial inputs of arbitrary forms. Unlike end-to-end shape completion methods based on point clouds or voxels our approach relies on a mesh-based generative network that is easy to optimize…

2024

GTPT: Group-based Token Pruning Transformer for Efficient Human Pose Estimation

ECCV 2024poster

"In recent years, 2D human pose estimation has made significant progress on public benchmarks. However, many of these approaches face challenges of less applicability in the industrial community due to the large number of parametric quantities and computational overhead. Efficient human pose estimat…

2024

Improving IP Geolocation With Target-Centric IP Graph (Student Abstract)

AAAI 2024technical

Accurate IP geolocation is indispensable for location-aware applications. While recent advances based on router-centric IP graphs are considered cutting-edge, one challenge remain: the prevalence of sparse IP graphs (14.24% with fewer than 10 nodes, 9.73% isolated) limits graph learning. To mitigate…

Cited by 0SourcePDFScholar
2024

Loop Structure-Aware Learning for Fully Automated Pulmonary Fissure Completeness Assessment

ICASSP 2024accepted

Pulmonary fissures are anatomical biomarkers used to evaluate the severity of chronic obstructive pulmonary disease. The completeness of the fissures is significantly associated with this disease. This work proposes a new fully automated fissure completeness assessment framework on the basis of deep…

Cited by 0SourceScholar
2024

On the Concept Trustworthiness in Concept Bottleneck Models

AAAI 2024technical

Concept Bottleneck Models (CBMs), which break down the reasoning process into the input-to-concept mapping and the concept-to-label prediction, have garnered significant attention due to their remarkable interpretability achieved by the interpretable concept bottleneck. However, despite the transpar…

2024

Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion Recognition

ICASSP 2024accepted

Swin-Transformer has demonstrated remarkable success in computer vision by leveraging its hierarchical feature representation based on Transformer. In speech signals, emotional information is distributed across different scales of speech features, e. g., word, phrase, and utterance. Drawing above in…

Cited by 0SourceScholar
2024

Training-Free Pretrained Model Merging

CVPR 2024poster

Recently model merging techniques have surfaced as a solution to combine multiple single-talent models into a single multi-talent model. However previous endeavors in this field have either necessitated additional training or fine-tuning processes or require that the models possess the same pre-trai…

2023

HDNet: Hierarchical Dynamic Network for Gait Recognition using Millimeter-wave radar

ICASSP 2023accepted

Gait recognition is widely used in diversified practical applications. Currently, the most prevalent approach is to recognize human gait from RGB images, owing to the progress of computer vision technologies. Nevertheless, the perception capability of RGB cameras deteriorates in rough circumstances,…

Cited by 0SourceScholar
2023

MetaPortrait: Identity-Preserving Talking Head Generation With Fast Personalized Adaptation

CVPR 2023poster

In this work, we propose an ID-preserving talking head generation framework, which advances previous methods in two aspects. First, as opposed to interpolating from sparse flow, we claim that dense landmarks are crucial to achieving accurate geometry-aware flow fields. Second, inspired by face-swapp…

2023

Optimized Covariance Design for AB Test on Social Network under Interference

NeurIPS 2023poster

Online A/B tests have become increasingly popular and important for social platforms. However, accurately estimating the global average treatment effect (GATE) has proven to be challenging due to network interference, which violates the Stable Unit Treatment Value Assumption (SUTVA) and poses great…

2023

TR-Rules: Rule-based Model for Link Forecasting on Temporal Knowledge Graph Considering Temporal Redundancy

EMNLP 2023long findings

Temporal knowledge graph (TKG) has been proved to be an effective way for modeling dynamic facts in real world. Many efforts have been devoted into predicting future events i.e. extrapolation, on TKGs. Recently, rule-based knowledge graph completion methods which are considered to be more interpreta…

Cited by 0SourceScholar
2022

A Transformer-based Threshold-Free Framework for Multi-Intent NLU

COLING 2022main

Multi-intent natural language understanding (NLU) has recently gained attention. It detects multiple intents in an utterance, which is better suited to real-world scenarios. However, the state-of-the-art joint NLU models mainly detect multiple intents on threshold-based strategy, resulting in one ma…

2022

StyleSwin: Transformer-Based GAN for High-Resolution Image Generation

CVPR 2022poster

Despite the tantalizing success in a broad of vision tasks, transformers have not yet demonstrated on-par ability as ConvNets in high-resolution image generative modeling. In this paper, we seek to explore using pure transformers to build a generative adversarial network for high-resolution image sy…

Cited by 318PDFcodeScholar
2022

XLM-D: Decorate Cross-lingual Pre-training Model as Non-Autoregressive Neural Machine Translation

EMNLP 2022main

Pre-training language models have achieved thriving success in numerous natural language understanding and autoregressive generation tasks, but non-autoregressive generation in applications such as machine translation has not sufficiently benefited from the pre-training paradigm. In this work, we es…

2021

Evolving Quantized Neural Networks for Image Classification Using A Multi-Objective Genetic Algorithm

ICASSP 2021accepted

Recently, many model quantization approaches have been investigated to reduce the model size and improve the inference speed of convolutional neural networks (CNNs). However, these approaches usually inevitably lead to a decrease in classification accuracy. To address this problem, this paper propos…

Cited by 0SourceScholar
2021

Prototypical Pseudo Label Denoising and Target Structure Learning for Domain Adaptive Semantic Segmentation

CVPR 2021poster

Self-training is a competitive approach in domain adaptive segmentation, which trains the network with the pseudo labels on the target domain. However inevitably, the pseudo labels are noisy and the target features are dispersed due to the discrepancy between source and target domains. In this paper…

Cited by 631PDFcodeScholar
2020

Lexical-Constraint-Aware Neural Machine Translation via Data Augmentation

IJCAI 2020poster

Leveraging lexical constraint is extremely significant in domain-specific machine translation and interactive machine translation. Previous studies mainly focus on extending beam search algorithm or augmenting the training corpus by replacing source phrases with the corresponding target translation.…

2019

Local Pose optimization with an Attention-based Neural Network

IROS 2019poster

In this paper, we propose a novel pose optimizer which can be inserted into either supervised or unsupervised end-to-end visual odometry for the purpose of local pose optimization. The pose optimizer is an analogue of the pose graph optimization used in traditional VSLAM algorithms. Local pose optim…

Cited by 2SourceScholar
2017

Design of an automated controller with collision-avoidance capability for in-vivo transportation of biological cells

IROS 2017poster

As the rapid development of precision medicine, in-vivo manipulation of micro/nano-scaled particles has attracted increasing attention in recent years. The collision is one of the main reasons that falls the in-vivo particle transportation fail. In this paper, we develop an in-vivo cell transportati…

Cited by 1SourceScholar
2016

Automated in-vivo transportation of biological cells with a disturbance compensation controller

IROS 2016poster

As rapid development of precision medicine, in vivo manipulation of micro/nano-scaled particles have attracted increasing attention in recent years. To accommodate complex in-vivo environment, robot-aided automated manipulation technology is highly demanded in trapping and controlling micro/nano-par…

Cited by 3SourceScholar
2015

Modelling and control of optical manipulation for cell rotation

ICRA 2015poster

Optical tweezers has become a powerful tool in automated cell transportation control and has been used in a variety of biological applications. The use of optical tweezers for cell surgery has great potential for various biomedical applications such as microinjection, organelle extraction and modifi…

Cited by 15SourceScholar