← Search

Wenxin Huang

14 accepted papers

2026

Beyond the Horizon: Decoupling Multi-View UAV Action Recognition via Partial Order Transfer

AAAI 2026technical

Action recognition using uncrewed aerial vehicles (UAVs) faces unique challenges due to substantial view variations along the vertical spatial axis. Unlike ground-based scenarios, UAVs capture actions from diverse altitudes, resulting in pronounced appearance discrepancies and reduced recognition ro

Cited by 0SourcePDFScholar
2026

Evolving Quantitative Reasoning through Self-Play in Digital Twin Markets

ICML 2026poster

Large Language Models (LLMs) exhibit strong capabilities in high-level semantic understanding and strategic planning, yet they suffer from persistent quantitative failure modes, such as imprecise computation and the illusion of quantitative coherence, which limit their reliability in high-stakes dec…

Cited by 0SourceScholar
2026

KeenKT: Knowledge Mastery-State Disambiguation for Knowledge Tracing

AAAI 2026technical

Knowledge Tracing (KT) aims to dynamically model a student’s mastery of knowledge concepts based on their historical learning interactions. Most current methods rely on single-point estimates, which cannot distinguish true ability from outburst or carelessness, creating ambiguity in judging mastery.

Cited by 0SourcePDFScholar
2025

Agent Trading Arena: A Study on Numerical Understanding in LLM-Based Agents

EMNLP 2025

Large language models (LLMs) have demonstrated remarkable capabilities in natural language tasks, yet their performance in dynamic, real-world financial environments remains underexplored. Existing approaches are confined to historical backtesting, where trading actions cannot influence market price

2025

StoryLLaVA: Enhancing Visual Storytelling with Multi-Modal Large Language Models

COLING 2025main

The rapid development of multimodal large language models (MLLMs) has positioned visual storytelling as a crucial area in content creation. However, existing models often struggle to maintain temporal, spatial, and narrative coherence across image sequences, and they frequently lack the depth and en…

Cited by 3SourcePDFScholar
2024

Diversity-Driven Synthesis: Enhancing Dataset Distillation through Directed Weight Adjustment

NeurIPS 2024spotlight

The sharp increase in data-related expenses has motivated research into condensing datasets while retaining the most informative features. Dataset distillation has thus recently come to the fore. This paradigm generates synthetic datasets that are representative enough to replace the original datase…

2023

Background Disturbance Mitigation for Video Captioning Via Entity-Action Relocation

ICASSP 2023accepted

Video captioning aims to generate sentences to accurately describe the video content, in which video background plays the role of prompts. State-of-the-art methods tend to explore richer video representations adequately, fusing with language to improve caption quality, which has shown great success.…

Cited by 0SourceScholar
2023

Background-Weakening Consistency Regularization for Semi-Supervised Video Action Detection

ICASSP 2023accepted

Consistency-based techniques have produced state-of-the-art results in semi-supervised action detection. When the model false detects the dynamic information in the background as an action, spatio-temporal consistency calculations can hardly reflect this false detection result. We consider weakening…

Cited by 0SourceScholar
2023

Bat: Bi-Alignment Based On Transformation in Multi-Target Domain Adaptation for Semantic Segmentation

ICASSP 2023accepted

While enlightening progress has been made recently in single-target domain adaptive semantic segmentation (ST-DASS), the multi-peak distributed multi-target domain cannot be directly aligned well with the single-peak distributed source domain. As a result, it is impossible for existing methods to ha…

Cited by 0SourceScholar
2023

Neighborhood Information-Based Label Refinement for Person Re-Identification with Label Noise

ICASSP 2023accepted

The existing excellent person re-identification (Re-ID) model is still affected by the samples with the incorrect labels. It is difficult to accurately annotate person images in the real scene, resulting in label noise. To avoid fitting to the noisy labels, a common solution in Re-ID is to replace t…

Cited by 0SourceScholar
2022

Rainy WCity: A Real Rainfall Dataset with Diverse Conditions for Semantic Driving Scene Understanding

IJCAI 2022poster

Scene understanding in adverse weather conditions (e.g. rainy and foggy days) has drawn increasing attention, arising some specific benchmarks and algorithms. However, scene segmentation under rainy weather is still challenging and under-explored due to the following limitations on the datasets and…

Cited by 32SourcePDFScholar
2022

VCD: View-Constraint Disentanglement for Action Recognition

ICASSP 2022accepted

Action recognition is a hot topic in computer vision due to its wide range of applications in urban surveillance. Although some methods are more advanced from an invariant view perspective, those approaches do not perform well for the viewpoint change. To address this issue, one possible solution is…

Cited by 0SourceScholar
2021

Part-Aligned Network with Background for Misaligned Person Search

ICASSP 2021accepted

Person search is a significant computer vision task that requires addressing person detection and re-identification simultaneously. Body parts are frequently misaligned due to variation poses, occlusions, and partial missing, leading to the unsatisfied results of person search. Existing methods usua…

Cited by 0SourceScholar
2020

Multi-Scale Residual Network for Image Classification

ICASSP 2020accepted

Multi-scale approach representing image objects at various levels-of-details has been applied to various computer vision tasks. Existing image classification approaches place more emphasis on multi-scale convolution kernels, and overlook multi-scale feature maps. As such, some shallower information…

Cited by 0SourceScholar