← Search

Kun Hu

32 accepted papers

2026

BoostSLT: Boosting Sign Language Translation via a Plug-and-Play Diffusion-Based Semantic Enhancer

CVPR 2026

Sign Language Translation (SLT) converts continuous sign videos into spoken language text, yet current models, whether gloss-based or gloss-free, struggle with long or discourse-level inputs. Recent architectures such as TwoStreamNetwork and CV-SLT have nearly saturated short-sentence accuracy, but

Cited by 0SourcecodeScholar
2026

CABTO: Context-Aware Behavior Tree Grounding for Robot Manipulation

AAAI 2026technical

Behavior Trees (BTs) offer a powerful paradigm for designing modular and reactive robot controllers. BT planning, an emerging field, provides theoretical guarantees for the automated generation of reliable BTs. However, BT planning typically assumes that a well-designed BT system is already grounded

Cited by 0SourcePDFScholar
2026

CoPE: A Framework for Optimizing Coordination between Planning and Execution in LLM-based Agents

ICML 2026poster

Fine-tuning Large Language Models (LLMs) as autonomous agents on domain-specific data has emerged as a promising paradigm for tackling interactive, real-world tasks. However, existing studies have overlooked the critical coordination between long-term planning and multi-step execution in optimizing …

Cited by 0SourceScholar
2026

DuoCast: Duo-Probabilistic Diffusion for Precipitation Nowcasting

AAAI 2026technical

Accurate short-term precipitation forecasting is critical for weather-sensitive decision-making in agriculture, transportation, and disaster response. Existing deep learning approaches often struggle to balance global structural consistency with local detail preservation, especially under complex me

Cited by 0SourcePDFScholar
2026

ECD: Evidence-guided Contrastive Decoding in Retrieval-Augmented Generation with Accurate Knowledge Reference Adjustment

AAAI 2026technical

Retrieval-Augmented Generation (RAG) enhances the quality of question answering by integrating external knowledge with internal knowledge. A robust RAG system needs to precisely regulate the dependence of the response on the two types of knowledge. The recently proposed context-aware contrastive dec

Cited by 0SourcePDFScholar
2026

F2Net: A Frequency-Fused Network for Ultra-High Resolution Remote Sensing Segmentation

CVPR 2026

Semantic segmentation of ultra-high-resolution (UHR) remote sensing imagery is critical for applications like environmental monitoring and urban planning but faces com- putational and optimization challenges. Conventional methods either lose fine details through downsampling or fragment global conte

Cited by 0SourcecodeScholar
2026

Pb4U-GNet: Resolution-Adaptive Garment Simulation via Propagation-before-Update Graph Network

AAAI 2026technical

Garment simulation is fundamental to various applications in computer vision and graphics, from virtual try-on to digital human modelling. However, conventional physics-based methods remain computationally expensive, hindering their application in time-sensitive scenarios. While graph neural network

Cited by 0SourcePDFScholar
2026

PhenoYieldNet: Learning Crop-Aware Phenological Responses for Multi-Crop Yield Prediction

CVPR 2026

Accurate crop yield prediction is crucial for sustainable agriculture and global food security. While existing methods are predominantly developed for single-crop prediction, they often struggle to generalize across diverse crop types, without addressing the unique crop phenological responses that a

Cited by 0SourcecodeScholar
2026

RebRL: Reinforcing Discrete Visual Diffusion Models with Rebalanced Timestep Credits

CVPR 2026

Discrete Diffusion Models (DDMs) have shown great potential in image generation, especially when equipped with reinforcement learning (RL) techniques.However, a fundamental yet overlooked limitation is revealed in our experiments: severe imbalance of credit assignment across timesteps during trainin

Cited by 0SourceScholar
2026

TWINFUZZ: Dual-Model Fuzzing for Robustness Generalization in Deep Learning

AAAI 2026technical

Deep learning (DL) models are increasingly deployed in safety-critical applications such as face recognition, autonomous driving, and medical diagnosis. Despite their impressive accuracy, they remain vulnerable to adversarial examples - subtle perturbations that can cause incorrect predictions, i.e.

Cited by 0SourcePDFScholar
2026

Unstitching the Chimera: Frame-Level Risk and Train-Free Mitigation for Video Hallucination

CVPR 2026

Hallucination limits the reliability of multimodal large language models (MLLMs), and it is particularly damaging in video where errors manifest as distorted narratives rather than single-frame mistakes. We introduce a frame-first study of **Chimera Hallucination**: model stitches visual segments th

Cited by 0SourceScholar
2025

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens

ICCV 2025poster

Recently, Vision Large Language Models (VLLMs) with integrated vision encoders have shown promising performance in vision understanding. They encode visual content into sequences of visual tokens, enabling joint processing of visual and textual data. However, understanding videos, especially long vi…

2025

CM-LIUW-Odometry: Robust and High-Precision LiDAR-Inertial-UWB-Wheel Odometry for Extreme Degradation Coal Mine Tunnels

IROS 2025

Simultaneous Localization and Mapping (SLAM) in large-scale, complex, and GPS-denied underground coal mine environments presents significant challenges. Sensors must contend with abnormal operating conditions: GPS unavailability impedes scene reconstruction and absolute geographic referencing, uneve

Cited by 1SourceScholar
2025

DC-PCN: Point Cloud Completion Network with Dual-Codebook Guided Quantization

AAAI 2025technical

Point cloud completion aims to reconstruct complete 3D shapes from partial 3D point clouds. With advancements in deep learning techniques, various methods for point cloud completion have been developed. Despite achieving encouraging results, a significant issue remains: these methods often overlook…

2025

LiDAR-IMU Fusion System with Adaptive Scanning for High-Resolution Deformation Monitoring of Underground Infrastructures

IROS 2025

A LiDAR-IMU fusion system utilizing adaptive scanning is developed for high-resolution deformation monitoring of underground coal mine infrastructure, such as sealed walls. The system integrates data from a LiDAR scanner and an IMU, employing a penalty function-based scanning strategy to optimize po

Cited by 0SourceScholar
2025

PUMPS: Skeleton-Agnostic Point-based Universal Motion Pre-Training for Synthesis in Human Motion Tasks

ICCV 2025poster

Motion skeletons drive 3D character animation by transforming bone hierarchies, but differences in proportions or structure make motion data hard to transfer across skeletons, posing challenges for data-driven motion synthesis. Temporal Point Clouds (TPCs) offer an unstructured, cross-compatible mot…

2025

RI-MAE: Rotation-Invariant Masked AutoEncoders for Self-Supervised Point Cloud Representation Learning

AAAI 2025technical

Masked point modeling methods have recently achieved great success in self-supervised learning for point cloud data. However, these methods are sensitive to rotations and often exhibit sharp performance drops when encountering rotational variations. In this paper, we propose a novel Rotation-Invaria…

2025

Reidentify: Context-Aware Identity Generation for Contextual Multi-Agent Reinforcement Learning

ICML 2025poster

Generalizing multi-agent reinforcement learning (MARL) to accommodate variations in problem configurations remains a critical challenge in real-world applications, where even subtle differences in task setups can cause pre-trained policies to fail. To address this, we propose Context-Aware Identity…

Cited by 0SourcePDFScholar
2025

Wave-wise Discriminative Tracking by Phase-Amplitude Separation, Augmentation and Mixture

IJCAI 2025

Distinguishing key features in complex visual tasks is challenging. A novel approach treats image patches (tokens) as waves. By using both phase and amplitude, it captures richer semantics and specific invariances compared to pixel-based methods, and allows for feature fusion across regions for a ho

Cited by 0SourcePDFScholar
2024

Autoregressive Omni-Aware Outpainting for Open-Vocabulary 360-Degree Image Generation

AAAI 2024technical

A 360-degree (omni-directional) image provides an all-encompassing spherical view of a scene. Recently, there has been an increasing interest in synthesising 360-degree images from conventional narrow field of view (NFoV) images captured by digital cameras and smartphones, for providing immersive ex…

2024

Sequential Fusion Based Multi-Granularity Consistency for Space-Time Transformer Tracking

AAAI 2024technical

Regarded as a template-matching task for a long time, visual object tracking has witnessed significant progress in space-wise exploration. However, since tracking is performed on videos with substantial time-wise information, it is important to simultaneously mine the temporal contexts which have no…

Cited by 7SourcePDFScholar
2024

SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation

AAAI 2024technical

The Segment Anything Model (SAM) is a powerful foundation model that has revolutionised image segmentation. To apply SAM to surgical instrument segmentation, a common approach is to locate precise points or boxes of instruments and then use them as prompts for SAM in a zero-shot manner. However, we…

2024

Terrain Diffusion Network: Climatic-Aware Terrain Generation with Geological Sketch Guidance

AAAI 2024technical

Sketch-based terrain generation seeks to create realistic landscapes for virtual environments in various applications such as computer games, animation and virtual reality. Recently, deep learning based terrain generation has emerged, notably the ones based on generative adversarial networks (GAN).…

2023

Continuous Intermediate Token Learning With Implicit Motion Manifold for Keyframe Based Motion Interpolation

CVPR 2023poster

Deriving sophisticated 3D motions from sparse keyframes is a particularly challenging problem, due to continuity and exceptionally skeletal precision. The action features are often derivable accurately from the full series of keyframes, and thus, leveraging the global context with transformers has b…

2023

Decomposition, Interaction, Reconstruction Meets Global Context Learning In Visual Tracking

ICASSP 2023accepted

Tensor decomposition and reconstruction attention is a promising global context learning approach because it can remain efficient while avoiding feature compression. To exploit its potential even further in visual tracking, we redesign a 3D tensor modeling paradigm, namely tensor Decomposition, Inte…

Cited by 0SourceScholar
2023

Enhanced Dcf Tracker Regularized by Reliable Sample Construction

ICASSP 2023accepted

Discriminative correlation filter (DCF) is a highly efficient tracking technique using the circulant shifted samples of search images to update the template, so the reliability of input samples determines template quality. In this paper, we rethink the reliability problem of input samples in advance…

Cited by 0SourceScholar
2023

ICD-Face: Intra-class Compactness Distillation for Face Recognition

ICCV 2023poster

Knowledge distillation is an effective model compression method to improve the performance of a lightweight student model by transferring the knowledge of a well-performed teacher model, which has been widely adopted in many computer vision tasks, including face recognition (FR). The current FR dist…

Cited by 6PDFScholar
2023

Multi-Scale Control Signal-Aware Transformer for Motion Synthesis without Phase

AAAI 2023technical

Synthesizing controllable motion for a character using deep learning has been a promising approach due to its potential to learn a compact model without laborious feature engineering. To produce dynamic motion from weak control signals such as desired paths, existing methods often require auxiliary…

Cited by 10SourcePDFScholar
2023

Progressive Perception Learning for Distribution Modulation in Siamese Tracking

ICASSP 2023accepted

We explore an innovative view on distribution modulation to boost Siamese trackers. Specially, we observed two cases of possible distribution inconsistency in Siamese tracking: 1) Two branches with different sizes may be in different distribution ranges after a shared backbone (including BN layers).…

Cited by 0SourceScholar
2023

Stochastic Feature Averaging for Learning with Long-Tailed Noisy Labels

IJCAI 2023poster

Deep neural networks have shown promising results on a wide variety of tasks using large-scale and well-annotated training datasets. However, data collected from real-world applications can suffer from two prevalent biases, i.e., long-tailed class distribution and label noise. Previous efforts on lo…

2022

OTExtSum: Extractive Text Summarisation with Optimal Transport

NAACL 2022findings

Extractive text summarisation aims to select salient sentences from a document to form a short yet informative summary. While learning-based methods have achieved promising results, they have several limitations, such as dependence on expensive training and lack of interpretability. Therefore, in th…