← Search

Kang Chen

26 accepted papers

2026

Adaptive Anisotropic Gaussian Splatting for Multi-contrast MRI Arbitrary-Scale Super-Resolution with Anatomy Guidance

CVPR 2026

Implicit neural representation (INR) based methods learn a continuous mapping from a low-resolution (LR) target magnetic resonance (MR) image and a high-resolution (HR) reference image to achieve arbitrary-scale super-resolution (SR). However, their inherent spectral bias favors learning low-frequen

Cited by 0SourcecodeScholar
2026

Do LLMs Signal When They’re Right? Evidence from Neuron Agreement

ICML 2026spotlight

Large language models (LLMs) commonly boost reasoning via sample-evaluate-ensemble decoders (e.g., majority voting), achieving label free gains without ground truth. However, prevailing strategies score candidates using only external outputs such as token probabilities, entropies, or self evaluation…

Cited by 0SourceScholar
2026

Error Amplification Limits ANN-to-SNN Conversion in Continuous Control

ICML 2026poster

Spiking Neural Networks (SNNs) can achieve competitive performance by converting already existing well-trained Artificial Neural Networks (ANNs), avoiding further costly training. This property is particularly attractive in Reinforcement Learning (RL), where training through environment interaction …

Cited by 0SourceScholar
2026

PHYS-DIFF: A PHYSICS-INSPIRED LATENT DIFFUSION MODEL FOR TROPICAL CYCLONE FORECASTING

ICASSP 2026poster

Tropical cyclone (TC) forecasting is critical for disaster warning and emergency response. Deep learning methods address computational challenges but often neglect physical relationships between TC attributes, resulting in predictions lacking physical consistency. To address this, we propose Phys-Di…

Cited by 0SourcePDFScholar
2026

RLux-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models

RSS 2026poster

Recent advances in vision-language-action (VLA) models have motivated the extension of their capabilities to embodied settings, where reinforcement learning (RL) offers a principled way to optimize task success through interaction. However, existing methods remain fragmented, lacking both a unified …

Cited by 0SourceScholar
2026

ROBUST ONLINE OVERDETERMINED INDEPENDENT VECTOR ANALYSIS BASED ON BILINEAR DECOMPOSITION

ICASSP 2026oral

Online blind source separation is essential for both speech communication and human-machine interaction. Among existing approaches, overdetermined independent vector analysis (OverIVA) delivers strong performance by exploiting the statistical independence of source signals and the orthogonality betw…

Cited by 0SourcePDFScholar
2025

A Multi-modal Hand Imitation Dataset for Dexterous Hand

IROS 2025

Multimodal data is indispensable for advancing imitation learning, particularly in the context of dexterous hands. However, existing datasets predominantly rely on single-modality inputs, such as RGB images, which inherently lack the capacity to capture the spatial and temporal dynamics essential fo

Cited by 0SourcecodeScholar
2025

Faster and Stronger: When ANN-SNN Conversion Meets Parallel Spiking Calculation

ICML 2025poster

Spiking Neural Network (SNN), as a brain-inspired and energy-efficient network, is currently facing the pivotal challenge of exploring a suitable and efficient learning framework. The predominant training methodologies, namely Spatial-Temporal Back-propagation (STBP) and ANN-SNN Conversion, are encu…

2025

Global Tropical Cyclone Intensity Forecasting with Multi-modal Multi-scale Causal Autoregressive Model

ICASSP 2025accepted

Accurate forecasting of tropical cyclone (TC) intensity is crucial for formulating disaster risk reduction strategies. Current methods predominantly rely on limited spatiotemporal information from ERA5 data and neglect the causal relationships between these physical variables, failing to fully captu…

Cited by 0SourceScholar
2025

MigGPT: Harnessing Large Language Models for Automated Migration of Out-of-Tree Linux Kernel Patches Across Versions

NeurIPS 2025spotlight

Out-of-tree kernel patches are essential for adapting the Linux kernel to new hardware or enabling specific functionalities. Maintaining and updating these patches across different kernel versions demands significant effort from experienced engineers. Large language models (LLMs) have shown remarkab…

Cited by 0SourceScholar
2025

On the Design of Low-Rank Differential Beamformers with Nonuniform Linear Microphone Arrays

ICASSP 2025accepted

Kronecker product beamforming is an effective technique for designing beamformers with nonuniform linear arrays (NULAs). However, current techniques are restricted to NULAs with specific configurations, where the steering vector of the array is represented as a Kronecker product of steering vectors…

Cited by 0SourceScholar
2025

Rethinking High-speed Image Reconstruction Framework with Spike Camera

AAAI 2025technical

Spike cameras, as innovative neuromorphic devices, generate continuous spike streams to capture high-speed scenes with lower bandwidth and higher dynamic range than traditional RGB cameras. However, reconstructing high-quality images from the spike input under low-light conditions remains challengin…

2025

SOTA: Spike-Navigated Optimal TrAnsport Saliency Region Detection in Composite-bias Videos

IJCAI 2025

Existing saliency detection methods struggle in real-world scenarios due to motion blur and occlusions. In contrast, spike cameras, with their high temporal resolution, significantly enhance visual saliency maps. However, the composite noise inherent to spike camera imaging introduces discontinuitie

2025

USP-Gaussian: Unifying Spike-based Image Reconstruction, Pose Correction and Gaussian Splatting

CVPR 2025highlight

Spike camera, as an innovative type of neuromorphic camera that captures scenes with 0-1 bit stream at 40 kHz, is increasingly being employed for the novel view synthesis task building on the techniques such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). Previous spike-based appr…

2025

VQLTI: Long-Term Tropical Cyclone Intensity Forecasting with Physical Constraints

AAAI 2025technical

Tropical cyclone (TC) intensity forecasting is crucial for early disaster warning and emergency decision-making. Numerous researchers have explored deep-learning methods to address computational and post-processing issues in operational forecasting. Regrettably, they exhibit subpar long-term forecas…

2024

Capturing Closely Interacted Two-Person Motions with Reaction Priors

CVPR 2024poster

In this paper we focus on capturing closely interacted two-person motions from monocular videos an important yet understudied topic. Unlike less-interacted motions closely interacted motions contain frequently occurring inter-human occlusions which pose significant challenges to existing capturing a…

Cited by 1SourcePDFScholar
2024

FNP: Fourier Neural Processes for Arbitrary-Resolution Data Assimilation

NeurIPS 2024poster

Data assimilation is a vital component in modern global medium-range weather forecasting systems to obtain the best estimation of the atmospheric state by combining the short-term forecast and observations. Recently, AI-based data assimilation approaches have attracted increasing attention for their…

2024

IPL: Leveraging Multimodal Large Language Models for Intelligent Product Listing

EMNLP 2024industry

Unlike professional Business-to-Consumer (B2C) e-commerce platforms (e.g., Amazon), Consumer-to-Consumer (C2C) platforms (e.g., Facebook marketplace) are mainly targeting individual sellers who usually lack sufficient experience in e-commerce. Individual sellers often struggle to compose proper desc…

Cited by 3SourcePDFScholar
2024

SpikeReveal: Unlocking Temporal Sequences from Real Blurry Inputs with Spike Streams

NeurIPS 2024spotlight

Reconstructing a sequence of sharp images from the blurry input is crucial for enhancing our insights into the captured scene and poses a significant challenge due to the limited temporal features embedded in the image. Spike cameras, sampling at rates up to 40,000 Hz, have proven effective in captu…

2024

Towards a Self-contained Data-driven Global Weather Forecasting Framework

ICML 2024poster

Data-driven weather forecasting models are advancing rapidly, yet they rely on initial states (i.e., analysis states) typically produced by traditional data assimilation algorithms. Four-dimensional variational assimilation (4DVar) is one of the most widely adopted data assimilation algorithms in nu…

Cited by 8SourcePDFScholar
2024

VTQA: Visual Text Question Answering via Entity Alignment and Cross-Media Reasoning

CVPR 2024poster

Achieving the optimal form of Visual Question Answering mandates a profound grasp of understanding grounding and reasoning within the intersecting domains of vision and language. Traditional VQA benchmarks have predominantly focused on simplistic tasks such as counting visual attributes and object d…

2023

CLIPVG: Text-Guided Image Manipulation Using Differentiable Vector Graphics

AAAI 2023technical

Considerable progress has recently been made in leveraging CLIP (Contrastive Language-Image Pre-Training) models for text-guided image manipulation. However, all existing works rely on additional generative models to ensure the quality of results, because CLIP alone cannot provide enough guidance in…

2023

Learning Analytical Posterior Probability for Human Mesh Recovery

CVPR 2023poster

Despite various probabilistic methods for modeling the uncertainty and ambiguity in human mesh recovery, their overall precision is limited because existing formulations for joint rotations are either not constrained to SO(3) or difficult to learn for neural networks. To address such an issue, we de…

2023

TODE-Trans: Transparent Object Depth Estimation with Transformer

ICRA 2023poster

Transparent objects are widely used in industrial automation and daily life. However, robust visual recognition and perception of transparent objects have always been a major challenge. Currently, most commercial-grade depth cameras are still not good at sensing the surfaces of transparent objects d…

Cited by 24SourcecodeScholar
2022

Face2Faceρ: Real-Time High-Resolution One-Shot Face Reenactment

ECCV 2022poster

"Existing one-shot face reenactment methods either present obvious artifacts in large pose transformations, or cannot well-preserve the identity information in the source images, or fail to meet the requirements of real-time applications due to the intensive amount of computation involved. In this p…

Cited by 37SourcePDFScholar