← Search

Hui Shen

16 accepted papers

2026

ATTS: Asynchronous Test-Time Scaling via Conformal Prediction

ICLR 2026poster

Large language models (LLMs) benefit from test-time scaling but are often hampered by high inference latency. Speculative decoding is a natural way to accelerate the scaling process; however, scaling along both the parallel and sequential dimensions poses significant challenges, including substantia…

Cited by 0SourcecodeScholar
2026

Autonomous Navigation in Unstructured Environments: A Probabilistic Approach for Generating Local Guidance With Limited Prior Information

RA-L 2026

Generating reliable local guidance information is crucial for autonomous navigation in unstructured environments with limited prior information. Conventional approaches often fuse navigation cues with different physical semantics into a single scalar objective through manually weighted cost terms, m

Cited by 0SourceScholar
2026

RealAppiance: Let High-fidelity Appliance Assets Controllable and Workable as Aligned Real Manauls

CVPR 2026

Existing appliance assets suffer from poor rendering, incomplete mechanisms, and misalignment with manuals, leading to simulation-reality gaps that hinder appliance manipulation development. In this work, we introduce the RealAppliance dataset, comprising 100 high-fidelity appliances with complete p

Cited by 0SourceScholar
2026

SWINGARENA: Adversarial Programming Arena for Long-context GitHub Issue Solving

ICLR 2026oral

We present \textsc{SwingArena}, a adversarial evaluation framework for Large Language Models (LLMs) that closely mirrors real-world software development workflows. Unlike traditional static benchmarks, \textsc{SwingArena} models the collaborative process of software iteration by pairing LLMs as \tex…

Cited by 0SourcecodeScholar
2026

TTAPFormer: Robust Arbitrary Point Tracking via Transient Asynchronous Fusion of Frames and Events

CVPR 2026

Tracking any point (TAP) is a fundamental yet challenging task in computer vision, requiring high precision and long-term motion reasoning. Recent attempts to combine RGB frames and event streams have shown promise, yet they typically rely on synchronous or non-adaptive fusion, leading to temporal m

Cited by 0SourcecodeScholar
2025

Argus: Benchmarking and Enhancing Vision-Language Models for 3D Radiology Report Generation

ACL 2025finding

Automatic radiology report generation holds significant potential to streamline the labor-intensive process of report writing by radiologists, particularly for 3D radiographs such as CT scans. While CT scans are critical for clinical diagnostics, they remain less explored compared to 2D radiographs.…

Cited by 0SourcePDFScholar
2025

Fully Spiking Neural Networks for Unified Frame-Event Object Tracking

NeurIPS 2025poster

The integration of image and event streams offers a promising approach for achieving robust visual object tracking in complex environments. However, current fusion methods achieve high performance at the cost of significant computational overhead and struggle to efficiently extract the sparse, async…

Cited by 0SourcecodeScholar
2025

MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference

NAACL 2025long

Long-context Multimodal Large Language Models (MLLMs) that incorporate long text-image and text-video modalities, demand substantial computational resources as their multimodal Key-Value (KV) cache grows with increasing input lengths, challenging memory and time efficiency. For multimodal scenarios,…

2025

MEIT: Multimodal Electrocardiogram Instruction Tuning on Large Language Models for Report Generation

ACL 2025finding

Electrocardiogram (ECG) is the primary non-invasive diagnostic tool for monitoring cardiac conditions and is crucial in assisting clinicians. Recent studies have concentrated on classifying cardiac conditions using ECG data but have overlooked ECG report generation, which is time-consuming and requi…

2025

SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning

NeurIPS 2025poster

Multimodal large language models (MLLMs) have shown promising capabilities in reasoning tasks, yet still struggle significantly with complex problems requiring explicit self-reflection and self-correction, especially compared to their unimodal text-based counterparts. Existing reflection methods are…

Cited by 0SourceScholar
2025

SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model Compression

NAACL 2025long

Despite significant advancements, the practical deployment of Large Language Models (LLMs) is often hampered by their immense sizes, highlighting the need for effective compression techniques. Singular Value Decomposition (SVD) emerges as a promising method for compressing LLMs. However, existing SV…

2025

Self-Supervised Traversability Learning With Online Prototype Adaptation for Off-Road Autonomous Driving

RA-L 2025

Achieving reliable and safe autonomous driving in off-road environments requires accurate and efficient terrain traversability analysis. However, this task faces several challenges, including the scarcity of large-scale datasets tailored for off-road scenarios, the high cost and potential errors of

Cited by 4SourceScholar
2025

Tracking Any Point with Frame-Event Fusion Network at High Frame Rate

IROS 2025

Tracking any point based on image frames is constrained by frame rates, leading to instability in high-speed scenarios and limited generalization in real-world applications. To overcome these limitations, we propose an image-event fusion point tracker, FE-TAP, which combines the contextual informati

Cited by 7SourceScholar
2022

CDX-NET: Cross-Domain Multi-Feature Fusion Modeling Via Deep Neural Networks for Multivariate Time Series Forecasting in AIOps

ICASSP 2022accepted

In the application of Artificial Intelligence for IT Operations (AIOps), monitoring data are usually modeled as MTS (Multivariate Time Series). The prediction of MTS has been widely studied and various models, including statistic algorithms and deep learning networks, have been proposed, which attem…

Cited by 0SourceScholar
2015

Low-latency list decoding of polar codes with double thresholding

ICASSP 2015accepted

For polar codes with short-to-medium code length, list successive cancellation decoding is used to achieve a good error-correcting performance. However, list pruning in the current list decoding is based on the sorting strategy and its timing complexity is high. This results in a long decoding laten…

Cited by 0SourceScholar