← Search

Sheng Liu

45 accepted papers

2026

A Tri-Axial FBG-Based Force Sensor at the Tool Tip of a Continuum Manipulator for Single-Port Access Surgery

ICRA 2026poster

Abstract— The absence of force feedback remains a major bottleneck in the development of robotic laparoendoscopic single-site (R-LESS) surgery, reducing the control precision of surgical instruments and increasing the risk of tissue damage. To address this challenge, we propose a miniature triaxial …

Cited by 0Scholar
2026

Focal-General Diffusion Model with Semantic Consistent Guidance for Sign Language Production

CVPR 2026

Sign Language Production (SLP) aims to translate spoken language into sign sequences, where the main challenge lies in generating coherent and natural poses from discrete glosses (G2P). Existing G2P methods typically treat each pose as an indivisible unit, limiting their ability to capture fine-grai

Cited by 0SourcecodeScholar
2026

In-The-Flow Agentic System Optimization for Effective Planning and Tool Use

ICLR 2026oral

Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full context; this scales poorly with long horizons and diverse tools and generalize…

Cited by 0SourcecodeScholar
2026

InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression

ICLR 2026oral

Accurate and efficient discrete video tokenization is essential for long video sequences processing. Yet, the inherent complexity and variable information density of videos present a significant bottleneck for current tokenizers, which rigidly compress all content at a fixed rate, leading to redunda…

Cited by 0SourcecodeScholar
2026

SRAM: Shape-Realism Alignment Metric for No Reference 3D Shape Evaluation

AAAI 2026technical

3D generation and reconstruction techniques have been widely used in computer games, film, and other content creation areas. As the application grows, there is a growing demand for 3D shapes that look truly realistic. Traditional evaluation methods rely on a ground truth to measure mesh fidelity. Ho

Cited by 0SourcePDFScholar
2026

SpatialLogic-Bench: A Diagnostic Benchmark for Task-Oriented Spatiotemporal Reasoning

AAAI 2026technical

Vision-Language Models (VLMs) have made significant progress in static perception, but their ability to understand dynamic task-oriented reasoning remains unclear. Existing benchmarks mainly focus on static spatial relationships and lack systematic assessment of dynamic reasoning capabilities. To th

Cited by 0SourcePDFScholar
2026

Textured Geometry Evaluation: Perceptual 3D Textured Shape Metric via 3D Latent-Geometry Network

AAAI 2026technical

Textured high-fidelity 3D models are crucial for games, AR/VR, and film, but human-aligned evaluation methods still fall behind despite recent advances in 3D reconstruction and generation. Existing metrics, such as Chamfer Distance, often fail to align with how humans evaluate the fidelity of 3D sha

Cited by 0SourcePDFScholar
2025

Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks

EMNLP 2025

Despite advances in improving large language model (LLM) to refuse to answer malicious instructions, widely used LLMs remain vulnerable to jailbreak attacks where attackers generate instructions with distributions differing from safety alignment corpora. New attacks expose LLMs’ inability to recogni

2025

Hierarchical Spatial-Temporal Enhancement Network For Continuous Sign Language Recognition

ICASSP 2025accepted

In continuous sign language recognition (CSLR), 2D-CNN-based extractors are often insufficiently trained for spatial capture and struggle with temporal modeling. This leads to incomplete spatial discrimination, hindering the understanding actions across frames. To address these limitations, we propo…

Cited by 0SourceScholar
2025

Improving Continuous Sign Language Recognition via Cross-Frame Interactions in Expanded Contextual Spaces

ICASSP 2025accepted

Current continuous sign language recognition (CSLR) methods typically rely on single or adjacent frames for calculations, which can overlook broader contextual information and result in lower accuracy. To address this issue, we introduce CVSign, which constructs an extended contextual space frame by…

Cited by 0SourceScholar
2025

MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine

ICLR 2025poster

This paper introduces MedTrinity-25M, a comprehensive, large-scale multimodal dataset for medicine, covering over 25 million images across 10 modalities with multigranular annotations for more than 65 diseases. These multigranular annotations encompass both global information, such as modality and o…

2025

More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models

NeurIPS 2025poster

Test-time compute has empowered multimodal large language models to generate extended reasoning chains, yielding strong performance on tasks such as multimodal math reasoning. However, we observe that this improved reasoning ability often comes with increased hallucination: as generations become lon…

Cited by 0SourceScholar
2025

OLMD: Orientation-aware Long-term Motion Decoupling for Continuous Sign Language Recognition

AAAI 2025technical

The primary challenge in continuous sign language recognition (CSLR) mainly stems from the presence of multi-orientational and long-term motions. However, current research overlooks these crucial aspects, significantly impacting accuracy. To tackle these issues, we propose a novel CSLR framework: Or…

Cited by 0SourcePDFScholar
2025

OTLRM: Orthogonal Learning-based Low-Rank Metric for Multi-Dimensional Inverse Problems

AAAI 2025technical

In real-world scenarios, complex data such as multispectral images and multi-frame videos inherently exhibit robust low-rank property. This property is vital for multi-dimensional inverse problems, such as tensor completion, spectral imaging reconstruction, and multispectral image denoising. Exist…

2025

Reducing Hallucinations in Large Vision-Language Models via Latent Space Steering

ICLR 2025spotlight

Hallucination poses a challenge to the deployment of large vision-language models (LVLMs) in applications. Unlike in large language models (LLMs), hallucination in LVLMs often arises from misalignments between visual inputs and textual outputs. This paper investigates the underlying mechanisms of ha…

Cited by 67SourcePDFScholar
2025

Single-View Reconstruction via Decoupled 3D Gaussian Splatting

ICASSP 2025accepted

Creating high-quality 3D object representations from a single-view image is challenging. Existing methods tend to infer the geometry and texture information simultaneously within a shared network. However, decoding geometry and texture from a unified network often leads to their entanglement, causin…

Cited by 0SourceScholar
2025

Two-dimensional Trajectory Tracking of a Magnetic Continuum Robot by Optimal Magnet Manipulation

IROS 2025

The steerability of catheter is critical to the success of interventional procedure. In this paper, a magnetic continuum robot is presumably mounted to the distal of a catheter to pull it in the narrow, bifurcate, tortuous pathways of the blood vessels. The continuum robot is actuated by a permanent

Cited by 1SourceScholar
2025

VISO-Grasp: Vision-Language Informed Spatial Object-centric 6-DoF Active View Planning and Grasping in Clutter and Invisibility

IROS 2025

We propose VISO-Grasp, a novel vision-language-informed system designed to systematically address visibility constraints for grasping in severely occluded environments. By leveraging Foundation Models (FMs) for spatial reasoning and active view planning, our framework constructs and updates an insta

Cited by 13SourcecodeScholar
2024

Design and Fabrication of a Novel Miniature Magnetic Gripper

ICRA 2024poster

Small-scale robots hold significant promise in the field of minimally invasive surgery (MIS). In this paper, we present a miniature magnetic gripper and develop a data-driven kinematic model. The gripper comprises four fingers, wherein each finger has a maximum size not exceeding 3mm, 4mm and 5.5mm…

Cited by 0SourceScholar
2024

In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering

ICML 2024poster

Large language models (LLMs) demonstrate emergent in-context learning capabilities, where they adapt to new tasks based on example demonstrations. However, in-context learning has seen limited effectiveness in many settings, is difficult to quantitatively control and takes up context window space. T…

2024

Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews

ICML 2024oral

We present an approach for estimating the fraction of text in a large corpus which is likely to be substantially modified or produced by a large language model (LLM). Our maximum likelihood model leverages expert-written and AI-generated reference texts to accurately and efficiently examine real-wor…

2024

POSE-HMR: Heuristic Transformer with Postural Prior Constraints for 3D Human Mesh Reconstruction

ICASSP 2024accepted

This paper proposes an efficient and lightweight model called PoseHMR to address the interference of irrelevant image features and the issues of model inefficiency in 3D human body mesh reconstruction. PoseHMR uses a transformer-decoder architecture and obtains holistic and regional prior constraint…

Cited by 0SourceScholar
2024

TFG: Unified Training-Free Guidance for Diffusion Models

NeurIPS 2024spotlight

Given an unconditional diffusion model and a predictor for a target property of interest (e.g., a classifier), the goal of training-free guidance is to generate samples with desirable target properties without additional training. Existing methods, though effective in various individual applications…

2023

Avoiding spurious correlations via logit correction

ICLR 2023poster

Empirical studies suggest that machine learning models trained with empirical risk minimization (ERM) often rely on attributes that may be spuriously correlated with the class labels. Such models typically lead to poor performance during inference for data lacking such correlations. In this work, we…

2023

IAST: Instance Association Relying on Spatio-Temporal Features for Video Instance Segmentation

ICASSP 2023accepted

Most offline video instance segmentation (VIS) methods lack consideration for multi-scale spatio-temporal features, which leads to unstable instance association across frames. To address this problem, we propose IAST that builds Instance Association relying on Spatio-Temporal features for video inst…

Cited by 0SourceScholar
2023

LEMaRT: Label-Efficient Masked Region Transform for Image Harmonization

CVPR 2023poster

We present a simple yet effective self-supervised pretraining method for image harmonization which can leverage large-scale unannotated image datasets. To achieve this goal, we first generate pre-training data online with our Label-Efficient Masked Region Transform (LEMaRT) pipeline. Given an image,…

Cited by 23SourcePDFScholar
2023

Multiple Instance Learning via Iterative Self-Paced Supervised Contrastive Learning

CVPR 2023poster

Learning representations for individual instances when only bag-level labels are available is a fundamental challenge in multiple instance learning (MIL). Recent works have shown promising results using contrastive self-supervised learning (CSSL), which learns to push apart representations correspon…

2023

PADDLES: Phase-Amplitude Spectrum Disentangled Early Stopping for Learning with Noisy Labels

ICCV 2023poster

Convolutional Neural Networks (CNNs) are powerful in learning patterns of different vision tasks, but they are sensitive to label noise and may overfit to noisy labels during training. The early stopping strategy averts updating CNNs during the early training phase and is widely employed in the pres…

Cited by 14PDFcodeScholar
2023

SQA: Strong Guidance Query with Self-Selected Attention for Human-Object Interaction Detection

ICASSP 2023accepted

The attention mechanism in Transformer-based HOI models plays important role in the comprehension of human and object interaction. However, most previous Transformer-based models ignore the guidance on the query and attention, which leads to a poor understanding of interaction behaviour. In this pap…

Cited by 0SourceScholar
2023

VLKP:Video Instance Segmentation with Visual-Linguistic Knowledge Prompts

ICASSP 2023accepted

Most video instance segmentation(VIS) models only focused on visual knowledge and ignored intrinsic linguistic knowledge. Based on the observation that incorporating linguistic knowledge can significantly improve the model’s contextual understanding of the video, in this paper, we present a Video In…

Cited by 0SourceScholar
2022

Adaptive Early-Learning Correction for Segmentation From Noisy Annotations

CVPR 2022oral

Deep learning in the presence of noisy annotations has been studied extensively in classification, but much less in segmentation tasks. In this work, we study the learning dynamics of deep segmentation networks trained on inaccurately-annotated data. We discover a phenomenon that has been previously…

Cited by 144PDFcodeScholar
2022

Are All Losses Created Equal: A Neural Collapse Perspective

NeurIPS 2022accept

While cross entropy (CE) is the most commonly used loss function to train deep neural networks for classification tasks, many alternative losses have been developed to obtain better empirical performance. Among them, which one is the best to use is still a mystery, because there seem to be multiple…

Cited by 67SourcePDFScholar
2022

Deep Probability Estimation

ICML 2022spotlight

Reliable probability estimation is of crucial importance in many real-world applications where there is inherent (aleatoric) uncertainty. Probability-estimation models are trained on observed outcomes (e.g. whether it has rained or not, or whether a patient has died or not), because the ground-truth…

Cited by 18SourcePDFScholar
2022

HMD-former: a Transformer-based Human Mesh Deformer with Inter-layer Semantic Consistency

ICRA 2022poster

We present a transformer-based network, Human Mesh Deformer (HMD-former), to tackle the problem of 3D human mesh reconstruction from a single RGB image. HMD-former applies a pre-trained CNN to extract image grid features and a transformer decoder to gradually warp the template 3D mesh to the deforme…

Cited by 1SourcecodeScholar
2022

OVIS: Open-Vocabulary Visual Instance Search via Visual-Semantic Aligned Representation Learning

AAAI 2022technical

We introduce the task of open-vocabulary visual instance search (OVIS). Given an arbitrary textual search query, Open-vocabulary Visual Instance Search (OVIS) aims to return a ranked list of visual instances, i.e., image patches, that satisfies the search intent from an image database. The term ``op…

Cited by 2SourcePDFScholar
2022

On Learning Contrastive Representations for Learning With Noisy Labels

CVPR 2022poster

Deep neural networks are able to memorize noisy labels easily with a softmax cross entropy (CE) loss. Previous studies attempted to address this issue focus on incorporating a noise-robust loss function to the CE loss. However, the memorization issue is alleviated but still remains due to the non-ro…

Cited by 81PDFcodeScholar
2022

PA-AWCNN: Two-stream Parallel Attention Adaptive Weight Network for RGB-D Action Recognition

ICRA 2022poster

Due to overly relying on appearance information or adopting direct static feature fusion, most of the existing action recognition methods based on multi-modality have poor robustness and insufficient consideration of modality differences. To address these problems, we propose a two-stream adaptive w…

Cited by 6SourcecodeScholar
2022

Robust Training under Label Noise by Over-parameterization

ICML 2022spotlight

Recently, over-parameterized deep networks, with increasingly more network parameters than training samples, have dominated the performances of modern machine learning. However, when the training data is corrupted, it has been well-known that over-parameterized networks tend to overfit and do not ge…

2022

Video Shadow Detection via Spatio-Temporal Interpolation Consistency Training

CVPR 2022poster

It is challenging to annotate large-scale datasets for supervised video shadow detection methods. Using a model trained on labeled images to the video frames directly may lead to high generalization error and temporal inconsistent results. In this paper, we address these challenges by proposing a Sp…

Cited by 21PDFcodeScholar
2021

Convolutional Normalization: Improving Deep Convolutional Network Robustness and Training

NeurIPS 2021poster

Normalization techniques have become a basic component in modern convolutional neural networks (ConvNets). In particular, many recent works demonstrate that promoting the orthogonality of the weights helps train deep models and improve robustness. For ConvNets, most existing methods are based on pen…

2020

Early-Learning Regularization Prevents Memorization of Noisy Labels

NeurIPS 2020poster

We propose a novel framework to perform classification via deep learning in the presence of noisy annotations. When trained on noisy labels, deep neural networks have been observed to first fit the training data with clean labels during an "early learning" phase, before eventually memorizing the exa…

2020

FCEM: A Novel Fast Correlation Extract Model For Real Time Steganalysis Of VoIP Stream Via Multi-Head Attention

ICASSP 2020accepted

Extracting correlation features between codes-words with high computational efficiency is crucial to steganalysis of Voice over IP (VoIP) streams. In this paper, we utilized attention mechanisms, which have recently attracted enormous interests due to their highly parallelizable computation and flex…

Cited by 0SourceScholar