← Search

Sheng Wang

63 accepted papers

2026

Let’s Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts

AAAI 2026technical

Recently, Interleaved-modal Chain-of-Thought (ICoT) reasoning has achieved remarkable success by leveraging both multimodal inputs and outputs, attracting increasing attention. While achieving promising performance, current ICoT methods still suffer from two major limitations: (1) Static Visual Thou

Cited by 0SourcePDFScholar
2026

Masked-Diffusion Autoencoders for 3D Medical Vision Representation Learning

CVPR 2026

Effective medical image analysis requires representations that capture both global anatomical structure and fine-grained tissue texture. Current self-supervised approaches exhibit limited capacity to address both requirements simultaneously. Invariance-based methods learn through augmentation consis

Cited by 0SourceScholar
2026

Mechanistic Analysis of Cable Tension Effects on the Stiffness of Cable-Driven Serpentine Manipulators

RA-L 2026

This paper presents a mechanistic analysis of stiffness in cable-driven serpentine manipulators (CDSMs), incorporating both cable tension and cable stiffness. First, we derive an analytical stiffness model based on robot statics, identifying cable tension and stiffness as the dominant factors govern

Cited by 0SourceScholar
2026

Mechanistic Analysis of Cable Tension Effects on the Stiffness of Cable-Driven Serpentine Manipulators

ICRA 2026poster

This paper presents a mechanistic analysis of stiffness in cable-driven serpentine manipulators (CDSMs), incorporating both cable tension and cable stiffness. First, we derive an analytical stiffness model based on robot statics, identifying cable tension and stiffness as the dominant factors govern…

Cited by 0SourceScholar
2025

A Study on Enhancing Wearer Adaptation Through Accurate Gait Phase Prediction and Gradual Increase in Assistive Force Magnitude in Exosuits

RA-L 2025

Human-exosuit adaptation is a bi-directional process: exosuit-to-human locomotion adaptation maximizes the benefits of exosuit assistance, while human-to-exosuit adaptation accelerates the wearer's access to these benefits. To promote bi-directional adaptation, we investigated precise gait phase pre

Cited by 2SourceScholar
2025

Demeter: A Parametric Model of Crop Plant Morphology from the Real World

ICCV 2025poster

Learning 3D parametric shape models of objects has gained popularity in vision and graphics and has showed broad utility in 3D reconstruction, generation, understanding, and simulation. While powerful models exist for humans and animals, equally expressive approaches for modeling plants are lacking.…

Cited by 0SourcePDFScholar
2025

Developing and Utilizing a Large-Scale Cantonese Dataset for Multi-Tasking in Large Language Models

EMNLP 2025

High-quality data resources play a crucial role in learning large language models (LLMs), particularly for low-resource languages like Cantonese. Despite having more than 85 million native speakers, Cantonese is still considered a low-resource language in the field of natural language processing (NL

2025

Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation

EMNLP 2025

End-to-end speech-to-speech (S2S) dialogue systems have recently garnered increasing research attention for their lower latency and more natural integration of nonverbal cues such as emotion and speaker identity. However, these systems face key challenges, particularly in incorporating external know

2025

Forewarned is Forearmed: Harnessing LLMs for Data Synthesis via Failure-induced Exploration

ICLR 2025poster

Large language models (LLMs) have significantly benefited from training on diverse, high-quality task-specific data, leading to impressive performance across a range of downstream applications. Current methods often rely on human-annotated data or predefined task templates to direct powerful LLMs in…

Cited by 0SourcePDFScholar
2025

GDTS: Goal-Guided Diffusion Model with Tree Sampling for Multi-Modal Pedestrian Trajectory Prediction

IROS 2025

Accurate prediction of pedestrian trajectories is crucial for improving the safety of autonomous driving. However, this task is generally nontrivial due to the inherent stochasticity of human motion, which naturally requires the predictor to generate multi-modal prediction. Previous works leverage v

Cited by 2SourceScholar
2025

Group Ligands Docking to Protein Pockets

ICLR 2025poster

Molecular docking is a key task in computational biology that has attracted increasing interest from the machine learning community. While existing methods have achieved success, they generally treat each protein-ligand pair in isolation. Inspired by the biochemical observation that ligands binding…

Cited by 1SourcePDFScholar
2025

HAMF: A Hybrid Attention-Mamba Framework for Joint Scene Context Understanding and Future Motion Representation Learning

IROS 2025

Motion forecasting represents a critical challenge in autonomous driving systems, requiring accurate prediction of surrounding agents’ future trajectories. While existing approaches predict future motion states with the extracted scene context feature from historical agent trajectories and road layo

Cited by 4SourceScholar
2025

Hotspot-Driven Peptide Design via Multi-Fragment Autoregressive Extension

ICLR 2025poster

Peptides, short chains of amino acids, interact with target proteins, making them a unique class of protein-based therapeutics for treating human diseases. Recently, deep generative models have shown great promise in peptide generation. However, several challenges remain in designing effective pepti…

2025

How Well Do LLMs Handle Cantonese? Benchmarking Cantonese Capabilities of Large Language Models

NAACL 2025findings

The rapid evolution of large language models (LLMs) has transformed the competitive landscape in natural language processing (NLP), particularly for English and other data-rich languages. However, underrepresented languages like Cantonese, spoken by over 85 million people, face significant developme…

2025

MITracker: Multi-View Integration for Visual Object Tracking

CVPR 2025highlight

Multi-view object tracking (MVOT) offers promising solutions to challenges such as occlusion and target loss, which are common in traditional single-view tracking. However, progress has been limited by the lack of comprehensive multi-view datasets and effective cross-view integration methods. To ove…

Cited by 0SourcePDFScholar
2025

MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models

ICLR 2025poster

Artificial Intelligence (AI) has demonstrated significant potential in healthcare, particularly in disease diagnosis and treatment planning. Recent progress in Medical Large Vision-Language Models (Med-LVLMs) has opened up new possibilities for interactive diagnostic tools. However, these models oft…

2025

MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization

ICML 2025poster

The advancement of Large Vision-Language Models (LVLMs) has propelled their application in the medical field. However, Medical LVLMs (Med-LVLMs) encounter factuality challenges due to modality misalignment, where the models prioritize textual knowledge over visual input, leading to hallucinations th…

2025

MUC: Mixture of Uncalibrated Cameras for Robust 3D Human Body Reconstruction

AAAI 2025technical

Multiple cameras can provide comprehensive multi-view video coverage of a person. Fusing this multi-view data is crucial for tasks like behavioral analysis, although it traditionally requires camera calibration—a process that is often complex. Moreover, previous studies have overlooked the challenge…

2025

MlingConf: A Comprehensive Study of Multilingual Confidence Estimation on Large Language Models

ACL 2025finding

The tendency of Large Language Models (LLMs) to generate hallucinations raises concerns regarding their reliability. Therefore, confidence estimations indicating the extent of trustworthiness of the generations become essential. However, current LLM confidence estimations in languages other than Eng…

2025

MoS: Unleashing Parameter Efficiency of Low-Rank Adaptation with Mixture of Shards

ICLR 2025poster

The rapid scaling of large language models necessitates more lightweight finetuning methods to reduce the explosive GPU memory overhead when numerous customized models are served simultaneously. Targeting more parameter-efficient low-rank adaptation (LoRA), parameter sharing presents a promising sol…

2025

ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and Wisdom

EMNLP 2025

Large vision-language models (LVLMs) have witnessed significant progress on visual understanding tasks. However, they often prioritize language knowledge over image information on visual reasoning tasks, incurring performance degradation. To tackle this issue, we first identify the drawbacks of exis

2025

QSpec: Speculative Decoding with Complementary Quantization Schemes

EMNLP 2025

Quantization is widely adopted to accelerate inference and reduce memory consumption in large language models (LLMs). While activation-weight joint quantization enables efficient low-precision decoding, it suffers substantial performance degradation on multi-step reasoning tasks. We propose QSPEC, a

2025

Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic Control

NeurIPS 2025poster

Despite recent advances in diffusion models, top-tier text-to-image (T2I) models still struggle to achieve precise spatial layout control, *i.e.* accurately generating entities with specified attributes and locations. Segmentation-mask-to-image (S2I) generation has emerged as a promising solution by…

Cited by 0SourceScholar
2025

TreeSynth: Synthesizing Diverse Data from Scratch via Tree-Guided Subspace Partitioning

NeurIPS 2025spotlight

Model customization necessitates high-quality and diverse datasets, but acquiring such data remains time-consuming and labor-intensive. Despite the great potential of large language models (LLMs) for data synthesis, current approaches are constrained by limited seed data, model biases and low-varia…

Cited by 0SourcecodeScholar
2025

UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models

ACL 2025long

Despite demonstrating impressive capabilities, Large Language Models (LLMs) still often struggle to accurately express the factual knowledge they possess, especially in cases where the LLMs’ knowledge boundaries are ambiguous. To improve LLMs’ factual expressions, we propose the UAlign framework, wh…

2024

A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery

EMNLP 2024main

In many scientific fields, large language models (LLMs) have revolutionized the way text and other modalities of data (e.g., molecules and proteins) are handled, achieving superior performance in various applications and augmenting the scientific discovery process. Nevertheless, previous surveys on…

2024

A Generic Trajectory Planning Method for Constrained All-Wheel-Steering Robots

IROS 2024poster

This paper presents a generic trajectory planning method for wheeled robots with fixed steering axes while the steering angle of each wheel is constrained. In the existing literatures, All-Wheel-Steering (AWS) robots, incorporating modes such as rotation-free translation maneuvers, in-situ rotationa…

Cited by 0SourcecodeScholar
2024

A Novel Friction Measuring Method and Its Application to Improve the Static Modeling Accuracy of Cable-Driven Continuum Manipulators

RA-L 2024

Cable-driven continuum manipulators exhibit high flexibility and dexterity, leading to their increased popularity in recent years. Friction analysis is a crucial problem for these manipulators. Previous research has introduced friction models that are applicable to dynamic states where the direction

Cited by 6SourceScholar
2024

Cross-Domain Contrastive Learning for Time Series Clustering

AAAI 2024technical

Most deep learning-based time series clustering models concentrate on data representation in a separate process from clustering. This leads to that clustering loss cannot guide feature extraction. Moreover, most methods solely analyze data from the temporal domain, disregarding the potential within…

2024

DHP-Mapping: A Dense Panoptic Mapping System with Hierarchical World Representation and Label Optimization Techniques

IROS 2024poster

Maps provide robots with crucial environmental knowledge, thereby enabling them to perform interactive tasks effectively. Easily accessing accurate abstract-to-detailed geometric and semantic concepts from maps is crucial for robots to make informed and efficient decisions. To comprehensively model…

Cited by 2SourcecodeScholar
2024

DragTraffic: Interactive and Controllable Traffic Scene Generation for Autonomous Driving

IROS 2024

Evaluating and training autonomous driving systems require diverse and scalable corner cases. However, most existing scene generation methods lack controllability, accuracy, and versatility, resulting in unsatisfactory generation results. Inspired by DragGAN in image generation, we propose DragTraff

Cited by 6SourcecodeScholar
2024

Improving Autonomous Driving Safety with POP: A Framework for Accurate Partially Observed Trajectory Predictions

ICRA 2024poster

Accurate trajectory prediction is crucial for safe and efficient autonomous driving, but handling partial observations presents significant challenges. To address this, we propose a novel trajectory prediction framework called Partial Observations Prediction (POP) for congested urban road scenarios.…

Cited by 6SourcecodeScholar
2024

LoRA Meets Dropout under a Unified Framework

ACL 2024findings

With the remarkable capabilities, large language models (LLMs) have emergedas essential elements in numerous NLP applications, while parameter-efficientfinetuning, especially LoRA, has gained popularity as a lightweight approachfor model customization. Meanwhile, various dropout methods, initially d…

Cited by 11SourcePDFScholar
2024

MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

ICLR 2024poster

Since the resurgence of deep learning, vision-language models (VLMs) enhanced by large language models (LLMs) have grown exponentially in popularity. However, while LLMs can utilize extensive background knowledge and task information with in-context learning, most VLMs still struggle with understan…

2024

Mining Gaze for Contrastive Learning toward Computer-Assisted Diagnosis

AAAI 2024technical

Obtaining large-scale radiology reports can be difficult for medical images due to ethical concerns, limiting the effectiveness of contrastive pre-training in the medical image domain and underscoring the need for alternative methods. In this paper, we propose eye-tracking as an alternative to text…

2024

PRoLoRA: Partial Rotation Empowers More Parameter-Efficient LoRA

ACL 2024long

With the rapid scaling of large language models (LLMs), serving numerouslow-rank adaptations (LoRAs) concurrently has become increasingly impractical,leading to unaffordable costs and necessitating more parameter-efficientfinetuning methods. In this work, we introduce Partially Rotation-enhanced Low…

2024

Physical Property Understanding from Language-Embedded Feature Fields

CVPR 2024poster

Can computers perceive the physical properties of objects solely through vision? Research in cognitive science and vision science has shown that humans excel at identifying materials and estimating their physical properties based purely on visual appearance. In this paper we present a novel approach…

Cited by 12SourcePDFScholar
2024

Retrieved Sequence Augmentation for Protein Representation Learning

EMNLP 2024main

Protein Language Models traditionally depend on Multiple Sequence Alignments (MSA) to incorporate evolutionary knowledge. However, MSA-based approaches suffer from substantial computational overhead and generally underperform in generalizing to de novo proteins. This study reevaluates the role of MS…

2023

A Cognitive Stimulation Dialogue System with Multi-source Knowledge Fusion for Elders with Cognitive Impairment

ACL 2023long

When communicating with elders with cognitive impairment, cognitive stimulation (CS) help to maintain the cognitive health of elders. Data sparsity is the main challenge in building CS-based dialogue systems, particularly in the Chinese language. To fill this gap, we construct a Chinese CS conversat…

2023

GraphPrompt: Graph-Based Prompt Templates for Biomedical Synonym Prediction

AAAI 2023technical

In the expansion of biomedical dataset, the same category may be labeled with different terms, thus being tedious and onerous to curate these terms. Therefore, automatically mapping synonymous terms onto the ontologies is desirable, which we name as biomedical synonym prediction task. Unlike biomedi…

2023

Robust One-Shot Segmentation of Brain Tissues via Image-Aligned Style Transformation

AAAI 2023technical

One-shot segmentation of brain tissues is typically a dual-model iterative learning: a registration model (reg-model) warps a carefully-labeled atlas onto unlabeled images to initialize their pseudo masks for training a segmentation model (seg-model); the seg-model revises the pseudo masks to enhanc…

2023

Self-Supervised Drivable Area Segmentation Using LiDAR's Depth Information for Autonomous Driving

IROS 2023poster

Drivable area segmentation is an essential component of the visual perception system for autonomous driving vehicles. Recent efforts in deep neural networks have sig-nificantly improved semantic segmentation performance for autonomous driving. However, most DNN-based methods need a large amount of d…

Cited by 9SourceScholar
2022

Antigen-Specific Antibody Design and Optimization with Diffusion-Based Generative Models for Protein Structures

NeurIPS 2022accept

Antibodies are immune system proteins that protect the host by binding to specific antigens such as viruses and bacteria. The binding between antibodies and antigens is mainly determined by the complementarity-determining regions (CDR) of the antibodies. In this work, we develop a deep generative mo…

Cited by 244SourcePDFScholar
2022

Contact-Distil: Boosting Low Homologous Protein Contact Map Prediction by Self-Supervised Distillation

AAAI 2022technical

Accurate protein contact map prediction (PCMP) is essential for precise protein structure estimation and further biological studies. Recent works achieve significant performance on this task with high quality multiple sequence alignment (MSA). However, the PCMP accuracy drops dramatically while only…

2022

DisenCite: Graph-Based Disentangled Representation Learning for Context-Specific Citation Generation

AAAI 2022technical

Citing and describing related literature are crucial to scientific writing. Many existing approaches show encouraging performance in citation recommendation, but are unable to accomplish the more challenging and onerous task of citation text generation. In this paper, we propose a novel disentangled…

2022

Distribution-Informed Neural Networks for Domain Adaptation Regression

NeurIPS 2022accept

In this paper, we study the problem of domain adaptation regression, which learns a regressor for a target domain by leveraging the knowledge from a relevant source domain. We start by proposing a distribution-informed neural network, which aims to build distribution-aware relationship of inputs and…

Cited by 18SourcePDFScholar
2022

MetaFill: Text Infilling for Meta-Path Generation on Heterogeneous Information Networks

EMNLP 2022main

Heterogeneous information network (HIN) is essential to study complicated networks containing multiple edge types and node types. Meta-path, a sequence of node types and edge types, is the core technique to embed HINs. Since manually curating meta-paths is time-consuming, there is a pressing need to…

2022

Pathway2Text: Dataset and Method for Biomedical Pathway Description Generation

NAACL 2022findings

Biomedical pathways have been extensively used to characterize the mechanism of complex diseases. One essential step in biomedical pathway analysis is to curate the description of a pathway based on its graph structure and node features. Neural text generation could be a plausible technique to circu…

2022

Seed-Guided Topic Discovery with Out-of-Vocabulary Seeds

NAACL 2022long

Discovering latent topics from text corpora has been studied for decades. Many existing topic models adopt a fully unsupervised setting, and their discovered topics may not cater to users’ particular interests due to their inability of leveraging user guidance. Although there exist seed-guided topic…

2022

Towards Accurate Active Camera Localization

ECCV 2022poster

"In this work, we tackle the problem of active camera localization, which controls the camera movements actively to achieve an accurate camera pose. The past solutions are mostly based on Markov Localization, which reduces the position-wise camera uncertainty for localization. These approaches local…

2022

Weakly Supervised Object Localization through Inter-class Feature Similarity and Intra-Class Appearance Consistency

ECCV 2022poster

"Weakly supervised object localization (WSOL) aims at detecting objects through only image-level labels. Class activation maps (CAMs) are the commonly used features for WSOL. However, existing CAM-based methods tend to excessively pursue discriminative features for object recognition and hence ignor…

Cited by 14SourcePDFScholar
2021

Adaptive Residue-wise Profile Fusion for Low Homologous Protein Secondary Structure Prediction Using External Knowledge

IJCAI 2021poster

Protein secondary structure prediction (PSSP) is essential for protein function analysis. However, for low homologous proteins, the PSSP suffers from insufficient input features. In this paper, we explicitly import external self-supervised knowledge for low homologous PSSP under the guidance of resi…

2021

Auto-Encoding Knowledge Graph for Unsupervised Medical Report Generation

NeurIPS 2021poster

Medical report generation, which aims to automatically generate a long and coherent report of a given medical image, has been receiving growing research interests. Existing approaches mainly adopt a supervised manner and heavily rely on coupled image-report pairs. However, in the medical domain, bui…

Cited by 135SourcePDFScholar
2021

Graphine: A Dataset for Graph-aware Terminology Definition Generation

EMNLP 2021main

Precisely defining the terminology is the first step in scientific communication. Developing neural text generation models for definition generation can circumvent the labor-intensity curation, further accelerating scientific discovery. Unfortunately, the lack of large-scale terminology definition d…

2021

InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds Through Instance Multi-Level Contextual Referring

ICCV 2021poster

Compared with the visual grounding on 2D images, the natural-language-guided 3D object localization on point clouds is more challenging. In this paper, we propose a new model, named InstanceRefer, to achieve a superior 3D visual grounding through the grounding-by-matching strategy. In practice, our…

Cited by 147PDFcodeScholar
2021

PSSM-Distil: Protein Secondary Structure Prediction (PSSP) on Low-Quality PSSM by Knowledge Distillation with Contrastive Learning

AAAI 2021technical

Protein secondary structure prediction (PSSP) is an essential task in computational biology. To achieve the accurate PSSP, the general and vital feature engineering is to use multiple sequence alignment (MSA) for Position-Specific Scoring Matrix (PSSM) extraction. However, when only low-quality PSSM…

2021

Shallow Feature Matters for Weakly Supervised Object Localization

CVPR 2021poster

Weakly supervised object localization (WSOL) aims to localize objects by only utilizing image-level labels. Class activation maps (CAMs) are the commonly used features to achieve WSOL. However, previous CAM-based methods did not take full advantage of the shallow features, despite their importance f…

Cited by 117PDFcodeScholar
2020

Label-Driven Reconstruction for Domain Adaptation in Semantic Segmentation

ECCV 2020poster

Unsupervised domain adaptation enables to alleviate the need for pixel-wise annotation in the semantic segmentation. One of the most common strategies is to translate images from the source domain to the target domain and then align their marginal distributions in the feature space using adversarial…

Cited by 113SourcePDFScholar
2020

PointASNL: Robust Point Clouds Processing Using Nonlocal Neural Networks With Adaptive Sampling

CVPR 2020poster

Raw point clouds data inevitably contains outliers or noise through acquisition from 3D sensors or reconstruction algorithms. In this paper, we present a novel end-to-end network for robust point clouds processing, named PointASNL, which can deal with point clouds with noise effectively. The key com…

Cited by 764PDFcodeScholar