← Search

Tao Liang

17 accepted papers

2026

A Kind of Flexible Forceps With High Clamping Force and Large Opening Angle for Endoscopic Surgery

RA-L 2026

In endoscopic submucosal dissection, endoscopic forceps is subject to stringent requirements regarding flexibility, miniaturization, and mechanical performance. However, most existing endoscopic forceps suffer from insufficient clamping force and limited opening angle. In this letter, we analyze the

Cited by 0SourceScholar
2026

Interpretable Reward Model via Sparse Autoencoder

AAAI 2026technical

Large language models (LLMs) have been widely deployed across numerous fields. Reinforcement Learning from Human Feedback (RLHF) leverages reward models (RMs) as proxies for human preferences to align LLM behaviors with human values, making the accuracy, reliability, and interpretability of RMs crit

Cited by 0SourcePDFScholar
2026

PreferThinker: Reasoning-based Personalized Image Preference Assessment

ICLR 2026poster

Personalized image preference assessment aims to evaluate an individual user's image preferences by relying only on a small set of reference images as prior information. Existing methods mainly focus on general preference assessment, training models with large-scale data to tackle well-defined task…

Cited by 0SourceScholar
2026

Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image Generation

CVPR 2026

Text-to-image generation has advanced rapidly, yet it still struggles to capture the nuanced user preferences. Existing approaches typically rely on multimodal large language models to infer user preferences, but the derived prompts or latent codes rarely reflect them faithfully, leading to suboptim

Cited by 0SourceScholar
2026

SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models

ICLR 2026poster

Erasing concepts from large-scale text-to-image (T2I) diffusion models has become increasingly crucial due to the growing concerns over copyright infringement, offensive content, and privacy violations. In scalable applications, fine-tuning-based methods are time-consuming to precisely erase multipl…

Cited by 0SourcecodeScholar
2025

Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment

ICCV 2025poster

Contemporary image generation systems have achieved high fidelity and superior aesthetic quality beyond basic text-image alignment. However, existing evaluation frameworks have failed to evolve in parallel. This study reveals that human preference reward models fine-tuned based on CLIP and BLIP arch…

2025

Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs

NeurIPS 2025poster

Large Vision-Language Models (LVLMs) are susceptible to hallucinations, where generated responses seem semantically plausible yet exhibit little or no relevance to the input image. Previous studies reveal that this issue primarily stems from LVLMs' over-reliance on language priors while disregarding…

Cited by 0SourceScholar
2025

Neuron-Level Sequential Editing for Large Language Models

ACL 2025long

This work explores sequential model editing in large language models (LLMs), a critical task that involves modifying internal knowledge within LLMs continuously through multi-round editing, each incorporating updates or corrections to adjust the model’s outputs without the need for costly retraining…

2025

PFDial: A Structured Dialogue Instruction Fine-tuning Method Based on UML Flowcharts

ACL 2025finding

Process-driven dialogue systems, which operate under strict predefined process constraints, are essential in customer service and equipment maintenance scenarios. Although Large Language Models (LLMs) have shown remarkable progress in dialogue and reasoning, they still struggle to solve these strict…

2025

Route Sparse Autoencoder to Interpret Large Language Models

EMNLP 2025

Mechanistic interpretability of large language models (LLMs) aims to uncover the internal processes of information propagation and reasoning. Sparse autoencoders (SAEs) have demonstrated promise in this domain by extracting interpretable and monosemantic features. However, prior works primarily focu

2024

A Novel Miniature Flexible Instrument With Unfolding and Decoupling Design for Endoscopic Surgery

RA-L 2024

Nowadays, gastrointestinal cancer has widely impacted people's health worldwide due to its high mortality rate. Early treatment of gastrointestinal cancer by endoscopic procedure can greatly increase survival rates of patients. Nevertheless, current flexible endoscopic instruments lack of degree of

Cited by 7SourceScholar
2024

A Novel Robot Platform With Decoupled Stiffness Control for Endoscopic Surgery

RA-L 2024

Endoscopic robot has garnered significant attention for its ability to offer auxiliary traction and precise maneuverability. However, the low stiffness of its insertion tube makes it susceptible to deformation, which poses great challenges for precise control. Existing variable stiffness technologie

Cited by 3SourceScholar
2022

Curriculum Knowledge Distillation for Emoji-supervised Cross-lingual Sentiment Analysis

EMNLP 2022main

Existing sentiment analysis models have achieved great advances with the help of sufficient sentiment annotations. Unfortunately, many languages do not have sufficient sentiment corpus. To this end, recent studies have proposed cross-lingual sentiment analysis to transfer sentiment analysis models f…

Cited by 6SourcePDFScholar
2022

Expanding Large Pre-Trained Unimodal Models With Multimodal Information Injection for Image-Text Multimodal Classification

CVPR 2022poster

Fine-tuning pre-trained models for downstream tasks is mainstream in deep learning. However, the pre-trained models are limited to be fine-tuned by data from a specific modality. For example, as a visual model, DenseNet cannot directly take the textual data as its input. Hence, although the large pr…

Cited by 45PDFScholar
2021

Attention Is Not Enough: Mitigating the Distribution Discrepancy in Asynchronous Multimodal Sequence Fusion

ICCV 2021poster

Videos flow as the mixture of language, acoustic, and vision modalities. A thorough video understanding needs to fuse time-series data of different modalities for prediction. Due to the variable receiving frequency for sequences from each modality, there usually exists inherent asynchrony across the…

Cited by 74PDFScholar
2020

Cross-Domain Semantic Segmentation via Domain-Invariant Interactive Relation Transfer

CVPR 2020poster

Exploiting photo-realistic synthetic data to train semantic segmentation models has received increasing attention over the past years. However, the domain mismatch between synthetic and real images will cause a significant performance drop when the model trained with synthetic images is directly app…

Cited by 117PDFScholar
2016

Secure performance analysis of buffer-aided cognitive relay networks under delay unconstraint case

ICASSP 2016accepted

This paper investigates the physical layer security for a buffer-aided cooperative cognitive radio network in the presence of an eavesdropper, wherein the relay is equipped with a buffer so that it can store packets received from secondary source. Multiple primary users (PUs) locate in the transmiss…

Cited by 0SourceScholar