← Search

Yan Hu

17 accepted papers

2026

Does Higher Interpretability Imply Better Utility? A Pairwise Analysis on Sparse Autoencoders

ICLR 2026poster

Sparse Autoencoders (SAEs) are widely used to steer large language models (LLMs), based on the assumption that their interpretable features naturally enable effective model behavior steering. Yet a fundamental question remains: does higher interpretability imply better steering utility? To answer th…

Cited by 0SourceScholar
2026

Open-Vocabulary Spatio-Temporal Scene Graph for Robot Perception and Teleoperation Planning

ICRA 2026poster

Teleoperation via natural-language reduces operator workload and enhances safety in high-risk or remote settings. However, in dynamic remote scenes, transmission latency during bidirectional communication creates gaps between remote perceived states and operator intent, leading to command misunderst…

2026

Real2Sim2Real: RetinalDepth-64K for Depth Estimation in Posterior Segment Ophthalmic Surgery

CVPR 2026

Accurate depth estimation is crucial for 3D reconstruction and precise navigation in posterior segment ophthalmic surgery. However, acquiring annotated data remains challenging due to the impracticality of depth sensors under surgical microscopes. To overcome this limitation, we introduce RetinalDep

Cited by 0SourceScholar
2025

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices

CVPR 2025poster

The emergence and growing popularity of multimodal large language models (MLLMs) have significant potential to enhance various aspects of daily life, from improving communication to facilitating learning and problem-solving. Mobile phones, as essential daily companions, represent the most effective…

2025

Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization

EMNLP 2025

Materials characterization is fundamental to acquiring materials information, revealing the processing-microstructure-property relationships that guide material design and optimization. While multimodal large language models (MLLMs) have recently shown promise in generative and predictive tasks with

2025

DRBO: Mitigating the Bottleneck Effect via Dynamic Reward Balancing in Multi-reward LLM Optimization

EMNLP 2025

In the current landscape of large language models (LLMs), many evaluation metrics have been developed and used as rewards during training to improve specific metrics. However, balancing these metrics and dynamically adjusting reward weights remains challenging, as current approaches often fail to en

2025

RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection

ACL 2025long

Large language models (LLMs) have demonstrated remarkable capabilities in various domains, including radiology report generation. Previous approaches have attempted to utilize multimodal LLMs for this task, enhancing their performance through the integration of domain-specific knowledge retrieval. H…

2025

Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis

ICML 2025poster

Fine-tuning significantly improves the performance of Large Language Models (LLMs), yet its underlying mechanisms remain poorly understood. This paper aims to provide an in-depth interpretation of the fine-tuning process through circuit analysis, a popular tool in *Mechanistic Interpretability (MI)*…

Cited by 0SourcePDFScholar
2025

TwinMarket: A Scalable Behavioral and Social Simulation for Financial Markets

NeurIPS 2025poster

The study of social emergence has long been a central focus in social science. Traditional modeling approaches, such as rule-based Agent-Based Models (ABMs), struggle to capture the diversity and complexity of human behavior, particularly the irrational factors emphasized in behavioral economics. Re…

Cited by 0SourcecodeScholar
2025

UCFE: A User-Centric Financial Expertise Benchmark for Large Language Models

NAACL 2025findings

This paper introduces the UCFE: User-Centric Financial Expertise benchmark, an innovative framework designed to evaluate the ability of large language models (LLMs) to handle complex real-world financial tasks. UCFE benchmark adopts a hybrid approach that combines human expert evaluations with dynam…

2024

ICON: Improving Inter-Report Consistency in Radiology Report Generation via Lesion-aware Mixup Augmentation

EMNLP 2024finding

Previous research on radiology report generation has made significant progress in terms of increasing the clinical accuracy of generated reports. In this paper, we emphasize another crucial quality that it should possess, i.e., inter-report consistency, which refers to the capability of generating c…

2024

Refining Airway Segmentation Through Breakage Filling and Leakage Reduction Using Point Clouds

IROS 2024poster

Bronchoscopy reveals air passages and internal tissues for accurate diagnosis of various lung diseases. Robot-assisted bronchoscopy using an airway tree model can help path planning before surgery and navigation during surgery. In airway tree modeling, though volumetric deep learning methods have ac…

Cited by 0SourceScholar
2023

GRI: Graph-based Relative Isomorphism of Word Embedding Spaces

EMNLP 2023long findings

Automated construction of bi-lingual dictionaries using monolingual embedding spaces is a core challenge in machine translation. The end performance of these dictionaries relies on the geometric similarity of individual spaces, i.e., their degree of isomorphism. Existing attempts aimed at controllin…

Cited by 0SourcecodeScholar
2023

Learning-Based Distortion Compensation for a Hybrid Simulator of Space Docking

RA-L 2023

By effectively utilizing the fidelity of a physical simulation and the flexibility of a numerical simulation, the hybrid simulation is applicable to test the complicated docking contact process of various kinds of spacecraft. However, the hybrid simulation of space docking often has a divergence or

Cited by 5SourceScholar