← Search

Yuxuan Zhou

30 accepted papers

2026

AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile Perception

ICLR 2026poster

Real-world contact-rich manipulation demands robots to perceive temporal tactile feedback, capture subtle surface deformations, and reason about object properties and force dynamics. Although optical tactile sensors are uniquely capable of providing such rich information, existing tactile datasets a…

Cited by 0SourcecodeScholar
2026

Clair Obscur: an Illumination-Aware Method for Real-World Image Vectorization

CVPR 2026

Image vectorization aims to convert raster images into editable, scalable vector representations while preserving visual fidelity. Existing vectorization methods struggle to represent complex real-world images, often producing fragmented shapes at the cost of semantic conciseness. In this paper, we

Cited by 0SourceScholar
2026

Critic–Adviser–Reviser Cyclic Refinement: Towards High-Quality EMR Corpus Generation with LLMs

ICLR 2026poster

Electronic medical records (EMRs) are vital for healthcare research, but their use is limited by privacy concerns. Synthetic EMR generation offers a promising alternative, yet most existing methods merely imitate real records without adhering to rigorous clinical quality principles. To address this,…

Cited by 0SourceScholar
2026

Overcoming Joint Intractability with Lossless Hierarchical Speculative Decoding

ICLR 2026oral

Verification is a key bottleneck in improving inference speed while maintaining distribution fidelity in Speculative Decoding. Recent work has shown that sequence-level verification leads to a higher number of accepted tokens compared to token-wise verification. However, existing solutions often rel…

Cited by 0SourcecodeScholar
2026

Uni-DocRobust: Universal Plug-and-Play Robustness Enhancement for Multi-modal LLMs via Feature Restoration

ICML 2026poster

Real-world degradations, such as noise, blur, and low resolution, significantly impair the performance of Multi-modal Large Language Models (MLLMs) in document understanding tasks. Despite recent advancements, progress in this field remains stifled by two critical bottlenecks: the scarcity of large-…

Cited by 0SourceScholar
2025

Balancing Diversity and Risk in LLM Sampling: How to Select Your Method and Parameter for Open-Ended Text Generation

ACL 2025long

Sampling-based decoding strategies have been widely adopted for Large Language Models (LLMs) in numerous applications, targeting a balance between diversity and quality via temperature tuning and tail truncation. Considering the strong dependency of the candidate next tokens on different prefixes, r…

2025

Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving

ICML 2025poster

Large language models (LLMs) have demonstrated remarkable performance on various medical benchmarks, but their capabilities across different cognitive levels remain underexplored. Inspired by Bloom's Taxonomy, we propose a multi-cognitive-level evaluation framework for assessing LLMs in the medical…

2025

GlyphSR: A Simple Glyph-Aware Framework for Scene Text Image Super-Resolution

AAAI 2025technical

The goal of scene text image super-resolution (STISR) is to enhance the clarity of text within line images, thereby improving readability and enabling more accurate text recognition. However, existing STISR methods often rely heavily on Text Prior (TP) derived from trained recognizers, which can be…

Cited by 0SourcePDFScholar
2025

Investigating and Mitigating Catastrophic Forgetting in Medical Knowledge Injection through Internal Knowledge Augmentation Learning

NeurIPS 2025poster

Large Language Models (LLMs) are expected to possess comprehensive medical knowledge to support real-world clinical applications. While domain-specific fine-tuning effectively injects medical knowledge into LLMs, it often causes catastrophic forgetting of previously acquired knowledge and instructio…

Cited by 0SourcecodeScholar
2025

LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability

EMNLP 2025

Generating long, informative, and factual outputs remains a major challenge for Large Language Models (LLMs). Existing benchmarks for long-form generation typically assess real-world queries with hard-to-verify metrics or use synthetic setups that ease evaluation but overlook real-world intricacies.

2025

MaxSup: Overcoming Representation Collapse in Label Smoothing

NeurIPS 2025oral

Label Smoothing (LS) is widely adopted to reduce overconfidence in neural network predictions and improve generalization. Despite these benefits, recent studies reveal two critical issues with LS. First, LS induces overconfidence in misclassified samples. Second, it compacts feature representations…

Cited by 0SourcecodeScholar
2025

Reliable and Diverse Evaluation of LLM Medical Knowledge Mastery

ICLR 2025poster

Mastering medical knowledge is crucial for medical-specific LLMs. However, despite the existence of medical benchmarks like MedQA, a unified framework that fully leverages existing knowledge bases to evaluate LLMs' mastery of medical knowledge is still lacking. We propose PretexEval, a novel framewo…

Cited by 0SourcePDFScholar
2025

Uni-MuMER: Unified Multi-Task Fine-Tuning of Vision-Language Model for Handwritten Mathematical Expression Recognition

NeurIPS 2025spotlight

Handwritten Mathematical Expression Recognition (HMER) remains a persistent challenge in Optical Character Recognition (OCR) due to the inherent freedom of symbol layouts and variability in handwriting styles. Prior methods have faced performance bottlenecks by proposing isolated architectural modif…

Cited by 0SourcecodeScholar
2025

iKalibr-RGBD: Partially-Specialized Target-Free Visual-Inertial Spatiotemporal Calibration for RGBDs via Continuous-Time Velocity Estimation

RA-L 2025

Visual-inertial systems have been widely studied and applied in the last two decades (from the early 2000 s to the present), mainly due to their low cost and power consumption, small footprint, and high availability. Such a trend simultaneously leads to a large amount of visual-inertial calibration

Cited by 3SourcecodeScholar
2024

BOTH2Hands: Inferring 3D Hands from Both Text Prompts and Body Dynamics

CVPR 2024poster

The recently emerging text-to-motion advances have spired numerous attempts for convenient and interactive human motion generation. Yet existing methods are largely limited to generating body motions only without considering the rich two-hand motions let alone handling various conditions like body d…

2024

BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition

CVPR 2024poster

Graph Convolutional Networks (GCNs) have long set the state-of-the-art in skeleton-based action recognition leveraging their ability to unravel the complex dynamics of human joint topology through the graph's adjacency matrix. However an inherent flaw has come to light in these cutting-edge models:…

2024

DBA-Fusion: Tightly Integrating Deep Dense Visual Bundle Adjustment With Multiple Sensors for Large-Scale Localization and Mapping

RA-L 2024

Visual simultaneous localization and mapping (VSLAM) has broad applications, with state-of-the-art methods leveraging deep neural networks for better robustness and applicability. However, there is a lack of research in fusing these learning-based methods with multi-sensor information, which could b

Cited by 16SourcecodeScholar
2024

Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions

NeurIPS 2024spotlight

Vision-and-Language Navigation (VLN) aims to develop embodied agents that navigate based on human instructions. However, current VLN frameworks often rely on static environments and optimal expert supervision, limiting their real-world applicability. To address this, we introduce Human-Aware Vision-…

2024

MI-Calib: An Open-Source Spatiotemporal Calibrator for Multiple IMUs Based on Continuous-Time Batch Optimization

RA-L 2024

The inertial measurement unit (IMU), as an interoceptive sensor typically providing high-frequency angular velocity and specific force measurements, has been widely exploited for accurate motion estimation in modern robotic applications, such as autonomous navigation and exploration. Recently, there

Cited by 4SourceScholar
2024

MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge

IJCAI 2024poster

Large language models (LLMs) have excelled across domains, also delivering notable performance on the medical evaluation benchmarks, such as MedQA. However, there still exists a significant gap between the reported performance and the practical effectiveness in real-world medical scenarios. In this…

2024

Recognition-Guided Diffusion Model for Scene Text Image Super-Resolution

ICASSP 2024accepted

Scene Text Image Super-Resolution (STISR) aims to enhance the resolution and legibility of text within low-resolution (LR) images, consequently elevating recognition accuracy in Scene Text Recognition (STR). Previous methods predominantly employ discriminative Convolutional Neural Networks (CNNs) au…

Cited by 0SourceScholar
2024

River: A Tightly-Coupled Radar-Inertial Velocity Estimator Based on Continuous-Time Optimization

RA-L 2024

Continuous and reliable ego-velocity information is significant for high-performance motion control and planning in a variety of robotic tasks, such as autonomous navigation and exploration. While linear velocities as first-order kinematics can be simultaneously estimated with other states or explic

Cited by 5SourceScholar
2023

A Magnetically Actuated Miniature Robotic Fish With the Flexible Tail Fin

RA-L 2023

Bionic robotic fish are of great importance in marine resource exploration, military applications and industrial production. However, existing bionic robotic fish often use motor-driven multi-link systems, which are complex and bulky. They are unable to perform narrow underwater operations and indus

Cited by 18SourceScholar
2023

Generative Action Description Prompts for Skeleton-based Action Recognition

ICCV 2023poster

Skeleton-based action recognition has recently received considerable attention. Current approaches to skeleton-based action recognition are typically formulated as one-hot classification tasks and do not fully exploit the semantic relations between actions. For example, "make victory sign" and "thum…

Cited by 65PDFcodeScholar
2022

Continuous and Precise Positioning in Urban Environments by Tightly Coupled Integration of GNSS, INS and Vision

RA-L 2022

Accurate, continuous and seamless state estimation is the fundamental module for intelligent navigation applications, such as self-driving cars and autonomous robots. However, it is often difficult for a standalone sensor to fulfill the demanding requirements of precise navigation in complex scenari

Cited by 42SourceScholar
2022

Table-based Fact Verification with Self-adaptive Mixture of Experts

ACL 2022findings

The table-based fact verification task has recently gained widespread attention and yet remains to be a very challenging problem. It inherently requires informative reasoning over natural language together with different numerical and logical reasoning on tables (e.g., count, superlative, comparativ…

2022

Visual Mapping and Localization System Based on Compact Instance-Level Road Markings With Spatial Uncertainty

RA-L 2022

High-definition (HD) map is crucial for intelligent vehicles to perform high-level localization and navigation. To improve the availability and usability of HD map, it is meaningful to investigate crowd-sourced mapping solutions and low-cost map-aided localization schemes which don't rely on high-en

Cited by 17SourceScholar
2020

Understanding Anomaly Detection with Deep Invertible Networks through Hierarchies of Distributions and Features

NeurIPS 2020poster

Deep generative networks trained via maximum likelihood on a natural image dataset like CIFAR10 often assign high likelihoods to images from datasets with different objects (e.g., SVHN). We refine previous investigations of this failure at anomaly detection for invertible generative networks and pr…