← Search

Yiming Chen

39 accepted papers

2026

CaFe-TeleVision: A Coarse-To-Fine Teleoperation System with Immersive Situated Visualization for Enhanced Ergonomics

ICRA 2026poster

Teleoperation presents a promising paradigm for remote control and robot proprioceptive data collection. Despite recent progress, current teleoperation systems still suffer from limitations in efficiency and ergonomics, particularly in challenging scenarios. In this paper, we propose CaFe-TeleVision…

2026

CaFe-TeleVision: A Coarse-to-Fine Teleoperation System With Immersive Situated Visualization for Enhanced Ergonomics

RA-L 2026

Teleoperation presents a promising paradigm for remote control and robot proprioceptive data collection. Despite recent progress, current teleoperation systems still suffer from limitations in efficiency and ergonomics, particularly in challenging scenarios. In this paper, we propose CaFe-TeleVision

Cited by 0SourcecodeScholar
2026

Reasoning in Space via Grounding in the World

ICLR 2026poster

In this paper, we claim that 3D visual grounding is the cornerstone of spatial reasoning and introduce the $\textit{Grounded-Spatial Reasoner (GS-Reasoner)}$ to explore the effective spatial representations that bridge the gap between them. Existing 3D LLMs suffer from the absence of a unified 3D re…

Cited by 0SourcecodeScholar
2026

SSR-SAM: Retrieval-Style Segment Anything Model for Semi-Supervised Ultra-High-Resolution Image Segmentation

AAAI 2026technical

Accurate segmentation of ultra-high-resolution (UHR) images, which often exceed tens of millions of pixels, is critically important in domains such as remote sensing and biomedical imaging. However, acquiring pixel-level annotations for such high-resolution images is prohibitively expensive and labo

Cited by 0SourcePDFScholar
2026

Submodel Extraction for Efficient and Personalized Federated Learning via Optimal Transport

CVPR 2026

Federated Learning (FL) enables collaborative model training while preserving data privacy, but its practical deployment is hampered by system and statistical heterogeneity. While federated network pruning offers a path to mitigate these issues, existing methods face a critical dilemma: server-side

Cited by 0SourceScholar
2025

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models

ICML 2025poster

The memory challenges associated with training Large Language Models (LLMs) have become a critical concern, particularly when using the Adam optimizer. To address this issue, numerous memory-efficient techniques have been proposed, with GaLore standing out as a notable example designed to reduce the…

Cited by 1SourcePDFScholar
2025

Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic Evaluation

EMNLP 2025

In the era of evaluating large language models (LLMs), data contamination has become an increasingly prominent concern. To address this risk, LLM benchmarking has evolved from a *static* to a *dynamic* paradigm. In this work, we conduct an in-depth analysis of existing *static* and *dynamic* benchma

2025

Clink! Chop! Thud! - Learning Object Sounds from Real-World Interactions

ICCV 2025poster

Can a model distinguish between the sound of a spoon hitting a hardwood floor versus a carpeted one? Everyday object interactions produce sounds unique to the objects involved. We introduce the sounding object detection task to evaluate a model's ability to link these sounds to the objects directly…

Cited by 0SourcePDFScholar
2025

Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization

ICASSP 2025accepted

Current emotional text-to-speech (TTS) models pre-dominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a single emotion per text-speech pair. These models only learn the correct emotional outputs without fully comprehending…

Cited by 0SourceScholar
2025

Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank Structures

ICLR 2025poster

Parameter-efficient fine-tuning (PEFT) significantly reduces memory costs when adapting large language models (LLMs) for downstream applications. However, traditional first-order (FO) fine-tuning algorithms incur substantial memory overhead due to the need to store activation values for back-propaga…

2025

FDPT: Federated Discrete Prompt Tuning for Black-Box Visual-Language Models

ICCV 2025poster

General-purpose Vision-Language Models (VLMs) have driven major advancements in multimodal AI. Fine-tuning these models with task-specific data enhances adaptability to various downstream tasks but suffers from privacy risks. While potential solutions like federated learning can address user data pr…

Cited by 0SourcePDFScholar
2025

HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class

NeurIPS 2025poster

Large language models (LLMs) have shown remarkable progress in mathematical problem-solving, but evaluation has largely focused on problems that have exact analytical solutions or involve formal proofs, often overlooking approximation-based problems ubiquitous in applied science and engineering. To…

Cited by 0SourcecodeScholar
2025

Integrating Ergonomics and Manipulability for Upper Limb Postural Optimization in Bimanual Human-Robot Collaboration

IROS 2025

This paper introduces an upper limb postural optimization method for enhancing physical ergonomics and force manipulability during bimanual human-robot co-carrying tasks. Existing research typically emphasizes human safety or manipulative efficiency, whereas our proposed method uniquely integrates b

Cited by 1SourceScholar
2025

OmniTry: Virtual Try-On Anything without Masks

NeurIPS 2025poster

Virtual Try-ON (VTON) is a practical and widely-applied task, for which most of existing works focus on clothes. This paper presents OmniTry, a unified framework that extends VTON beyond garment to encompass any wearable objects, e.g., jewelries and accessories, with mask-free setting for more pract…

Cited by 0SourcecodeScholar
2024

A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue Evaluators

AAAI 2024technical

Automatic evaluation is an integral aspect of dialogue system research. The traditional reference-based NLG metrics are generally found to be unsuitable for dialogue assessment. Consequently, recent studies have suggested various unique, reference-free neural metrics that better align with human eva…

2024

Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models

EMNLP 2024finding

Various audio-LLMs (ALLMs) have been explored recently for tackling different audio tasks simultaneously using a single, unified model. While existing evaluations of ALLMs primarily focus on single-audio tasks, real-world applications often involve processing multiple audio streams simultaneously. T…

2024

DreamMesh4D: Video-to-4D Generation with Sparse-Controlled Gaussian-Mesh Hybrid Representation

NeurIPS 2024poster

Recent advancements in 2D/3D generative techniques have facilitated the generation of dynamic 3D objects from monocular videos. Previous methods mainly rely on the implicit neural radiance fields (NeRF) or explicit Gaussian Splatting as the underlying representation, and struggle to achieve satisfac…

Cited by 10SourcePDFScholar
2024

Preliminary Result of Cury: A Backdrivable Leg Design Using Linear Actuators

IROS 2024poster

This paper reports the design, simulation, and experiment of a robotic leg prototype named Cury, which has the potential to achieve minimal clearance and excellent backdrivability. Inspired by human walking data, the actuator design incorporates four-bar linkages and ball screws and is further optim…

Cited by 1SourceScholar
2024

Progressive Poisoned Data Isolation for Training-Time Backdoor Defense

AAAI 2024technical

Deep Neural Networks (DNN) are susceptible to backdoor attacks where malicious attackers manipulate the model's predictions via data poisoning. It is hence imperative to develop a strategy for training a clean model using a potentially poisoned dataset. Previous training-time defense mechanisms typi…

2024

Self-Supervised Multi-Scale Hierarchical Refinement Method for Joint Learning of Optical Flow and Depth

ICASSP 2024accepted

Recurrently refining the optical flow based on a single high-resolution feature demonstrates high performance. We exploit the strength of this strategy to build a novel architecture for the joint learning of optical flow and depth. Our pro-posed architecture is improved to work in the case of traini…

Cited by 0SourceScholar
2024

Unveiling the Achilles’ Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models

ACL 2024findings

The automatic evaluation of natural language generation (NLG) systems presents a long-lasting challenge. Recent studies have highlighted various neural metrics that align well with human evaluations. Yet, the robustness of these evaluators against adversarial perturbations remains largely under-expl…

2023

Dynamic Transformers Provide a False Sense of Efficiency

ACL 2023long

Despite much success in natural language processing (NLP), pre-trained language models typically lead to a high computational cost during inference. Multi-exit is a mainstream approach to address this issue by making a trade-off between efficiency and accuracy, where the saving of computation comes…

2023

Effective Ambiguity Attack Against Passport-Based DNN Intellectual Property Protection Schemes Through Fully Connected Layer Substitution

CVPR 2023poster

Since training a deep neural network (DNN) is costly, the well-trained deep models can be regarded as valuable intellectual property (IP) assets. The IP protection associated with deep models has been receiving increasing attentions in recent years. Passport-based method, which replaces normalizatio…

Cited by 16SourcePDFScholar
2022

Analyzing and Evaluating Faithfulness in Dialogue Summarization

EMNLP 2022main

Dialogue summarization is abstractive in nature, making it suffer from factual errors. The factual correctness of summaries has the highest priority before practical applications. Many efforts have been made to improve faithfulness in text summarization. However, there is a lack of systematic study…

2022

Generate, Discriminate and Contrast: A Semi-Supervised Sentence Representation Learning Framework

EMNLP 2022main

Most sentence embedding techniques heavily rely on expensive human-annotated sentence pairs as the supervised signals. Despite the use of large-scale unlabeled data, the performance of unsupervised methods typically lags far behind that of the supervised counterparts in most downstream tasks. In thi…

2022

Lower Bounds and Nearly Optimal Algorithms in Distributed Learning with Communication Compression

NeurIPS 2022accept

Recent advances in distributed optimization and learning have shown that communication compression is one of the most effective means of reducing communication. While there have been many results for convergence rates with compressed communication, a lower bound is still missing. Analyses of algori…

Cited by 30SourcePDFScholar
2022

Revisiting Optimal Convergence Rate for Smooth and Non-convex Stochastic Decentralized Optimization

NeurIPS 2022accept

While numerous effective decentralized algorithms have been proposed with theoretical guarantees and empirical successes, the performance limits in decentralized optimization, especially the influence of network topology and its associated weight matrix on the optimal convergence rate, have not been…

Cited by 23SourcePDFScholar
2021

Accelerating Gossip SGD with Periodic Global Averaging

ICML 2021spotlight

Communication overhead hinders the scalability of large-scale distributed training. Gossip SGD, where each node averages only with its neighbors, is more communication-efficient than the prevalent parallel SGD. However, its convergence rate is reversely proportional to quantity $1-\beta$ which measu…

Cited by 48SourcePDFScholar
2021

DecentLaM: Decentralized Momentum SGD for Large-Batch Deep Training

ICCV 2021poster

The scale of deep learning nowadays calls for efficient distributed training algorithms. Decentralized momentum SGD (DmSGD), in which each node averages only with its neighbors, is more communication efficient than vanilla Parallel momentum SGD that incurs global average across all computing nodes.…

Cited by 58PDFcodeScholar
2021

DynaEval: Unifying Turn and Dialogue Level Evaluation

ACL 2021long

A dialogue is essentially a multi-turn interaction among interlocutors. Effective evaluation metrics should reflect the dynamics of such interaction. Existing automatic metrics are focused very much on the turn-level quality, while ignoring such dynamics. To this end, we propose DynaEval, a unified…

2021

Exponential Graph is Provably Efficient for Decentralized Deep Training

NeurIPS 2021poster

Decentralized SGD is an emerging training method for deep learning known for its much less (thus faster) communication per iteration, which relaxes the averaging step in parallel SGD to inexact averaging. The less exact the averaging is, however, the more the total iterations the training needs to t…

2021

IMU Data Processing For Inertial Aided Navigation: A Recurrent Neural Network Based Approach

ICRA 2021poster

In this work, we propose a novel method for performing inertial aided navigation, by using deep neural net-works (DNNs). To date, most DNN inertial navigation methods focus on the task of inertial odometry, by taking gyroscope and accelerometer readings as input and regressing for integrated IMU pos…

Cited by 62SourceScholar
2021

Revisiting Self-training for Few-shot Learning of Language Model

EMNLP 2021main

As unlabeled data carry rich task-relevant information, they are proven useful for few-shot learning of language model. The question is how to effectively make use of such data. In this work, we revisit the self-training technique for language model fine-tuning and present a state-of-the-art prompt-…

2021

Road Mapping and Localization Using Sparse Semantic Visual Features

RA-L 2021

We present a novel method for visual mapping and localization for autonomous vehicles, by extracting, modeling, and optimizing semantic road elements. Specifically, our method integrates cascaded deep models to detect standardized road elements instead of traditional point features, to seek for impr

Cited by 31SourceScholar
2020

A Lightweight and Accurate Localization Algorithm Using Multiple Inertial Measurement Units

RA-L 2020

This paper proposes a novel inertial-aided localization approach by fusing information from multiple inertial measurement units (IMUs) and exteroceptive sensors. IMU is a low-cost motion sensor which provides measurements on angular velocity and gravity compensated linear acceleration of a moving pl

Cited by 64SourceScholar
2019

Perception System Design for Low-Cost Commercial Ground Robots: Sensor Configurations, Calibration, Localization and Mapping

IROS 2019poster

For commercially successful ground robots, high degree of autonomy, low manufacturing and maintenance cost, as well as minimized deployment limitations in different environments are essential attributes. To deliver an `anywhere deployable' product, it is impractical to rely on one single sensor or o…

Cited by 8SourceScholar