← Search

Yao He

15 accepted papers

2026

Balanced Knowledge Distillation for Large Language Models with Mix-of-Experts

AAAI 2026technical

Mixture-of-Experts (MoE) architectures have recently become a more prevalent choice for large language models (LLMs) than dense architectures due to their superior performance. However, billions of parameters bring MoE LLMs a huge cost for deployment and inference. To address these issues, knowledge

Cited by 0SourcePDFScholar
2026

Benchmarking and Enhancing Relational Diagrams Reasoning for Multimodal Large Language Models

IJCAI 2026

Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision--language tasks. However, their capabilities remain unclear in relational-diagram (RD) reasoning, where correct answers must satisfy diagram-defined constraints such as directed dependencies, branchin

Cited by 0Scholar
2026

Feed-forward Human Performance Capture via Progressive Canonical Space Updates

ICLR 2026poster

We present a feed-forward human performance capture method that renders novel views of a performer from a monocular RGB stream. A key challenge in this setting is the lack of sufficient observations, especially for unseen regions. Assuming the subject moves continuously over time, we take advantage…

Cited by 0SourceScholar
2026

Optimizing LoRA Allocation of MoE with the Alignment of Topic Correlation

AAAI 2026technical

Mixture of experts (MoE) dynamically routes inputs to specialized expert networks to scale model capacity with low inference overhead. However, the excessive parameter growth in MoE models poses challenges in low-resource settings. To address these issues, MoE with parameter-efficient fine-tuning (P

Cited by 0SourcePDFScholar
2026

PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model

ICLR 2026poster

Vision-Language-Action models (VLAs) are emerging as powerful tools for learning generalizable visuomotor control policies. However, current VLAs are mostly trained on large-scale image–text–action data and remain limited in two key ways: (i) they struggle with pixel-level scene understanding, and (…

Cited by 0SourceScholar
2026

S2FT: Parameter-Efficient Fine-Tuning in Sparse Spectrum Domain

CVPR 2026

Parameter Efficient Fine-Tuning (PEFT) is a key technique for adapting a large pretrained model to downstream tasks by fine-tuning only a small number of parameters. Recent methods based on Fourier transforms have further reduced the fine-tuned parameters scale by only fine-tuning a few spectral coe

Cited by 0SourceScholar
2025

Demonstrating ViSafe: Vision-enabled Safety for High-speed Detect and Avoid

RSS 2025poster

Maintaining visual separation is crucial to achieving safe and seamless high-density operation of airborne vehicles in shared airspace, where pilots currently shoulder this responsibility. To automate this, we present ViSafe, a high-speed airborne vision-only collision avoidance system. Designed un…

Cited by 0PDFScholar
2025

FIReStereo: Forest InfraRed Stereo Dataset for UAS Depth Perception in Visually Degraded Environments

RA-L 2025

Robust depth perception in visually-degraded environments is crucial for autonomous aerial systems. Thermal imaging cameras, which capture infrared radiation, are robust to visual degradation. However, due to lack of a large-scale dataset, the use of thermal cameras for uncrewed aerial system (UAS)

Cited by 8SourceScholar
2025

Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids

CoRL 2025poster

Simulation-based reinforcement learning (RL) has significantly advanced humanoid locomotion tasks, yet direct real-world RL from scratch or starting from pretrained policies remains rare, limiting the full potential of humanoid robots. Real-world training, despite being crucial for overcoming the si…

Cited by 0SourceScholar
2025

Robust Visual Odometry Using Rigidly-Bundled Arbitrarily-Arranged Multi-Cameras

RA-L 2025

Making multi-camera visual SLAM systems easier to set up and more robust to the environment is attractive for vision robots. Existing monocular and binocular vision SLAM systems have narrow sensing Field-of-View (FoV), resulting in degenerated accuracy and limited robustness in textureless environme

Cited by 0SourcecodeScholar
2025

Stem-OB: Generalizable Visual Imitation Learning with Stem-Like Convergent Observation through Diffusion Inversion

ICLR 2025spotlight

Visual imitation learning methods demonstrate strong performance, yet they lack generalization when faced with visual input perturbations like variations in lighting and textures. This limitation hampers their practical application in real-world settings. To address this, we propose ***Stem-OB*** th…

2025

Towards Improved Text-Aligned Codebook Learning: Multi-Hierarchical Codebook-Text Alignment with Long Text

CVPR 2025highlight

Image quantization is a crucial technique in image generation, aimed at learning a codebook that encodes an image into a discrete token sequence. Recent advancements have seen researchers exploring learning multi-modal codebook (i.e., text-aligned codebook) by utilizing image caption semantics, aimi…

Cited by 0SourcePDFScholar
2024

Codebook Transfer with Part-of-Speech for Vector-Quantized Image Modeling

CVPR 2024poster

Vector-Quantized Image Modeling (VQIM) is a fundamental research problem in image synthesis which aims to represent an image with a discrete token sequence. Existing studies effectively address this problem by learning a discrete codebook from scratch and in a code-independent manner to quantize con…

Cited by 11SourcePDFScholar
2022

Towards Robust Visual-Inertial Odometry with Multiple Non-Overlapping Monocular Cameras

IROS 2022poster

We present a Visual-Inertial Odometry (VIO) algorithm with multiple non-overlapping monocular cameras aiming at improving the robustness of the VIO algorithm. An initialization scheme and tightly-coupled bundle adjustment for multiple non-overlapping monocular cameras are proposed. With more stable…

Cited by 7SourceScholar