← Search

Yun Zhang

22 accepted papers

2026

EnerGS: Energy-Based Gaussian Splatting under Partial Geometric Observability

ICML 2026poster

3D Gaussian Splatting (3DGS) has been widely adopted for scene reconstruction, where training inherently constitutes a highly coupled and non-convex optimization problem. Recent works commonly incorporate geometric priors, such as LiDAR measurements, either for initialization or as training constrai…

Cited by 0SourceScholar
2026

RelMap: Enhancing Online Map Construction with Class-Aware Spatial Relation and Semantic Priors

ICRA 2026poster

Online high-definition (HD) map construction is crucial for scaling autonomous driving systems. While Transformer-based methods have become prevalent in online HD map construction, most existing approaches overlook the inherent spatial dependencies and semantic relationships among map elements, whic…

2026

TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments

ICML 2026poster

Robots in dynamic, human-centric environments must follow language instructions while maintaining real-time reactive control. Vision-language-action (VLA) models offer a promising framework, but they assume temporally aligned reasoning and control, despite semantic inference being inherently delayed…

Cited by 0SourceScholar
2025

AugKD: Ingenious Augmentations Empower Knowledge Distillation for Image Super-Resolution

ICLR 2025poster

Knowledge distillation (KD) compresses deep neural networks by transferring task-related knowledge from cumbersome pre-trained teacher models to more compact student models. However, vanilla KD for image super-resolution (SR) networks yields only limited improvements due to the inherent nature of SR…

Cited by 0SourcePDFScholar
2025

AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning

NeurIPS 2025poster

Recent advancements in Vision-Language-Action (VLA) models have shown promise for end-to-end autonomous driving by leveraging world knowledge and reasoning capabilities. However, current VLA models often struggle with physically infeasible action outputs, complex model structures, or unnecessarily l…

Cited by 0SourcecodeScholar
2025

CBQ: Cross-Block Quantization for Large Language Models

ICLR 2025spotlight

Post-training quantization (PTQ) has played a pivotal role in compressing large language models (LLMs) at ultra-low costs. Although current PTQ methods have achieved promising results by addressing outliers and employing layer- or block-wise loss optimization techniques, they still suffer from signi…

Cited by 13SourcePDFScholar
2025

Knowledge Distillation with Multi-granularity Mixture of Priors for Image Super-Resolution

ICLR 2025spotlight

Knowledge distillation (KD) is a promising yet challenging model compression approach that transmits rich learning representations from robust but resource-demanding teacher models to efficient student models. Previous methods for image super-resolution (SR) are often tailored to specific teacher-st…

Cited by 4SourcePDFScholar
2025

MeMoTune: A Measure and Moment-Driven Fine-Tuning Framework for Quantized Large Language Models

ACL 2025finding

Quantizing large language models (LLMs) is essential for reducing memory and computational costs in natural language processing. Existing methods combine quantization with parameter-efficient fine-tuning but often fail to meet practical performance requirements. This paper introduces MeMoTune, a nov…

2025

PANTHER: Generative Pretraining Beyond Language for Sequential User Behavior Modeling

NeurIPS 2025poster

Large language models (LLMs) have shown that generative pretraining can distill vast world knowledge into compact token representations. While LLMs encapsulate extensive world knowledge, they remain limited in modeling the behavioral knowledge contained within user interaction histories. User behavi…

Cited by 0SourceScholar
2025

PC-SRIF: Preconditioned Cholesky-based Square Root Information Filter for Vision-aided Inertial Navigation

IROS 2025

In this paper, we introduce a novel estimator for vision-aided inertial navigation systems (VINS), the Preconditioned Cholesky-based Square Root Information Filter (PC-SRIF). When solving linear systems, employing Cholesky decomposition offers superior efficiency but can compromise numerical stabili

Cited by 2SourceScholar
2025

V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and Prediction

ICCV 2025poster

Vehicle-to-everything (V2X) technologies offer a promising paradigm to mitigate the limitations of constrained observability in single-vehicle systems. Prior work primarily focuses on single-frame cooperative perception, which fuses agents' information across different spatial locations but ignores…

2024

Eliminating Warping Shakes for Unsupervised Online Video Stitching

ECCV 2024poster

"In this paper, we retarget video stitching to an emerging issue, named warping shake, when extending image stitching to video stitching. It unveils the temporal instability of warped content in non-overlapping regions, despite image stitching having endeavored to preserve the natural structures. Th…

2024

Soften to Defend: Towards Adversarial Robustness via Self-Guided Label Refinement

CVPR 2024poster

Adversarial training (AT) is currently one of the most effective ways to obtain the robustness of deep neural networks against adversarial attacks. However most AT methods suffer from robust overfitting i.e. a significant generalization gap in adversarial robustness between the training and testing…

Cited by 3SourcePDFScholar
2022

Jet-HR2: A Flying Bipedal Robot Based on Thrust Vector Control

RA-L 2022

Achieving short-distance flight helps improve the efficiency of bipedal robots moving in complex environments (e.g., crossing large obstacles or reaching high places) for rapid emergency missions. This study proposes a design of a flying bipedal robot named Jet-HR2 ( <xref ref-type="fig" rid="fig1"

Cited by 10SourceScholar
2021

Biomimetic Flip-and-Flap Strategy of Flying Objects for Perching on Inclined Surfaces

RA-L 2021

Animals can use the maneuver of a flipping body and flapping wings to reduce the normal rebound force of impact during landing, decreasing the adsorption force required by the contact point. This capability aids aerial vehicles with landing not only on vertical surfaces, but also on inclined surface

Cited by 11SourceScholar
2021

DymSLAM: 4D Dynamic Scene Reconstruction Based on Geometrical Motion Segmentation

RA-L 2021

Most SLAM (Simultaneous Localization and Mapping) algorithms are based on the assumption that the scene is static. However, in practice, most real scenes usually contain moving objects. In this letter, we introduce DymSLAM, a dynamic stereo visual SLAM system being capable of reconstructing a 4D (3D

Cited by 50SourceScholar
2020

Three-Dimensional Posture Optimization for Biped Robot Stepping over Large Ditch Based on a Ducted-Fan Propulsion System

IROS 2020poster

The recent progress of an ongoing project utilizing a ducted-fan propulsion system to improve a humanoid robot's ability to step over large ditches is reported. A novel method (GAS) based on the genetic algorithm with smoothness constraint can effectively minimize the thrust by optimizing the robot'…

Cited by 9SourceScholar
2019

Interactive Subjective Study on Picture-level Just Noticeable Difference of Compressed Stereoscopic Images

ICASSP 2019accepted

The Just Noticeable Difference (JND) reveals the minimum distortion that the Human Visual System (HVS) can perceive. Traditional studies on JND mainly focus on background luminance adaptation and contrast masking. However, the HVS does not perceive visual content based on individual pixels or blocks…

Cited by 0SourceScholar
2018

Connecting Gaze, Scene, and Attention: Generalized Attention Estimation via Joint Modeling of Gaze and Scene Saliency

ECCV 2018poster

This paper addresses the challenging problem of estimating the general visual attention of people in images. Our proposed method is designed to work across multiple naturalistic social scenarios and provides a full picture of the subject’s attention and gaze. In contrast, earlier works on gaze and a…

2018

Jet-HR1: Stepping Posture Optimization for Bipedal Robot Over Large Ditch Based on a Ducted-fan Propulsion System

IROS 2018poster

This paper reports the latest progress of an ongoing project utilizing a ducted-fan propulsion system to improve a humanoid robot's ability to step over a broad ditch with a height difference between the two sides. This work focuses on the methods of calculating the boundary and optimizing stepping…

Cited by 9SourceScholar
2015

Multi-task rank learning for image quality assessment

ICASSP 2015accepted

In practice, multiple types of distortions are associated with an image quality degradation process. The existing machine learning (ML) based image quality assessment (IQA) approaches generally established a unified model for all distortion types, or each model is trained independently for each dist…

Cited by 0SourceScholar