← Search

Tong Nie

5 accepted papers

2026

E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving

CVPR 2026

End-to-end autonomous driving (AD) systems increasingly adopt vision-language-action (VLA) models, yet they ignore the passenger's emotional state, which is central to comfort and AD acceptance. We introduce Open-Domain End-to-End (OD-E2E) AD, where an autonomous vehicle must interpret free-form nat

Cited by 0SourceScholar
2026

Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization

ICML 2026poster

Efficient Distillation (EDistill) compresses large language models (LLMs) by structured pruning parameters and tuning lightweight modules with high training efficiency. Although these EDistilled LLMs achieve state-of-the-art (SOTA) performance on general ability benchmarks relative to similarly size…

Cited by 0SourceScholar
2026

Steerable Adversarial Scenario Generation through Test-Time Preference Alignment

ICLR 2026poster

Adversarial scenario generation is a cost-effective approach for safety assessment of autonomous driving systems. However, existing methods are often constrained to a single, fixed trade-off between competing objectives such as adversariality and realism. This yields behavior-specific models that c…

Cited by 0SourcecodeScholar
2025

Geolocation Representation from Large Language Models Are Generic Enhancers for Spatio-Temporal Learning

AAAI 2025technical

In the geospatial domain, universal representation models are significantly less prevalent than their extensive use in natural language processing and computer vision. This discrepancy arises primarily from the high costs associated with the input of existing representation models, which often requi…

Cited by 7SourcePDFScholar
2023

Category-Level 6D Pose Estimation Using Geometry-Guided Instance-Aware Prior and Multi-Stage Reconstruction

RA-L 2023

Category-level object 6D pose estimation is essential for robotic manipulation, augmented reality and 3D scene understanding. It aims to accurately predict the translation and rotation of arbitrary shape instances from a given set of object classes without models of each instance. However, such esti

Cited by 5SourceScholar