← Search

Jiawei Yang

22 accepted papers

2026

SEAP: Sparse Expert Activation Pruning Unlocks the Brainpower of Large Language Models

AAAI 2026technical

Pruning is a promising approach to reduce the high inference cost of large language models (LLMs), but it often comes at the expense of performance. Motivated by the "functional localization" theory in neuroscience, we hypothesize that LLMs contain task-specific expert activation paths, where specif

Cited by 0SourcePDFScholar
2025

InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models

ICCV 2025poster

We present InfiniCube, a scalable and controllable method to generate unbounded and dynamic 3D driving scenes with high fidelity.Previous methods for scene generation are constrained either by their applicability to indoor scenes or by their lack of controllability.In contrast, we take advantage of…

Cited by 0SourcePDFScholar
2025

OmniRe: Omni Urban Scene Reconstruction

ICLR 2025spotlight

We introduce OmniRe, a comprehensive system for efficiently creating high-fidelity digital twins of dynamic real-world scenes from on-device logs. Recent methods using neural fields or Gaussian Splatting primarily focus on vehicles, hindering a holistic framework for all dynamic foregrounds demanded…

2025

Retrieval-Augmented Multilingual Citation Generation

ICASSP 2025accepted

Retrieval-augmented citation generation (RACG) helps users trust the large language model output by retrieving evidence from reliable sources. However, most current RACG research focuses on single-language tasks, particularly in English, and overlooks the need for cross-lingual evidence retrieval an…

Cited by 0SourceScholar
2025

STORM: Spatio-TempOral Reconstruction Model For Large-Scale Outdoor Scenes

ICLR 2025poster

We present STORM, a spatio-temporal reconstruction model designed for reconstructing dynamic outdoor scenes from sparse observations. Existing dynamic reconstruction methods often rely on per-scene optimization, dense observations across space and time, and strong motion supervision, resulting in le…

2025

SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model

ACL 2025long

The indexing-retrieval-generation paradigm of retrieval-augmented generation (RAG) has been highly successful in solving knowledge-intensive tasks by integrating external knowledge into large language models (LLMs). However, the incorporation of external and unverified knowledge increases the vulner…

2025

SampleMix: A Sample-wise Pre-training Data Mixing Strategy by Coordinating Data Quality and Diversity

EMNLP 2025

Existing pretraining data mixing methods for large language models (LLMs) typically follow a domain-wise methodology, a top-down process that first determines domain weights and then performs uniform data sampling across each domain. However, these approaches neglect significant inter-domain overlap

Cited by 0SourcePDFScholar
2025

When Sparse Graph Representation Learning Falls into Domain Shift: Feature Augmentation for Cross-Domain Graph Meta-Learning

ICASSP 2025accepted

Graph Meta-learning methods have improved the performance of few-shot node classification by means of applying meta-learning to the data in non-Euclidean domains. However, most works focus on adopting a single domain, ignoring the fact that tasks in various domains may be distinct, which can cause o…

Cited by 0SourceScholar
2024

DistillNeRF: Perceiving 3D Scenes from Single-Glance Images by Distilling Neural Fields and Foundation Model Features

NeurIPS 2024poster

We propose DistillNeRF, a self-supervised learning framework addressing the challenge of understanding 3D environments from limited 2D observations in outdoor autonomous driving scenes. Our method is a generalizable feedforward model that predicts a rich neural scene representation from sparse, sing…

2024

EmerNeRF: Emergent Spatial-Temporal Scene Decomposition via Self-Supervision

ICLR 2024poster

We present EmerNeRF, a simple yet powerful approach for learning spatial-temporal representations of dynamic driving scenes. Grounded in neural fields, EmerNeRF simultaneously captures scene geometry, appearance, motion, and semantics via self-bootstrapping. EmerNeRF hinges upon two core components:…

2024

Parallelized Spatiotemporal Slot Binding for Videos

ICML 2024poster

While modern best practices advocate for scalable architectures that support long-range interactions, object-centric models are yet to fully embrace these architectures. In particular, existing object-centric models for handling sequential inputs, due to their reliance on RNN-based implementation, s…

Cited by 0SourcePDFScholar
2024

PreSight: Enhancing Autonomous Vehicle Perception with City-Scale NeRF Priors

ECCV 2024poster

"Autonomous vehicles rely extensively on perception systems to navigate and interpret their surroundings. Despite significant advancements in these systems recently, challenges persist under conditions like occlusion, extreme lighting, or in unfamiliar urban areas. Unlike these systems, humans do no…

2023

FreeNeRF: Improving Few-Shot Neural Rendering With Free Frequency Regularization

CVPR 2023poster

Novel view synthesis with sparse inputs is a challenging problem for neural radiance fields (NeRF). Recent efforts alleviate this challenge by introducing external supervision, such as pre-trained models and extra depth signals, or by using non-trivial patch-based rendering. In this paper, we presen…

2022

ConCL: Concept Contrastive Learning for Dense Prediction Pre-training in Pathology Images

ECCV 2022poster

"Detecting and segmenting objects within whole slide images is essential in computational pathology workflow. Self-supervised learning (SSL) is appealing to such annotation-heavy tasks. Despite the extensive benchmarks in natural images for dense tasks, such studies are, unfortunately, absent in cur…

2022

Research on Target Tracking for Robotic Fish Based on Low-Cost Scarce Sensing Information Fusion

RA-L 2022

Target tracking for underwater robots is always challenging, due to low-quality sensing information, information interference, and environmental disturbances. Traditional sensing methods for underwater target tracking include vision-based tracking, acoustic-based tracking, etc., but most of the adop

Cited by 21SourceScholar
2022

Towards Better Understanding and Better Generalization of Low-shot Classification in Histology Images with Contrastive Learning

ICLR 2022poster

Few-shot learning is an established topic in natural images for years, but few work is attended to histology images, which is of high clinical value since well-labeled datasets and rare abnormal samples are expensive to collect. Here, we facilitate the study of few-shot learning in histology images…

Cited by 34SourcePDFScholar
2022

TreeMoCo: Contrastive Neuron Morphology Representation Learning

NeurIPS 2022accept

Morphology of neuron trees is a key indicator to delineate neuronal cell-types, analyze brain development process, and evaluate pathological changes in neurological diseases. Traditional analysis mostly relies on heuristic features and visual inspections. A quantitative, informative, and comprehensi…

2021

A Versatile Pneumatic Actuator Based on Scissor Mechanisms: Design, Modeling, and Experiments

RA-L 2021

A new versatile pneumatic actuator called a scissor-mechanism-based vacuum powered artificial muscle (SMVAM) is proposed in this letter. The actuator is composed of a soft airtight skin and a scissor structural skeleton, and based on the types of scissor units in the skeleton, the actuator can be de

Cited by 27SourceScholar
2021

Oral-3D: Reconstructing the 3D Structure of Oral Cavity from Panoramic X-ray

AAAI 2021technical

Panoramic X-ray (PX) provides a 2D picture of the patient's mouth in a panoramic view to help dentists observe the invisible disease inside the gum. However, it provides limited 2D information compared with cone-beam computed tomography (CBCT), another dental imaging method that generates a 3D pictu…

Cited by 41SourcePDFScholar