← Search

Yanan Wang

8 accepted papers

2026

LILAC: Language-Conditioned Object-Centric Optical Flow for Open-Loop Trajectory Generation

RA-L 2026

We address language-conditioned robotic manipulation using flow-based trajectory generation, which enables training on human and web videos of object manipulation and requires only minimal embodiment-specific data. This task is challenging, as object trajectory generation from pre-manipulation image

Cited by 3SourceScholar
2026

MultiKD: Backdoor Defense in Federated Graph Learning via Attention-Guided Multi-Teacher Distillation

AAAI 2026technical

Backdoor attacks pose a severe threat to federated graph learning (FGL), where malicious clients can inject hidden triggers into the global model without being detected. Defending against such attacks is particularly challenging due to the complex graph structures and the stealthy nature of trigger

Cited by 0SourcePDFScholar
2026

RCPU: Rotation-Constrained Error Compensation for Structured Pruning of a Large Language Model

ICLR 2026poster

In this paper, we propose a rotation-constrained compensation method to address the errors introduced by structured pruning of large language models (LLMs). LLMs are trained on massive datasets and accumulate rich semantic knowledge in their representation space. In contrast, pruning is typically…

Cited by 0SourcecodeScholar
2026

Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Artificial Intelligence

ICML 2026poster

The impressive performance of generalist large language models (LLMs) such as GPT-4 and Claude in healthcare raises a critical question: will domain-specific medical specialist models become obsolete? We argue that the future of medical artificial intelligence (AI) lies not in building monolithic me…

Cited by 0SourceScholar
2023

VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question Answering

ICCV 2023poster

Visual question answering (VQA) requires systems to perform concept-level reasoning by unifying unstructured (e.g., the context in question and answer; "QA context") and structured (e.g., knowledge graph for the QA context and scene; "concept graph") multimodal knowledge. Existing works typically co…

Cited by 42PDFScholar
2022

Cloning Outfits From Real-World Images to 3D Characters for Generalizable Person Re-Identification

CVPR 2022poster

Recently, large-scale synthetic datasets are shown to be very useful for generalizable person re-identification. However, synthesized persons in existing datasets are mostly cartoon-like and in random dress collocation, which limits their performance. To address this, in this work, an automatic appr…

Cited by 36PDFcodeScholar
2022

DRG-SLAM: A Semantic RGB-D SLAM using Geometric Features for Indoor Dynamic Scene

IROS 2022poster

Visual SLAM methods based on point features have achieved acceptable results in texture-rich static scenes, but they often suffer from a deficiency of texture and the existence of dynamic objects in real indoor scenes, which limits the application of these methods. In this paper, we have presented D…

Cited by 20SourceScholar