← Search

Yang Jiao

23 accepted papers

2026

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning

CVPR 2026

Embodied Visual Reasoning (EVR) seeks to follow complex, free-form instructions based on egocentric video, enabling semantic understanding and spatiotemporal reasoning in dynamic environments. Despite its promising potential, EVR encounters significant challenges stemming from the diversity of compl

Cited by 0SourcecodeScholar
2026

Identity-Aware Vision-Language Model for Explainable Face Forgery Detection

AAAI 2026technical

Recent advances in generative artificial intelligence have enabled the creation of highly realistic image forgeries, raising significant concerns about digital media authenticity. While existing detection methods demonstrate promising results on benchmark datasets, they face critical limitations in

Cited by 0SourcePDFScholar
2026

Learning by Analogy: A Causal Framework for Compositional Generalization

CVPR 2026

Compositional generalization -- the ability to understand and generate novel combinations of learned concepts -- enables models to extend their capabilities beyond limited experiences. While effective, the data structures and principles that enable this crucial capability remain poorly understood. W

Cited by 0SourceScholar
2026

Pix2Key: Controllable Open-Vocabulary Retrieval with Semantic Decomposition and Self-Supervised Visual Dictionary Learning

ICML 2026poster

Composed image retrieval uses a reference image plus a natural-language edit to retrieve images that apply the requested change while preserving other relevant visual content. Classic fusion pipelines typically rely on supervised triplets and can lose fine-grained cues, while recent zero-shot approa…

Cited by 0SourceScholar
2026

SIGMA-PPG: Statistical-prior Informed Generative Masking Architecture for PPG Foundation Model

ICML 2026poster

Current foundation model for photoplethysmography (PPG) signals is challenged by the intrinsic redundancy and noise of the signal. Standard masked modeling often yields trivial solutions while contrastive methods lack morphological precision. To address these limitations, we propose a Statistical-pr…

Cited by 0SourceScholar
2025

A Magnetically-Actuated Ultrasound Capsule Endoscope (MUSCE) for Endoluminal Imaging in Tubular Environments

RA-L 2025

Endoscopic ultrasound (EUS) has the ability to image tissue in and beyond the wall of the gastrointestinal (GI) tract, assisting in the early diagnosis of digestive diseases. However, traditional EUS based on flexible endoscopes could make the operation procedure traumatic and intolerable to patient

Cited by 9SourceScholar
2025

ATLAS: Autoformalizing Theorems through Lifting, Augmentation, and Synthesis of Data

NeurIPS 2025poster

Autoformalization, the automatic translation of mathematical content from natural language into machine-verifiable formal languages, has seen significant progress driven by advances in large language models (LLMs). Nonetheless, a primary barrier to further improvements is the limited availability of…

Cited by 0SourcecodeScholar
2025

DTZO: Distributed Trilevel Zeroth Order Learning with Provable Non-Asymptotic Convergence

ICML 2025poster

Trilevel learning (TLL) with zeroth order constraints is a fundamental problem in machine learning, arising in scenarios where gradient information is inaccessible due to data privacy or model opacity, such as in federated learning, healthcare, and financial systems. These problems are notoriously d…

Cited by 0SourcePDFScholar
2025

Divide and Orthogonalize: Efficient Continual Learning with Local Model Space Projection

UAI 2025

Continual learning (CL) has gained increasing interest in recent years due to the need for models that can continuously learn new tasks while retaining knowledge from previous ones. However, existing CL methods often require either computationally expensive layer-wise gradient projections or large-s

Cited by 0SourcePDFScholar
2025

Simultaneous 6-DOF localization and scanning angle detection of magnetic ultrasound capsule endoscope (MUSCE) with internal sensors

IROS 2025

Localization of magnetically actuated capsule endoscope (MCE) is essential for accurate actuation. Despite extensive progress in pose estimation using internal magnetic field sensors and external magnetic sources, it remains challenging to achieve localization when a time-varying internal magnetic f

Cited by 0SourceScholar
2024

Instance-Aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting Learning

AAAI 2024technical

Camera-based bird-eye-view (BEV) perception paradigm has made significant progress in the autonomous driving field. Under such a paradigm, accurate BEV representation construction relies on reliable depth estimation for multi-camera images. However, existing approaches exhaustively predict depths fo…

2024

Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models

NeurIPS 2024poster

Large Multimodal Model (LMM) is a hot research topic in the computer vision area and has also demonstrated remarkable potential across multiple disciplinary fields. A recent trend is to further extend and enhance the perception capabilities of LMMs. The current methods follow the paradigm of adaptin…

2024

NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving Scenario

AAAI 2024technical

We introduce a novel visual question answering (VQA) task in the context of autonomous driving, aiming to answer natural language questions based on street-view clues. Compared to traditional VQA tasks, VQA in autonomous driving scenario presents more challenges. Firstly, the raw visual data are mul…

2024

Provably Convergent Federated Trilevel Learning

AAAI 2024technical

Trilevel learning, also called trilevel optimization (TLO), has been recognized as a powerful modelling tool for hierarchical decision process and widely applied in many machine learning applications, such as robust neural architecture search, hyperparameter optimization, and domain adaptation. Tack…

Cited by 6SourcePDFScholar
2024

Tri-Level Navigator: LLM-Empowered Tri-Level Learning for Time Series OOD Generalization

NeurIPS 2024poster

Out-of-Distribution (OOD) generalization in machine learning is a burgeoning area of study. Its primary goal is to enhance the adaptability and resilience of machine learning models when faced with new, unseen, and potentially adversarial data that significantly diverges from their original training…

Cited by 4SourcePDFScholar
2023

Asynchronous Distributed Bilevel Optimization

ICLR 2023poster

Bilevel optimization plays an essential role in many machine learning tasks, ranging from hyperparameter optimization to meta-learning. Existing studies on bilevel optimization, however, focus on either centralized or synchronous distributed setting. The centralized bilevel optimization approaches r…

2023

Learning Attribute and Class-Specific Representation Duet for Fine-Grained Fashion Analysis

CVPR 2023poster

Fashion representation learning involves the analysis and understanding of various visual elements at different granularities and the interactions among them. Existing works often learn fine-grained fashion representations at the attribute-level without considering their relationships and inter-depe…

Cited by 12SourcePDFScholar
2023

MSMDFusion: Fusing LiDAR and Camera at Multiple Scales With Multi-Depth Seeds for 3D Object Detection

CVPR 2023poster

Fusing LiDAR and camera information is essential for accurate and reliable 3D object detection in autonomous driving systems. This is challenging due to the difficulty of combining multi-granularity geometric and semantic features from two drastically different modalities. Recent approaches aim at e…

2022

Fine-Grained Fashion Representation Learning by Online Deep Clustering

ECCV 2022poster

"Fashion designs are rich in visual details associated with various visual attributes at both global and local levels. As a result, effective modeling and analyzing fashion requires fine-grained representations for individual attributes. In this work, we present a deep learning based online clusteri…

Cited by 19SourcePDFScholar
2022

MORE: Multi-Order RElation Mining for Dense Captioning in 3D Scenes

ECCV 2022poster

"3D dense captioning is a recently-proposed novel task, where point clouds contain more geometric information than the 2D counterpart. However, it is also more challenging due to the higher complexity and wider variety of inter-object relations contained in point clouds. Existing methods only treat…

2021

EffiScene: Efficient Per-Pixel Rigidity Inference for Unsupervised Joint Learning of Optical Flow, Depth, Camera Pose and Motion Segmentation

CVPR 2021poster

This paper addresses the challenging unsupervised scene flow estimation problem by jointly learning four low-level vision sub-tasks: optical flow F, stereo-depth D, camera pose P and motion segmentation S. Our key insight is that the rigidity of the scene shares the same inherent geometrical structu…

Cited by 47PDFScholar
2020

Automated High-Productivity Microinjection System for Adherent Cells

RA-L 2020

Automated microinjection systems for suspension cells have been studied for years. Nevertheless, microinjection systems for adherent cells still suffer from laborious manual operations and low productivity. This paper presents a new automated microinjection system with high productivity for adherent

Cited by 27SourceScholar