← Search

Yuewei Lin

13 accepted papers

2026

FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics

ICLR 2026poster

Large language models have revolutionized artificial intelligence by enabling large, generalizable models trained through self-supervision. This paradigm has inspired the development of scientific foundation models (FMs). However, applying this capability to experimental particle physics is challeng…

Cited by 0SourceScholar
2025

Efficient and Accurate Low-Resolution Transformer Tracking

IROS 2025

High-performance Transformer trackers have exhibited excellent results, yet they often bear a heavy computational load. Observing that a smaller input can immediately and conveniently reduce computations without changing the model, an easy solution is to adopt a low-resolution input for efficient Tr

Cited by 0SourcecodeScholar
2025

Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding

ICLR 2025oral

Transformer has attracted increasing interest in spatio-temporal video grounding, or STVG, owing to its end-to-end pipeline and promising result. Existing Transformer-based STVG approaches often leverage a set of object queries, which are initialized simply using zeros and then gradually learn targe…

2025

VAPO: Visibility-Aware Keypoint Localization for Efficient 6DoF Object Pose Estimation

IROS 2025

Localizing predefined 3D keypoints in a 2D image is an effective way to establish 3D-2D correspondences for instance-level 6DoF object pose estimation. However, unreliable localization results of invisible keypoints degrade the quality of correspondences. In this paper, we address this issue by loca

Cited by 2SourcecodeScholar
2024

AesFA: An Aesthetic Feature-Aware Arbitrary Neural Style Transfer

AAAI 2024technical

Neural style transfer (NST) has evolved significantly in recent years. Yet, despite its rapid progress and advancement, existing NST methods either struggle to transfer aesthetic information from a style effectively or suffer from high computational costs and inefficiencies in feature disentanglemen…

2024

CLIPCEIL: Domain Generalization through CLIP via Channel rEfinement and Image-text aLignment

NeurIPS 2024poster

Domain generalization (DG) is a fundamental yet challenging topic in machine learning. Recently, the remarkable zero-shot capabilities of the large pre-trained vision-language model (e.g., CLIP) have made it popular for various downstream tasks. However, the effectiveness of this capacity often degr…

2024

Efficient Temporal Action Segmentation via Boundary-aware Query Voting

NeurIPS 2024poster

Although the performance of Temporal Action Segmentation (TAS) has been improved in recent years, achieving promising results often comes with a high computational cost due to dense inputs, complex model structures, and resource-intensive post-processing requirements. To improve the efficiency while…

2021

AGKD-BML: Defense Against Adversarial Attack by Attention Guided Knowledge Distillation and Bi-Directional Metric Learning

ICCV 2021poster

While deep neural networks have shown impressive performance in many tasks, they are fragile to carefully designed adversarial attacks. We propose a novel adversarial training-based model by Attention Guided Knowledge Distillation and Bi-directional Metric Learning (AGKD-BML). The attention knowledg…

Cited by 22PDFcodeScholar
2021

Transparent Object Tracking Benchmark

ICCV 2021poster

Visual tracking has achieved considerable progress in recent years. However, current research in the field mainly focuses on tracking of opaque objects, while little attention is paid to transparent object tracking. In this paper, we make the first attempt in exploring this problem by proposing a Tr…

Cited by 33PDFcodeScholar
2017

Learning View-Invariant Features for Person Identification in Temporally Synchronized Videos Taken by Wearable Cameras

ICCV 2017poster

In this paper, we study the problem of Cross-View Person Identification (CVPI), which aims at identifying the same person from temporally synchronized videos taken by different wearable cameras. Our basic idea is to utilize the human motion consistency for CVPI, where human motion can be computed by…

Cited by 28PDFScholar
2016

Groupwise Tracking of Crowded Similar-Appearance Targets From Low-Continuity Image Sequences

CVPR 2016spotlight

Automatic tracking of large-scale crowded targets are of particular importance in many applications, such as crowded people/vehicle tracking in video surveillance, fiber tracking in materials science, and cell tracking in biomedical imaging. This problem becomes very challenging when the targets sho…

Cited by 37PDFScholar
2015

Co-Interest Person Detection From Multiple Wearable Camera Videos

ICCV 2015poster

Wearable cameras, such as Google Glass and Go Pro, enable video data collection over larger areas and from different views. In this paper, we tackle a new problem of locating the co-interest person (CIP), i.e., the one who draws attention from most camera wearers, from temporally synchronized videos…

Cited by 29PDFScholar
2015

Combining Local Appearance and Holistic View: Dual-Source Deep Neural Networks for Human Pose Estimation

CVPR 2015poster

We propose a new learning-based method for estimating 2D human pose from a single image, using Dual-Source Deep Convolutional Neural Networks (DS-CNN). Recently, many methods have been developed to estimate human pose by using pose priors that are estimated from physiologically inspired graphical mo…

Cited by 299SourcePDFScholar