← Search

Jin Gao

29 accepted papers

2026

Credit-Budgeted ICPC-Style Coding: When LLM Agents Must Pay for Every Decision

ICLR 2026poster

Contemporary coding-agent benchmarks applaud “first correct answer,” silently assuming infinite tokens, container minutes, and developer patience. In production, every LLM call, test re-run, and rollback incurs hard cost; agents that cannot budget these resources are dead on arrival. We close the ga…

Cited by 0SourcecodeScholar
2026

HDGS: Hierarchical Dynamic Gaussian Splatting for Urban Driving Scenes

AAAI 2026technical

This paper tackles the challenging task of achieving storage-efficient yet high-fidelity motion representation in large-scale dynamic 3D Gaussian Splatting. Our motivation stems from the truth that existing urban-scale methods, which rely on massive and unstructured individual Gaussians for scene mo

Cited by 0SourcePDFScholar
2026

Integrating Diverse Assignment Strategies into DETRs

AAAI 2026technical

Label assignment is a critical component in object detectors, particularly within DETR-style frameworks where the one-to-one matching strategy, despite its end-to-end elegance, suffers from slow convergence due to sparse supervision. While recent works have explored one-to-many assignments to enrich

Cited by 0SourcePDFScholar
2026

TadABench-1M: A Large-Scale Wet-Lab Protein Benchmark For Rigorous OOD Evaluation

ICML 2026poster

Existing benchmarks for biological language models (BLMs) inadequately capture the challenges of real-world applications, often lacking realistic out-of-distribution (OOD) scenarios, evolutionary depth, and consistency in measurement. To address this, we introduce TadABench-1M, a new benchmark based…

Cited by 0SourceScholar
2025

DeepTAGE: Deep Temporal-Aligned Gradient Enhancement for Optimizing Spiking Neural Networks

ICLR 2025poster

Spiking Neural Networks (SNNs), with their biologically inspired spatio-temporal dynamics and spike-driven processing, are emerging as a promising low-power alternative to traditional Artificial Neural Networks (ANNs). However, the complex neuronal dynamics and non-differentiable spike communication…

Cited by 0SourcePDFScholar
2025

Each Complexity Deserves a Pruning Policy

NeurIPS 2025poster

The established redundancy in visual tokens within large vision–language models (LVLMs) allows for pruning to effectively reduce their substantial computational demands. Empirical evidence from previous works indicates that visual tokens in later decoder stages receive less attention than shallow la…

Cited by 0SourcecodeScholar
2025

Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection

ICCV 2025poster

We present the first work demonstrating that a pure Mamba block can achieve efficient Dense Global Fusion, meanwhile guaranteeing top performance for camera-LiDAR multi-modal 3D object detection. Our motivation stems from the observation that existing fusion strategies are constrained by their inabi…

2025

Online Segment Any 3D Thing as Instance Tracking

NeurIPS 2025poster

Online, real-time, and fine-grained 3D segmentation constitutes a fundamental capability for embodied intelligent agents to perceive and comprehend their operational environments. Recent advancements employ predefined object queries to aggregate semantic information from Vision Foundation Models (VF…

Cited by 0SourcecodeScholar
2025

SSTrack: Sample-interval Scheduling for Lightweight Visual Object Tracking

IJCAI 2025

In recent years, CPU real-time object tracking has gained significant attention due to its broad applications such as UAV-tracking. To maintain computational efficiency, most existing CPU real-time object trackers rely on lightweight backbones and employ a single initial template image without inter

2025

SynCL: A Synergistic Training Strategy with Instance-Aware Contrastive Learning for End-to-End Multi-Camera 3D Tracking

NeurIPS 2025poster

While existing query-based 3D end-to-end visual trackers integrate detection and tracking via the *tracking-by-attention* paradigm, these two chicken-and-egg tasks encounter optimization difficulties when sharing the same parameters. Our findings reveal that these difficulties arise due to two inher…

Cited by 0SourcecodeScholar
2025

Towards More Discriminative Feature Learning in SNNs with Temporal-Self-Erasing Supervision

AAAI 2025technical

Spiking Neural Networks (SNNs) are biologically inspired models that process visual inputs over multiple time steps. However, they often struggle with limited feature discrimination along the temporal dimension due to inherent spatiotemporal invariance. This limitation arises from the redundant acti…

Cited by 0SourcePDFScholar
2024

A-Teacher: Asymmetric Network for 3D Semi-Supervised Object Detection

CVPR 2024poster

This work proposes the first online asymmetric semi-supervised framework namely A-Teacher for LiDAR-based 3D object detection. Our motivation stems from the observation that 1) existing symmetric teacher-student methods for semi-supervised 3D object detection have characterized simplicity but impede…

Cited by 2SourcePDFScholar
2024

Animate3D: Animating Any 3D Model with Multi-view Video Diffusion

NeurIPS 2024poster

Recent advances in 4D generation mainly focus on generating 4D content by distilling pre-trained text or single-view image conditioned models. It is inconvenient for them to take advantage of various off-the-shelf 3D assets with multi-view attributes, and their results suffer from spatiotemporal inc…

Cited by 13SourcePDFScholar
2024

BEV2PR: BEV-Enhanced Visual Place Recognition with Structural Cues

IROS 2024

In this paper, we propose a new image-based visual place recognition (VPR) framework by exploiting the structural cues in bird’s-eye view (BEV) from a single monocular camera. The motivation arises from two key observations about place recognition methods based on both appearance and structure: 1) F

Cited by 4SourcecodeScholar
2024

Consistent4D: Consistent 360° Dynamic Object Generation from Monocular Video

ICLR 2024poster

In this paper, we present Consistent4D, a novel approach for generating 4D dynamic objects from uncalibrated monocular videos. Uniquely, we cast the 360-degree dynamic object reconstruction as a 4D generation problem, eliminating the need for tedious multi-view data collection and camera calibration…

2024

Dissecting Dissonance: Benchmarking Large Multimodal Models Against Self-Contradictory Instructions

ECCV 2024poster

"Large multimodal models (LMMs) excel in adhering to human instructions. However, self-contradictory instructions may arise due to the increasing trend of multimodal interaction and context length, which is challenging for language beginners and vulnerable populations. We introduce the Self-Contradi…

2024

VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization

NeurIPS 2024poster

Bird's-eye-view (BEV) map layout estimation requires an accurate and full understanding of the semantics for the environmental elements around the ego car to make the results coherent and realistic. Due to the challenges posed by occlusion, unfavourable imaging conditions and low resolution, \emph{g…

2023

A Closer Look at Self-Supervised Lightweight Vision Transformers

ICML 2023poster

Self-supervised learning on large-scale Vision Transformers (ViTs) as pre-training methods has achieved promising downstream performance. Yet, how much these pre-training paradigms promote lightweight ViTs' performance is considerably less studied. In this work, we develop and benchmark several self…

2023

Back to the Source: Diffusion-Driven Adaptation To Test-Time Corruption

CVPR 2023poster

Test-time adaptation harnesses test inputs to improve the accuracy of a model trained on source data when tested on shifted target data. Most methods update the source model by (re-)training on each target domain. While re-training can help, it is sensitive to the amount and order of the data and th…

Cited by 121SourcePDFScholar
2023

Multi-Correlation Siamese Transformer Network With Dense Connection for 3D Single Object Tracking

RA-L 2023

Point cloud-based 3D object tracking is an important task in autonomous driving. Though great advances regarding Siamese-based 3D tracking have been made recently, it remains challenging to learn the correlation between the template and search branches effectively with the sparse LIDAR point cloud d

Cited by 10SourcecodeScholar
2023

PolarFormer: Multi-Camera 3D Object Detection with Polar Transformer

AAAI 2023technical

3D object detection in autonomous driving aims to reason “what” and “where” the objects of interest present in a 3D world. Following the conventional wisdom of previous 2D object detection, existing methods often adopt the canonical Cartesian coordinate system with perpendicular axis. However, we co…

2023

ZoomTrack: Target-aware Non-uniform Resizing for Efficient Visual Tracking

NeurIPS 2023spotlight

Recently, the transformer has enabled the speed-oriented trackers to approach state-of-the-art (SOTA) performance with high-speed thanks to the smaller input size or the lighter feature extraction backbone, though they still substantially lag behind their corresponding performance-oriented versions.…

2022

Open-Vocabulary One-Stage Detection With Hierarchical Visual-Language Knowledge Distillation

CVPR 2022poster

Open-vocabulary object detection aims to detect novel object categories beyond the training set. The advanced open-vocabulary two-stage detectors employ instance-level visual-to-visual knowledge distillation to align the visual space of the detector with the semantic space of the Pre-trained Visual-…

Cited by 106PDFcodeScholar
2021

Graspness Discovery in Clutters for Fast and Accurate Grasp Detection

ICCV 2021poster

Efficient and robust grasp pose detection is vital for robotic manipulation. For general 6 DoF grasping, conventional methods treat all points in a scene equally and usually adopt uniform sampling to select grasp candidates. However, we discover that ignoring where to grasp greatly harms the speed a…

Cited by 120PDFcodeScholar
2018

Learning Attentions: Residual Attentional Siamese Network for High Performance Online Visual Tracking

CVPR 2018poster

Offline training for object tracking has recently shown great potentials in balancing tracking accuracy and speed. However, it is still difficult to adapt an offline trained model to a target tracked online. This work presents a Residual Attentional Siamese Network (RASNet) for high performance obje…

2018

Visual Tracking via Spatially Aligned Correlation Filters Network

ECCV 2018poster

Correlation filters based trackers rely on a periodic assumption of the search sample to efficiently distinguish the target from the background. This assumption however yields undesired boundary effects and restricts aspect ratios of search samples. To handle these issues, an end-to-end deep archite…