← Search

Siyu Zhang

8 accepted papers

2026

Decoupling Continual Semantic Segmentation

AAAI 2026technical

Continual Semantic Segmentation (CSS) requires learning new classes without forgetting previously acquired knowledge, addressing the fundamental challenge of catastrophic forgetting in dense prediction tasks. However, existing CSS methods typically employ single-stage encoder-decoder architectures w

Cited by 0SourcePDFScholar
2026

MOBO: A Merging-Oriented Bi-Level Optimization Framework for Class Incremental Learning

IJCAI 2026

Class-Incremental Learning (CIL) aims to enable models to sequentially learn new tasks while retaining knowledge from previous ones. Recently, merging-based pre-trained CIL methods have gained significant attention due to their competitive performance and high inference efficiency. However, most exi

Cited by 0Scholar
2026

Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought

ICML 2026poster

Multimodal models for text-to-image generation have achieved strong visual fidelity, yet they remain brittle under compositional structural constraints—notably generative numeracy, attribute binding, and part-level relations. To address these challenges, we propose **Shape-of-Thought (SoT)**, a visu…

Cited by 0SourceScholar
2025

A Multi-Modal Fusion-Based 3D Multi-Object Tracking Framework With Joint Detection

RA-L 2025

In the classical tracking-by-detection (TBD) paradigm, detection and tracking are separately and sequentially conducted, and data association must be properly performed to achieve satisfactory tracking performance. In this letter, a new multi-object tracking framework is proposed, which integrates o

Cited by 18SourceScholar
2025

MCTrack: A Unified 3D Multi-Object Tracking Framework for Autonomous Driving

IROS 2025

This paper introduces MCTrack, a new 3D multi-object tracking method that achieves performance across KITTI, nuScenes, and Waymo datasets. Addressing the gap in existing tracking paradigms, which often perform well on specific datasets but lack generalizability, MCTrack offers a unified solution. Ad

Cited by 27SourcecodeScholar
2022

OnePose: One-Shot Object Pose Estimation Without CAD Models

CVPR 2022poster

We propose a new method named OnePose for object pose estimation. Unlike existing instance-level or category-level methods, OnePose does not rely on CAD models and can handle objects in arbitrary categories without instance- or category-specific network training. OnePose draws the idea from visual l…

Cited by 174PDFcodeScholar
2021

You Don't Only Look Once: Constructing Spatial-Temporal Memory for Integrated 3D Object Detection and Tracking

ICCV 2021poster

Humans are able to continuously detect and track surrounding objects by constructing a spatial-temporal memory of the objects when looking around. In contrast, 3D object detectors in existing tracking-by-detection systems often search for objects in every new video frame from scratch, without fully…

Cited by 13PDFcodeScholar
2020

Disp R-CNN: Stereo 3D Object Detection via Shape Prior Guided Instance Disparity Estimation

CVPR 2020poster

In this paper, we propose a novel system named Disp R-CNN for 3D object detection from stereo images. Many recent works solve this problem by first recovering a point cloud with disparity estimation and then apply a 3D detector. The disparity map is computed for the entire image, which is costly and…

Cited by 145PDFcodeScholar