← Search

Chu-Song Chen

20 accepted papers

2026

Submodular Optimization for Minimal Augmentation in Robust Language Model Alignment

ICML 2026poster

Safety alignment of large language models is fragile: even small fine-tuning perturbations elastically revert behaviors toward those of the pre-training, with degradation inversely proportional to the size of the alignment set. We ask how to achieve safety alignment with \emph{minimal augmentation}.…

Cited by 0SourceScholar
2025

PDSeg: Patch-Wise Distillation and Controllable Image Generation for Weakly-Supervised Histopathology Tissue Segmentation

ICASSP 2025accepted

Weakly-supervised semantic segmentation, which achieves pixel-wise segmentation using image-level labels, has emerged as an alternative to fully supervised methods by reducing the need for detailed annotations. Inspired by the recent success of the teacher-student strategy in various vision tasks, w…

Cited by 0SourceScholar
2025

Relation-Rich Visual Document Generator for Visual Information Extraction

CVPR 2025poster

Despite advances in Large Language Models (LLMs) and Multimodal LLMs (MLLMs) for visual document understanding (VDU), visual information extraction (VIE) from relation-rich documents remains challenging due to the layout diversity and limited training data. While existing synthetic document generato…

2025

Safety Depth in Large Language Models: A Markov Chain Perspective

NeurIPS 2025poster

Large Language Models (LLMs) are increasingly adopted in high-stakes scenarios, yet their safety mechanisms often remain fragile. Simple jailbreak prompts or even benign fine-tuning can bypass internal safeguards, underscoring the need to understand the failure modes of current safety strategies. R…

Cited by 0SourceScholar
2024

ACCEPT: Adaptive Codebook for Composite and Efficient Prompt Tuning

EMNLP 2024finding

Prompt Tuning has been a popular Parameter-Efficient Fine-Tuning method attributed to its remarkable performance with few updated parameters on various large-scale pretrained Language Models (PLMs). Traditionally, each prompt has been considered indivisible and updated independently, leading the par…

2023

Class-incremental Continual Learning for Instance Segmentation with Image-level Weak Supervision

ICCV 2023poster

Instance segmentation requires labor-intensive manual labeling of the contours of complex objects in images for training. The labels can also be provided incrementally in practice to balance the human labor in different time steps. However, research on incremental learning for instance segmentation…

Cited by 12PDFcodeScholar
2023

Continual Cell Instance Segmentation of Microscopy Images

ICASSP 2023accepted

A continual cell instance segmenter aims to continually learn to segment new objects while preserving the ability to localize and distinguish old objects without access to previous data. Besides catastrophic forgetting, background shift, where the background class could contain objects in the old an…

Cited by 0SourceScholar
2023

D4AM: A General Denoising Framework for Downstream Acoustic Models

ICLR 2023poster

The performance of acoustic models degrades notably in noisy environments. Speech enhancement (SE) can be used as a front-end strategy to aid automatic speech recognition (ASR) systems. However, existing training objectives of SE methods are not fully effective at integrating speech-text and noise-c…

2023

Hearing and Seeing Abnormality: Self-Supervised Audio-Visual Mutual Learning for Deepfake Detection

ICASSP 2023accepted

The recent development of deepfakes has resulted in serious threats to society, such as spreading misinformation, defamation, etc. Although recent deepfake detection methods are capable of achieving satisfactory results for seen forgeries, the performance drops significantly for unseen ones. With pr…

Cited by 0SourceScholar
2023

Scalable Spatial Memory for Scene Rendering and Navigation

AAAI 2023technical

Neural scene representation and rendering methods have shown promise in learning the implicit form of scene structure without supervision. However, the implicit representation learned in most existing methods is non-expandable and cannot be inferred online for novel scenes, which makes the learned r…

Cited by 1SourcePDFScholar
2022

Continual Learning for Visual Search With Backward Consistent Feature Embedding

CVPR 2022poster

In visual search, the gallery set could be incrementally growing and added to the database in practice. However, existing methods rely on the model trained on the entire dataset, ignoring the continual updating of the model. Besides, as the model updates, the new model must re-extract features for t…

Cited by 28PDFcodeScholar
2021

3D Video Stabilization With Depth Estimation by CNN-Based Optimization

CVPR 2021poster

Video stabilization is an essential component of visual quality enhancement. Early methods rely on feature tracking to recover either 2D or 3D frame motion, which suffer from the robustness of local feature extraction and tracking in shaky videos. Recently, learning-based methods seek to find frame…

Cited by 40PDFScholar
2021

STR-GQN: Scene Representation and Rendering for Unknown Cameras Based on Spatial Transformation Routing

ICCV 2021poster

Geometry-aware modules are widely applied in recent deep learning architectures for scene representation and rendering. However, these modules require intrinsic camera information that might not be obtained accurately. In this paper, we propose a Spatial Transformation Routing (STR) mechanism to mod…

Cited by 5PDFScholar
2019

Compacting, Picking and Growing for Unforgetting Continual Learning

NeurIPS 2019poster

Continual lifelong learning is essential to many applications. In this paper, we propose a simple but effective approach to continual deep learning. Our approach leverages the principles of deep model compression, critical weights selection, and progressive networks expansion. By enforcing their int…

2017

Learning and inferring human actions with temporal pyramid features based on conditional random fields

ICASSP 2017accepted

Finding an effective way to represent human actions is yet an open problem because it usually requires taking evidences extracted from various temporal resolutions into account. A conventional way of representing an action employs temporally ordered fine-grained movements, e.g., key poses or subtle…

Cited by 0SourceScholar