← Search

Bin Kang

14 accepted papers

2026

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech

ICML 2026poster

While existing text-to-speech (TTS) models exhibit high expressiveness, fine-grained control over composite instructions remains challenging due to the structural mismatch between discrete textual intents and continuous acoustic realizations. Inspired by human cognitive decoupling, we introduce Agen…

Cited by 0SourceScholar
2026

Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation

ICML 2026poster

Text-to-Image (T2I) models and Unified Multimodal Models (UMMs) have achieved remarkable progress in visual generation. However, their reliance on a single-pass generation paradigm limits their ability to handle complex prompts requiring iterative refinement. To enable multi-round Reflective Visual …

Cited by 0SourceScholar
2026

LongHorizonUI: A Unified Framework for Robust long-horizon Task Automation of GUI Agent

ICLR 2026poster

Although agents based on multimodal large language models (MLLMs) demonstrate proficiency in general short-term graphical user interface (GUI) tasks, their robustness remains a significant challenge for handling complex long-horizon tasks in dynamic environments . In response, the LongHorizonUI fram…

Cited by 0SourcecodeScholar
2025

DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception

CVPR 2025poster

Dense visual prediction tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbounded. While Vision-Language Models (VLMs) like CLIP have shown promise in open-vocabulary tasks, their direct applicatio…

2025

Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior

NeurIPS 2025poster

Recent advances in Video Large Language Models (VLLMs) have achieved remarkable video understanding capabilities, yet face critical efficiency bottlenecks due to quadratic computational growth with lengthy visual token sequences of long videos. While existing keyframe sampling methods can improve te…

Cited by 0SourceScholar
2025

OV-DQUO: Open-Vocabulary DETR with Denoising Text Query Training and Open-World Unknown Objects Supervision

AAAI 2025technical

Open-vocabulary detection aims to detect objects from novel categories beyond the base categories on which the detector is trained. However, existing open-vocabulary detectors trained on base category data tend to assign higher confidence to trained categories and confuse novel categories with the b…

2023

ParaFormer: Parallel Attention Transformer for Efficient Feature Matching

AAAI 2023technical

Heavy computation is a bottleneck limiting deep-learning-based feature matching algorithms to be applied in many real-time applications. However, existing lightweight networks optimized for Euclidean data cannot address classical feature matching tasks, since sparse keypoint based descriptors are ex…

Cited by 18SourcePDFScholar
2022

PointTAD: Multi-Label Temporal Action Detection with Learnable Query Points

NeurIPS 2022accept

Traditional temporal action detection (TAD) usually handles untrimmed videos with small number of action instances from a single label (e.g., ActivityNet, THUMOS). However, this setting might be unrealistic as different classes of actions often co-occur in practice. In this paper, we focus on the ta…

2021

Cross Scene Video Foreground Segmentation Via Co-Occurrence Probability Oriented Supervised and Unsupervised Model Interaction

ICASSP 2021accepted

Using only one deep model for cross scene video foreground segmentation is still very challenging because existing methods are scene-dependent, which restricts the consistent segmentation. In this paper, we propose a cross scene video foreground segmentation framework to extend the generalization ca…

Cited by 0SourceScholar
2020

FDDWNet: A Lightweight Convolutional Neural Network for Real-Time Semantic Segmentation

ICASSP 2020accepted

This paper introduces a lightweight convolutional neural network, called FDDWNet, for real-time accurate semantic segmentation. In contrast to recent advances of lightweight networks that prefer to utilize shallow structure, FDDWNet makes an effort to design more deeper network architecture, while m…

Cited by 0SourceScholar
2020

TEA: Temporal Excitation and Aggregation for Action Recognition

CVPR 2020poster

Temporal modeling is key for action recognition in videos. It normally considers both short-range motions and long-range aggregations. In this paper, we propose a Temporal Excitation and Aggregation (TEA) block, including a motion excitation (ME) module and a multiple temporal aggregation (MTA) modu…

Cited by 638PDFScholar
2019

Grayscale-thermal Tracking via Canonical Correlation Analysis Based Inverse Sparse Representation

ICASSP 2019accepted

The grayscale-thermal tracking has attracted increasing attention due to the fact that it can make thermal information complement with grayscale information. Since there exists a large gap between the grayscale and the thermal video sequences, how to exploit the intrinsic relation between the graysc…

Cited by 0SourceScholar
2019

Score-specific Non-maximum Suppression and Coexistence Prior for Multi-scale Face Detection

ICASSP 2019accepted

Face detection is an ultimate component to support various visual facial related tasks. However, detecting faces with extremely low resolution or high occlusion is still an open problem. In this paper, we propose a two-step general approach to refine the performance of modern face detectors accordin…

Cited by 0SourceScholar