← Search

Bin Luo

32 accepted papers

2026

Beyond Graph Model: Reliable VLM Fine-Tuning via Random Graph Adapter

CVPR 2026

Textual adapter-based tuning methods have shown significant potential in transferring knowledge from pre-trained Vision-Language Models (VLMs) to downstream tasks. Existing works generally employ the deterministic textual feature adapter to refine each category textual representation. However, due t

Cited by 0SourceScholar
2026

CHASE: Contextual History for Adaptive and Simple Exploitation in Large Language Model Jailbreaking

AAAI 2026technical

We propose Contextual History for Adaptive and Simple Exploitation (CHASE), a novel multi-turn method for Large Language Model (LLM) jailbreaking. Rather than directly attack an LLM that may be difficult to jailbreak, CHASE first collects jailbroken histories from an easy-to-jailbreak LLM and then t

Cited by 0SourcePDFScholar
2026

NESTOR: A Nested MOE-based Neural Operator for Large-Scale PDE Pre-Training

CVPR 2026

Neural operators have emerged as an efficient paradigm for solving PDEs, overcoming the limitations of traditional numerical methods and significantly improving computational efficiency. However, due to the diversity and complexity of PDE systems, existing neural operators typically rely on a single

Cited by 0SourcecodeScholar
2026

Semantic-Driven Visual Progressive Refinement for Aerial-Ground Person ReID: A Challenging Large-Scale Benchmark

AAAI 2026technical

Aerial-Ground Person Re-IDentification (AGPReID) aims to extract identity-discriminative representations from heterogeneous perspectives across different platforms in complex real-world environments. However, existing methods primarily focus on visual appearance modeling and make insufficient use of

Cited by 0SourcePDFScholar
2026

Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

CVPR 2026

Missing modalities in RGBT tracking often lead to incomplete and unstable multimodal feature representations that greatly degrade the performance. Existing methods typically attempt to recover missing modalities from available ones, but the quality of data generated in challenging scenarios might be

Cited by 0SourceScholar
2026

Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models

CVPR 2026

Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in multimodal tasks.Despite their impressive performance, MLLMs suffer from the modality imbalance issue, where visual information is often underutilized compared to textual representations in deeper layers, leading to

Cited by 0SourcecodeScholar
2025

Alignment-Free RGB-T Salient Object Detection: A Large-Scale Dataset and Progressive Correlation Network

AAAI 2025technical

Alignment-free RGB-Thermal (RGB-T) salient object detection (SOD) aims to achieve robust performance in complex scenes by directly leveraging the complementary information from unaligned visible-thermal image pairs, without requiring manual alignment. However, the labor-intensive process of collecti…

2025

MetaDesigner: Advancing Artistic Typography through AI-Driven, User-Centric, and Multilingual WordArt Synthesis

ICLR 2025poster

MetaDesigner introduces a transformative framework for artistic typography synthesis, powered by Large Language Models (LLMs) and grounded in a user-centric design paradigm. Its foundation is a multi-agent system comprising the Pipeline, Glyph, and Texture agents, which collectively orchestrate the…

Cited by 2SourcePDFScholar
2025

Quantum Algorithms for Finite-horizon Markov Decision Processes

ICML 2025poster

In this work, we design quantum algorithms that are more efficient than classical algorithms to solve time-dependent and finite-horizon Markov Decision Processes (MDPs) in two distinct settings: (1) In the exact dynamics setting, where the agent has full knowledge of the environment's dynamics (i.e.…

Cited by 0SourcePDFScholar
2025

RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba

AAAI 2025technical

Existing RGBT tracking methods often design various interaction models to perform cross-modal fusion of each layer, but can not execute the feature interactions among all layers, which plays a critical role in robust multimodal representation, due to large computational burden. To address this issue…

Cited by 2SourcePDFScholar
2024

Building Bridges across Spatial and Temporal Resolutions: Reference-Based Super-Resolution via Change Priors and Conditional Diffusion Model

CVPR 2024poster

Reference-based super-resolution (RefSR) has the potential to build bridges across spatial and temporal resolutions of remote sensing images. However existing RefSR methods are limited by the faithfulness of content reconstruction and the effectiveness of texture transfer in large scaling factors. C…

2024

DCPT: Darkness Clue-Prompted Tracking in Nighttime UAVs

ICRA 2024poster

Existing nighttime unmanned aerial vehicle (UAV) trackers follow an "Enhance-then-Track" architecture - first using a light enhancer to brighten the nighttime video, then employing a daytime tracker to locate the object. This separate enhancement and tracking fails to build an end-to-end trainable v…

Cited by 17SourcecodeScholar
2024

LJPCheck: Functional Tests for Legal Judgment Prediction

ACL 2024findings

Legal Judgment Prediction (LJP) refers to the task of automatically predicting judgment results (e.g., charges, law articles and term of penalty) given the fact description of cases. While SOTA models have achieved high accuracy and F1 scores on public datasets, existing datasets fail to evaluate sp…

Cited by 0SourcePDFScholar
2024

Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception

CVPR 2024highlight

Multimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs with visual perceiving and grounding capabilities. However there still remains a gap in providing fine-grained pixel-level…

2023

Backdooring Neural Code Search

ACL 2023long

Reusing off-the-shelf code snippets from online repositories is a common practice, which significantly enhances the productivity of software developers. To find desired code snippets, developers resort to code search engines through natural language queries. Neural code search models are hence behin…

2023

Communication Resources Constrained Hierarchical Federated Learning for End-to-End Autonomous Driving

IROS 2023poster

While federated learning (FL) improves the generalization of end-to-end autonomous driving by model aggregation, the conventional single-hop FL (SFL) suffers from slow convergence rate due to long-range communications among vehicles and cloud server. Hierarchical federated learning (HFL) overcomes s…

Cited by 20SourcecodeScholar
2023

DAMO-StreamNet: Optimizing Streaming Perception in Autonomous Driving

IJCAI 2023poster

In the realm of autonomous driving, real-time perception or streaming perception remains under-explored. This research introduces DAMO-StreamNet, a novel framework that merges the cutting-edge elements of the YOLO series with a detailed examination of spatial and temporal perception techniques. DAMO…

2023

HDFormer: High-order Directed Transformer for 3D Human Pose Estimation

IJCAI 2023poster

Human pose estimation is a challenging task due to its structured data sequence nature. Existing methods primarily focus on pair-wise interaction of body joints, which is insufficient for scenarios involving overlapping joints and rapidly changing poses. To overcome these issues, we introduce a nove…

2023

Longshortnet: Exploring Temporal and Semantic Features Fusion In Streaming Perception

ICASSP 2023accepted

Streaming perception is a fundamental task in autonomous driving that requires a careful balance between the latency and accuracy of the autopilot system. However, current methods for streaming perception are limited as they rely only on the current and adjacent two frames to learn movement patterns…

Cited by 0SourceScholar
2023

Procontext: Exploring Progressive Context Transformer for Tracking

ICASSP 2023accepted

Existing Visual Object Tracking (VOT) only takes the target area in the first frame as a template. This causes tracking to inevitably fail in fast-changing and crowded scenes, as it cannot account for changes in object appearance between frames. To this end, we revamped the tracking framework with P…

Cited by 0SourceScholar
2023

Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance Learning

ICCV 2023oral

Depth-aware panoptic segmentation is an emerging topic in computer vision which combines semantic and geometric understanding for more robust scene interpretation. Recent works pursue unified frameworks to tackle this challenge but mostly still treat it as two individual learning tasks, which limits…

Cited by 12PDFcodeScholar
2023

Unbiased Multiple Instance Learning for Weakly Supervised Video Anomaly Detection

CVPR 2023poster

Weakly Supervised Video Anomaly Detection (WSVAD) is challenging because the binary anomaly label is only given on the video level, but the output requires snippet-level predictions. So, Multiple Instance Learning (MIL) is prevailing in WSVAD. However, MIL is notoriously known to suffer from many fa…

2022

Deep Learning Meets Software Engineering: A Survey on Pre-Trained Models of Source Code

IJCAI 2022poster

Recent years have seen the successful application of deep learning to software engineering (SE). In particular, the development and use of pre-trained models of source code has enabled state-of-the-art results to be achieved on a wide variety of SE tasks. This paper provides an overview of this rapi…

Cited by 54SourcePDFScholar
2022

ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding

EMNLP 2022finding

Recent years have witnessed the rise and success of pre-training techniques in visually-rich document understanding. However, most existing methods lack the systematic mining and utilization of layout-centered knowledge, leading to sub-optimal performances. In this paper, we propose ERNIE-Layout, a…

2021

Don’t Miss the Potential Customers! Retrieving Similar Ads to Improve User Targeting

EMNLP 2021finding

User targeting is an essential task in the modern advertising industry: given a package of ads for a particular category of products (e.g., green tea), identify the online users to whom the ad package should be targeted. A (ad package specific) user targeting model is typically trained using histori…

Cited by 1SourcePDFScholar
2021

DymSLAM: 4D Dynamic Scene Reconstruction Based on Geometrical Motion Segmentation

RA-L 2021

Most SLAM (Simultaneous Localization and Mapping) algorithms are based on the assumption that the scene is static. However, in practice, most real scenes usually contain moving objects. In this letter, we introduce DymSLAM, a dynamic stereo visual SLAM system being capable of reconstructing a 4D (3D

Cited by 50SourceScholar
2018

SINT++: Robust Visual Tracking via Adversarial Positive Instance Generation

CVPR 2018poster

Existing visual trackers are easily disturbed by occlusion,blurandlargedeformation. Inthechallengesofocclusion, motion blur and large object deformation, the performance of existing visual trackers may be limited due to the followingissues: i)Adoptingthedensesamplingstrategyto generate positive exam…

Cited by 152SourcePDFScholar
2016

Reversible data hiding in encrypted image based on block histogram shifting

ICASSP 2016accepted

Since there is good potential for practical applications such as encrypted image authentication, content owner identification and privacy protection, reversible data hiding in encrypted image (RDHEI) has attracted increasing attention in recent years. In this paper, we propose and evaluate a new sep…

Cited by 0SourceScholar