← Search

Jianbin Jiao

33 accepted papers

2026

Agentic Reinforcement Learning with Implicit Step Rewards

ICLR 2026poster

Large language models (LLMs) are increasingly developed as autonomous agents using reinforcement learning (agentic RL) that reason and act in interactive environments. However, sparse and sometimes unverifiable rewards make it extremely challenging to assign credit when training LLM agents that serv…

Cited by 0SourceScholar
2026

Balancing Understanding and Generation in Discrete Diffusion Models

ICML 2026spotlight

In discrete generative modeling, two dominant paradigms demonstrate divergent capabilities: Masked Diffusion Language Models (MDLM) excel at semantic understanding and zero-shot generalization, whereas Uniform-noise Diffusion Language Models (UDLM) achieve strong few-step generation quality, yet nei…

Cited by 0SourceScholar
2026

HeroGS: Hierarchical Guidance for Robust 3D Gaussian Splatting under Sparse Views

CVPR 2026

3D Gaussian Splatting (3DGS) has recently emerged as a promising approach in novel view synthesis, combining photorealistic rendering with real-time efficiency. However, its success heavily relies on dense camera coverage; under sparse-view conditions, insufficient supervision leads to irregular Gau

Cited by 0SourceScholar
2025

Adaptive Keyframe Sampling for Long Video Understanding

CVPR 2025poster

Multimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language models (LLMs) as contexts. However, when the visual input changes from a single image to a long video, the above paradigm encounters difficulty because…

2025

EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning

ACL 2025long

Large Language Models (LLMs) have shown impressive reasoning capabilities in well-defined problems with clear solutions, such as mathematics and coding. However, they still struggle with complex real-world scenarios like business negotiations, which require strategic reasoning—an ability to navigate…

2024

BadRL: Sparse Targeted Backdoor Attack against Reinforcement Learning

AAAI 2024technical

Backdoor attacks in reinforcement learning (RL) have previously employed intense attack strategies to ensure attack success. However, these methods suffer from high attack costs and increased detectability. In this work, we propose a novel approach, BadRL, which focuses on conducting highly sparse b…

2024

P2Seg: Pointly-supervised Segmentation via Mutual Distillation

ICLR 2024poster

Point-level Supervised Instance Segmentation (PSIS) aims to enhance the applicability and scalability of instance segmentation by utilizing low-cost yet instance-informative annotations. Existing PSIS methods usually rely on positional information to distinguish objects, but predicting precise bound…

Cited by 2SourcePDFScholar
2024

Position: Foundation Agents as the Paradigm Shift for Decision Making

ICML 2024poster

Decision making demands intricate interplay between perception, memory, and reasoning to discern optimal policies. Conventional approaches to decision making face challenges related to low sample efficiency and poor generalization. In contrast, foundation models in language and vision have showcased…

2024

Semantic-aware SAM for Point-Prompted Instance Segmentation

CVPR 2024highlight

Single-point annotation in visual tasks with the goal of minimizing labeling costs is becoming increasingly prominent in research. Recently visual foundation models such as Segment Anything (SAM) have gained widespread usage due to their robust zero-shot capabilities and exceptional annotation perfo…

2024

VMamba: Visual State Space Model

NeurIPS 2024spotlight

Designing computationally efficient network architectures remains an ongoing necessity in computer vision. In this paper, we adapt Mamba, a state-space language model, into VMamba, a vision backbone with linear time complexity. At the core of VMamba is a stack of Visual State-Space (VSS) blocks with…

2023

Integrally Pre-Trained Transformer Pyramid Networks

CVPR 2023poster

In this paper, we present an integral pre-training framework based on masked image modeling (MIM). We advocate for pre-training the backbone and neck jointly so that the transfer gap between MIM and downstream recognition tasks is minimal. We make two technical contributions. First, we unify the rec…

2023

Spatial Self-Distillation for Object Detection with Inaccurate Bounding Boxes

ICCV 2023poster

Object detection via inaccurate bounding box supervision has boosted a broad interest due to the expensive high-quality annotation data or the occasional inevitability of low annotation quality (e.g. tiny objects). The previous works usually utilize multiple instance learning (MIL), which highly dep…

Cited by 18PDFcodeScholar
2021

Anti-Aliasing Semantic Reconstruction for Few-Shot Semantic Segmentation

CVPR 2021poster

Encouraging progress in few-shot semantic segmentation has been made by leveraging features learned upon base classes with sufficient training data to represent novel classes with few-shot examples. However, this feature sharing mechanism inevitably causes semantic aliasing between novel classes whe…

Cited by 63PDFcodeScholar
2021

Beyond Bounding-Box: Convex-Hull Feature Adaptation for Oriented and Densely Packed Object Detection

CVPR 2021poster

Detecting oriented and densely packed objects remains challenging for spatial feature aliasing caused by the intersection of reception fields between objects. In this paper, we propose a convex-hull feature adaptation (CFA) approach for configuring convolutional features in accordance with oriented…

Cited by 222PDFcodeScholar
2021

Conformer: Local Features Coupling Global Representations for Visual Recognition

ICCV 2021poster

Within Convolutional Neural Network (CNN), the convolution operations are good at extracting local features but experience difficulty to capture global representations. Within visual transformer, the cascaded self-attention modules can capture long-distance feature dependencies but unfortunately det…

Cited by 890PDFcodeScholar
2021

Self-Motivated Communication Agent for Real-World Vision-Dialog Navigation

ICCV 2021poster

Vision-Dialog Navigation (VDN) requires an agent to ask questions and navigate following the human responses to find target objects. Conventional approaches are only allowed to ask questions at predefined locations, which are built upon expensive dialogue annotations, and inconvenience the real-word…

Cited by 35PDFScholar
2020

Learning Saliency Propagation for Semi-Supervised Instance Segmentation

CVPR 2020poster

Instance segmentation is a challenging task for both modeling and annotation. Due to the high annotation cost, modeling becomes more difficult because of the limited amount of supervision. We aim to improve the accuracy of the existing instance segmentation models by utilizing a large amount of dete…

Cited by 41PDFcodeScholar
2020

Prototype Mixture Models for Few-shot Semantic Segmentation

ECCV 2020poster

Few-shot segmentation is challenging because objects within the support and query images could significantly differ in appearance and pose. Using a single prototype acquired directly from the support image to segment the query image causes semantic ambiguity. In this paper, we propose prototype mixt…

2020

Vision-Dialog Navigation by Exploring Cross-Modal Memory

CVPR 2020poster

Vision-dialog navigation posed as a new holy-grail task in vision-language disciplinary targets at learning an agent endowed with the capability of constant conversation for help with natural language and navigating according to human responses. Besides the common challenges faced in visual language…

Cited by 55PDFcodeScholar
2019

C-MIL: Continuation Multiple Instance Learning for Weakly Supervised Object Detection

CVPR 2019oral

Weakly supervised object detection (WSOD) is a challenging task when provided with image category supervision but required to simultaneously learn object locations and object detectors. Many WSOD approaches adopt multiple instance learning (MIL) and have non-convex loss functions which are prone to…

Cited by 299PDFcodeScholar
2019

DANet: Divergent Activation for Weakly Supervised Object Localization

ICCV 2019poster

Weakly supervised object localization remains a challenge when learning object localization models from image category labels. Optimizing image classification tends to activate object parts and ignore the full object extent, while expanding object parts into full object extent could deteriorate the…

Cited by 242PDFcodeScholar
2019

Learning Instance Activation Maps for Weakly Supervised Instance Segmentation

CVPR 2019poster

Discriminative region responses residing inside an object instance can be extracted from networks trained with image-level label supervision. However, learning the full extent of pixel-level instance response in a weakly supervised manner remains unexplored. In this work, we tackle this challenging…

Cited by 95PDFScholar
2019

SIXray: A Large-Scale Security Inspection X-Ray Benchmark for Prohibited Item Discovery in Overlapping Images

CVPR 2019poster

In this paper, we present a large-scale dataset and establish a baseline for prohibited item discovery in Security Inspection X-ray images. Our dataset, named SIXray, consists of 1,059,231 X-ray images, in which 6 classes of 8,929 prohibited items are manually annotated. It raises a brand new challe…

Cited by 359PDFcodeScholar
2019

Selective Sparse Sampling for Fine-Grained Image Recognition

ICCV 2019poster

Fine-grained recognition poses the unique challenge of capturing subtle inter-class differences under considerable intra-class variances (e.g., beaks for bird species). Conventional approaches crop local regions and learn detailed representation from those regions, but suffer from the fixed number o…

Cited by 316PDFcodeScholar
2018

Image-Image Domain Adaptation With Preserved Self-Similarity and Domain-Dissimilarity for Person Re-Identification

CVPR 2018poster

Person re-identification (re-ID) models trained on one domain often fail to generalize well to another. In our attempt, we present a ``learning via translation'' framework. In the baseline, we translate the labeled images from source to target domain in an unsupervised manner. We then train re-ID mo…

Cited by 1224SourcePDFScholar
2018

Min-Entropy Latent Model for Weakly Supervised Object Detection

CVPR 2018poster

Weakly supervised object detection is a challenging task when provided with image category supervision but required to learn, at the same time, object locations and object detectors. The inconsistency between the weak supervision and learning objectives introduces randomness to object locations and…

2018

Weakly Supervised Instance Segmentation Using Class Peak Response

CVPR 2018poster

Weakly supervised instance segmentation with image-level labels, instead of expensive pixel-level masks, remains unexplored. In this paper, we tackle this challenging problem by exploiting class peak responses to enable a classification network for instance mask extraction. With image labels supervi…

Cited by 357SourcePDFScholar
2017

A scalable convolutional neural network for task-specified scenarios via knowledge distillation

ICASSP 2017accepted

In this paper, we explore the redundancy in convolutional neural network, which scales with the complexity of vision tasks. Considering that many front-end visual systems are interested in only a limited range of visual targets, the removing of task-specified network redundancy can promote a wide ra…

Cited by 0SourceScholar
2017

SRN: Side-output Residual Network for Object Symmetry Detection in the Wild

CVPR 2017oral

In this paper, we establish a baseline for object symmetry detection in complex backgrounds by presenting a new benchmark and an end-to-end deep learning approach, opening up a promising direction for symmetry detection in the wild. The new benchmark, named Sym-PASCAL, spans challenges including obj…

Cited by 120PDFcodeScholar
2017

Soft Proposal Networks for Weakly Supervised Object Localization

ICCV 2017poster

Weakly supervised object localization remains challenging, where only image labels instead of bounding boxes are available during training. Object proposal is an effective component in localization, but often computationally expensive and incapable of joint optimization with some of the remaining mo…

Cited by 181PDFScholar
2015

Pedestrian detection via PCA filters based convolutional channel features

ICASSP 2015accepted

In this paper, we propose a kind of image representation, named PCA filters based convolutional channel features (PCA-CCF) for pedestrian detection. The motivation is to use the convolutional network architecture with orthogonal PCA filters to enhance the state-of-the-art aggregate channel features…

Cited by 0SourceScholar