← Search

Feng Li

58 accepted papers

2026

BDRP: A Binary Divisive Recursive Planner for Path Planning

RA-L 2026

Narrow passage scenarios pose significant challenges for path planning, especially for tasks requiring real-time performance. Traditional asymptotically converging sampling-based planners (SBPs) often exhibit poor initial path quality and slow convergence, limiting their ability to efficiently const

Cited by 0SourceScholar
2026

CATNet: Collaborative Alignment and Transformation Network for Cooperative Perception

CVPR 2026

Cooperative perception significantly enhances scene understanding by integrating complementary information from diverse agents. However, existing research often overlooks critical challenges inherent in real-world multi-source data integration, specifically high temporal latency and multi-source noi

Cited by 0SourceScholar
2026

Can Large Language Models Grasp 3D Medical Anatomy Shapes? (Student Abstract)

AAAI 2026technical

What if the next generation of human-computer interaction is not a screen... but a conversation? Large Language Models (LLMs) offer a new paradigm for interacting with computers through text, but they lack shape reasoning capabilities. We introduce Textual Anatomy Encoding (TAE), a workflow that con

Cited by 0SourcePDFScholar
2026

Geometry-Aware Visual Odometry for Bronchoscopic Navigation Via High-Gain Observer Fusion

ICRA 2026poster

Navigational bronchoscopy is critical for pulmonary interventions, yet current platforms depend heavily on pre-operative CT or external sensors, limiting their use in critical care and resource-constrained settings. Vision-only navigation offers a scalable alternative, but conventional visual odomet…

Cited by 0Scholar
2026

Layer Consistency Matters: Elegant Latent Transition Discrepancy for Generalizable Synthetic Image Detection

CVPR 2026

Recent rapid advancement of generative models has significantly improved the fidelity and accessibility of AI-generated synthetic images. While enabling various innovative applications, the unprecedented realism of these synthetics makes them increasingly indistinguishable from authentic photographs

Cited by 0SourcecodeScholar
2026

ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation

ICLR 2026poster

Unified multimodal models (UMMs) have shown remarkable advances in jointly understanding and generating text and images. However, prevailing evaluations treat these abilities in isolation, such that tasks with multimodal inputs and outputs are scored primarily through unimodal reasoning: textual ben…

Cited by 0SourcecodeScholar
2026

SAM 3: Segment Anything with Concepts

ICLR 2026poster

We present Segment Anything Model (SAM) 3, a unified model that detects, segments, and tracks objects in images and videos based on concept prompts, which we define as either short noun phrases (e.g., “yellow school bus”), image exemplars, or a combination of both. Promptable Concept Segmentation (P…

Cited by 687SourcecodeScholar
2026

Surgical Workflow Prediction via Visual Information in Laparoscopic and Robot-Assisted Surgery

RA-L 2026

Surgical workflow prediction is critical for enhancing safety and providing real-time guidance in Computer-Assisted Surgery (CAS), particularly in laparoscopic and Robot-Assisted Surgery (RAS). We propose a novel visual information-based method for predicting surgical workflows at fine temporal scal

Cited by 0SourceScholar
2026

VQ-VA World: Towards High-Quality Visual Question-Visual Answering

CVPR 2026

This paper studies Visual Question-Visual Answering (VQ-VA): generating an image, rather than text, in response to a visual question---an ability that has recently emerged in proprietary systems such as NanoBanana and GPT-Image. To also bring this capability to open-source models, we introduce VQ-VA

Cited by 0SourcecodeScholar
2025

ATP-TTS: Adaptive Thresholding Pseudo-Labeling for Low-Resource Multi-Speaker Text-to-Speech

ICASSP 2025accepted

To address the challenge of high annotation costs in text-to-speech (TTS) generation, this paper introduces a semi-supervised learning framework specifically designed for low-resource TTS scenarios. The framework incorporates adaptive thresholding to select appropriate pseudo-labels and leverages au…

Cited by 0SourceScholar
2025

Attend and Enrich: Enhanced Visual Prompt for Zero-Shot Learning

AAAI 2025technical

Zero-shot learning (ZSL) endeavors to transfer knowledge from the seen categories to recognize unseen categories, which mostly relies on the semantic-visual interactions between image and attribute tokens. Recently, the prompt learning has emerged in ZSL and demonstrated significant potential as it…

2025

CLEAR: A Clinically Grounded Tabular Framework for Radiology Report Evaluation

EMNLP 2025

Existing metrics often lack the granularity and interpretability to capture nuanced clinical differences between candidate and ground-truth radiology reports, resulting in suboptimal evaluation. We introduce a **Cl**inically grounded tabular framework with **E**xpert-curated labels and **A**ttribute

2025

EvEnhancer: Empowering Effectiveness, Efficiency and Generalizability for Continuous Space-Time Video Super-Resolution with Events

CVPR 2025highlight

Continuous space-time video super-resolution (C-STVSR) endeavors to upscale videos simultaneously at arbitrary spatial and temporal scales, which has recently garnered increasing interest. However, prevailing methods struggle to yield satisfactory videos at out-of-distribution spatial and temporal s…

2025

Human-Like Walking Motion Generation for Self-Balancing Lower Limb Rehabilitation Exoskeletons

ICRA 2025

Self-balancing lower limb rehabilitation exoskeletons (SLLREs) allow individuals with lower limb dysfunction to walk without the use of crutches. Stable and human-like walking motions are crucial for SLLREs because achieving a close imitation of healthy human walking is a key goal in rehabilitation

Cited by 0SourceScholar
2025

LLaVA-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

ICLR 2025spotlight

Visual instruction tuning has made considerable strides in enhancing the capabilities of Large Multimodal Models (LMMs). However, existing open LMMs largely focus on single-image tasks, their applications to multi-image scenarios remains less explored. Additionally, prior LMM research separately tac…

Cited by 0SourcePDFScholar
2025

MDD-5k: A New Diagnostic Conversation Dataset for Mental Disorders Synthesized via Neuro-Symbolic LLM Agents

AAAI 2025technical

The clinical diagnosis of most mental disorders primarily relies on the conversations between psychiatrist and patient. The creation of such diagnostic conversation datasets is promising to boost the AI mental healthcare community. However, directly collecting the conversations in real diagnosis sce…

2025

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

ICLR 2025poster

Comprehensive evaluation of Multimodal Large Language Models (MLLMs) has recently garnered widespread attention in the research community. However, we observe that existing benchmarks present several common barriers that make it difficult to measure the significant challenges that models face in the…

Cited by 41SourcePDFScholar
2025

Once-for-All: Controllable Generative Image Compression with Dynamic Granularity Adaptation

ICLR 2025poster

Although recent generative image compression methods have demonstrated impressive potential in optimizing the rate-distortion-perception trade-off, they still face the critical challenge of flexible rate adaptation to diverse compression necessities and scenarios. To overcome this challenge, this pa…

Cited by 1SourcePDFScholar
2025

Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production

AAAI 2025technical

Sign Language Production (SLP) aims to generate semantically consistent sign videos from textual statements, where the conversion from textual glosses to sign poses (G2P) is a crucial step. Existing G2P methods typically treat sign poses as discrete three-dimensional coordinates and directly fit the…

2025

Thinking in Granularity: Dynamic Quantization for Image Super-Resolution by Intriguing Multi-Granularity Clues

AAAI 2025technical

Dynamic quantization has attracted rising attention in image super-resolution (SR) as it expands the potential of heavy SR models onto mobile devices while preserving competitive performance. Most current methods explore layer-to-bit configuration upon varying local regions, adaptively allocating th…

2024

A Closed-loop Control for Lower Limb Exoskeleton Considering Overall Deformations: A Simple and Direct Application Method

IROS 2024

In this paper, considering overall deformations of the exoskeleton, we couple deformations relationship network (DRN) with fractional order viscoelastic (FOV) controller, proposing a novel DRN-FOV closed-loop control method, endowing exoskeleton with stable dynamic walking ability. Simply by utilizi

Cited by 0SourceScholar
2024

Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

ECCV 2024poster

"In this paper, we develop an open-set object detector, called Grounding DINO, by marrying Transformer-based detector DINO with grounded pre-training, which can detect arbitrary objects with human inputs such as category names or referring expressions. The key solution of open-set object detection i…

2024

Interfacing Foundation Models' Embeddings

NeurIPS 2024poster

Foundation models possess strong capabilities in reasoning and memorizing across modalities. To further unleash the power of foundation models, we present FIND, a generalized interface for aligning foundation models' embeddings with unified image and dataset-level understanding spanning modality and…

2024

LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models

ECCV 2024poster

"With the recent significant advancements in large multimodal models (LMMs), the importance of their grounding capability in visual chat is increasingly recognized. Despite recent efforts to enable LMMs to support grounding, their capabilities for grounding and chat are usually separate, and their c…

2024

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

ECCV 2024poster

"This paper presents (), a general-purpose multimodal assistant trained using an end-to-end approach that systematically expands the capabilities of large multimodal models (LMMs). maintains a skill repository that contains a wide range of vision and vision-language pre-trained models (tools), and i…

2024

T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy

ECCV 2024poster

"We present , a highly practical model for open-set object detection. Previous open-set object detection methods relying on text prompts effectively encapsulate the abstract concept of common objects, but struggle with rare or complex object representation due to data scarcity and descriptive limita…

2024

TAPTR: Tracking Any Point with Transformers as Detection

ECCV 2024poster

"In this paper, we propose a simple yet effective approach for Tracking Any Point with TRansformers (). Based on the observation that point tracking bears a great resemblance to object detection and tracking, we borrow designs from DETR-like algorithms to address the task of TAP. In , in each video…

Cited by 20SourcePDFScholar
2024

TAPTRv2: Attention-based Position Update Improves Tracking Any Point

NeurIPS 2024poster

In this paper, we present TAPTRv2, a Transformer-based approach built upon TAPTR for solving the Tracking Any Point (TAP) task. TAPTR borrows designs from DEtection TRansformer (DETR) and formulates each tracking point as a point query, making it possible to leverage well-studied operations in DETR-…

Cited by 7SourcePDFScholar
2023

A Simple Framework for Open-Vocabulary Segmentation and Detection

ICCV 2023poster

In this work, we present OpenSeeD, a simple Open-vocabulary Segmentation and Detection framework that learns from different segmentation and detection datasets. To bridge the gap of vocabulary and annotation granularity, we first introduce a pretrained text encoder to encode all the visual concepts…

Cited by 176PDFcodeScholar
2023

DFA3D: 3D Deformable Attention For 2D-to-3D Feature Lifting

ICCV 2023poster

In this paper, we propose a new operator, called 3D DeFormable Attention (DFA3D), for 2D-to-3D feature lifting, which transforms multi-view 2D image features into a unified 3D space for 3D object detection. Existing feature lifting approaches, such as Lift-Splat-based and 2D attention-based, either…

Cited by 44PDFcodeScholar
2023

DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

ICLR 2023poster

We present DINO (DETR with Improved deNoising anchOr boxes), a strong end-to-end object detector. DINO improves over previous DETR-like models in performance and efficiency by using a contrastive way for denoising training, a look forward twice scheme for box prediction, and a mixed query selection…

2023

DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and Grounding

AAAI 2023technical

In this paper, we study the problem of visual grounding by considering both phrase extraction and grounding (PEG). In contrast to the previous phrase-known-at-test setting, PEG requires a model to extract phrases from text and locate objects from image simultaneously, which is a more practical setti…

2023

Detection Transformer with Stable Matching

ICCV 2023poster

This paper is concerned with the matching stability problem across different decoder layers in DEtection TRansformers (DETR). We point out that the unstable matching in DETR is caused by a multi-optimization path problem, which is highlighted by the one-to-one matching design in DETR. To address thi…

Cited by 46PDFcodeScholar
2023

Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation

ICLR 2023poster

This paper presents a novel end-to-end framework with Explicit box Detection for multi-person Pose estimation, called ED-Pose, where it unifies the contextual learning between human-level (global) and keypoint-level (local) information. Different from previous one-stage methods, ED-Pose re-considers…

2023

Lite DETR: An Interleaved Multi-Scale Encoder for Efficient DETR

CVPR 2023poster

Recent DEtection TRansformer-based (DETR) models have obtained remarkable performance. Its success cannot be achieved without the re-introduction of multi-scale feature fusion in the encoder. However, the excessively increased tokens in multi-scale features, especially for about 75% of low-level fea…

2023

MP-Former: Mask-Piloted Transformer for Image Segmentation

CVPR 2023poster

We present a mask-piloted Transformer which improves masked-attention in Mask2Former for image segmentation. The improvement is based on our observation that Mask2Former suffers from inconsistent mask predictions between consecutive decoder layers, which leads to inconsistent optimization goals and…

2023

Mask DINO: Towards a Unified Transformer-Based Framework for Object Detection and Segmentation

CVPR 2023poster

In this paper we present Mask DINO, a unified object detection and segmentation framework. Mask DINO extends DINO (DETR with Improved Denoising Anchor Boxes) by adding a mask prediction branch which supports all image segmentation tasks (instance, panoptic, and semantic). It makes use of the query e…

2023

Progressive Semantic-Visual Mutual Adaption for Generalized Zero-Shot Learning

CVPR 2023highlight

Generalized Zero-Shot Learning (GZSL) identifies unseen categories by knowledge transferred from the seen domain, relying on the intrinsic interactions between visual and semantic information. Prior works mainly localize regions corresponding to the sharing attributes. When various visual appearance…

2023

Segment Everything Everywhere All at Once

NeurIPS 2023poster

In this work, we present SEEM, a promotable and interactive model for segmenting everything everywhere all at once in an image. In SEEM, we propose a novel and versatile decoding mechanism that enables diverse prompting for all types of segmentation tasks, aiming at a universal interface that behave…

Cited by 621SourcePDFScholar
2022

APG: Adaptive Parameter Generation Network for Click-Through Rate Prediction

NeurIPS 2022accept

In many web applications, deep learning-based CTR prediction models (deep CTR models for short) are widely adopted. Traditional deep CTR models learn patterns in a static manner, i.e., the network parameters are the same across all the instances. However, such a manner can hardly characterize each…

Cited by 37SourcePDFScholar
2022

An Accelerated Rank-(L, L, 1, 1) Block Term Decomposition Of Multi-Subject Fmri Data Under Spatial Orthonormality Constraint

ICASSP 2022accepted

The decomposition of multi-subject fMRI data using rank-(L,L,1,1) block term decomposition (BTD) can preserve higher-way data structure and is more robust to noise effects by decomposing shared spatial maps (SMs) into a product of two rank-L loading matrices. However, since the number of whole-brain…

Cited by 4SourceScholar
2022

DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR

ICLR 2022poster

We present in this paper a novel query formulation using dynamic anchor boxes for DETR (DEtection TRansformer) and offer a deeper understanding of the role of queries in DETR. This new formulation directly uses box coordinates as queries in Transformer decoders and dynamically updates them layer by…

2022

DN-DETR: Accelerate DETR Training by Introducing Query DeNoising

CVPR 2022oral

We present in this paper a novel denoising training method to speedup DETR (DEtection TRansformer) training and offer a deepened understanding of the slow convergence issue of DETR-like methods. We show that the slow convergence results from the instability of bipartite graph matching which causes i…

Cited by 921PDFcodeScholar
2022

Learn To Remember: Transformer with Recurrent Memory for Document-Level Machine Translation

NAACL 2022findings

The Transformer architecture has led to significant gains in machine translation. However, most studies focus on only sentence-level translation without considering the context dependency within documents, leading to the inadequacy of document-level coherence. Some recent research tried to mitigate…

Cited by 20SourcePDFScholar
2021

Deep Texture Recognition via Exploiting Cross-Layer Statistical Self-Similarity

CVPR 2021poster

In recent years, convolutional neural networks (CNNs) have become a prominent tool for texture recognition. The key of existing CNN-based approaches is aggregating the convolutional features into a robust yet discriminative description. This paper presents a novel feature aggregation module called C…

Cited by 49PDFScholar
2021

Encoding Spatial Distribution of Convolutional Features for Texture Representation

NeurIPS 2021poster

Existing convolutional neural networks (CNNs) often use global average pooling (GAP) to aggregate feature maps into a single representation. However, GAP cannot well characterize complex distributive patterns of spatial features while such patterns play an important role in texture-oriented applicat…

2021

SARG: A Novel Semi Autoregressive Generator for Multi-turn Incomplete Utterance Restoration

AAAI 2021technical

Dialogue systems in open domain have achieved great success due to the easily obtained single-turn corpus and the development of deep learning, but the multi-turn scenario is still a challenge because of the frequent coreference and information omission. In this paper, we investigate the incomplete…

2021

Towards Complete Scene and Regular Shape for Distortion Rectification by Curve-Aware Extrapolation

ICCV 2021poster

The wide-angle lens gains increasing attention since it can capture a wide field-of-view scene (FoV). However, the obtained image is contaminated with radial distortion, making the scene not realistic. Previous distortion rectification methods rectify the image in a rectangle or invagination, failin…

Cited by 8PDFScholar
2021

Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline

CVPR 2021poster

Depth maps obtained by commercial depth sensors are always in low-resolution, making it difficult to be used in various computer vision tasks. Thus, depth map super-resolution (SR) is a practical and valuable task, which upscales the depth map into high-resolution (HR) space. However, limited by the…

Cited by 100PDFScholar
2020

Deep Interleaved Network for Single Image Super-Resolution with Asymmetric Co-Attention

IJCAI 2020poster

Recently, Convolutional Neural Networks (CNN) based image super-resolution (SR) have shown significant success in the literature. However, these methods are implemented as single-path stream to enrich feature maps from the input for the final prediction, which fail to fully incorporate former low-le…

Cited by 0SourcePDFScholar
2018

Learning Spatial-Temporal Regularized Correlation Filters for Visual Tracking

CVPR 2018poster

Discriminative Correlation Filters (DCF) are efficient in visual tracking but suffer from unwanted boundary effects. Spatially Regularized DCF (SRDCF) has been suggested to resolve this issue by enforcing spatial penalty on DCF coefficients, which, inevitably, improves the tracking performance at th…