← Search

Jingyu Li

25 accepted papers

2026

CollectiveKV: Decoupling and Sharing Collaborative Information in Sequential Recommendation

ICLR 2026poster

Sequential recommendation models are widely used in applications, yet they face stringent latency requirements. Mainstream models leverage the Transformer attention mechanism to improve performance, but its computational complexity grows with the sequence length, leading to a latency challenge for…

Cited by 0SourceScholar
2026

DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object Detection

CVPR 2026

Multimodal tiny object detection plays a critical role in real-world applications, yet remains highly challenging due to weak target representations and complex cross-modal interference. Existing frequency-domain methods for tiny object detection are still largely limited to the visible modality and

Cited by 0SourceScholar
2026

EAKV: An Entropy-Driven Adaptive KV Compression Framework for Long Video Understanding

ICML 2026poster

Although Multimodal Large Language Models have made remarkable progress, they still struggle with long-video understanding due to the massive memory footprint of KV Caches. Exsiting methods often resort to disjoint retrieval or attention-based static reduction to achieve compression. However, these …

Cited by 0SourceScholar
2026

GeoTeacher: Geometry-Guided Semi-Supervised 3D Object Detection

ICRA 2026poster

Semi-supervised 3D object detection (SS3D), aiming to explore unlabeled data for boosting 3D object detectors, has emerged as an active research area in recent years. Some previous methods have shown substantial improvements by either employing heterogeneous teacher models to provide high-quality ps…

2026

ImagiDrive: A Unified Imagination-And-Planning Framework for Autonomous Driving

ICRA 2026poster

Autonomous driving requires rich contextual comprehension and precise predictive reasoning to navigate dynamic and complex environments safely. Vision-Language Models (VLMs) and Driving World Models (DWMs) have independently emerged as powerful recipes addressing different aspects of this challenge.…

2026

Length-Adaptive Interest Network for Balancing Long and Short Sequence Modeling in CTR Prediction

AAAI 2026technical

User behavior sequences in modern recommendation systems exhibit significant length heterogeneity, ranging from sparse short-term interactions to rich long-term histories. While longer sequences provide more context, we observe that increasing the maximum input sequence length in existing CTR models

Cited by 0SourcePDFScholar
2026

Perception in Plan: Coupled Perception and Planning for End-to-End Autonomous Driving

AAAI 2026technical

End-to-end autonomous driving has achieved remarkable advancements in recent years. Existing methods primarily follow a perception–planning paradigm, where perception and planning are executed sequentially within a fully differentiable framework for planning-oriented optimization. We further advance

Cited by 0SourcePDFScholar
2026

SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Driving

CVPR 2026

Recent end-to-end autonomous driving approaches have leveraged Vision-Language Models (VLMs) to enhance planning capabilities in complex driving scenarios. However, VLMs are inherently trained as generalist models, lacking specialized understanding of driving-specific reasoning in 3D space and time.

Cited by 0SourcecodeScholar
2025

Contrastive Representation for Interactive Recommendation

AAAI 2025technical

Interactive Recommendation (IR) has gained significant attention recently for its capability to quickly capture dynamic interest and optimize both short and long term objectives. IR agents are typically implemented through Deep Reinforcement Learning (DRL), because DRL is inherently compatible with…

2025

DH-Set: Improving Vision-Language Alignment with Diverse and Hybrid Set-Embeddings Learning

CVPR 2025poster

Vision-Language (VL) alignment across image and text modalities is a challenging task due to the inherent semantic ambiguity of data with multiple possible meanings. Existing methods typically solve it by learning multiple sub-representation spaces to encode each input data as a set of embeddings, a…

Cited by 0SourcePDFScholar
2025

Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution

NeurIPS 2025poster

End-to-end autonomous driving methods aim to directly map raw sensor inputs to future driving actions such as planned trajectories, bypassing traditional modular pipelines. While these approaches have shown promise, they often operate under a one-shot paradigm that relies heavily on the current scen…

Cited by 0SourcecodeScholar
2025

MiCEval: Unveiling Multimodal Chain of Thought’s Quality via Image Description and Reasoning Steps

NAACL 2025long

**Multimodal Chain of Thought (MCoT)** is a popular prompting strategy for improving the performance of multimodal large language models (MLLMs) across a range of complex reasoning tasks. Despite its popularity, there is a notable absence of automated methods for evaluating the quality of reasoning…

2025

Towards Irreversible Attack: Fooling Scene Text Recognition via Multi-Population Coevolution Search

NeurIPS 2025poster

Recent work has shown that scene text recognition (STR) models are vulnerable to adversarial examples. Different from non-sequential vision tasks, the output sequence of STR models contains rich information. However, existing adversarial attacks against STR models can only lead to a few incorrect c…

Cited by 0SourcecodeScholar
2025

UniMotion: A Unified Motion Framework for Simulation, Prediction and Planning

NeurIPS 2025poster

Motion simulation, prediction and planning are foundational tasks in autonomous driving, each essential for modeling and reasoning about dynamic traffic scenarios. While often addressed in isolation due to their differing objectives, such as generating diverse motion states or estimating optimal tra…

Cited by 0SourceScholar
2024

Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction Loss

ICASSP 2024accepted

This research is about the creation of personalized synthetic voices for head and neck cancer survivors. It is focused particularly on tongue cancer patients whose speech might exhibit severe articulation impairment. Our goal is to restore normal articulation in the synthesized speech, while maximal…

Cited by 0SourceScholar
2024

Efficient Black-Box Speaker Verification Model Adaptation With Reprogramming And Backend Learning

ICASSP 2024accepted

The development of deep neural networks (DNN) has significantly enhanced the performance of speaker verification (SV) systems in recent years. However, a critical issue that persists when applying DNN-based SV systems in practical applications is domain mismatch. To mitigate the performance degradat…

Cited by 0SourceScholar
2023

A Simple Vision Transformer for Weakly Semi-supervised 3D Object Detection

ICCV 2023poster

Advanced 3D object detection methods usually rely on large-scale, elaborately labeled datasets to achieve good performance. However, labeling the bounding boxes for the 3D objects is difficult and expensive. Although semi-supervised (SS3D) and weakly-supervised 3D object detection (WS3D) methods can…

Cited by 29PDFScholar
2023

Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification

ICASSP 2023accepted

Deep convolutional neural networks (CNNs) have been applied to extracting speaker embeddings with significant success in speaker verification. Incorporating the attention mechanism has shown to be effective in improving the model performance. This paper presents an efficient two-dimensional convolut…

Cited by 0SourceScholar
2023

DDS3D: Dense Pseudo-Labels with Dynamic Threshold for Semi-Supervised 3D Object Detection

ICRA 2023poster

In this paper, we present a simple yet effective semi-supervised 3D object detector named DDS3D. Our main contributions have two-fold. On the one hand, different from previous works using Non-Maximal Suppression (NMS) or its variants for obtaining the sparse pseudo labels, we propose a dense pseudo-…

Cited by 19SourcecodeScholar
2023

SOOD: Towards Semi-Supervised Oriented Object Detection

CVPR 2023poster

Semi-Supervised Object Detection (SSOD), aiming to explore unlabeled data for boosting object detectors, has become an active task in recent years. However, existing SSOD approaches mainly focus on horizontal objects, leaving multi-oriented objects that are common in aerial images unexplored. This p…

2022

ER-SAN: Enhanced-Adaptive Relation Self-Attention Network for Image Captioning

IJCAI 2022poster

Image captioning (IC), bringing vision to language, has drawn extensive attention. Precisely describing visual relations between image objects is a key challenge in IC. We argue that the visual relations, that is geometric positions (i.e., distance and size) and semantic interactions (i.e., actions…

2019

Differentiable Learning-to-Group Channels via Groupable Convolutional Neural Networks

ICCV 2019poster

Group convolution, which divides the channels of ConvNets into groups, has achieved impressive improvement over the regular convolution operation. However, existing models, e.g. ResNext, still suffers from the sub-optimal performance due to manually defining the number of groups as a constant over a…

Cited by 49PDFScholar
2019

Differentiable Learning-to-Normalize via Switchable Normalization

ICLR 2019poster

We address a learning-to-normalize problem by proposing Switchable Normalization (SN), which learns to select different normalizers for different normalization layers of a deep neural network. SN employs three distinct scopes to compute statistics (means and variances) including a channel, a layer,…

Cited by 262SourcePDFScholar
2019

SSN: Learning Sparse Switchable Normalization via SparsestMax

CVPR 2019poster

Normalization methods improve both optimization and generalization of ConvNets. To further boost performance, the recently-proposed switchable normalization (SN) provides a new perspective for deep learning: it learns to select different normalizers for different convolution layers of a ConvNet. How…

Cited by 71PDFcodeScholar
2016

AnalogCast: Full linear coding and pseudo analog transmission for satellite remote-sensing images

ICASSP 2016accepted

In this paper, we propose a novel image coding and transmission scheme called AnalogCast, which is a pseudo analog coding system for transmitting satellite remote-sensing images to large number of receivers. AnalogCast follows the idea originally developed for SoftCast [1-3] but with two special tec…

Cited by 0SourceScholar