← Search

Qin Li

10 accepted papers

2026

A Deep Learning Model of Mental Rotation Informed by Interactive VR Experiments

ICML 2026poster

Mental rotation—the ability to compare objects seen from different viewpoints—is a fundamental example of mental simulation and spatial world modeling in humans. Here we propose a mechanistic model of human mental rotation, leveraging recent advances in deep, equivariant, and neuro-symbolic learning…

Cited by 0SourceScholar
2026

Maximizing Schatten-p Norm Regularization Toward Balance

AAAI 2026technical

The Schatten-p norm, as a class of structure-inducing norms based on singular values, has been widely used to enhance model low-rankness and representation capability due to its flexibility in structural modeling and favorable mathematical properties. However, its potential in cluster distribution m

Cited by 0SourcePDFScholar
2026

SeD-UD: An Influence-Driven and Hierarchically-Decoupled Information Bottleneck for Multimodal Intent Recognition

CVPR 2026

Multimodal intent recognition (MIR) is hindered by substantial redundancy and noise originating from text, speech, and visual inputs, which weakens feature distinctiveness and ultimately harms recognition performance. Although recent approaches based on the information bottleneck (IB) principle miti

Cited by 0SourcecodeScholar
2026

Tensorized Label Learning via Balanced Tensor Regression

AAAI 2026technical

The multi-view clustering methods based on tensor regression can make full use of the potential structural information between views and achieve data-level fusion. However, existing tensor regression-based approaches for anchor graph often overlook the probabilistic nature of anchor graph, focusing

Cited by 0SourcePDFScholar
2025

SURE: Safety Understanding and Reasoning Enhancement for Multimodal Large Language Models

EMNLP 2025

Multimodal large language models (MLLMs) demonstrate impressive capabilities by integrating visual and textual information. However, the incorporation of visual modalities also introduces new and complex safety risks, rendering even the most advanced models vulnerable to sophisticated jailbreak atta

2024

Beyond MOT: Semantic Multi-Object Tracking

ECCV 2024poster

"Current multi-object tracking (MOT) aims to predict trajectories of targets (, “where”) in videos. Yet, knowing merely “where” is insufficient in many crucial applications. In comparison, semantic understanding such as fine-grained behaviors, interactions, and overall summarized captions (, “what”)…

2021

I2UV-HandNet: Image-to-UV Prediction Network for Accurate and High-Fidelity 3D Hand Mesh Modeling

ICCV 2021poster

Reconstructing a high-precision and high-fidelity 3D human hand from a color image plays a central role in replicating a realistic virtual hand in human-computer interaction and virtual reality applications. Current methods are lacking in accuracy and fidelity due to various hand poses and severe oc…

Cited by 73PDFScholar