← Search

Dan Meng

16 accepted papers

2026

Aligning Cross-View Visual Geometries in LVLMs Through Human-Like Reasoning Learning

AAAI 2026technical

Spatial understanding is a critical capability for LVLMs (Large Vision-Language Models) to advance embodied AI applications. Existing works primarily focus on enhancing spatial understanding within a single frame, i.e., injecting 3D spatial concepts into LVLMs under single coordinate system. However

Cited by 0SourcePDFScholar
2026

JoPPO: Hierarchical Photography Assessment via Contrastive Joint Conditional Probabilistic Reinforcement Learning

CVPR 2026

With the advancement of Vision-Language Models (VLMs), employing VLM-as-a-Judge for visual evaluation has become a widely adopted metric in vision research. However, existing VLM-as-a-Judge approaches suffer from biased scoring outcomes with low discrimination and lack the capacity for unified multi

Cited by 0SourcecodeScholar
2026

Sketch-Based Low-Rank Model Merging with Shared Circulant Transforms

ICML 2026poster

Merging multiple low-rank adapters (LoRA) provides a practical route to scaling multi-task learning and deployment more efficiently than full-model weight merging, while avoiding reliance on task-specific training data. However, most existing approaches either treat LoRA updates as dense weight delt…

Cited by 0SourceScholar
2025

Attention to Trajectory: Trajectory-Aware Open-Vocabulary Tracking

ICCV 2025poster

Open-Vocabulary Multi-Object Tracking (OV-MOT) aims to enable approaches to track objects without being limited to a predefined set of categories. Current OV-MOT methods typically rely primarily on instance-level detection and association, often overlooking trajectory information that is unique and…

2025

Jack of All Trades, Master of None: PMP-Guided Adaptive Multi-Teacher Distillation with Meta-Learning

ICASSP 2025accepted

To enhance the robustness and accuracy of the small model, existing approaches combine adversarial training with knowledge distillation, introducing a comprehensive single-teacher model to improve the performance of the student model (small model). However, due to the limited knowledge of a teacher…

Cited by 0SourceScholar
2025

LoRATEE: A Secure and Efficient Inference Framework for Multi-Tenant LoRA LLMs Based on TEE

ICASSP 2025accepted

Low-Rank Adaptation (LoRA) is a parameter-efficient fine-tuning approach that adaptes pre-trained Large Language Models (LLMs) to multi-tenant tasks by generating a variety of LoRA adapters. However, this approach faces significant security challenges and is particularly susceptible to malicious ser…

Cited by 0SourceScholar
2025

MsRAG: Knowledge Augumented Image Captioning with Object-level Multi-source RAG

IJCAI 2025

Language-Visual Large Models (LVLMs) have made significant strides in enhancing visual understanding capabilities. However, these models often struggle with knowledge-based visual tasks due to constrains in their pre-training data scope and timeliness. Existing Retrieval-Augmented Generation (RAG) m

Cited by 0SourcePDFScholar
2025

RanDoctor: System-Level Ransomware Detection with ProbSparse Self-Attention

ICASSP 2025accepted

Ransomware attacks pose significant threats and have caused substantial economic losses across various industries worldwide. Existing defense mechanisms typically focus on detecting ransomware in environments free from interference by other legitimate programs. However, in real-world applications, r…

Cited by 0SourceScholar
2024

An Evaluation Mechanism of LLM-based Agents on Manipulating APIs

EMNLP 2024finding

LLM-based agents can greatly extend the abilities of LLMs and thus attract sharply increased studies. An ambitious vision – serving users by manipulating massive API-based tools – has been proposed and explored. However, we find a widely accepted evaluation mechanism for generic agents is still miss…

2024

Stealthy Backdoor Attack Towards Federated Automatic Speaker Verification

ICASSP 2024accepted

Automatic speech verification (ASV) authenticates individuals based on distinct vocal patterns, playing a pivotal role in many applications such as voice-based unlocking systems for devices. The ASV system comprises three stages: training, registration, and validation. The model refines using voice…

Cited by 0SourceScholar
2023

UltraRE: Enhancing RecEraser for Recommendation Unlearning via Error Decomposition

NeurIPS 2023poster

With growing concerns regarding privacy in machine learning models, regulations have committed to granting individuals the right to be forgotten while mandating companies to develop non-discriminatory machine learning systems, thereby fueling the study of the machine unlearning problem. Our attentio…

2022

Randomized Sketches for Clustering: Fast and Optimal Kernel $k$-Means

NeurIPS 2022accept

Kernel $k$-means is arguably one of the most common approaches to clustering. In this paper, we investigate the efficiency of kernel $k$-means combined with randomized sketches in terms of both statistical analysis and computational requirements. More precisely, we propose a unified randomized sketc…

Cited by 3SourcePDFScholar