← Search

Lin Sun

29 accepted papers

2026

AgentPO: Enhancing Multi-Agent Collaboration via Reinforcement Learning

ICLR 2026poster

Multi-Agent Systems (MAS) offer a powerful paradigm for solving complex problems through distributed reasoning and collaboration. However, their effectiveness is often hindered by the challenge of optimizing interactions among agents. To address this, we introduce AgentPO, a novel framework that dir…

Cited by 0SourceScholar
2026

Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training

AAAI 2026technical

Current methods for content safety in Large Language Models (LLMs), such as Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often rely on multi-stage training pipelines and lack fine-grained, post-deployment controllability. To address these limitations, we propos

Cited by 0SourcePDFScholar
2026

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

ICLR 2026poster

Temporal context is essential for robotic manipulation because such tasks are inherently non-Markovian, yet mainstream VLA models typically overlook it and struggle with long-horizon, temporally dependent tasks. Cognitive science suggests that humans rely on working memory to buffer short-lived repr…

Cited by 0SourcecodeScholar
2025

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation

ICCV 2025poster

Contrastive Language-Image Pre-training (CLIP) exhibits strong zero-shot classification ability on image-level tasks, leading to the research to adapt CLIP for open-vocabulary semantic segmentation without training. The key is to improve spatial representation of image-level CLIP, such as replacing…

2025

Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision

EMNLP 2025

Recent advances in Large Language Models (LLMs) have highlighted the challenge of handling long-context tasks, where models need to reason over extensive input contexts to aggregate target information. While Chain-of-Thought (CoT) prompting has shown promise for multi-step reasoning, its effectivene

Cited by 0SourcePDFScholar
2025

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules

AAAI 2025technical

Distinguishing spatial relations is a basic part of human cognition which requires fine-grained perception on cross-instance. Although benchmarks like MME, MMBench and SEED comprehensively have evaluated various capabilities which already include visual spatial reasoning(VSR). There is still a la…

2025

Large Language Models Badly Generalize across Option Length, Problem Types, and Irrelevant Noun Replacements

EMNLP 2025

In this paper, we propose a “Generalization Stress Test” to assess Large Language Models’ (LLMs) generalization ability under slight and controlled perturbations, including option length, problem types, and irrelevant noun replacements. We achieve novel and significant findings that, despite high be

Cited by 0SourcePDFScholar
2025

LongAttn: Selecting Long-context Training Data via Token-level Attention

ACL 2025finding

With the development of large language models (LLMs), there has been an increasing need for significant advancements in handling long contexts. To enhance long-context capabilities, constructing high-quality training data with **long-range dependencies** is crucial. Existing methods to select long-c…

2025

TACLR: A Scalable and Efficient Retrieval-based Method for Industrial Product Attribute Value Identification

ACL 2025long

Product Attribute Value Identification (PAVI) involves identifying attribute values from product profiles, a key task for improving product search, recommendation, and business analytics on e-commerce platforms.However, existing PAVI methods face critical challenges, such as inferring implicit value…

2024

Aegis:An Advanced LLM-Based Multi-Agent for Intelligent Functional Safety Engineering

EMNLP 2024industry

Functional safety is a critical aspect of automotive engineering, encompassing all phases of a vehicle’s lifecycle, including design, development, production, operation, and decommissioning. This domain involves highly knowledge-intensive tasks. This paper introduces Aegis: An Advanced LLM-Based Mul…

Cited by 1SourcePDFScholar
2024

Multi-View Point Cloud Registration Based on Improved NDT Algorithm and ODM Optimization Method

RA-L 2024

The acquisition of targets' complete point cloud model is crucial for tasks such as 3D reconstruction and disordered grasping. Shooting targets from multiple perspectives and registering point clouds from different perspectives can obtain a relatively complete point cloud model. However, small scene

Cited by 8SourceScholar
2024

PathAsst: A Generative Foundation AI Assistant towards Artificial General Intelligence of Pathology

AAAI 2024technical

As advances in large language models (LLMs) and multimodal techniques continue to mature, the development of general-purpose multimodal large language models (MLLMs) has surged, offering significant applications in interpreting natural images. However, the field of pathology has largely remained unt…

2024

UMIE: Unified Multimodal Information Extraction with Instruction Tuning

AAAI 2024technical

Multimodal information extraction (MIE) gains significant attention as the popularity of multimedia content increases. However, current MIE methods often resort to using task-specific model structures, which results in limited generalizability across tasks and underutilizes shared knowledge across M…

2023

PAGE: A Position-Aware Graph-Based Model for Emotion Cause Entailment in Conversation

ICASSP 2023accepted

Conversational Causal Emotion Entailment (C<inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</inf>E<inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</inf>) is a task that aims at recognizing the causes corr…

Cited by 0SourceScholar
2022

Stochastic Consensus: Enhancing Semi-Supervised Learning with Consistency of Stochastic Classifiers

ECCV 2022poster

"Semi-supervised learning (SSL) has achieved new progress recently with the emerging framework of self-training deep networks, where the criteria for selection of unlabeled samples with pseudo labels play a key role in the empirical success. In this work, we propose such a new criterion based on con…

Cited by 7SourcePDFScholar
2022

VISTA: Boosting 3D Object Detection via Dual Cross-VIew SpaTial Attention

CVPR 2022poster

Detecting objects from LiDAR point clouds is of tremendous significance in autonomous driving. In spite of good progress, accurate and reliable 3D detection is yet to be achieved due to the sparsity and irregularity of LiDAR point clouds. Among existing strategies, multi-view methods have shown grea…

Cited by 100PDFcodeScholar
2021

Improving NER in Social Media via Entity Type-Compatible Unknown Word Substitution

ICASSP 2021accepted

Named entity recognition (NER) is a fundamental task for information extraction (IE), and current state-of-the-art methods try to address this issue and achieve high performance on clean text (e.g., newswire genres). However, most of these algorithms do not generalize well when they transit to the n…

Cited by 0SourceScholar
2021

RpBERT: A Text-image Relation Propagation-based BERT Model for Multimodal NER

AAAI 2021technical

Recently multimodal named entity recognition (MNER) has utilized images to improve the accuracy of NER in tweets. However, most of the multimodal methods use attention mechanisms to extract visual clues regardless of whether the text and image are relevant. Practically, the irrelevant text-image pai…

2020

Every View Counts: Cross-View Consistency in 3D Object Detection with Hybrid-Cylindrical-Spherical Voxelization

NeurIPS 2020poster

Recent voxel-based 3D object detectors for autonomous vehicles learn point cloud representations either from bird eye view (BEV) or range view (RV, a.k.a. the perspective view). However, each view has its own strengths and weaknesses. In this paper, we present a novel framework to unify and leverage…

Cited by 129SourcePDFScholar
2020

Grasp Proposal Networks: An End-to-End Solution for Visual Learning of Robotic Grasps

NeurIPS 2020poster

Learning robotic grasps from visual observations is a promising yet challenging task. Recent research shows its great potential by preparing and learning from large-scale synthetic datasets. For the popular, 6 degree-of-freedom (6-DOF) grasp setting of parallel-jaw gripper, most of existing methods…

2020

Object as Hotspots: An Anchor-Free 3D Object Detection Approach via Firing of Hotspots

ECCV 2020poster

Accurate 3D object detection in LiDAR based point clouds suffers from the challenges of data sparsity and irregularities. Existing methods strive to organize the points regularly, e.g. voxelize, pass them through a designed 2D/3D neural network, and then define object-level anchors that predict offs…

Cited by 211SourcePDFScholar
2020

Probabilistic Multi-modal Trajectory Prediction with Lane Attention for Autonomous Vehicles

IROS 2020poster

Trajectory prediction is crucial for autonomous vehicles. The planning system not only needs to know the current state of the surrounding objects but also their possible states in the future. As for vehicles, their trajectories are significantly influenced by the lane geometry and how to effectively…

Cited by 98SourceScholar
2020

RIVA: A Pre-trained Tweet Multimodal Model Based on Text-image Relation for Multimodal NER

COLING 2020main

Multimodal named entity recognition (MNER) for tweets has received increasing attention recently. Most of the multimodal methods used attention mechanisms to capture the text-related visual information. However, unrelated or weakly related text-image pairs account for a large proportion in tweets. V…

Cited by 35SourcePDFScholar
2017

Lattice Long Short-Term Memory for Human Action Recognition

ICCV 2017poster

Human actions captured in video sequences are three-dimensional signals characterizing visual appearance and motion dynamics. To learn action patterns, existing methods adopt Convolutional and/or Recurrent Neural Networks (CNNs and RNNs). CNN based methods are effective in learning spatial appearanc…

Cited by 232PDFScholar
2015

Human Action Recognition Using Factorized Spatio-Temporal Convolutional Networks

ICCV 2015poster

Human actions in video sequences are three-dimensional (3D) spatio-temporal signals characterizing both the visual appearance and motion dynamics of the involved humans and objects. Inspired by the success of convolutional neural networks (CNN) for image classification, recent attempts have been mad…

Cited by 745PDFScholar