← Search

Li Jin

24 accepted papers

2026

CoSMo3D: Open-World Promptable 3D Semantic Segmentation through LLM-Guided Canonical Spatial Modeling

CVPR 2026

Open-world promptable 3D semantic segmentation remains brittle as semantics are inferred in the input sensor coordinates. Yet, humans, in contrast, interpret parts via functional roles in a canonical space -- wings extend laterally, handles protrude to the side, and legs support from below. Psychoph

Cited by 0SourcecodeScholar
2026

Rectify Evaluation Preference: Improving LLMs’ Critique on Math Reasoning via Perplexity-aware Reinforcement Learning

AAAI 2026technical

To improve Multi-step Mathematical Reasoning (MsMR) of Large Language Models (LLMs), it is crucial to obtain scalable supervision from the corpus by automatically critiquing mistakes in the reasoning process of MsMR and rendering a final verdict of the problem-solution. Most existing methods rely on

Cited by 0SourcePDFScholar
2026

Secure Multi-agent Reinforcement Learning for Service Systems with Affinity and Byzantine Nodes: Stability Analysis and Protection Design

ICML 2026poster

We study decentralized multi-agent reinforcement learning (MARL) for networked service systems with affinity in the presence of Byzantine nodes. The way that a server processes a job depends on an affinity state that captures the correlation between the job and the server. Each node learns a local c…

Cited by 0SourceScholar
2025

A Visual Servo System for Robotic on-Orbit Servicing Based on 3D Perception of Non-Cooperative Satellite

ICRA 2025

The 3D perception of satellites, including both their shape and pose, is a key foundation for robotic on-orbit servicing. However, the demanding space environment-such as intense and dim illumination-presents significant challenges. Previous non-cooperative methods focus on specific geometric featur

Cited by 0SourceScholar
2025

Chain-of-Specificity: Enhancing Task-Specific Constraint Adherence in Large Language Models

COLING 2025main

Large Language Models (LLMs) exhibit remarkable generative capabilities, enabling the generation of valuable information. Despite these advancements, previous research found that LLMs sometimes struggle with adhering to specific constraints, such as being in a specific place or at a specific time, a…

Cited by 1SourcePDFScholar
2025

HyperMixer: Specializable Hypergraph Channel Mixing for Long-term Multivariate Time Series Forecasting

AAAI 2025technical

Long-term Multivariate Time Series (LMTS) forecasting aims to predict extended future trends based on channel-interrelated historical data. Considering the elusive channel correlations, most existing methods compromise by treating channels as independent or tentatively modeling pairwise channel int…

Cited by 0SourcePDFScholar
2025

One-shot 3D Object Canonicalization based on Geometric and Semantic Consistency

CVPR 2025highlight

3D object canonicalization is a fundamental task, essential for various downstream tasks. Existing methods rely on either cumbersome manual processes or priors learned from extensive, per-category training samples. Real-world datasets, however, often exhibit long-tail distributions, challenging exis…

2025

PIPER: Benchmarking and Prompting Event Reasoning Boundary of LLMs via Debiasing-Distillation Enhanced Tuning

ACL 2025long

While Large Language Models (LLMs) excel in diverse domains, their validity in event reasoning remains underexplored. Most existing works merely stagnate at assessing LLMs’ event reasoning with a single event relational type or reasoning format, failing to conduct a complete evaluation and provide a…

Cited by 0SourcePDFScholar
2025

Prior-free 3D Object Tracking

CVPR 2025highlight

In this paper, we introduce a novel, truly prior-free 3D object tracking method that operates without given any model or training priors. Unlike existing methods that typically require pre-defined 3D models or specific training datasets as priors, which limit their applicability, our method is free…

2025

P²Net: Parallel Pointer-based Network for Key Information Extraction with Complex Layouts

ACL 2025finding

Key Information Extraction (KIE) is a challenging multimodal task aimed at extracting structured value entities from visually rich documents. Despite recent advancements, two major challenges remain. First, existing datasets typically feature fixed layouts and a limited set of entity categories, whi…

Cited by 0SourcePDFScholar
2024

CAMEL: Capturing Metaphorical Alignment with Context Disentangling for Multimodal Emotion Recognition

AAAI 2024technical

Understanding the emotional polarity of multimodal content with metaphorical characteristics, such as memes, poses a significant challenge in Multimodal Emotion Recognition (MER). Previous MER researches have overlooked the phenomenon of metaphorical alignment in multimedia content, which involves n…

Cited by 9SourcePDFScholar
2024

GOME: Grounding-based Metaphor Binding With Conceptual Elaboration For Figurative Language Illustration

EMNLP 2024main

The illustration or visualization of figurative language, such as linguistic metaphors, is an emerging challenge for existing Large Language Models (LLMs) and multimodal models. Due to their comparison of seemingly unrelated concepts in metaphors, existing LLMs have a tendency of over-literalization…

Cited by 0SourcePDFScholar
2024

Implicit Coarse-to-Fine 3D Perception for Category-level Object Pose Estimation from Monocular RGB Image

ICRA 2024poster

Category-level object pose estimation demonstrates robust generalization capabilities that benefit robotics applications. However, exclusive reliance on RGB images without leveraging any 3D information introduces ambiguity in the translation and size of objects, leading to suboptimal performance. In…

Cited by 0SourceScholar
2024

Rethinking the Reversal Curse of LLMs: a Prescription from Human Knowledge Reversal

EMNLP 2024main

Large Language Models (LLMs) have exhibited exceptional performance across diverse domains. However, recent studies reveal that LLMs are plagued by the “reversal curse”. Most existing methods rely on aggressive sample permutation and pay little attention to delving into the underlying reasons for th…

Cited by 4SourcePDFScholar
2024

Semi-Decoupled 6D Pose Estimation via Multi-Modal Feature Fusion

ICASSP 2024accepted

The existing methods for 6D pose estimation based on RGB-D employ RGB images and observed point cloud derived from depth maps as input, then concurrently predicting both rotation and translation. However, rotation and translation possess distinct characteristics and scale ranges, and their simultane…

Cited by 0SourceScholar
2024

V2X-Real: a Largs-Scale Dataset for Vehicle-to-Everything Cooperative Perception

ECCV 2024poster

"Recent advancements in Vehicle-to-Everything (V2X) technologies have enabled autonomous vehicles to share sensing information to see through occlusions, greatly boosting the perception capability. However, there are no real-world datasets to facilitate the real V2X cooperative perception research –…

2024

Video Event Extraction with Multi-View Interaction Knowledge Distillation

AAAI 2024technical

Video event extraction (VEE) aims to extract key events and generate the event arguments for their semantic roles from the video. Despite promising results have been achieved by existing methods, they still lack an elaborate learning strategy to adequately consider: (1) inter-object interaction, whi…

Cited by 2SourcePDFScholar
2023

Event Causality Extraction via Implicit Cause-Effect Interactions

EMNLP 2023long main

Event Causality Extraction (ECE) aims to extract the cause-effect event pairs from the given text, which requires the model to possess a strong reasoning ability to capture event causalities. However, existing works have not adequately exploited the interactions between the cause and effect event th…

Cited by 0SourceScholar
2023

Guide the Many-to-One Assignment: Open Information Extraction via IoU-aware Optimal Transport

ACL 2023long

Open Information Extraction (OIE) seeks to extract structured information from raw text without the limitations of close ontology. Recently, the detection-based OIE methods have received great attention from the community due to their parallelism. However, as the essential step of those models, how…

Cited by 14SourcePDFScholar
2023

Narrative Order Aware Story Generation via Bidirectional Pretraining Model with Optimal Transport Reward

EMNLP 2023long findings

To create a captivating story, a writer often plans a sequence of logically coherent events and ingeniously manipulates the narrative order to generate flashback in place. However, existing storytelling systems suffer from both insufficient understanding of event correlations and inadequate awarenes…

Cited by 0SourceScholar
2023

Online Hand-Eye Calibration with Decoupling by 3D Textureless Object Tracking

ICRA 2023poster

Hand-eye calibration estimates the pose of a camera relative to a robot, which is a fundamental problem for visually guided robots, especially for dynamic object grasping. Most methods use 2D fiducial markers with distinctive visual features and require pre-calibration for accurate calibration, whic…

Cited by 2SourceScholar
2023

TOT:Topology-Aware Optimal Transport for Multimodal Hate Detection

AAAI 2023technical

Multimodal hate detection, which aims to identify the harmful content online such as memes, is crucial for building a wholesome internet environment. Previous work has made enlightening exploration in detecting explicit hate remarks. However, most of their approaches neglect the analysis of implicit…

Cited by 14SourcePDFScholar
2022

Assist Non-native Viewers: Multimodal Cross-Lingual Summarization for How2 Videos

EMNLP 2022main

Multimodal summarization for videos aims to generate summaries from multi-source information (videos, audio transcripts), which has achieved promising progress. However, existing works are restricted to monolingual video scenarios, ignoring the demands of non-native video viewers to understand the c…

2021

Trigger is Not Sufficient: Exploiting Frame-aware Knowledge for Implicit Event Argument Extraction

ACL 2021long

Implicit Event Argument Extraction seeks to identify arguments that play direct or implicit roles in a given event. However, most prior works focus on capturing direct relations between arguments and the event trigger. The lack of reasoning ability brings many challenges to the extraction of implici…

Cited by 76SourcePDFScholar