← Search

hang li

53 accepted papers

2026

CoLoR: The Devil is in Scene Coordinate Regression for Large-Scale Visual Localization

CVPR 2026

Scene Coordinate Regression (SCR) has emerged as a memory-efficient paradigm for visual localization. While SCR has demonstrated performance comparable to classic feature matching based approaches in small-scale scenes, it has consistently underperformed in large-scale environments. Large-scale loca

Cited by 0SourceScholar
2026

ReVeal: Self-Evolving Code Agents via Reliable Self-Verification

ICLR 2026poster

Reinforcement learning with verifiable rewards (RLVR) has advanced the reasoning capabilities of large language models. Howerer, existing methods rely solely on outcome rewards, without explicitly optimizing verification or leveraging reliable signals from realistic environments, leading to unreliab…

Cited by 0SourceScholar
2026

Scaling Zero-Shot Reference-to-Video Generation

CVPR 2026

Reference-to-video (R2V) generation aims to synthesize videos that align with a text prompt while preserving the subject identity from reference images. However, current R2V methods are hindered by the reliance on explicit reference image-video-text triplets, whose construction is highly expensive a

Cited by 0SourcecodeScholar
2026

Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory

ICLR 2026poster

We introduce M3-Agent, a novel multimodal agent framework equipped with long-term memory. Like humans, M3-Agent can process real-time visual and auditory inputs to build and update episodic and semantic memories, gradually accumulating world knowledge. Its memory is organized in an entity-centric, m…

Cited by 0SourcecodeScholar
2026

TacTape: Real-Time High-Accuracy Tactile Fiducial System with Structured 3D Texture for Vision-Based Tactile Sensors

ICRA 2026poster

Vision-based tactile sensors enable high-resolution tactile perception by capturing image-based contact data. However, their utility in tactile localization is limited by their inherently small and local sensing area, as well as their dependence on distinct object surface features. We propose TacTap…

Cited by 0Scholar
2026

Towards Universal Gene Regulatory Network Inference: Unlocking Generalizable Regulatory Knowledge in Single-cell Foundation Models

ICML 2026poster

Gene Regulatory Network (GRN) inference is essential for understanding complex cellular mechanisms, rendered tractable through single-cell transcriptomic data. With the emergence of single-cell Foundation Models (scFMs), enhanced transcriptomic encoding is widely expected to revolutionize GRN infere…

Cited by 0SourceScholar
2025

Ask-Before-Detection: Identifying and Mitigating Conformity Bias in LLM-Powered Error Detector for Math Word Problem Solutions

ACL 2025long

The rise of large language models (LLMs) offers new opportunities for automatic error detection in education, particularly for math word problems (MWPs). While prior studies demonstrate the promise of LLMs as error detectors, they overlook the presence of multiple valid solutions for a single MWP. O…

Cited by 0SourcePDFScholar
2025

FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models

CVPR 2025poster

One-Shot Federated Learning (OSFL), a special decentralized machine learning paradigm, has recently gained significant attention. OSFL requires only a single round of client data or model upload, which reduces communication costs and mitigates privacy threats compared to traditional FL. Despite thes…

2025

Knowledge Tagging with Large Language Model Based Multi-Agent System

AAAI 2025technical

Knowledge tagging for questions is vital in modern intelligent educational applications, including learning progress diagnosis, practice question recommendations, and course content organization. Traditionally, these annotations have been performed by pedagogical experts, as the task demands not onl…

Cited by 1SourcePDFScholar
2025

Language-Guided Object-Centric Diffusion Policy for Generalizable and Collision-Aware Manipulation

ICRA 2025

Learning from demonstrations faces challenges in generalizing beyond the training data and often lacks collision awareness. This paper introduces Lan-o3dp, a language-guided object-centric diffusion policy framework that can adapt to unseen situations such as cluttered scenes, shifting camera views,

Cited by 8SourceScholar
2025

Learning Flow Fields in Attention for Controllable Person Image Generation

CVPR 2025poster

Controllable person image generation aims to generate a person image conditioned on reference images, allowing precise control over the person's appearance or pose.However, prior methods often distort fine-grained textural details from the reference image, despite achieving high overall image qualit…

2025

MIMO: A Medical Vision Language Model with Visual Referring Multimodal Input and Pixel Grounding Multimodal Output

CVPR 2025poster

Currently, medical vision language models are widely used in medical vision question answering tasks. However, existing models are confronted with two issues: for input, the model only relies on text instructions and lacks direct understanding of visual clues in the image; for output, the model only…

2025

Multi-Scale Conditional Generative Adversarial Networks for Wind Speed Data Imputation in Earthen Ruins Protection

ICASSP 2025accepted

Time-series data are vital for preserving earthen ruins and evaluating wind erosion effects. Harsh conditions at these sites often lead to sensor degradation and significant data gaps. To tackle wind speed data imputation for such environments, we introduce a Multi-Scale Conditional Generative Adver…

Cited by 0SourceScholar
2025

Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation

RA-L 2025

In real-world scenarios, many robotic manipulation tasks are hindered by occlusions and limited fields of view, posing significant challenges for passive observation-based models that rely on fixed or wrist-mounted cameras. In this letter, we investigate the problem of robotic manipulation under lim

Cited by 12SourceScholar
2025

PaSa: An LLM Agent for Comprehensive Academic Paper Search

ACL 2025long

We introduce PaSa, an advanced Paper Search agent powered by large language models. PaSa can autonomously make a series of decisions, including invoking search tools, reading papers, and selecting relevant references, to ultimately obtain comprehensive and accurate results for complex scholar querie…

2025

Robust Multi-bit Text Watermark with LLM-based Paraphrasers

ICML 2025poster

We propose an imperceptible multi-bit text watermark embedded by paraphrasing with LLMs. We fine-tune a pair of LLM paraphrasers that are designed to behave differently so that their paraphrasing difference reflected in the text semantics can be identified by a trained decoder. To embed our multi-bi…

2025

Toward Optimal LLM Alignments Using Two-Player Games

EMNLP 2025

Alignment of large language models (LLM) is a process that ensures the model’s responses to user prompts align with human intentions and social values. This optimization typically relies on pre-collected prompts. The collection of these prompts often either requires careful human interventions or pr

2024

AGILE: A Novel Reinforcement Learning Framework of LLM Agents

NeurIPS 2024poster

We introduce a novel reinforcement learning framework of LLM agents named AGILE (AGent that Interacts and Learns from Environments) designed to perform complex conversational tasks with users, leveraging LLMs, memory, tools, and interactions with experts. The agent possesses capabilities beyond con…

2024

Are Large Language Models (LLMs) Good Social Predictors?

EMNLP 2024finding

With the recent advancement of Large Language Models (LLMs), efforts have been made to leverage LLMs in crucial social science study methods, including predicting human features of social life such as presidential voting. Existing works suggest that LLMs are capable of generating human-like response…

Cited by 9SourcePDFScholar
2024

Boximator: Generating Rich and Controllable Motions for Video Synthesis

ICML 2024poster

Generating rich and controllable motion is a pivotal challenge in video synthesis. We propose *Boximator*, a new approach for fine-grained motion control. Boximator introduces two constraint types: *hard box* and *soft box*. Users select objects in the conditional frame using hard boxes and then use…

Cited by 50SourcePDFScholar
2024

Flight Structure Optimization of Modular Reconfigurable UAVs

IROS 2024poster

This paper presents a Genetic Algorithm (GA) designed to reconfigure a large group of modular Unmanned Aerial Vehicles (UAVs), each with different weights and inertia parameters, into an over-actuated flight structure with improved dynamic properties. Previous research efforts either utilized expert…

Cited by 9SourceScholar
2024

MLeVLM: Improve Multi-level Progressive Capabilities based on Multimodal Large Language Model for Medical Visual Question Answering

ACL 2024findings

Medical visual question answering (MVQA) requires in-depth understanding of medical images and questions to provide reliable answers. We summarize multi-level progressive capabilities that models need to focus on in MVQA: recognition, details, diagnosis, knowledge, and reasoning. Existing MVQA model…

2024

Make Pixels Dance: High-Dynamic Video Generation

CVPR 2024poster

Creating high-dynamic videos such as motion-rich actions and sophisticated visual effects poses a significant challenge in the field of artificial intelligence. Unfortunately current state-of-the-art video generation methods primarily focusing on text-to-video generation tend to produce video clips…

Cited by 102SourcePDFScholar
2024

ReFT: Reasoning with Reinforced Fine-Tuning

ACL 2024long

One way to enhance the reasoning capability of Large Language Models (LLMs) is to conduct Supervised Fine-Tuning (SFT) using Chain-of-Thought (CoT) annotations. This approach does not show sufficiently strong generalization ability, however, because the training only relies on the given CoT data. In…

2024

Real-time Dynamic-consistent Motion Planning for Over-actuated UAVs

ICRA 2024poster

Existing motion planning approaches for over-actuated unmanned aerial vehicle (UAV) platforms can achieve online planning without considering dynamics. However, in many envisioned application areas such as aerial manipulation, payload delivery, and moving target tracking, it is critical to ensure dy…

Cited by 4SourceScholar
2024

Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation

CVPR 2024poster

Diffusion-based models have gained significant popularity for text-to-image generation due to their exceptional image-generation capabilities. A risk with these models is the potential generation of inappropriate content such as biased or harmful images. However the underlying reasons for generating…

2024

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

ICLR 2024poster

Generative pre-trained models have demonstrated remarkable effectiveness in language and vision domains by learning useful representations. In this paper, we extend the scope of this effectiveness by showing that visual robot manipulation can significantly benefit from large-scale video generative p…

2024

Vision-Language Foundation Models as Effective Robot Imitators

ICLR 2024spotlight

Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipulation. We seek a straightforward way of making use of existing vision-language models (VLMs) with simple fine-tuning on…

Cited by 133SourcePDFScholar
2023

Aggregating Single-Wheeled Mobile Robots for Omnidirectional Movements

IROS 2023poster

This paper presents a novel modular robot system that can self-reconfigure to achieve omnidirectional movements for collaborative object transportation. Each robotic module is equipped with a steerable omni-wheel for navigation and is shaped as a regular icositetragon with a permanent magnet install…

Cited by 2SourceScholar
2023

Generative Diffusion Models on Graphs: Methods and Applications

IJCAI 2023poster

Diffusion models, as a novel generative paradigm, have achieved remarkable success in various image generation tasks such as image inpainting, image-to-text translation, and video generation. Graph generation is a crucial computational task on graphs with numerous real-world applications. It aims to…

2023

Sequential Manipulation Planning for Over-Actuated Unmanned Aerial Manipulators

IROS 2023poster

We investigate the sequential manipulation planning problem for unmanned aerial manipulators (UAMs). Unlike prior work that primarily focuses on one-step manipulation tasks, sequential manipulations require coordinated motions of a UAM's floating base, the manipulator, and the object being manipulat…

Cited by 17SourceScholar
2023

Uncertainty-Aware Instance Reweighting for Off-Policy Learning

NeurIPS 2023poster

Off-policy learning, referring to the procedure of policy optimization with access only to logged feedback data, has shown importance in various important real-world applications, such as search engines and recommender systems. While the ground-truth logging policy is usually unknown, previous work…

Cited by 7SourcePDFScholar
2023

Weak Proxies are Sufficient and Preferable for Fairness with Missing Sensitive Attributes

ICML 2023poster

Evaluating fairness can be challenging in practice because the sensitive attributes of data are often inaccessible due to privacy constraints. The go-to approach that the industry frequently adopts is using off-the-shelf proxy models to predict the missing sensitive attributes, e.g. Meta (Alao et al…

2022

A Neural-Symbolic Approach to Natural Language Understanding

EMNLP 2022finding

Deep neural networks, empowered by pre-trained language models, have achieved remarkable results in natural language understanding (NLU) tasks. However, their performances can drastically deteriorate when logical reasoning is needed. This is because NLU in principle depends on not only analogical re…

2022

Directed Acyclic Transformer for Non-Autoregressive Machine Translation

ICML 2022spotlight

Non-autoregressive Transformers (NATs) significantly reduce the decoding latency by generating all tokens in parallel. However, such independent predictions prevent NATs from capturing the dependencies between the tokens for generating multiple possible translations. In this paper, we propose Direct…

2022

Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

ICML 2022spotlight

Most existing methods in vision language pre-training rely on object-centric features extracted through object detection and make fine-grained alignments between the extracted features and texts. It is challenging for these methods to learn relations among multiple objects. To this end, we propose a…

2022

Self-Supervised Audio-and-Text Pre-training with Extremely Low-Resource Parallel Data

AAAI 2022technical

Multimodal pre-training for audio-and-text has recently been proved to be effective and has significantly improved the performance of many downstream speech understanding tasks. However, these state-of-the-art pre-training audio-text models work well only when provided with large amount of parallel…

2021

CTAL: Pre-training Cross-modal Transformer for Audio-and-Language Representations

EMNLP 2021main

Existing audio-language task-specific predictive approaches focus on building complicated late-fusion mechanisms. However, these models are facing challenges of overfitting with limited labels and low model generalization abilities. In this paper, we present a Cross-modal Transformer for Audio-and-L…

2021

Disentangled Contrastive Learning on Graphs

NeurIPS 2021poster

Recently, self-supervised learning for graph neural networks (GNNs) has attracted considerable attention because of their notable successes in learning the representation of graph-structure data. However, the formation of a real-world graph typically arises from the highly complex interaction of man…

Cited by 114SourcePDFScholar
2021

Mathematical Word Problem Generation from Commonsense Knowledge Graph and Equations

EMNLP 2021main

There is an increasing interest in the use of mathematical word problem (MWP) generation in educational assessment. Different from standard natural question generation, MWP generation needs to maintain the underlying mathematical operations between quantities and variables, while at the same time en…

2021

Secoco: Self-Correcting Encoding for Neural Machine Translation

EMNLP 2021finding

This paper presents Self-correcting Encoding (Secoco), a framework that effectively deals with noisy input for robust neural machine translation by introducing self-correcting predictors. Different from previous robust approaches, Secoco enables NMT to explicitly correct noisy inputs and delete spec…

2020

Multimodal Learning for Classroom Activity Detection

ICASSP 2020accepted

Classroom activity detection (CAD) focuses on accurately classifying whether the teacher or student is speaking and recording both the length of individual utterances during a class. A CAD solution helps teachers get instant feedback on their pedagogical instructions. This greatly improves educators…

Cited by 0SourceScholar
2018

Streaming Influence Maximization in Social Networks Based on Multi-Action Credit Distribution

ICASSP 2018accepted

In a social network, influence maximization is the problem of identifying a set of users that own the maximum influence ability across the network. In this paper, a novel credit distribution (CD) based model, termed as the multi-action CD (mCD) model, is introduced to quantify the influence ability…

Cited by 0SourceScholar
2017

Coupling Distributed and Symbolic Execution for Natural Language Queries

ICML 2017poster

Building neural networks to query a knowledge base (a table) with natural language is an emerging research topic in deep learning. An executor for table querying typically requires multiple steps of execution because queries may have complicated structures. In previous studies, researchers have deve…

Cited by 52SourcePDFScholar
2016

Dropped pronoun generation for dialogue machine translation

ICASSP 2016accepted

Dropped pronoun (DP) is a common problem in dialogue machine translation, in which pronouns are frequently dropped in the source sentence and thus are missing in its translation. In response to this problem, we propose a novel approach to improve the translation of DPs for dialogue machine translati…

Cited by 0SourceScholar