← Search

Liang Zhao

110 accepted papers

2026

CasMoE: A Cascaded Framework for Efficient MoE Inference on Resource-constrained Devices

AAAI 2026technical

The Mixture-of-Experts (MoE) architecture has emerged as a key enabler for scaling large language models (LLMs), empowering increased model capacity with minimal computational overhead through gating-based dynamic expert activation. However, due to the memory demands introduced by expert modules, Mo

Cited by 0SourcePDFScholar
2026

Deep Identification of Propagation Trees in Graph Diffusion

IJCAI 2026

Understanding how information or influence propagates through a network, such as during an epidemic outbreak or the spread of misinformation, is a fundamental yet challenging problem. While prior works have focused on cascade prediction (forecasting future infected nodes), network inference (recover

Cited by 0Scholar
2026

KNNDA: A New Perspective of Alignment Recovery for Partially View-Aligned Clustering

AAAI 2026technical

In multi-view clustering (MVC), complementary and consistent information from multiple views is integrated to improve clustering performance. However, inter-view sample correspondences may be partially missing in practice, making it difficult to learn cross-view consistency, which leads to the parti

Cited by 0SourcePDFScholar
2026

MedVCoT: Bridging the Modality Gap in Medical VQA Through Latent Visual Reasoning

IJCAI 2026

With the rising demand for trustworthy AI in clinical practice, strong interpretability is now a critical requirement as well as accuracy. However, the modality gap for medical visual question answering is quite severe when continuous visual signals are forcibly projected into discrete text space fo

Cited by 0Scholar
2026

Non-Rigid Structure-From-Motion Via Differential Geometry with Recoverable Conformal Scale

ICRA 2026poster

Non-rigid structure-from-motion (NRSfM), a promising technique for addressing the mapping challenges in monocular visual deformable simultaneous localization and mapping (SLAM), has attracted growing attention. We introduce a novel method, called Con-NRSfM, for NRSfM under conformal deformations, en…

2026

PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra

ICLR 2026poster

Current methods for personality control in Large Language Models rely on static prompting or expensive fine-tuning, failing to capture the dynamic and compositional nature of human traits. We introduce PERSONA, a training-free framework that achieves fine-tuning level performance through direct mani…

Cited by 0SourceScholar
2026

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

ICML 2026poster

We introduce the Perception Rubric Benchmark (PRB), a rubric-based evaluation framework for Multimodal Large Language Models (MLLMs) that addresses the growing gap between benchmark scores and human-perceived quality. While standard perception metrics approach saturation, they produce compressed ran…

Cited by 0SourceScholar
2026

RiemanLine: Riemannian Manifold Representation of 3D Lines for Factor Graph Optimization

AAAI 2026technical

Minimal parametrization of 3D lines plays a critical role in camera localization and structural mapping. Existing representations in robotics and computer vision predominantly handle independent lines, overlooking structural regularities such as sets of parallel lines that are pervasive in man-made

Cited by 0SourcePDFScholar
2026

Robust Selective Activation with Randomized Temporal K-Winner-Take-All in Spiking Neural Networks for Continual Learning

ICLR 2026poster

The human brain exhibits remarkable efficiency in processing sequential information, a capability deeply rooted in the temporal selectivity and stochastic competition of neuronal activation. Current continual learning in spiking neural networks (SNNs) faces a critical challenge: balancing task-speci…

Cited by 0SourceScholar
2026

Sample Weighted Incomplete Multimodal Clustering Based on Graph Coarsening Label Extraction

AAAI 2026technical

Multimodal data is typically collected through heterogeneous sensors and processing pipelines. However, due to variations in acquisition environments, device capabilities, and feature extraction methods, such data often suffers from incompleteness and inconsistent quality across modalities. To addre

Cited by 0SourcePDFScholar
2026

Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers

ICML 2026poster

Reinforcement learning (RL) has emerged as a crucial approach for enhancing the capabilities of large language models. However, in Mixture-of-Experts (MoE) models, the routing mechanism often introduces instability, even leading to catastrophic RL training collapse. We analyze the training-inference…

Cited by 0SourceScholar
2026

mHC: Manifold-Constrained Hyper-Connections

ICML 2026spotlight

Recently, studies exemplified by Hyper-Connections (HC) have extended the ubiquitous residual connection paradigm established over the past decade by expanding the residual stream width and diversifying connectivity patterns. While yielding substantial performance gains, this diversification fundame…

Cited by 0SourceScholar
2025

Alleviating Hallucinations from Knowledge Misalignment in Large Language Models via Selective Abstention Learning

ACL 2025long

Large language models (LLMs) are known to suffer from severe hallucination issues. One of the main causes lies in the knowledge misalignment between the pre-training stage and the supervised fine-tuning stage. The unfamiliar knowledge encountered during fine-tuning may encourage LLMs to generate fac…

2025

CLaw: Benchmarking Chinese Legal Knowledge in Large Language Models - A Fine-grained Corpus and Reasoning Analysis

EMNLP 2025

Large Language Models (LLMs) are increasingly tasked with analyzing legal texts and citing relevant statutes, yet their reliability is often compromised by general pre-training that ingests legal texts without specialized focus, obscuring the true depth of their legal knowledge. This paper introduce

Cited by 0SourcePDFScholar
2025

CSF-GAN: Cross-modal Semantic Fusion-based Generative Adversarial Network for Text-guided Image Inpainting

IJCAI 2025

Most visual-guided image inpainting methods based on generative adversarial networks (GANs) struggle when the missing region has weak correlations with the surrounding visual context. Recently, diffusion-based methods guided by textual context have been proposed to address this limitation by leverag

Cited by 0SourcePDFScholar
2025

Consistency-Aware Padding for Incomplete Multi-Modal Alignment Clustering Based on Self-Repellent Greedy Anchor Search

IJCAI 2025

Multi-modal representation is faithful and highly effective in describing real-world data samples' characteristics by describing their complementary information. However, the collected data often exhibits incomplete and misaligned characteristics due to factors such as inconsistent sensor frequencie

2025

Correspondence-Free Multiview Point Cloud Registration via Depth-Guided Joint Optimisation

IROS 2025

Multiview point cloud registration is a fundamental task for constructing globally consistent 3D models. Existing approaches typically rely on feature extraction and data association across multiple point clouds. However, these processes are challenging to obtain global optimal solution in complex e

Cited by 0SourceScholar
2025

Dual Robust Unbiased Multi-View Clustering for Incomplete and Unpaired Information

IJCAI 2025

Recently, multi-view data has gradually attracted attention. However, real-world applications often face Partial View-aligned Problem (PVP) and Partially Sample-missing Problem (PSP) due to data loss or corruption. Existing methods addressing PVP typically focus only on learning from the information

Cited by 0SourcePDFScholar
2025

EDeformNet: Estimating Fishing Net Deformations from Sparse Observations

IROS 2025

This paper introduces EDeformNet, a novel method for real-time 3D reconstruction of fishing nets using sparse positional measurements. Currently, net deployment during large-scale fishing operations is challenging as the submerged lattice deformations that occur in response to the various environmen

Cited by 1SourceScholar
2025

EchoGPT: An Interactive Cardiac Function Assessment Model for Echocardiogram Videos

IJCAI 2025

With the development of wearable cardiac ultrasound devices, it is no longer sufficient to solely rely on doctors for diagnosing long-term echocardiogram videos. Automated diagnosis of echocardiogram videos has now become a research hotspot. Existing studies only analyze echocardiogram video through

2025

FedSpaLLM: Federated Pruning of Large Language Models

NAACL 2025long

Large Language Models (LLMs) achieve state-of-the-art performance but are challenging to deploy due to their high computational and storage demands. Pruning can reduce model size, yet existing methods assume public access to calibration data, which is impractical for privacy-sensitive applications.…

2025

From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems

EMNLP 2025

Research is a fundamental process driving the advancement of human civilization, yet it demands substantial time and effort from researchers. In recent years, the rapid development of artificial intelligence (AI) technologies has inspired researchers to explore how AI can accelerate and enhance rese

Cited by 0SourcePDFScholar
2025

GRAG: Graph Retrieval-Augmented Generation

NAACL 2025findings

Naive Retrieval-Augmented Generation (RAG) focuses on individual documents during retrieval and, as a result, falls short in handling networked documents which are very popular in many applications such as citation graphs, social media, and knowledge graphs. To overcome this limitation, we introduce…

2025

GraphNarrator: Generating Textual Explanations for Graph Neural Networks

ACL 2025long

Graph representation learning has garnered significant attention due to its broad applications in various domains, such as recommendation systems and social network analysis. Despite advancements in graph learning methods, challenges still remain in explainability when graphs are associated with sem…

Cited by 0SourcePDFScholar
2025

Incomplete and Unpaired Multi-View Graph Clustering with Cross-View Feature Fusion

AAAI 2025technical

Due to its effectiveness and efficiency, graph-based multi-view clustering has recently attracted much attention. However, the multi-view data are often incomplete and unpaired in real-world applications as a consequence of data loss or corruption. Although efforts have been made through a series of…

Cited by 0SourcePDFScholar
2025

JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

CVPR 2025poster

We present JanusFlow, a powerful framework that unifies image understanding and generation in a single model.JanusFlow introduces a minimalist architecture that integrates autoregressive language models with rectified flow, a state-of-the-art method in generative modeling.Our key finding demonstrate…

2025

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

ACL 2025long

Long-context modeling is crucial for next-generation language models, yet the high computational cost of standard attention mechanisms poses significant computational challenges. Sparse attention offers a promising direction for improving efficiency while maintaining model capabilities. We present N…

Cited by 0SourcePDFScholar
2025

Neuron Similarity-Based Neural Network Verification via Abstraction and Refinement

IJCAI 2025

Deep neural networks (DNNs) have become integral to numerous safety-critical applications, necessitating rigorous verification of their trustworthiness. However, the problem of verifying DNNs has high computational complexity, and existing techniques have limited efficiency, insufficient to deal wit

Cited by 0SourcePDFScholar
2025

Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning

NeurIPS 2025poster

The remarkable reasoning capability of large language models (LLMs) stems from cognitive behaviors that emerge through reinforcement with verifiable rewards. This work investigates how to transfer this principle to Multimodal LLMs (MLLMs) to unlock advanced visual reasoning. We introduce a two-stage…

Cited by 0SourceScholar
2025

PL-VIWO: A Lightweight and Robust Point-Line Monocular Visual Inertial Wheel Odometry

IROS 2025

This paper presents a novel tightly coupled Filter-based monocular visual-inertial-wheel odometry (VIWO) system for ground robots, designed to deliver accurate and robust localization in long-term complex outdoor navigation scenarios. As an external sensor, the camera enhances localization performan

Cited by 2SourcecodeScholar
2025

Partial-to-Full Registration based on Gradient-SDF for Computer-Assisted Orthopedic Surgery

ICRA 2025

In computer-assisted orthopedic surgery (CAOS), accurate pre-operative to intra-operative bone registration is an essential and critical requirement for providing navigational guidance. This registration process is challenging since the intra-operative 3D points are sparse, only partially overlapped

Cited by 2SourceScholar
2025

Perception-R1: Pioneering Perception Policy with Reinforcement Learning

NeurIPS 2025poster

Inspired by the success of DeepSeek-R1, we explore the potential of rule-based reinforcement learning (RL) in MLLM post-training for perception policy learning. While promising, our initial experiments reveal that incorporating a thinking process through RL does not consistently lead to performance…

Cited by 0SourcecodeScholar
2025

PolyhedronNet: Representation Learning for Polyhedra with Surface-attributed Graph

ICLR 2025poster

Ubiquitous geometric objects can be precisely and efficiently represented as polyhedra. The transformation of a polyhedron into a vector, known as polyhedra representation learning, is crucial for manipulating these shapes with mathematical and statistical tools for tasks like classification, cluste…

2025

Probabilistic Person-in-Bed Detection Using Accelerometer Signals

ICASSP 2025accepted

Using accelerometer data in smart bed systems offers a cost-effective solution for person-in-bed detection. In this work, we propose a lightweight probabilistic model for this task. The accelerometer time series is first divided into multiple patches, with high-frequency noise filtered through a com…

Cited by 0SourceScholar
2025

Self-supervised 3D Reconstruction of Tibia and Fibula from Biplanar X-rays

IROS 2025

With the growing number of patients experiencing knee-related conditions, total knee arthroplasty (TKA) has become a common procedure, where a 3D visualisation of the patient’s tibia and fibula is essential for preoperative planning. Traditional imaging techniques, such as computed tomography (CT),

Cited by 0SourcecodeScholar
2025

Unhackable Temporal Reward for Scalable Video MLLMs

ICLR 2025poster

In the pursuit of superior video-processing MLLMs, we have encountered a perplexing paradox: the “anti-scaling law”, where more data and larger models lead to worse performance. This study unmasks the culprit: “temporal hacking”, a phenomenon where models shortcut by fixating on select frames, missi…

Cited by 0SourcePDFScholar
2024

Advancing Large Language Model Attribution through Self-Improving

EMNLP 2024main

Teaching large language models (LLMs) to generate text with citations to evidence sources can mitigate hallucinations and enhance verifiability in information-seeking systems. However, improving this capability requires high-quality attribution data, which is costly and labor-intensive. Inspired by…

Cited by 6SourcePDFScholar
2024

ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

IJCAI 2024poster

Human-AI interactivity is a critical aspect that reflects the usability of Multimodal Large Language Models (MLLMs). However, existing end-to-end MLLMs only allow users to interact with them through language instructions, leading to the limitation of the interactive accuracy and efficiency. In this…

2024

DreamLLM: Synergistic Multimodal Comprehension and Creation

ICLR 2024spotlight

This paper presents DreamLLM, a learning framework that first achieves versatile Multimodal Large Language Models (MLLMs) empowered with frequently overlooked synergy between multimodal comprehension and creation. DreamLLM operates on two fundamental principles. The first focuses on the generative m…

2024

ELAD: Explanation-Guided Large Language Models Active Distillation

ACL 2024findings

The deployment and application of Large Language Models (LLMs) is hindered by their memory inefficiency, computational demands, and the high costs of API inferences. Traditional distillation methods, which transfer the capabilities of LLMs to smaller models, often fail to determine whether the knowl…

Cited by 6SourcePDFScholar
2024

Grid-based Submap Joining: An Efficient Algorithm for Simultaneously Optimizing Global Occupancy Map and Local Submap Frames

IROS 2024poster

Optimizing robot poses and the map simultaneously has been shown to provide more accurate SLAM results. However, for non-feature based SLAM approaches, directly optimizing all the robot poses and the whole map will greatly increase the computational cost, making SLAM problems difficult to solve in l…

Cited by 0SourceScholar
2024

Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models

ACL 2024long

Though advanced in understanding visual information with human languages, Large Vision-Language Models (LVLMs) still suffer from multimodal hallucinations. A natural concern is that during multimodal interaction, the generated hallucinations could influence the LVLMs’ subsequent generation. Thus, we…

2024

Length Extrapolation of Transformers: A Survey from the Perspective of Positional Encoding

EMNLP 2024finding

Built upon the Transformer, large language models (LLMs) have captured worldwide attention due to their remarkable abilities. Nevertheless, all Transformer-based models including LLMs suffer from a preset length limit and can hardly generalize from short training sequences to longer inference ones,…

Cited by 21SourcePDFScholar
2024

LibMOON: A Gradient-based MultiObjective OptimizatioN Library in PyTorch

NeurIPS 2024poster

Multiobjective optimization problems (MOPs) are prevalent in machine learning, with applications in multi-task learning, learning under fairness or robustness constraints, etc. Instead of reducing multiple objective functions into a scalar objective, MOPs aim to optimize for the so-called Pareto opt…

2024

MIM-Reasoner: Learning with Theoretical Guarantees for Multiplex Influence Maximization

AISTATS 2024poster

Multiplex influence maximization (MIM) asks us to identify a set of seed users such as to maximize the expected number of influenced users in a multiplex network. MIM has been one of central research topics, especially in nowadays social networking landscape where users participate in multiple onlin…

2024

Merlin: Empowering Multimodal LLMs with Foresight Minds

ECCV 2024poster

"Humans can foresee the future based on present observations, a skill we term as foresight minds. However, this capability remains under-explored within existing MLLMs, hindering their capacity to understand intentions behind subjects. To address this, we integrate the future modeling into MLLMs. By…

2024

Open-Structure: Structural Benchmark Dataset for SLAM Algorithms

RA-L 2024

This letter presents Open-Structure, a novel benchmark dataset for evaluating visual odometry and SLAM methods. Compared to existing public datasets that primarily offer raw images, Open-Structure provides direct access to point and line measurements, correspondences, structural associations, and co

Cited by 5SourcecodeScholar
2024

SparseLLM: Towards Global Pruning of Pre-trained Language Models

NeurIPS 2024poster

The transformative impact of large language models (LLMs) like LLaMA and GPT on natural language processing is countered by their prohibitive computational demands. Pruning has emerged as a pivotal compression strategy, introducing sparsity to enhance both memory and computational efficiency. Yet, t…

2024

TEG-DB: A Comprehensive Dataset and Benchmark of Textual-Edge Graphs

NeurIPS 2024poster

Text-Attributed Graphs (TAGs) augment graph structures with natural language descriptions, facilitating detailed depictions of data and their interconnections across various real-world settings. However, existing TAG datasets predominantly feature textual information only at the nodes, with edges ty…

2024

Uncertainty Quantification for In-Context Learning of Large Language Models

NAACL 2024long

In-context learning has emerged as a groundbreaking ability of Large Language Models (LLMs) and revolutionized various fields by providing a few task-relevant demonstrations in the prompt. However, trustworthy issues with LLM’s response, such as hallucination, have also been actively discussed. Exis…

2024

Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

ECCV 2024poster

"Most Large Vision-Language Models (LVLMs) enjoy the same vision vocabulary, i.e., CLIP, for common vision tasks. However, for some special task that needs dense and fine-grained perception, the CLIP-style vocabulary may encounter low efficiency in tokenizing corresponding vision knowledge and even…

2024

Visual Attention Prompted Prediction and Learning

IJCAI 2024poster

Visual explanation (attention)-guided learning uses not only labels but also explanations to guide the model reasoning process. While visual attention-guided learning has shown promising results, it requires a large number of explanation annotations that are time-consuming to prepare. However, in ma…

2023

3D Reconstruction of Tibia and Fibula using One General Model and Two X-ray Images

ICRA 2023poster

The 3D reconstruction of patient specific bone models plays a crucial role in orthopaedic surgery for clinical evaluation, surgical planning and precise implant design or selection. This paper considers the problem of reconstructing a patient-specific 3D tibia and fibula model from only two 2D X-ray…

Cited by 2SourceScholar
2023

An End-to-End Framework for Partial View-Aligned Clustering with Graph Structure

ICASSP 2023accepted

Over the last decade, many multi-view clustering (MVC) methods have achieved promising results with intact and completely correct correspondence of multi-view data, which is hard to satisfy in practice leading to the problem of partially view-aligned clustering. In this paper, we propose a novel met…

Cited by 0SourceScholar
2023

Curriculum Learning for Graph Neural Networks: Which Edges Should We Learn First

NeurIPS 2023poster

Graph Neural Networks (GNNs) have achieved great success in representing data with dependencies by recursively propagating and aggregating messages along the edges. However, edges in real-world graphs often have varying degrees of difficulty, and some edges may even be noisy to the downstream tasks.…

2023

Deep Graph Representation Learning and Optimization for Influence Maximization

ICML 2023poster

Influence maximization (IM) is formulated as selecting a set of initial users from a social network to maximize the expected number of influenced users. Researchers have made great progresses to design various traditional methods, yet both theoretical design and performance gain are close to their l…

2023

Dialogue Context Modelling for Action Item Detection: Solution for ICASSP 2023 Mug Challenge Track 5

ICASSP 2023accepted

Action item detection aims at recognizing sentences containing information about actionable tasks, which can help people quickly grasp core tasks in the meeting without going through the redundant meeting contents. Therefore, in this paper, we thoroughly describe our carefully designed solution for…

Cited by 0SourceScholar
2023

Domain Adaptation on Point Clouds for 6D Pose Estimation in Bin-Picking Scenarios

IROS 2023poster

Training with simulated data is a common approach in pose estimation research. However, a sim-to-real gap between clean simulated data and noisy real data will seriously weaken the generalization ability of the algorithm, especially for point clouds. To address this problem, this paper proposes a do…

Cited by 4SourceScholar
2023

Enhanced Coprime Array Configuration for DoA Estimation of Non-Circular Signals

ICASSP 2023accepted

Recently, sparse arrays have received considerable attention owing to their capability of achieving increased degrees of freedom (DoFs) by exploiting the virtual sensors resulting from their difference or sum-difference coarrays. Mutual coupling is another factor that attracts interest to these kind…

Cited by 0SourceScholar
2023

Open-ended Commonsense Reasoning with Unrestricted Answer Candidates

EMNLP 2023long findings

Open-ended Commonsense Reasoning is defined as solving a commonsense question without providing 1) a short list of answer candidates and 2) a pre-defined answer scope. Conventional ways of formulating the commonsense question into a question-answering form or utilizing external knowledge to learn re…

Cited by 0SourceScholar
2023

Temporal Domain Generalization with Drift-Aware Dynamic Neural Networks

ICLR 2023top-5%

Temporal domain generalization is a promising yet extremely challenging area where the goal is to learn models under temporally changing data distributions and generalize to unseen data distributions following the trends of the change. The advancement of this area is challenged by: 1) characterizing…

2023

Unrestricted Anchor Graph Based GCN for Incomplete Multi-View Clustering

ICASSP 2023accepted

In recent years, the task of multi-view clustering(MVC) has attracted more and more attention. Meanwhile, the graph convolution network(GCN) based MVC method has made consistent achievements in processing graph-structured data. However, real world data often suffers from missing some instances in ea…

Cited by 0SourceScholar
2022

A New Coprime-Array-based Configuration with Augmented Degrees of Freedom and Reduced Mutual Coupling

ICASSP 2022accepted

In this paper, a new type of coprime-array-based structure, named AtCADiS, is proposed to achieve increased degrees of freedom (DoFs) and reduced mutual coupling. The closed-form expressions for the sensor positions and the number of uniform DoFs (uDoFs) of AtCADiS are provided. Specifically, AtCADi…

Cited by 0SourceScholar
2022

A Right Invariant Extended Kalman Filter for Object Based SLAM

RA-L 2022

With the recent advance of deep learning based object recognition and estimation, it is possible to consider object level SLAM where the pose of each object is estimated in the SLAM process. In this letter, based on a novel Lie group structure, a right invariant extended Kalman filter (RI-EKF) for o

Cited by 37SourceScholar
2022

Adaptive Kernel Graph Neural Network

AAAI 2022technical

Graph neural networks (GNNs) have demonstrated great success in representation learning for graph-structured data. The layer-wise graph convolution in GNNs is shown to be powerful at capturing graph topology. During this process, GNNs are usually guided by pre-defined kernels such as Laplacian matri…

2022

Background-Insensitive Scene Text Recognition with Text Semantic Segmentation

ECCV 2022poster

"Scene Text Recognition (STR) has many important applications in computer vision. Complex backgrounds continue to be a big challenge for STR because they interfere with text feature extraction. Many existing methods use attentional regions, bounding boxes or polygons to reduce such interference. How…

Cited by 19SourcePDFScholar
2022

Disentangled Spatiotemporal Graph Generative Models

AAAI 2022technical

Spatiotemporal graph represents a crucial data structure where the nodes and edges are embedded in a geometric space and their attribute values can evolve dynamically over time. Nowadays, spatiotemporal graph data is becoming increasingly popular and important, ranging from microscale (e.g. protein…

Cited by 29SourcePDFScholar
2022

Multi-objective Deep Data Generation with Correlated Property Control

NeurIPS 2022accept

Developing deep generative models has been an emerging field due to the ability to model and generate complex data for various purposes, such as image synthesis and molecular design. However, the advance of deep generative models is limited by the challenges to generate objects that possess multiple…

Cited by 12SourcePDFScholar
2022

Occupancy-SLAM: Simultaneously Optimizing Robot Poses and Continuous Occupancy Map

RSS 2022poster

In this paper, we propose an optimization based SLAM approach to simultaneously optimize the robot trajectory and the occupancy map using 2D laser scans (and odometry) information. The key novelty is that the robot poses and the occupancy map are optimized together, which is significantly different…

2022

OneEE: A One-Stage Framework for Fast Overlapping and Nested Event Extraction

COLING 2022main

Event extraction (EE) is an essential task of information extraction, which aims to extract structured event information from unstructured text. Most prior work focuses on extracting flat events while neglecting overlapped or nested ones. A few models for overlapped and nested EE includes several su…

2022

Zero-Shot Cross-Lingual Machine Reading Comprehension via Inter-sentence Dependency Graph

AAAI 2022technical

We target the task of cross-lingual Machine Reading Comprehension (MRC) in the direct zero-shot setting, by incorporating syntactic features from Universal Dependencies (UD), and the key features we use are the syntactic relations within each sentence. While previous work has demonstrated effective…

2021

2D Laser SLAM With Closed Shape Features: Fourier Series Parameterization and Submap Joining

RA-L 2021

One of the valuable directions in feature based SLAM is to parameterize and estimate features accurately. In the real world, closed shape features are especially common. It is necessary to study the feature based SLAM problem on closed shape features. The main contribution of this letter is a 2D las

Cited by 15SourceScholar
2021

3D Reconstruction of Deformable Colon Structures based on Preoperative Model and Deep Neural Network

ICRA 2021poster

In colonoscopy procedures, it is important to rebuild and visualize the colonic surface to minimize the missing regions and reinspect for abnormalities. Due to the fast camera motion and deformation of the colon in standard forward-viewing colonoscopies, traditional simultaneous localization and map…

Cited by 10SourceScholar
2021

Deep Graph Spectral Evolution Networks for Graph Topological Evolution

AAAI 2021technical

Characterizing the underlying mechanism of graph topological evolution from a source graph to a target graph has attracted fast increasing attention in the deep graph learning domain. However, it is very challenging to build expressive and efficient models that can handle global and local evolution…

2021

Direct Bundle Adjustment for 3D Image Fusion with Application to Transesophageal Echocardiography

IROS 2021poster

In this paper, we propose a novel algorithm for fusing a sequence of 3D images, named as Direct Bundle Adjustment (DBA). This algorithm simultaneously optimizes the global pose parameters of image frames and the intensity values of the fused global image using the 3D image data directly (without ext…

Cited by 2SourceScholar
2021

GraphGT: Machine Learning Datasets for Graph Generation and Transformation

NeurIPS 2021poster

Graph generation has shown great potential in applications like network design and mobility synthesis and is one of the fastest-growing domains in machine learning for graphs. Despite the success of graph generation, the corresponding real-world datasets are few and limited to areas such as molecule…

Cited by 54SourcecodeScholar
2021

KNAS: Green Neural Architecture Search

ICML 2021spotlight

Many existing neural architecture search (NAS) solutions rely on downstream training for architecture evaluation, which takes enormous computations. Considering that these computations bring a large carbon footprint, this paper aims to explore a green (namely environmental-friendly) NAS solution tha…

2021

Long-term, Short-term and Sudden Event: Trading Volume Movement Prediction with Graph-based Multi-view Modeling

IJCAI 2021poster

Trading volume movement prediction is the key in a variety of financial applications. Despite its importance, there is few research on this topic because of its requirement for comprehensive understanding of information from different sources. For instance, the relation between mult…

2021

Property Controllable Variational Autoencoder via Invertible Mutual Dependence

ICLR 2021poster

Deep generative models have made important progress towards modeling complex, high dimensional data via learning latent representations. Their usefulness is nevertheless often limited by a lack of control over the generative process or a poor understanding of the latent representation. To overcome t…

2021

Some Research Questions for SLAM in Deformable Environments

IROS 2021poster

SLAM in deformable environments is a very challenging research topic. Some research works have been presented by different research groups in the past few years. However, there are still some challenging research questions remaining unanswered. This paper discusses some of these research questions f…

Cited by 6SourcecodeScholar
2020

Analysis of Minima for Geodesic and Chordal Cost for a Minimal 2-D Pose-Graph SLAM Problem

RA-L 2020

In this letter, we show that for a minimal 2D pose-graph SLAM problem, even in the ideal case of perfect measurements and spherical covariance, using geodesic distance (in 2D, the “wrap function”) to compare angles results in multiple suboptimal local minima. We numerically estimate regions of attra

Cited by 1SourceScholar
2020

Aortic 3D Deformation Reconstruction using 2D X-ray Fluoroscopy and 3D Pre-operative Data for Endovascular Interventions

ICRA 2020poster

Current clinical endovascular interventions rely on 2D guidance for catheter manipulation. Although an aortic 3D surface is available from the pre-operative CT/MRI imaging, it cannot be used directly as a 3D intra-operative guidance since the vessel will deform during the procedure. This paper aims…

Cited by 7SourceScholar
2020

Broadcast Your Weaknesses: Cooperative Active Pose-Graph SLAM for Multiple Robots

RA-L 2020

In this letter, we propose a low-cost, high-efficiency framework for cooperative active pose-graph simultaneous localization and mapping (SLAM) for multiple robots in three-dimensional (3D) environments based on graph topology. Based on the selection of weak connections in pose graphs, this method a

Cited by 27SourceScholar
2020

Dense Isometric Non-Rigid Shape-From-Motion Based on Graph Optimization and Edge Selection

RA-L 2020

In this letter, we propose a novel framework for dense isometric non-rigid shape-from-motion (Iso-NRSfM) based on graph topology and edge selection. A weighted undirected graph, of which nodes, edges, and weighted values are respectively the images, the image warps, and the number of the common feat

Cited by 3SourceScholar
2020

Efficient two step optimization for large embedded deformation graph based SLAM

ICRA 2020poster

Embedded deformation graph is a widely used technique in deformable geometry and graphical problems. Although the technique has been transmitted to stereo (or RGB-D) camera based SLAM applications, it remains challenging to compromise the computational cost as the model grows. In practice, the proce…

Cited by 3SourceScholar
2019

On-line 3D active pose-graph SLAM based on key poses using graph topology and sub-maps

ICRA 2019poster

In this paper, we present an on-line active pose-graph simultaneous localization and mapping (SLAM) frame-work for robots in three-dimensional (3D) environments using graph topology and sub-maps. This framework aims to find the best trajectory for loop-closure by re-visiting old poses based on the T…

Cited by 18SourceScholar
2019

Robust Global Structure From Motion Pipeline With Parallax on Manifold Bundle Adjustment and Initialization

RA-L 2019

In this letter, we present a novel global structure from motion (SfM) pipeline that is particularly effective in dealing with low-parallax scenes and camera motion collinear with the features that represent the environment structure. It is therefore particularly suitable in Urban SLAM, in which freq

Cited by 8SourceScholar
2018

Dynamic Reconstruction of Deformable Soft-Tissue With Stereo Scope in Minimal Invasive Surgery

RA-L 2018

In minimal invasive surgery, it is important to rebuild and visualize the latest deformed shape of soft-tissue surfaces to mitigate tissue damages. This letter proposes an innovative Simultaneous localization and mapping (SLAM) algorithm for deformable dense reconstruction of surfaces using a sequen

Cited by 69SourceScholar
2018

MIS-SLAM: Real-Time Large-Scale Dense Deformable SLAM System in Minimal Invasive Surgery Based on Heterogeneous Computing

RA-L 2018

Real-time simultaneous localization and dense mapping is very helpful for providing virtual reality and augmented reality for surgeons or even surgical robots. In this letter, we propose MIS-SLAM: A complete real-time large-scale dense deformable SLAM system with stereoscope in minimal invasive surg

Cited by 113SourceScholar
2018

Occlusion Aware Unsupervised Learning of Optical Flow

CVPR 2018poster

It has been recently shown that a convolutional neural network can learn optical flow estimation with unsuper- vised learning. However, the performance of the unsuper- vised methods still has a relatively large gap compared to its supervised counterpart. Occlusion and large motion are some of the ma…

Cited by 376SourcePDFScholar
2017

Theoretical Properties for Neural Networks with Weight Matrices of Low Displacement Rank

ICML 2017poster

Recently low displacement rank (LDR) matrices, or so-called structured matrices, have been proposed to compress large-scale neural networks. Empirical results have shown that neural networks with weight matrices of LDR matrices, referred as LDR neural networks, can achieve significant reduction in s…

Cited by 79SourcePDFScholar
2016

SCEM+: Real-Time Robust Simultaneous Catheter and Environment Modeling for Endovascular Navigation

RA-L 2016

Endovascular procedures are characterised by significant challenges mainly due to the complexity in catheter control and navigation. Real-time recovery of the 3-D structure of the vasculature is necessary to visualise the interaction between the catheter and its surrounding environment to facilitate

Cited by 17SourceScholar