← Search

Ming Liu

192 accepted papers

2026

AD-BTS: Adaptive Dual-Branch Token Sparsification via Spatial Information Density

ICML 2026poster

High-resolution visual encoders in multimodal large language models (MLLMs) substantially improve fine-grained perception, yet incur prohibitive computational costs.Existing token pruning methods are effective on natural images but struggle with spatially sparse structured inputs (e.g., charts), whe…

Cited by 0SourceScholar
2026

AgentTailor: A Semantic-Aware LLM-Based Multi-Agent System with Actor-Critic Structure

ICML 2026poster

Large Language Model (LLM)-based multi-agent systems often suffer from high communication cost due to redundant interactions, as existing methods optimize communication structures without explicitly measuring whether exchanged messages contribute to the final decision. To better utilize the semantic…

Cited by 0SourceScholar
2026

CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation

AAAI 2026technical

As Large Language Models (LLMs) are increasingly popularized in the multilingual world, ensuring hallucination-free factuality becomes markedly crucial. However, existing benchmarks for evaluating the reliability of Multimodal Large Language Models (MLLMs) predominantly focus on textual or visual m

Cited by 0SourcePDFScholar
2026

CE-GOCD: Central Entity-Guided Graph Optimization for Community Detection to Augment LLM Scientific Question Answering

ICASSP 2026poster

Large Language Models (LLMs) are increasingly used for question answering over scientific research papers. Existing retrieval augmentation methods often rely on isolated text chunks or concepts, but overlook deeper semantic connections between papers. This impairs the LLM's comprehension of scientif…

Cited by 0SourcePDFScholar
2026

Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining

CVPR 2026

Large-scale video-language pretraining enables strong generalization across multimodal tasks but often incurs prohibitive computational costs. Although recent advances in masked visual modeling help mitigate this issue, they still suffer from two fundamental limitations: severe visual information lo

Cited by 0SourcecodeScholar
2026

From Sampling to Cognition: Modeling Internal Cognitive Confidence in Language Models for Robust Uncertainty Calibration

AAAI 2026technical

Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of tasks, yet they generally lack self-awareness, often displaying overconfidence when confronted with questions beyond their knowledge boundaries. This limitation severely hinders their trustworthiness in high

Cited by 0SourcePDFScholar
2026

PreferThinker: Reasoning-based Personalized Image Preference Assessment

ICLR 2026poster

Personalized image preference assessment aims to evaluate an individual user's image preferences by relying only on a small set of reference images as prior information. Existing methods mainly focus on general preference assessment, training models with large-scale data to tackle well-defined task…

Cited by 0SourceScholar
2026

RefSTAR: Blind Face Image Restoration with Reference Selection, Transfer, and Reconstruction

AAAI 2026technical

Introducing high-quality references can largely alleviate the uncertainty in blind face image restoration tasks, yet the equivocal utilization of reference priors makes it still a struggle to well preserve the human identity. We attribute the identity inconsistency to two deficiencies of existing re

Cited by 0SourcePDFScholar
2026

Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved notable success in enhancing translation performance by integrating multimodal information. However, existing research primarily focuses on image-guided methods, whose applicability is constrained by the scarcity of multilingual image-text pairs…

Cited by 0SourceScholar
2026

Semantic-LiDAR-Inertial-Wheel Odometry Fusion for Robust Localization in Large-Scale Dynamic Environments

ICRA 2026poster

Reliable, drift-free global localization presents significant challenges yet remains crucial for autonomous navigation in large-scale dynamic environments. In this paper, we introduce a tightly-coupled Semantic-LiDAR-Inertial-Wheel Odometry fusion framework, which is specifically designed to provide…

2026

Smarter Not Harder: Generative Process Evaluation with Intrinsic-Signal Driving and Ability‑Adaptive Reward Shaping

ICLR 2026poster

Large reasoning models (LRMs) have shown strong performance in complex mathematical reasoning when optimized via reinforcement learning (RL). However, conventional outcome-only reward provides sparse feedback, leading to inefficient optimization. In this work, we investigate whether generative proce…

Cited by 0SourceScholar
2025

AnRe: Analogical Replay for Temporal Knowledge Graph Forecasting

ACL 2025long

Temporal Knowledge Graphs (TKGs) are vital for event prediction, yet current methods face limitations. Graph neural networks mainly depend on structural information, often overlooking semantic understanding and requiring high computational costs. Meanwhile, Large Language Models (LLMs) support zero-…

Cited by 0SourcePDFScholar
2025

Breaking the Reasoning Barrier A Survey on LLM Complex Reasoning through the Lens of Self-Evolution

ACL 2025finding

The release of OpenAI’s O1 and subsequent projects like DeepSeek R1 has significantly advanced research on complex reasoning in LLMs. This paper systematically analyzes existing reasoning studies from the perspective of self-evolution, structured into three components: data evolution, model evolutio…

Cited by 0SourcePDFScholar
2025

CFSP: An Efficient Structured Pruning Framework for LLMs with Coarse-to-Fine Activation Information

COLING 2025main

The colossal parameters and computational overhead of Large Language Models (LLMs) challenge their real-world applications. Network pruning, which targets unstructured or structured sparsity by removing redundant parameters, has recently been explored for LLM acceleration. Existing LLM pruning works…

2025

Can Language Models Capture Human Writing Preferences for Domain-Specific Text Summarization?

ACL 2025finding

With the popularity of large language models and their high-quality text generation capabilities, researchers are using them as auxiliary tools for text summary writing. Although summaries generated by these large language models are smooth and capture key information sufficiently, the quality of th…

2025

DuLoc: Life-Long Dual-Layer Localization in Changing and Dynamic Expansive Scenarios

IROS 2025

LiDAR-based localization serves as a critical component in autonomous systems, yet existing approaches face persistent challenges in balancing repeatability, accuracy, and environmental adaptability. Traditional point cloud registration methods relying solely on offline maps often exhibit limited ro

Cited by 0SourceScholar
2025

Effective Interplay between Sparsity and Quantization: From Theory to Practice

ICLR 2025spotlight

The increasing size of deep neural networks (DNNs) necessitates effective model compression to reduce their computational and memory footprints. Sparsity and quantization are two prominent compression methods that have been shown to reduce DNNs' computational and memory footprints significantly whil…

Cited by 5SourcePDFScholar
2025

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models

ACL 2025long

Large Vision-Language Models (LVLMs) have achieved remarkable success, yet their significant computational demands hinder practicaldeployment. While efforts to improve LVLM efficiency are growing, existing methods lack comprehensive evaluation across diverse backbones, benchmarks, and metrics. In th…

Cited by 0SourcePDFScholar
2025

Exploring and Detecting Self-disclosure in Multi-modal posts on Chinese Social Media

EMNLP 2025

Self-disclosure can provide psychological comfort and social support, but it also carries the risk of unintentionally revealing sensitive information, leading to serious privacy concerns. Research on self-disclosure in Chinese multimodal contexts remains limited, lacking high-quality corpora, analys

2025

FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation

ICCV 2025poster

Layout-to-image (L2I) generation has exhibited promising results in natural domains, but suffers from limited generative fidelity and weak alignment with user-provided layouts when applied to degraded scenes (i.e., low-light, underwater). We primarily attribute these limitations to the "contextual i…

Cited by 0SourcePDFScholar
2025

FisheyeDepth: A Real Scale Self-Supervised Depth Estimation Model for Fisheye Camera

ICRA 2025

Accurate depth estimation is crucial for 3D scene comprehension in robotics and autonomous vehicles. Fisheye cameras, known for their wide field of view, have inherent geometric benefits. However, their use in depth estimation is restricted by a scarcity of ground truth data and image distortions. W

Cited by 7SourcecodeScholar
2025

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities

ACL 2025finding

To tackle complex tasks in real-world scenarios, more researchers are focusing on Omni-MLLMs, which aim to achieve omni-modal understanding and generation. Beyond the constraints of any specific non-linguistic modality, Omni-MLLMs map various non-linguistic modalities into the embedding space of LLM…

2025

GDTS: Goal-Guided Diffusion Model with Tree Sampling for Multi-Modal Pedestrian Trajectory Prediction

IROS 2025

Accurate prediction of pedestrian trajectories is crucial for improving the safety of autonomous driving. However, this task is generally nontrivial due to the inherent stochasticity of human motion, which naturally requires the predictor to generate multi-modal prediction. Previous works leverage v

Cited by 2SourceScholar
2025

GraCoRe: Benchmarking Graph Comprehension and Complex Reasoning in Large Language Models

COLING 2025main

Evaluating the graph comprehension and reasoning abilities of Large Language Models (LLMs) is challenging and often incomplete. Existing benchmarks focus primarily on pure graph understanding, lacking a comprehensive evaluation across all graph types and detailed capability definitions. This paper p…

2025

How do Language Models Reshape Entity Alignment? A Survey of LM-Driven EA Methods: Advances, Benchmarks, and Future

EMNLP 2025

Entity alignment (EA), critical for knowledge graph (KG) integration, identifies equivalent entities across different KGs. Traditional methods often face challenges in semantic understanding and scalability. The rise of language models (LMs), particularly large language models (LLMs), has provided p

Cited by 0SourcePDFScholar
2025

Improved Diffusion-based Generative Model with Better Adversarial Robustness

ICLR 2025poster

Diffusion Probabilistic Models (DPMs) have achieved significant success in generative tasks. However, their training and sampling processes suffer from the issue of distribution mismatch. During the denoising process, the input data distributions differ between the training and inference stages, pot…

2025

Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem Solving

ICCV 2025poster

Current large vision-language models (LVLMs) typically employ a connector module to link visual features with text embeddings of large language models (LLMs) and use end-to-end training to achieve multi-modal understanding in a unified process. Well alignment needs high-quality pre-training data and…

2025

Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency

ACL 2025long

Large Multimodal Models (LMMs) have recently demonstrated impressive performance on general video comprehension benchmarks. Nevertheless, for broader applications, the robustness of their temporal analysis capability needs to be thoroughly investigated yet predominantly ignored. Motivated by this, w…

Cited by 0SourcePDFScholar
2025

MA-GTS: A Multi-Agent Framework for Solving Complex Graph Problems in Real-World Applications

EMNLP 2025

Graph-theoretic problems arise in real-world applications like logistics, communication networks, and traffic optimization. These problems are often complex, noisy, and irregular, posing challenges for traditional algorithms. Large language models offer potential solutions but face several challenge

2025

Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning

ACL 2025long

Multimodal Large Language Models (MLLMs) have achieved significant success in Speech-to-Text Translation (S2TT) tasks. While most existing research has focused on English-centric translation directions, the exploration of many-to-many translation is still limited by the scarcity of parallel data. To…

2025

On Fairness of Unified Multimodal Large Language Model for Image Generation

NeurIPS 2025poster

Unified multimodal large language models (U-MLLMs) have demonstrated impressive performance in end-to-end visual understanding and generation tasks. However, compared to generation-only systems (e.g., Stable Diffusion), the unified architecture of U-MLLMs introduces new risks of propagating demograp…

Cited by 0SourceScholar
2025

Ontology-Guided Reverse Thinking Makes Large Language Models Stronger on Knowledge Graph Question Answering

ACL 2025long

Large language models (LLMs) have shown remarkable capabilities in natural language processing. However, in knowledge graph question answering tasks (KGQA), there remains the issue of answering questions that require multi-hop reasoning. Existing methods rely on entity vector matching, but the purpo…

Cited by 0SourcePDFScholar
2025

Profiling LLM’s Copyright Infringement Risks under Adversarial Persuasive Prompting

EMNLP 2025

Large Language Models (LLMs) have demonstrated impressive capabilities in text generation but raise concerns regarding potential copyright infringement. While prior research has explored mitigation strategies like content filtering and alignment, the impact of adversarial persuasion techniques in el

2025

Self-Critique Guided Iterative Reasoning for Multi-hop Question Answering

ACL 2025finding

Although large language models (LLMs) have demonstrated remarkable reasoning capabilities, they still face challenges in knowledge-intensive multi-hop reasoning. Recent work explores iterative retrieval to address complex problems. However, the absence of intermediate guidance often leads to inaccur…

2025

Sequence Knowledge Enhancement Distillation Framework for Ultra-Fast Image Deraining

ICASSP 2025accepted

Traditional knowledge distillation techniques are aimed at compressing models and speeding up inference, but they often fail to maintain the superior capabilities of complex models in simpler ones. To address this issue, this paper focus on the deraining task and introduces the Sequential Knowledge-…

Cited by 0SourceScholar
2025

Simulation-Free Hierarchical Latent Policy Planning for Proactive Dialogues

AAAI 2025technical

Recent advancements in proactive dialogues have garnered significant attention, particularly for more complex objectives (e.g. emotion support and persuasion). Unlike traditional task-oriented dialogues, proactive dialogues demand advanced policy planning and adaptability, requiring rich scenarios a…

Cited by 1SourcePDFScholar
2025

TSCLIP: Robust CLIP Fine-Tuning for Worldwide Cross-Regional Traffic Sign Recognition

ICRA 2025

Traffic sign is a critical map feature for navigation and traffic control. Nevertheless, current methods for traffic sign recognition rely on traditional deep learning models, which typically suffer from significant performance degradation considering the variations in data distribution across diffe

Cited by 6SourcecodeScholar
2025

Towards Faithful Multi-step Reasoning through Fine-Grained Causal-aware Attribution Reasoning Distillation

COLING 2025main

Despite the remarkable reasoning capabilities demonstrated by large language models (LLM), the substantial computational overhead limits their practices. Some efforts have been directed toward distilling multi-step reasoning capabilities into smaller models through chain-of-thought (CoT). While CoT…

2025

Triad: Empowering LMM-based Anomaly Detection with Expert-guided Region-of-Interest Tokenizer and Manufacturing Process

ICCV 2025poster

Although recent methods have tried to introduce large multimodal models (LMMs) into industrial anomaly detection (IAD), their generalization in the IAD field is far inferior to that for general purposes. We summarize the main reasons for this gap into two aspects. On one hand, general-purpose LMMs l…

2025

UltraFastCrackSeg: A Lightweight Real-Time Crack Segmentation Model with Task-Oriented Pretraining

ICRA 2025

Crack segmentation is pivotal for structural health monitoring, enabling the timely maintenance of critical infrastructure such as bridges and roads. However, existing deep learning models are often too computationally intensive for deployment on resource-constrained devices. To address this limitat

Cited by 0SourcecodeScholar
2024

A Generic Trajectory Planning Method for Constrained All-Wheel-Steering Robots

IROS 2024poster

This paper presents a generic trajectory planning method for wheeled robots with fixed steering axes while the steering angle of each wheel is constrained. In the existing literatures, All-Wheel-Steering (AWS) robots, incorporating modes such as rotation-free translation maneuvers, in-situ rotationa…

Cited by 0SourcecodeScholar
2024

Accurate Prior-centric Monocular Positioning with Offline LiDAR Fusion

ICRA 2024poster

Unmanned vehicles usually rely on Global Positioning System (GPS) and Light Detection and Ranging (LiDAR) sensors to achieve high-precision localization results for navigation purpose. However, this combination with their associated costs and infrastructure demands, poses challenges for widespread a…

Cited by 3SourceScholar
2024

An Image Acquisition Scheme for Visual Odometry based on Image Bracketing and Online Attribute Control

ICRA 2024poster

Visual odometry (VO) system is challenged by complex illumination environments. Image quality and its consistency in the time domain directly determine feature detection and tracking performance, which further affect the robustness and accuracy of the entire system. In this paper, an image acquisiti…

Cited by 2SourceScholar
2024

Annotating Chinese Word Senses with English WordNet: A Practice on OntoNotes Chinese Sense Inventories

COLING 2024main

In this paper, we present our exploration of annotating Chinese word senses using English WordNet synsets, with examples extracted from OntoNotes Chinese sense inventories. Given a target word along with the example that contains it, the annotators select a WordNet synset that best describes the mea…

Cited by 0SourcePDFScholar
2024

BeamAggR: Beam Aggregation Reasoning over Multi-source Knowledge for Multi-hop Question Answering

ACL 2024long

Large language models (LLMs) have demonstrated strong reasoning capabilities.Nevertheless, they still suffer from factual errors when tackling knowledge-intensive tasks.Retrieval-augmented reasoning represents a promising approach.However, significant challenges still persist, including inaccurate a…

2024

BeautyMap: Binary-Encoded Adaptable Ground Matrix for Dynamic Points Removal in Global Maps

RA-L 2024

Global point clouds that correctly represent the static environment features can facilitate accurate localization and robust path planning. However, dynamic objects introduce undesired <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">‘ghost’</i> track

Cited by 21SourcecodeScholar
2024

Car-Studio: Learning Car Radiance Fields From Single-View and Unlimited In-the-Wild Images

RA-L 2024

Compositional neural scene graph studies have shown that radiance fields can be an efficient tool in an editable autonomous driving simulator. However, previous studies learned within a sequence of autonomous driving datasets, resulting in unsatisfactory blurring when rotating the car in the simulat

Cited by 6SourceScholar
2024

CoLRIO: LiDAR-Ranging-Inertial Centralized State Estimation for Robotic Swarms

ICRA 2024poster

Collaborative state estimation using different heterogeneous sensors is a fundamental prerequisite for robotic swarms operating in GPS-denied environments, posing a significant research challenge. In this paper, we introduce a centralized system to facilitate collaborative LiDAR-ranging-inertial sta…

Cited by 3SourcecodeScholar
2024

CogGPT: Unleashing the Power of Cognitive Dynamics on Large Language Models

EMNLP 2024finding

Cognitive dynamics, which refer to the evolution in human cognitive processes, are pivotal to advance human understanding of the world. Recent advancements in large language models (LLMs) highlight their potential for cognitive simulation. However, these LLM-based cognitive studies primarily focus o…

2024

DHP-Mapping: A Dense Panoptic Mapping System with Hierarchical World Representation and Label Optimization Techniques

IROS 2024poster

Maps provide robots with crucial environmental knowledge, thereby enabling them to perform interactive tasks effectively. Easily accessing accurate abstract-to-detailed geometric and semantic concepts from maps is crucial for robots to make informed and efficient decisions. To comprehensively model…

Cited by 2SourcecodeScholar
2024

Decompose, Prioritize, and Eliminate: Dynamically Integrating Diverse Representations for Multimodal Named Entity Recognition

COLING 2024main

Multi-modal Named Entity Recognition, a fundamental task for multi-modal knowledge graph construction, requires integrating multi-modal information to extract named entities from text. Previous research has explored the integration of multi-modal representations at different granularities. However,…

Cited by 1SourcePDFScholar
2024

Divide-and-Conquer Meets Consensus: Unleashing the Power of Functions in Code Generation

NeurIPS 2024oral

Despite recent progress made by large language models in code generation, they still struggle with programs that meet complex requirements. Recent work utilizes plan-and-solve decomposition to decrease the complexity and leverage self-tests to refine the generated program. Yet, planning deep-inside…

Cited by 3SourcePDFScholar
2024

Every Dataset Counts: Scaling up Monocular 3D Object Detection with Joint Datasets Training

IROS 2024poster

Monocular 3D object detection is essential for autonomous driving. However, current monocular 3D detection algorithms rely on expensive 3D labels from LiDAR scans, making it difficult to use in new datasets and unfamiliar environments. This study explores training a monocular 3D object detection mod…

Cited by 4SourceScholar
2024

GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension

IJCAI 2024poster

There are substantial instructional videos on the Internet, which provide us tutorials for completing various tasks. Existing instructional video datasets only focus on specific steps at the video level, lacking experiential guidelines at the task level, which can lead to beginners struggling to lea…

Cited by 1SourcePDFScholar
2024

Improving Autonomous Driving Safety with POP: A Framework for Accurate Partially Observed Trajectory Predictions

ICRA 2024poster

Accurate trajectory prediction is crucial for safe and efficient autonomous driving, but handling partial observations presents significant challenges. To address this, we propose a novel trajectory prediction framework called Partial Observations Prediction (POP) for congested urban road scenarios.…

Cited by 6SourcecodeScholar
2024

Infrared-LLaVA: Enhancing Understanding of Infrared Images in Multi-Modal Large Language Models

EMNLP 2024finding

Expanding the understanding capabilities of multi-modal large language models (MLLMs) for infrared modality is a challenge due to the single-modality nature and limited amount of training data. Existing methods typically construct a uniform embedding space for cross-modal alignment and leverage abun…

Cited by 1SourcePDFScholar
2024

MGCBS: An Optimal and Efficient Algorithm for Solving Multi-Goal Multi-Agent Path Finding Problem

IJCAI 2024poster

With the expansion of the scale of robotics applications, the multi-goal multi-agent pathfinding (MG-MAPF) problem began to gain widespread attention. This problem requires each agent to visit pre-assigned multiple goal points at least once without conflict. Some previous methods have been proposed…

2024

Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future

ACL 2024long

Reasoning, a fundamental cognitive process integral to human intelligence, has garnered substantial interest within artificial intelligence.Notably, recent studies have revealed that chain-of-thought prompting significantly enhances LLM’s reasoning capabilities, which attracts widespread attention f…

2024

OmniColor: A Global Camera Pose Optimization Approach of LiDAR-360Camera Fusion for Colorizing Point Clouds

ICRA 2024poster

A Colored point cloud, as a simple and efficient 3D representation, has many advantages in various fields, including robotic navigation and scene reconstruction. This representation is now commonly used in 3D reconstruction tasks relying on cameras and LiDARs. However, fusing data from these two typ…

Cited by 4SourcecodeScholar
2024

Planning Like Human: A Dual-process Framework for Dialogue Planning

ACL 2024long

In proactive dialogue, the challenge lies not just in generating responses but in steering conversations toward predetermined goals, a task where Large Language Models (LLMs) typically struggle due to their reactive nature. Traditional approaches to enhance dialogue planning in LLMs, ranging from el…

2024

RELEAD: Resilient Localization with Enhanced LiDAR Odometry in Adverse Environments

ICRA 2024poster

LiDAR-based localization is valuable for applications like mining surveys and underground facility maintenance. However, existing methods can struggle when dealing with uninformative geometric structures in challenging scenarios. This paper presents RELEAD, a LiDAR-centric solution designed to addre…

Cited by 3SourceScholar
2024

Relational Graph-Bridged Image-Text Interaction: A Novel Method for Multi-Modal Relation Extraction

ICASSP 2024accepted

Multi-modal relation extraction (MRE) requires the integration of multi-modal information to identify relationships between entities. Although fine-grained correlations between visual objects and textual words have the potential to improve cross-modal interaction, they are typically modeled implicit…

Cited by 0SourceScholar
2024

Rethinking Imitation-based Planners for Autonomous Driving

ICRA 2024poster

In recent years, imitation-based driving planners have reported considerable success. However, due to the absence of a standardized benchmark, the effectiveness of various designs remains unclear. The newly released nuPlan addresses this issue by offering a large-scale real-world dataset and a stand…

Cited by 45SourcecodeScholar
2024

SmartTrim: Adaptive Tokens and Attention Pruning for Efficient Vision-Language Models

COLING 2024main

Despite achieving remarkable performance on various vision-language tasks, Transformer-based Vision-Language Models (VLMs) suffer from redundancy in inputs and parameters, significantly hampering their efficiency in real-world applications. Moreover, the degree of redundancy in token representations…

2024

SumSurvey: An Abstractive Dataset of Scientific Survey Papers for Long Document Summarization

ACL 2024findings

With the popularity of large language models (LLMs) and their ability to handle longer input documents, there is a growing need for high-quality long document summarization datasets. Although many models already support 16k input, current lengths of summarization datasets are inadequate, and salient…

2024

TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

ACL 2024long

Grasping the concept of time is a fundamental facet of human cognition, indispensable for truly comprehending the intricacies of the world.Previous studies typically focus on specific aspects of time, lacking a comprehensive temporal reasoning benchmark.To address this, we propose TimeBench, a compr…

2024

Towards Benchmarking Situational Awareness of Large Language Models:Comprehensive Benchmark, Evaluation and Analysis

EMNLP 2024finding

Situational awareness refers to the capacity to perceive and comprehend the present context and anticipate forthcoming events, which plays a critical role in aiding decision-making, anticipating potential issues, and adapting to dynamic circumstances. Nevertheless, the situational awareness capabili…

Cited by 0SourcePDFScholar
2023

Beyond Image Borders: Learning Feature Extrapolation for Unbounded Image Composition

ICCV 2023poster

For improving image composition and aesthetic quality, most existing methods modulate the captured images by striking out redundant content near the image borders. However, such image cropping methods are limited in the range of image views. Some methods have been suggested to extrapolate the images…

Cited by 2PDFcodeScholar
2023

Black-Box Tuning of Vision-Language Models with Effective Gradient Approximation

EMNLP 2023long findings

Parameter-efficient fine-tuning (PEFT) methods have provided an effective way for adapting large vision-language models to specific tasks or scenarios. Typically, they learn a very small scale of parameters for pre-trained models in a white-box formulation, which assumes model architectures to be kn…

Cited by 0SourcecodeScholar
2023

CenterLineDet: CenterLine Graph Detection for Road Lanes with Vehicle-mounted Sensors by Transformer for HD Map Generation

ICRA 2023poster

With the fast development of autonomous driving technologies, there is an increasing demand for high-definition (HD) maps, which provide reliable and robust prior information about the static part of the traffic environments. As one of the important elements in HD maps, road lane centerline is criti…

Cited by 18SourcecodeScholar
2023

Closed-Loop Feedback Control of Human Step Width During Walking by Mediolaterally Acting Robotic Hip Exoskeleton

IROS 2023poster

Maintaining balance during gait in the mediolateral direction requires more active motor control than in the anteroposterior direction. Step width modulation is a key strategy used by healthy individuals to achieve mediolateral walking balance, but it can be disrupted in populations with poor sensor…

Cited by 2SourceScholar
2023

Completely Rational $\text{SO}(n)$ Orthonormalization

ICRA 2023poster

The rotation orthonormalization on the special orthogonal group \text{SO}(n)\text{SO}(n), also known as the high dimensional nearest rotation problem, has been revisited. A new generalized simple iterative formula has been proposed that solves this problem in a completely rational manner. Rational o…

Cited by 0SourceScholar
2023

Development and Online Validation of an Intrinsic Fault Detector for a Powered Robotic Knee Prosthesis

IROS 2023poster

Robotic prosthetic legs have the potential to significantly improve the quality of life for lower limb amputees to perform locomotion in various environments and task conditions. However, these devices lack the capability to recover from internal intrinsic control faults, which can lead to harmful c…

Cited by 1SourceScholar
2023

Enhanced Multi-Relationships Integration Graph Convolutional Network for Inferring Substitutable and Complementary Items

AAAI 2023technical

Understanding the relationships between items can improve the accuracy and interpretability of recommender systems. Among these relationships, the substitute and complement relationships attract the most attention in e-commerce platforms. The substitutable items are interchangeable and might be comp…

Cited by 8SourcePDFScholar
2023

FDLNet: Boosting Real-time Semantic Segmentation by Image-size Convolution via Frequency Domain Learning

ICRA 2023poster

This paper proposes a novel real-time semantic segmentation network via frequency domain learning, called FDLNet, which revisits the segmentation task from two critical perspectives: spatial structure description and multilevel feature fusion. We first devise an image-size convolution (IS-Conv) as a…

Cited by 6SourcecodeScholar
2023

Forecast-MAE: Self-supervised Pre-training for Motion Forecasting with Masked Autoencoders

ICCV 2023poster

This study explores the application of self-supervised learning (SSL) to the task of motion forecasting, an area that has not yet been extensively investigated despite the widespread success of SSL in computer vision and natural language processing. To address this gap, we introduce Forecast-MAE, an…

Cited by 77PDFcodeScholar
2023

GTR: A Grafting-Then-Reassembling Framework for Dynamic Scene Graph Generation

IJCAI 2023poster

Dynamic scene graph generation aims to identify visual relationships (subject-predicate-object) in frames based on spatio-temporal contextual information in the video. Previous work implicitly models the spatio-temporal interaction simultaneously, which leads to entanglement of spatio-temporal conte…

Cited by 2SourcePDFScholar
2023

Human Guided Ground-Truth Generation for Realistic Image Super-Resolution

CVPR 2023poster

How to generate the ground-truth (GT) image is a critical issue for training realistic image super-resolution (Real-ISR) models. Existing methods mostly take a set of high-resolution (HR) images as GTs and apply various degradations to simulate their low-resolution (LR) counterparts. Though great pr…

2023

MTGER: Multi-view Temporal Graph Enhanced Temporal Reasoning over Time-Involved Document

EMNLP 2023long findings

The facts and time in the document are intricately intertwined, making temporal reasoning over documents challenging. Previous work models time implicitly, making it difficult to handle such complex relationships. To address this issue, we propose MTGER, a novel Multi-view Temporal Graph Enhanced Re…

Cited by 0SourceScholar
2023

MetaF2N: Blind Image Super-Resolution by Learning Efficient Model Adaptation from Faces

ICCV 2023poster

Due to their highly structured characteristics, faces are easier to recover than natural scenes for blind image super-resolution. Therefore, we can extract the degradation representation of an image from the low-quality and recovered face pairs. Using the degradation representation, realistic low-qu…

Cited by 7PDFcodeScholar
2023

Physics-Guided ISO-Dependent Sensor Noise Modeling for Extreme Low-Light Photography

CVPR 2023poster

Although deep neural networks have achieved astonishing performance in many vision tasks, existing learning-based methods are far inferior to the physical model-based solutions in extreme low-light sensor noise modeling. To tap the potential of learning-based sensor noise modeling, we investigate th…

2023

RNGDet++: Road Network Graph Detection by Transformer With Instance Segmentation and Multi-Scale Features Enhancement

RA-L 2023

The road network graph is a critical component for downstream tasks in autonomous driving, such as global route planning and navigation. In the past years, road network graphs are usually annotated by human experts manually, which is time-consuming and labor-intensive. To annotate road network graph

Cited by 50SourceScholar
2023

Real-Time Neural Dense Elevation Mapping for Urban Terrain With Uncertainty Estimations

RA-L 2023

Having good knowledge of terrain information is essential for improving the performance of various downstream tasks on complex terrains, especially for the locomotion and navigation of legged robots. We present a novel framework for neural urban terrain reconstruction with uncertainty estimations. I

Cited by 21SourceScholar
2023

Self-Supervised Drivable Area Segmentation Using LiDAR's Depth Information for Autonomous Driving

IROS 2023poster

Drivable area segmentation is an essential component of the visual perception system for autonomous driving vehicles. Recent efforts in deep neural networks have sig-nificantly improved semantic segmentation performance for autonomous driving. However, most DNN-based methods need a large amount of d…

Cited by 9SourceScholar
2022

360ST-Mapping: An Online Semantics-Guided Topological Mapping Module for Omnidirectional Visual SLAM

IROS 2022poster

As an abstract representation of the environment structure, a topological map has advantageous properties for path-planning and navigation. Here we proposed an online topological mapping method, 360ST-Mapping, using omnidirectional vision. The 360° field-of-view allows the agent to obtain consistent…

Cited by 7SourceScholar
2022

An Online Interactive Approach for Crowd Navigation of Quadrupedal Robots

IROS 2022poster

Robot navigation in human crowds remains the challenge of understanding human behaviors in different scenarios. We present an approach for interactive and human-friendly crowd navigation in complex static environments. The planner models the online interactions among the robot, humans, and the stati…

Cited by 7SourceScholar
2022

Characterizing Prosthesis Control Fault During Human-Prosthesis Interactive Walking Using Intrinsic Sensors

RA-L 2022

The physical interactions between wearable lower limb robots and humans have been investigated to inform effective robot design for walking augmentation. However, human-robot interactions when internal faults occur within robots have not been systematically reported, but it is essential to improve t

Cited by 8SourceScholar
2022

Cost Ensemble with Gradient Selecting for GANs

IJCAI 2022poster

Generative Adversarial Networks(GANs) are powerful generative models on numerous tasks and datasets but are also known for their training instability and mode collapse. The latter is because the optimal transportation map is discontinuous, but DNNs can only approximate continuous ones. One way to so…

Cited by 0SourcePDFScholar
2022

Design of EMG-driven Musculoskeletal Model for Volitional Control of a Robotic Ankle Prosthesis

IROS 2022poster

Existing robotic lower-limb prostheses use autonomous control to address cyclic, locomotive tasks, but are inadequate in adapting to variations in non-cyclic and unpredictable tasks. This study aims to address this challenge by designing a novel electromyography (EMG)-driven musculoskeletal model fo…

Cited by 10SourceScholar
2022

Distilled Dual-Encoder Model for Vision-Language Understanding

EMNLP 2022main

On vision-language understanding (VLU) tasks, fusion-encoder vision-language models achieve superior results but sacrifice efficiency because of the simultaneous encoding of images and text. On the contrary, the dual encoder model that separately encodes images and text has the advantage in efficien…

2022

Efficient Speed Planning for Autonomous Driving in Dynamic Environment With Interaction Point Model

RA-L 2022

Safely interacting with other traffic participants is one of the core requirements for autonomous driving, especially in intersections and occlusions. Most existing approaches are designed for particular scenarios and require significant human labor in parameter tuning to be applied to different sit

Cited by 16SourcecodeScholar
2022

FusionPortable: A Multi-Sensor Campus-Scene Dataset for Evaluation of Localization and Mapping Accuracy on Diverse Platforms

IROS 2022poster

Combining multiple sensors enables a robot to maximize its perceptual awareness of environments and enhance its robustness to external disturbance, crucial to robotic navigation. This paper proposes the FusionPortable benchmark, a complete multi-sensor dataset with a diverse set of sequences for mob…

Cited by 39SourceScholar
2022

HGCN-GJS: Hierarchical Graph Convolutional Network with Groupwise Joint Sampling for Trajectory Prediction

IROS 2022poster

Pedestrian trajectory prediction is of great importance for downstream tasks, such as autonomous driving and mobile robot navigation. Realistic models of the social interactions within the crowd is crucial for accurate pedestrian trajectory prediction. However, most existing methods do not capture g…

Cited by 16SourceScholar
2022

HoloSeg: An Efficient Holographic Segmentation Network for Real-time Scene Parsing

ICRA 2022poster

Real-time semantic segmentation is a crucial but challenging dense prediction task for scene parsing. However, the existing CNN-based methods commonly bias the model in favor of speed-boosting compromising spatial resolution due to business requirements and hardware constrains, which impedes the hig…

Cited by 8SourcecodeScholar
2022

How Far are We from Robust Long Abstractive Summarization?

EMNLP 2022main

Abstractive summarization has made tremendous progress in recent years. In this work, we perform fine-grained human annotations to evaluate long document abstractive summarization systems (i.e., models and metrics) with the aim of implementing them to generate reliable summaries. For long document a…

2022

Open-World Semantic Segmentation for LIDAR Point Clouds

ECCV 2022poster

"Classical LIDAR semantic segmentation is not robust for real-world applications, e.g., autonomous driving, since it is closed-set and static. The closed-set network is only able to output labels of trained classes, even for objects never seen before, while a static network cannot update its knowled…

2022

Point Cloud Compression with Range Image-Based Entropy Model for Autonomous Driving

ECCV 2022poster

"For autonomous driving systems, the storage cost and transmission speed of the large-scale point clouds become an important bottleneck because of their large volume. In this paper, we propose a range image-based three-stage framework to compress the scanning LiDAR’s point clouds using the entropy m…

2022

Real-Time Trajectory Planning for Autonomous Driving with Gaussian Process and Incremental Refinement

ICRA 2022poster

Real-time kinodynamic trajectory planning in dy-namic environments is critical yet challenging for autonomous driving. In this paper, we propose an efficient trajectory plan-ning system for autonomous driving in complex dynamic sce-narios through iterative and incremental path-speed optimization. Ex…

Cited by 51SourcecodeScholar
2022

Self-Supervised Image Restoration with Blurry and Noisy Pairs

NeurIPS 2022accept

When taking photos under an environment with insufficient light, the exposure time and the sensor gain usually require to be carefully chosen to obtain images with satisfying visual quality. For example, the images with high ISO usually have inescapable noise, while the long-exposure ones may be blu…

2022

UnDAF: A General Unsupervised Domain Adaptation Framework for Disparity or Optical Flow Estimation

ICRA 2022poster

Disparity and optical flow estimation are respectively 1D and 2D dense correspondence matching (DCM) tasks in nature. Unsupervised domain adaptation (UDA) is crucial for their success in new and unseen scenarios, enabling networks to draw inferences across different domains without manually-labeled…

Cited by 8SourceScholar
2022

Why-So-Deep: Towards Boosting Previously Trained Models for Visual Place Recognition

RA-L 2022

Deep learning-based image retrieval techniques for the loop closure detection demonstrate satisfactory performance. However, it is still challenging to achieve high-level performance based on previously trained models in different geographical regions. This letter addresses the problem of their depl

Cited by 8SourcecodeScholar
2022

csBoundary: City-Scale Road-Boundary Detection in Aerial Images for High-Definition Maps

RA-L 2022

High-Definition (HD) maps can provide precise geometric and semantic information of static traffic environments for autonomous driving. Road-boundary is one important information presented in HD maps since it distinguishes between road areas and off-road areas, which can guide vehicles to drive with

Cited by 37SourceScholar
2021

3D Surfel Map-Aided Visual Relocalization with Learned Descriptors

ICRA 2021poster

In this paper, we introduce a method for visual relocalization using the geometric information from a 3D surfel map. A visual database is first built by global indices from the 3D surfel map rendering, which provides associations between image points and 3D surfels. Surfel reprojection constraints a…

Cited by 2SourceScholar
2021

A Data-Driven Reinforcement Learning Solution Framework for Optimal and Adaptive Personalization of a Hip Exoskeleton

ICRA 2021poster

Robotic exoskeletons are exciting technologies for augmenting human mobility. However, designing such a device for seamless integration with the human user and to assist human movement still is a major challenge. This paper aims at developing a novel data-driven solution framework based on reinforce…

Cited by 37SourceScholar
2021

A Powered Prosthetic Ankle Designed for Task Variability – A Concept Validation

IROS 2021poster

Ankle joints play key roles in everyday locomotion, such as walking, stair climbing, and sit-to-stand. Despite the achievement in designing powered prosthetic ankles, engineers still face challenges to duplicate the full mechanics of ankle joints, including high torque, large range of motion (ROM),…

Cited by 2SourceScholar
2021

AVGCN: Trajectory Prediction using Graph Convolutional Networks Guided by Human Attention

ICRA 2021poster

Pedestrian trajectory prediction is a critical yet challenging task especially for crowded scenes. We suggest that introducing an attention mechanism to infer the importance of different neighbors is critical for accurate trajectory prediction in scenes with varying crowd size. In this work, we prop…

Cited by 35SourceScholar
2021

CP-loss: Connectivity-preserving Loss for Road Curb Detection in Autonomous Driving with Aerial Images

IROS 2021poster

Road curb detection is important for autonomous driving. It can be used to determine road boundaries to constrain vehicles on roads, so that potential accidents could be avoided. Most of the current methods detect road curbs online using vehicle-mounted sensors, such as cameras or 3-D Lidars. Howeve…

Cited by 13SourceScholar
2021

DiGNet: Learning Scalable Self-Driving Policies for Generic Traffic Scenarios with Graph Neural Networks

IROS 2021poster

Traditional decision and planning frameworks for self-driving vehicles (SDVs) scale poorly in new scenarios, thus they require tedious hand-tuning of rules and parameters to maintain acceptable performance in all foreseeable cases. Recently, self-driving methods based on deep learning have shown pro…

Cited by 19SourcecodeScholar
2021

DiTNet: End-to-End 3D Object Detection and Track ID Assignment in Spatio-Temporal World

RA-L 2021

End-to-end 3D object detection and tracking based on point clouds is receiving more and more attention in many robotics applications, such as autonomous driving. Compared with 2D images, 3D point clouds do not have enough texture information for data association. Thus, we propose an end-to-end point

Cited by 22SourceScholar
2021

Differential Information Aided 3-D Registration for Accurate Navigation and Scene Reconstruction

ICRA 2021poster

A novel 3-dimensional (3-D) alignment method for point-cloud registration is proposed where the time-differential information of the measured points is employed. The new problem turns out to be a novel multi-dimensional optimization. Analytical solution to this optimization is then obtained, which s…

Cited by 2SourceScholar
2021

Greedy-Based Feature Selection for Efficient LiDAR SLAM

ICRA 2021poster

Modern LiDAR-SLAM (L-SLAM) systems have shown excellent results in large-scale, real-world scenarios. However, they commonly have a high latency due to the expensive data association and nonlinear optimization. This paper demonstrates that actively selecting a subset of features significantly improv…

Cited by 50SourceScholar
2021

In Defense of Knowledge Distillation for Task Incremental Learning and Its Application in 3D Object Detection

RA-L 2021

Making robots learn skills incrementally is an efficient way to design real intelligent agents. To achieve this, researchers adopt knowledge distillation to transfer old-task knowledge from old models to new ones. However, when the length of the task sequence increases, the effectiveness of knowledg

Cited by 23SourceScholar
2021

Learning Interpretable End-to-End Vision-Based Motion Planning for Autonomous Driving with Optical Flow Distillation

ICRA 2021poster

Recently, deep-learning based approaches have achieved impressive performance for autonomous driving. However, end-to-end vision-based methods typically have limited interpretability, making the behaviors of the deep networks difficult to explain. Hence, their potential applications could be limited…

Cited by 57SourceScholar
2021

Learning RAW-to-sRGB Mappings With Inaccurately Aligned Supervision

ICCV 2021poster

Learning RAW-to-sRGB mapping has drawn increasing attention in recent years, wherein an input raw image is trained to imitate the target sRGB image captured by another camera. However, the severe color inconsistency makes it very challenging to generate well-aligned training pairs of input raw and t…

Cited by 52PDFcodeScholar
2021

Less Is More: Domain Adaptation with Lottery Ticket for Reading Comprehension

EMNLP 2021finding

In this paper, we propose a simple few-shot domain adaptation paradigm for reading comprehension. We first identify the lottery subnetwork structure within the Transformer-based source domain model via gradual magnitude pruning. Then, we only fine-tune the lottery subnetwork, a small fraction of the…

2021

Leveraging Information Bottleneck for Scientific Document Summarization

EMNLP 2021finding

This paper presents an unsupervised extractive approach to summarize scientific long documents based on the Information Bottleneck principle. Inspired by previous work which uses the Information Bottleneck principle for sentence compression, we extend it to document level summarization with two sepa…

Cited by 20SourcePDFScholar
2021

On Bundle Adjustment for Multiview Point Cloud Registration

RA-L 2021

Multiview registration is used to estimate Rigid Body Transformations (RBTs) from multiple frames and reconstruct a scene with corresponding scans. Despite the success of pairwise registration and pose synchronization, the concept of Bundle Adjustment (BA) has been proven to better maintain global c

Cited by 24SourcecodeScholar
2021

PointMoSeg: Sparse Tensor-Based End-to-End Moving-Obstacle Segmentation in 3-D Lidar Point Clouds for Autonomous Driving

RA-L 2021

Moving-obstacle segmentation is an essential capability for autonomous driving. For example, it can serve as a fundamental component for motion planning in dynamic traffic environments. Most of the current 3-D Lidar-based methods use road segmentation to find obstacles, and then employ ego-motion co

Cited by 31SourceScholar
2021

Real-time Optimal Navigation Planning Using Learned Motion Costs

ICRA 2021poster

Navigation on challenging terrain topographies requires the understanding of robots’ locomotion capabilities to produce optimal solutions. We present an integrated framework for real-time autonomous navigation of mobile robots based on elevation maps. The framework performs rapid global path plannin…

Cited by 37SourceScholar
2021

S2P2: Self-Supervised Goal-Directed Path Planning Using RGB-D Data for Robotic Wheelchairs

ICRA 2021poster

Path planning is a fundamental capability for autonomous navigation of robotic wheelchairs. With the impressive development of deep-learning technologies, imitation learning-based path planning approaches have achieved effective results in recent years. However, the disadvantages of these approaches…

Cited by 4SourceScholar
2021

SNE-RoadSeg+: Rethinking Depth-Normal Translation and Deep Supervision for Freespace Detection

IROS 2021poster

Freespace detection is a fundamental component of autonomous driving perception. Recently, deep convolutional neural networks (DCNNs) have achieved impressive performance for this task. In particular, SNE-RoadSeg, our previously proposed method based on a surface normal estimator (SNE) and a data-fu…

Cited by 67SourceScholar
2021

Three-Filters-to-Normal: An Accurate and Ultrafast Surface Normal Estimator

RA-L 2021

This letter proposes three-filters-to-normal (3F2N), an accurate and ultrafast surface normal estimator (SNE), which is designed for structured range sensor data, e.g., depth/disparity images. 3F2N SNE computes surface normals by simply performing three filtering operations (two image gradient filte

Cited by 44SourcecodeScholar
2021

Topo-Boundary: A Benchmark Dataset on Topological Road-Boundary Detection Using Aerial Images for Autonomous Driving

RA-L 2021

Road-boundary detection is important for autonomous driving. It can be used to constrain autonomous vehicles running on road areas to ensure driving safety. Compared with online road-boundary detection using on-vehicle cameras/Lidars, offline detection using aerial images could alleviate the severe

Cited by 49SourcecodeScholar
2021

Transformer over Pre-trained Transformer for Neural Text Segmentation with Enhanced Topic Coherence

EMNLP 2021finding

This paper proposes a transformer over transformer framework, called Transformerˆ2, to perform neural text segmentation. It consists of two components: bottom-level sentence encoders using pre-trained transformers, and an upper-level transformer-based segmentation model based on the sentence embeddi…

Cited by 47SourcePDFScholar
2021

User Controlled Interface for Tuning Robotic Knee Prosthesis

IROS 2021poster

The tuning process for a robotic prosthesis is a challenging and time-consuming task both for users and clinicians. An automatic tuning approach using reinforcement learning (RL) has been developed for a knee prosthesis to address the challenges of manual tuning methods. The algorithm tunes the opti…

Cited by 16SourceScholar
2021

Vision-Based Autonomous Car Racing Using Deep Imitative Reinforcement Learning

RA-L 2021

Autonomous car racing is a challenging task in the robotic control area. Traditional modular methods require accurate mapping, localization and planning, which makes them computationally inefficient and sensitive to environmental changes. Recently, deep-learning-based end-to-end systems have shown p

Cited by 78SourcecodeScholar
2021

iCurb: Imitation Learning-Based Detection of Road Curbs Using Aerial Images for Autonomous Driving

RA-L 2021

Detection of road curbs is an essential capability for autonomous driving. It can be used for autonomous vehicles to determine drivable areas on roads. Usually, road curbs are detected on-line using vehicle-mounted sensors, such as video cameras and 3-D Lidars. However, on-line detection using video

Cited by 52SourcecodeScholar
2020

AMLN: Adversarial-based Mutual Learning Network for Online Knowledge Distillation

ECCV 2020poster

Online knowledge distillation has attracted increasing interest recently, which jointly learns teacher and student models or an ensemble of student models simultaneously and collaboratively. On the other hand, existing works focus more on outcome-driven learning according to knowledge like classific…

Cited by 18SourcePDFScholar
2020

Applying Surface Normal Information in Drivable Area and Road Anomaly Detection for Ground Mobile Robots

IROS 2020poster

The joint detection of drivable areas and road anomalies is a crucial task for ground mobile robots. In recent years, many impressive semantic segmentation networks, which can be used for pixel-level drivable area and road anomaly detection, have been developed. However, the detection accuracy still…

Cited by 75SourcecodeScholar
2020

CoT-AMFlow: Adaptive Modulation Network with Co-Teaching Strategy for Unsupervised Optical Flow Estimation

CoRL 2020

The interpretation of ego motion and scene change is a fundamental task for mobile robots. Optical flow information can be employed to estimate motion in the surroundings. Recently, unsupervised optical flow estimation has become a research hotspot. However, unsupervised approaches are often easy to

Cited by 0SourcePDFScholar
2020

Federated Imitation Learning: A Novel Framework for Cloud Robotic Systems With Heterogeneous Sensor Data

RA-L 2020

Humans are capable of learning a new behavior by observing others to perform the skill. Similarly, robots can also implement this by imitation learning. Furthermore, if with external guidance, humans can master the new behavior more efficiently. So, how can robots achieve this? To address the issue,

Cited by 80SourceScholar
2020

LINS: A Lidar-Inertial State Estimator for Robust and Efficient Navigation

ICRA 2020poster

We present LINS, a lightweight lidar-inertial state estimator, for real-time ego-motion estimation. The proposed method enables robust and efficient navigation for ground vehicles in challenging environments, such as feature-less scenes, via fusing a 6-axis IMU and a 3D lidar in a tightly-coupled sc…

Cited by 370SourceScholar
2020

Learning Flow-based Feature Warping for Face Frontalization with Illumination Inconsistent Supervision

ECCV 2020poster

Despite recent advances in deep learning-based face frontalization methods, photo-realistic and illumination preserving frontal face synthesis is still challenging due to large pose and illumination discrepancy during training. We propose a novel Flow-based Feature Warping Model (FFWM) which can lea…

2020

MLOD: Awareness of Extrinsic Perturbation in Multi-LiDAR 3D Object Detection for Autonomous Driving

IROS 2020poster

Extrinsic perturbation always exists in multiple sensors. In this paper, we focus on the extrinsic uncertainty in multi-LiDAR systems for 3D object detection. We first analyze the influence of extrinsic perturbation on geometric tasks with two basic examples. To minimize the detrimental effect of ex…

Cited by 15SourceScholar
2020

Molweni: A Challenge Multiparty Dialogues-based Machine Reading Comprehension Dataset with Discourse Structure

COLING 2020main

Research into the area of multiparty dialog has grown considerably over recent years. We present the Molweni dataset, a machine reading comprehension (MRC) dataset with discourse structure built over multiparty dialog. Molweni’s source samples from the Ubuntu Chat Corpus, including 10,000 dialogs co…

2020

Monocular Visual Odometry using Learned Repeatability and Description

ICRA 2020poster

Robustness and accuracy for monocular visual odometry (VO) under challenging environments are widely concerned. In this paper, we present a monocular VO system leveraging learned repeatability and description. In a hybrid scheme, the camera pose is initially tracked on the predicted repeatability ma…

Cited by 12SourceScholar
2020

PointTrackNet: An End-to-End Network For 3-D Object Detection and Tracking From Point Clouds

RA-L 2020

Recent machine learning-based multi-object tracking (MOT) frameworks are becoming popular for 3-D point clouds. Most traditional tracking approaches use filters (e.g., Kalman filter or particle filter) to predict object locations in a time sequence, however, they are vulnerable to extreme motion con

Cited by 61SourceScholar
2020

Probabilistic End-to-End Vehicle Navigation in Complex Dynamic Environments With Multimodal Sensor Fusion

RA-L 2020

All-day and all-weather navigation is a critical capability for autonomous driving, which requires proper reaction to varied environmental conditions and complex agent behaviors. Recently, with the rise of deep learning, end-to-end control for autonomous vehicles has been well studied. However, most

Cited by 82SourceScholar
2020

Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point

NeurIPS 2020poster

In this paper, we explore the limits of Microsoft Floating Point (MSFP), a new class of datatypes developed for production cloud-scale inferencing on custom hardware. Through the co-evolution of hardware design and algorithms, MSFP achieves accuracy comparable to or better than industry standards Bf…

2020

Robot Navigation in Crowds by Graph Convolutional Networks With Attention Learned From Human Gaze

RA-L 2020

Safe and efficient crowd navigation for mobile robot is a crucial yet challenging task. Previous work has shown the power of deep reinforcement learning frameworks to train efficient policies. However, their performance deteriorates when the crowd size grows. We suggest that this can be addressed by

Cited by 144SourceScholar
2020

Robust Pedestrian Tracking in Crowd Scenarios Using an Adaptive GMM-based Framework

IROS 2020poster

In this paper, we address the issue of pedestrian tracking in crowd scenarios. People in close social relationships tend to act as a group which is a great challenge to individually discriminate and track pedestrians on a LiDAR system. In this paper, we integrally model groups of people and track th…

Cited by 4SourceScholar
2020

SNE-RoadSeg: Incorporating Surface Normal Information into Semantic Segmentation for Accurate Freespace Detection

ECCV 2020poster

Freespace detection is an essential component of visual perception for self-driving cars. The recent efforts made in data-fusion convolutional neural networks (CNNs) have significantly improved semantic driving scene segmentation. Freespace can be hypothesized as a ground plane, on which the points…

2020

See the Future: A Semantic Segmentation Network Predicting Ego-Vehicle Trajectory With a Single Monocular Camera

RA-L 2020

Ego-vehicle trajectory prediction is important for autonomous vehicles to detect collisions and accordingly avoid accidents. Recent approaches employ prior-known or on-line acquired road topology or geometries as motion constraints for their predictive models. However, the prior-known information (e

Cited by 31SourceScholar
2020

Smart-Inspect: Micro Scale Localization and Classification of Smartphone Glass Defects for Industrial Automation

IROS 2020poster

The presence of any type of defect on the glass screen of smart devices has a great impact on their quality. We present a robust semi-supervised learning framework for intelligent micro-scaled localization and classification of defects on a 16K pixel image of smartphone glass. Our model features the…

Cited by 8SourceScholar
2019

A GPS-aided Omnidirectional Visual-Inertial State Estimator in Ubiquitous Environments

IROS 2019poster

The visual-inertial navigation system (VINS) has been a practical approach for state estimation in recent years. In this paper, we propose a general GPS-aided omnidirectional visual-inertial state estimator capable of operating in ubiquitous environments and platforms. Our system consists of two par…

Cited by 59SourceScholar
2019

Automatic Calibration of Multiple 3D LiDARs in Urban Environments

IROS 2019poster

Multiple LiDARs have progressively emerged on autonomous vehicles for rendering a rich view and dense measurements. However, the lack of precise calibration negatively affects their potential applications. In this paper, we propose a novel system that enables automatic multi-LiDAR calibration method…

Cited by 64SourceScholar
2019

Gaze Training by Modulated Dropout Improves Imitation Learning

IROS 2019poster

Imitation learning by behavioral cloning is a prevalent method that has achieved some success in vision-based autonomous driving. The basic idea behind behavioral cloning is to have the neural network learn from observing a human expert's behavior. Typically, a convolutional neural network learns to…

Cited by 27SourceScholar
2019

Lifelong Federated Reinforcement Learning: A Learning Architecture for Navigation in Cloud Robotic Systems

IROS 2019poster

This paper was motivated by the problem of how to make robots fuse and transfer their experience so that they can effectively use prior knowledge and quickly adapt to new environments. To address the problem, we present a learning architecture for navigation in cloud robotic systems: Lifelong Federa…

Cited by 325SourceScholar
2019

Lifelong Federated Reinforcement Learning: A Learning Architecture for Navigation in Cloud Robotic Systems

RA-L 2019

This letter was motivated by the problem of how to make robots fuse and transfer their experience so that they can effectively use prior knowledge and quickly adapt to new environments. To address the problem, we present a learning architecture for navigation in cloud robotic systems: Lifelong Feder

Cited by 272SourceScholar
2019

STGAN: A Unified Selective Transfer Network for Arbitrary Image Attribute Editing

CVPR 2019poster

Arbitrary attribute editing generally can be tackled by incorporating encoder-decoder and generative adversarial networks. However, the bottleneck layer in encoder-decoder usually gives rise to blurry and low quality editing result. And adding skip connections improves image quality at the cost of w…

Cited by 427PDFcodeScholar
2019

Self-Supervised Drivable Area and Road Anomaly Segmentation Using RGB-D Data For Robotic Wheelchairs

RA-L 2019

The segmentation of drivable areas and road anomalies are critical capabilities to achieve autonomous navigation for robotic wheelchairs. The recent progress of semantic segmentation using deep learning techniques has presented effective results. However, the acquisition of large-scale datasets with

Cited by 70SourcecodeScholar
2019

VR-Goggles for Robots: Real-to-Sim Domain Adaptation for Visual Control

RA-L 2019

In this letter, we deal with the reality gap from a novel perspective, targeting transferring deep reinforcement learning (DRL) policies learned in simulated environments to the real-world domain for visual control tasks. Instead of adopting the common solutions to the problem by increasing the visu

Cited by 133SourceScholar
2019

Visual-based Autonomous Driving Deployment from a Stochastic and Uncertainty-aware Perspective

IROS 2019poster

End-to-end visual-based imitation learning has been widely applied in autonomous driving. When deploying the trained visual-based driving policy, a deterministic command is usually directly applied without considering the uncertainty of the input data. Such kind of policies may bring dramatical dama…

Cited by 28SourcecodeScholar
2018

Indoor Mapping and Localization for Pedestrians using Opportunistic Sensing with Smartphones

IROS 2018poster

Indoor localization for pedestrians has gained increasing popularity among the rich body of literature for the last decade. In this paper, a low-cost indoor mapping and localization solution is proposed using the opportunistic signals from ambient indoor environments with a smartphone. It is compose…

Cited by 11SourceScholar
2018

Learning Warped Guidance for Blind Face Restoration

ECCV 2018poster

This paper studies the problem of blind face restoration from an unconstrained blurry, noisy, low-resolution, or compressed image (i.e., degraded observation). For better recovery of fine facial details, we modify the problem setting by taking both the degraded observation and a high-quality guided…

2018

Plugo: A Scalable Visible Light Communication System Towards Low-Cost Indoor Localization

IROS 2018poster

Indoor localization is critical to many location-aware applications, however, a low-cost solution with guaranteed accuracies has not yet come. Visible Light Communication (VLC-) based localization techniques are very promising to fill this gap. In this paper, we propose Plugo, a novel VLC system wit…

Cited by 14SourceScholar
2018

Socially Compliant Navigation Through Raw Depth Inputs with Generative Adversarial Imitation Learning

ICRA 2018poster

We present an approach for mobile robots to learn to navigate in dynamic environments with pedestrians via raw depth inputs, in a socially compliant manner. To achieve this, we adopt a generative adversarial imitation learning (GAIL) strategy, which improves upon a pre-trained behavior cloning polic…

Cited by 239SourcecodeScholar
2017

A unified leader-follower scheme for mobile robots with uncalibrated on-board camera

ICRA 2017poster

This paper studies the problem of image-based leader-follower formation control for mobile robots, where the controller is designed independently of the leader's motion. An adaptive control scheme, which is suitable for both omnidirectional and perspective cameras, is proposed. The proposed approach…

Cited by 14SourceScholar
2017

Virtual-to-real deep reinforcement learning: Continuous control of mobile robots for mapless navigation

IROS 2017poster

We present a learning-based mapless motion planner by taking the sparse 10-dimensional range findings and the target position with respect to the mobile robot coordinate frame as input and the continuous steering commands as output. Traditional motion planners for mobile ground robots with a laser r…

Cited by 951SourceScholar
2015

Asynchronous blind signal decomposition using tiny-length code for Visible Light Communication-based indoor localization

ICRA 2015poster

Indoor localization is a fundamental capability for service robots and indoor applications on mobile devices. To realize that, the cost and performance are of great concern. In this paper, we introduce a lightweight signal encoding and decomposition method for a low-cost and low-power Visible Light…

Cited by 10SourceScholar