← Search

Shuo Wang

160 accepted papers

2026

Accelerating Controllable Generation via Hybrid-grained Cache

AAAI 2026technical

Controllable generative models have been widely used to improve the realism of synthetic visual content. However, such models must handle control conditions and content generation computational requirements, resulting in generally low generation efficiency. To address this issue, we propose a Hybrid

Cited by 0SourcePDFScholar
2026

Adapting Point Cloud Analysis via Multimodal Bayesian Distribution Learning

CVPR 2026

Large multimodal 3D vision-language models show strong generalization across diverse 3D tasks, but their performance still degrades under domain shifts. This has motivated recent studies on test-time adaptation (TTA), which enables models to adapt online using test-time data. Among existing TTA meth

Cited by 0SourceScholar
2026

Adaptive Action Chunking at Inference-time for Vision-Language-Action Models

CVPR 2026

In Vision-Language-Action (VLA) models, action chunking (i.e., executing a sequence of actions without intermediate replanning) is a key technique to improve robotic manipulation abilities. However, a large chunk size reduces the model's responsiveness to new information, while a small one increases

Cited by 0SourcecodeScholar
2026

Advancing Off-Road Autonomous Driving: The Large-Scale ORAD-3D Dataset and Comprehensive Benchmarks

ICRA 2026poster

A major bottleneck in off-road autonomous driving research lies in the scarcity of large-scale, high-quality datasets and benchmarks. To bridge this gap, we present ORAD-3D, which, to the best of our knowledge, is the largest dataset specifically curated for off-road autonomous driving. ORAD-3D cove…

2026

Best of Sim and Real: Decoupled Visuomotor Manipulation Via Learning Control in Simulation and Perception in Real

ICRA 2026poster

Sim-to-real transfer remains a fundamental challenge in robot manipulation due to the entanglement of perception and control in end-to-end learning. We present a decoupled framework that learns each component where it is most reliable: control policies are trained in simulation with privileged state…

2026

Beyond Endpoints: Path-Centric Reasoning for Vectorized Off-Road Network Extraction

CVPR 2026

Deep learning has advanced vectorized road extraction in urban settings, yet off-road environments remain underexplored and challenging. A significant domain gap causes advanced models to fail in wild terrains due to two key issues: lack of large-scale vectorized datasets and structural weakness in

Cited by 0SourcecodeScholar
2026

CompTrack: Information Bottleneck-Guided Low-Rank Dynamic Token Compression for Point Cloud Tracking

AAAI 2026technical

3D single object tracking (SOT) in LiDAR point clouds is a critical task in computer vision and autonomous driving. Despite great success having been achieved, the inherent sparsity of point clouds introduces a dual-redundancy challenge that limits existing trackers: (1) vast spatial redundancy from

Cited by 0SourcePDFScholar
2026

CurvZO: Adaptive Curvature-Guided Sparse Zeroth-Order Optimization for Efficient LLM Fine-Tuning

ICML 2026poster

Fine-tuning large language models (LLMs) with backpropagation achieves high performance but incurs substantial memory overhead, limiting scalability on resource-constrained hardware. Zeroth-order (ZO) optimization provides a memory-efficient alternative by relying solely on forward passes, yet it ty…

Cited by 0SourceScholar
2026

DiffSemanticFusion: Semantic Raster BEV Fusion for Autonomous Driving via Online Map Diffusion

RA-L 2026

Autonomous driving requires accurate scene understanding, including road geometry, traffic agents, and their semantic relationships. In online HD map generation scenarios, raster-based representations are well-suited to vision models but lack geometric precision, while graph-based representations re

Cited by 1SourcecodeScholar
2026

Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns

ICLR 2026poster

Recent progress in large reasoning models for challenging mathematical reasoning has been driven by reinforcement learning (RL). Incorporating long chain-of-thought (CoT) data during mid-training has also been shown to substantially improve reasoning depth. However, current approaches often utiliz…

Cited by 0SourceScholar
2026

FlowSight: Vision-Based Artificial Lateral Line Sensor for Water Flow Perception

ICRA 2026poster

This article presents a novel vision-based artificial lateral line (ALL) sensor, FlowSight, enhancing the perception capabilities of underwater robots. Through an autonomous vision system, FlowSight allows for simultaneous sensing the speed and direction of local water flow without relying on extern…

Cited by 0SourceScholar
2026

ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution

ICML 2026poster

Recently, large language models (LLMs) have shown remarkable reasoning abilities by producing long reasoning traces. However, as the sequence length grows, the key-value (KV) cache expands linearly, incurring significant memory and computation costs. Existing KV cache eviction methods mitigate this …

Cited by 0SourceScholar
2026

Graph Domain Adaptation via Homophily-Agnostic Reconstructing Structure

AAAI 2026technical

Graph Domain Adaptation (GDA) transfers knowledge from labeled source graphs to unlabeled target graphs, addressing the challenge of label scarcity. However, existing GDA methods typically assume that both source and target graphs exhibit homophily, leading existing methods to perform poorly when he

Cited by 0SourcePDFScholar
2026

GuardAlign: Robust Safety Alignment in Multimodal Large Language Models

ICLR 2026poster

Multimodal large language models (MLLMs) have achieved remarkable progress in vision–language reasoning tasks, yet ensuring their safety remains a critical challenge. Recent input-side defenses detect unsafe images with CLIP and prepend safety prefixes to prompts, but they still suffer from inaccura…

Cited by 0SourceScholar
2026

HEV Generative Sandbox: A Framework for Assessing Domain-Specific Social Risks Through Human-LLM Simulation

AAAI 2026technical

Deploying Large Language Models (LLMs) in specialized domains introduces significant societal and compliance risks, including bias amplification, misinformation propagation, and privacy violations. These risks predominantly emerge from the dynamic interactions between LLMs and humans in specific con

Cited by 0SourcePDFScholar
2026

Hierarchical Semantic Alignment for Image Clustering

AAAI 2026technical

Image clustering is a classic problem in computer vision, which categorizes images into different groups. Recent studies utilize nouns as external semantic knowledge to improve clustering performance. However, these methods often overlook the inherent ambiguity of nouns, which can distort semantic r

Cited by 0SourcePDFScholar
2026

History to Future: Evolving Agent with Experience and Thought for Zero-shot Vision-and-Language Navigation

CVPR 2026

Vision-and-Language Navigation in Continuous Environment (VLN-CE) requires an agent to follow language instructions to navigate the target destination. With the advancement of large language models (LLMs), recent efforts have explored adapting them for zero-shot VLN-CE, offering a promising solution

Cited by 0SourceScholar
2026

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

RSS 2026poster

Recent robot foundation models largely rely on large-scale behavior cloning, which imitates expert actions but discards transferable dynamics knowledge embedded in heterogeneous embodied data. While the Unified World Model (UWM) formulation has the potential to leverage such diverse data, existing i…

Cited by 0SourceScholar
2026

LEGAL∆: ENHANCING LEGAL REASONING IN LLMS VIA REINFORCEMENT LEARNING WITH CHAIN-OF-THOUGHT GUIDED INFORMATION GAIN

ICASSP 2026poster

Legal Artificial Intelligence (LegalAI) has achieved notable advances in automating judicial decision-making with the support of Large Language Models (LLMs). However, existing legal LLMs still struggle to generate reliable and interpretable reasoning processes. They often default to fast-thinking b…

Cited by 0SourcePDFScholar
2026

Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos

CVPR 2026

Pre-training on large-scale videos to improve reinforcement learning efficiency is promising yet remains challenging. Existing methods typically treat the agent as an indivisible entity, modeling motion patterns globally. Such global modeling is tightly coupled with the morphology, hindering transfe

Cited by 0SourceScholar
2026

Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation

ICLR 2026poster

Multimodal large language models (MLLMs) have achieved remarkable progress in vision–language reasoning, yet they remain vulnerable to hallucination, where generated content deviates from the visual evidence. Existing mitigation strategies either demand costly supervision during training or introduc…

Cited by 0SourceScholar
2026

MapDream: Task-Driven Map Learning for Vision-Language Navigation

ICML 2026poster

Vision-Language Navigation (VLN) requires agents to follow natural language instructions in partially observed 3D environments, motivating map representations that aggregate spatial context beyond local perception. However, most existing approaches rely on hand-crafted maps constructed independently…

Cited by 0SourceScholar
2026

Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction

AAAI 2026technical

Reconstructing dense geometry for dynamic scenes from a monocular video is a critical yet challenging task. Recent memory-based methods enable efficient online reconstruction, but they fundamentally suffer from a Memory Demand Dilemma: The memory representation faces an inherent conflict be

Cited by 0SourcePDFScholar
2026

MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming

AAAI 2026technical

Vision-Language Navigation (VLN) tasks often leverage panoramic RGB and depth inputs to provide rich spatial cues for action planning, but these sensors can be costly or less accessible in real-world deployments. Recent approaches based on Vision-Language Action (VLA) models achieve strong results w

Cited by 0SourcePDFScholar
2026

PGP-DOR: A Point-Grid-Point Scheme for Efficient Dynamic Object Removal

ICRA 2026poster

In the field of autonomous driving, constructing high-precision maps, typically represented as 3D point cloud maps or bird's-eye view (BEV) grid maps, is essential for both offline and online applications. However, the presence of dynamic objects within a scene can introduce artifacts and noise that…

Cited by 0SourceScholar
2026

PMSPO: Progressive Matching and Semantic-Aware Policy Optimization for Camouflaged Object Detection

ICML 2026poster

Reinforcement learning-based Multimodal Large Language Models (MLLMs) provide new perspectives for visual grounding, yet face significant challenges in Camouflaged Object Detection (COD) where objects blend seamlessly with backgrounds. This stems primarily from: difficulties in multi-object matching…

Cited by 0SourceScholar
2026

Principled SVD-based Delta Compression via Quantization Error Minimization

ICML 2026poster

Supervised Fine-Tuning (SFT) empowers Large Language Models (LLMs) with exceptional performance on specialized tasks, but it yields dense, high-dimensional delta parameters that pose severe storage and distribution challenges. Singular Value Decomposition (SVD)-based compression offers a compact rep…

Cited by 0SourceScholar
2026

Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language Models

CVPR 2026

As vision-language models (VLMs) are increasingly deployed in open-world scenarios, they can be easily induced by visual jailbreak attacks to generate harmful content, posing serious risks to model safety and trustworthy usage.Recent activation steering methods inject directional vectors into model

Cited by 0SourceScholar
2026

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation

CVPR 2026

Vision-Language Navigation requires agents to act coherently over long horizons by understanding not only local visual context but also how far they have advanced within a multi-step instruction.However, recent Vision-Language-Action models focus on direct action prediction and earlier progress meth

Cited by 0SourceScholar
2026

Robustifying Vision-Language Models via Test-Time Prompt Adaptation

ICML 2026poster

Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing test-time adaptation methods typically rely on sample-level confidence heuristics, overlooking the intrinsic distributional…

Cited by 0SourceScholar
2026

Run, Ruminate, and Regulate: A Dual-process Thinking System for Vision-and-Language Navigation

AAAI 2026technical

Vision-and-Language Navigation (VLN) requires an agent to dynamically explore complex 3D environments following human instructions. Recent research underscores the potential of harnessing large language models (LLMs) for VLN, given their commonsense knowledge and general reasoning capabilities. Desp

Cited by 0SourcePDFScholar
2026

SAGA: Structural Aggregation Guided Alignment with Dynamic View and Neighborhood Order Selection for Multiview Graph Domain Adaptation

ICLR 2026poster

Graph domain adaptation (GDA) transfers knowledge from a labeled source graph to an unlabeled target graph to alleviate label scarcity. In multi-view graphs, the challenge of mitigating domain shift is constrained by structural information across various views. Moreover, within each view, structures…

Cited by 0SourcecodeScholar
2026

TGPO: Efficient Policy Optimization through Sequence Anchor and Information Gating

ICML 2026poster

Reinforcement learning from verifiable rewards (RLVR) has become an important paradigm for enhancing the reasoning capabilities of large language models, while it also involves a persistent tradeoff between optimization stability and learning efficiency. Token-level importance weighting supports fin…

Cited by 0SourceScholar
2026

TacFlex: Multi-Mode Tactile Imprints Simulation for Visuotactile Sensors with Coating Patterns

ICRA 2026poster

Visuotactile sensors can provide rich contact information for robots. However, how to build a high-fidelity visuotactile simulator that supports multi-mode tactile imprints and various sensor configurations remains a challenging problem. In this paper, we present TacFlex, a flexible simulator for vi…

Cited by 0SourceScholar
2026

UniUncer: Unified Dynamic–Static Uncertainty for End-To-End Driving

ICRA 2026poster

End-to-end (E2E) driving has become a cornerstone of both industry deployment and academic research, offering a single learnable pipeline that maps multi-sensor inputs to actions while avoiding hand-engineered modules. However, the reliability of such pipelines strongly depends on how well they hand…

2026

Unified Map Prior Encoder for Mapping and Planning

ICRA 2026poster

Online mapping and end-to-end (E2E) planning in autonomous driving are still largely sensor-centric, leaving rich map priors—HD/SD vector maps, rasterized SD maps, and satellite imagery—underused due to heterogeneity, pose drift, and inconsistent availability at test time. We present emph{UMPE}, a U…

2026

VGDM: Visual Localization-Guided 3D Dental Segmentation via Extrinsic–Intrinsic Bridging

IJCAI 2026

3D dental segmentation is a key task in digital dentistry. In real intraoral scans data (IOS), occlusion, scanning noise, and reconstruction artifacts often break down the geometric separation structure between teeth, resulting in adjacent teeth being incorrectly merged or a single tooth being over-

Cited by 0Scholar
2025

A*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource Settings

NeurIPS 2025poster

Large Reasoning Models (LRMs) achieve superior performance by extending the thought length. However, a lengthy thinking trajectory leads to reduced efficiency. Most of the existing methods are stuck in the assumption of overthinking and attempt to reason efficiently by compressing the Chain-of-Thoug…

Cited by 0SourcecodeScholar
2025

Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language Navigation

NeurIPS 2025poster

Vision-Language Navigation is a critical task for developing embodied agents that can follow natural language instructions to navigate in complex real-world environments. Recent advances by finetuning large pretrained models have significantly improved generalization and instruction grounding compa…

Cited by 0SourceScholar
2025

COAST: Enhancing the Code Debugging Ability of LLMs through Communicative Agent Based Data Synthesis

NAACL 2025findings

Code debugging is a vital stage of software development, essential for ensuring the reliability and performance of Large Language Models (LLMs) in the code generation task. Human debugging typically follows a multi-stage process, which includes Bug Localization, Bug Identification, Code Repair, and…

2025

ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation

ACL 2025long

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in chart understanding tasks. However, interpreting charts with textual descriptions often leads to information loss, as it fails to fully capture the dense information embedded in charts. In contrast, parsing charts…

2025

CoDe: Communication Delay-Tolerant Multi-Agent Collaboration via Dual Alignment of Intent and Timeliness

AAAI 2025technical

Communication has been widely employed to enhance multi-agent collaboration. Previous research has typically assumed delay-free communication, a strong assumption that is challenging to meet in practice. However, real-world agents suffer from channel delays, receiving messages sent at different time…

Cited by 0SourcePDFScholar
2025

Cooperation of Experts: Fusing Heterogeneous Information with Large Margin

ICML 2025poster

Fusing heterogeneous information remains a persistent challenge in modern data analysis. While significant progress has been made, existing approaches often fail to account for the inherent heterogeneity of object patterns across different semantic spaces. To address this limitation, we propose the…

Cited by 0SourcePDFScholar
2025

DAEA: Enhancing Entity Alignment in Real-World Knowledge Graphs Through Multi-Source Domain Adaptation

COLING 2025main

Entity Alignment (EA) is a critical task in Knowledge Graph (KG) integration, aimed at identifying and matching equivalent entities that represent the same real-world objects. While EA methods based on knowledge representation learning have shown strong performance on synthetic benchmark datasets su…

2025

DAMA: Data- and Model-aware Alignment of Multi-modal LLMs

ICML 2025poster

Direct Preference Optimization (DPO) has shown effectiveness in aligning multi-modal large language models (MLLM) with human preferences. However, existing methods exhibit an imbalanced responsiveness to the data of varying hardness, tending to overfit on the easy-to-distinguish data while underfit…

Cited by 0SourcePDFScholar
2025

DCAD-2000: A Multilingual Dataset across 2000+ Languages with Data Cleaning as Anomaly Detection

NeurIPS 2025poster

The rapid development of multilingual large language models (LLMs) highlights the need for high-quality, diverse, and well-curated multilingual datasets. In this paper, we introduce DCAD-2000 (Data Cleaning as Anomaly Detection), a large-scale multilingual corpus constructed from newly extracted Com…

Cited by 0SourcecodeScholar
2025

Distance between Relevant Information Pieces Causes Bias in Long-Context LLMs

ACL 2025finding

Positional bias in large language models hinders their ability to effectively process long inputs. A prominent example is the “lost in the middle” phenomenon, where LLMs struggle to utilize relevant information situated in the middle of the input. While prior research primarily focuses on single pie…

2025

Double-Feedback: Enhancing Large Language Models Reasoning in Robotic Tasks by Knowledge Graphs

RA-L 2025

Large language models (LLMs) have demonstrated remarkable reasoning capabilities. However, in real-world robotic tasks, LLMs face grounding issues and lack precise feedback, resulting in the generated solutions deviating from the actual situation. In this paper, we propose Double-Feedback, a method

Cited by 0SourceScholar
2025

Dynamic Multimodal Prototype Learning in Vision-Language Models

ICCV 2025poster

With the increasing attention to pre-trained vision-language models (VLMs), e.g., CLIP, substantial efforts have been devoted to many downstream tasks, especially in test-time adaptation (TTA). However, previous works focus on learning prototypes only in the textual modality while overlooking the am…

Cited by 0SourcePDFScholar
2025

EA-Vit: Efficient Adaptation for Elastic Vision Transformer

ICCV 2025poster

Vision Transformers (ViTs) have emerged as a foundational model in computer vision, excelling in generalization and adaptation to downstream tasks. However, deploying ViTs to support diverse resource constraints typically requires retraining multiple, size-specific ViTs, which is both time-consuming…

2025

Enhancing Adversarial Transferability with Checkpoints of a Single Model's Training

CVPR 2025poster

Adversarial attacks threaten the integrity of deep neural networks (DNNs), particularly in high-stakes applications. In this paper, we present a novel black-box adversarial attack that leverages the diverse checkpoints generated during a single model's training trajectory. Unlike conventional ensemb…

2025

Enhancing CLIP Robustness via Cross-Modality Alignment

NeurIPS 2025spotlight

Vision-language models (VLMs) such as CLIP demonstrate strong generalization in zero-shot classification but remain highly vulnerable to adversarial perturbations. Existing methods primarily focus on adversarial fine-tuning or prompt optimization, they often overlook the gaps in CLIP’s encoded featu…

Cited by 0SourceScholar
2025

Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation

CVPR 2025poster

Weakly Supervised Semantic Segmentation (WSSS) with image-level labels aims to achieve pixel-level predictions using Class Activation Maps (CAMs). Recently, Contrastive Language-Image Pre-training (CLIP) has been introduced in WSSS. However, recent methods primarily focus on image-text alignment for…

2025

Exploring the Impact of Personality Traits on LLM Bias and Toxicity

EMNLP 2025

With the different roles that AI is expected to play in human life, imbuing large language models (LLMs) with different personalities has attracted increasing research interest. While the “personification” enhances human experiences of interactivity and adaptability of LLMs, it gives rise to critica

Cited by 0SourcePDFScholar
2025

Fine-Grained and Efficient Self-Unlearning with Layered Iteration

IJCAI 2025

As machine learning models become widely deployed in data-driven applications, ensuring compliance with the 'right to be forgotten' as required by many privacy regulations is vital for safeguarding user privacy. To forget the given data, existing re-labeling based unlearning methods employ a single-

2025

From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora

EMNLP 2025

Continued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages. However, the unaligned nature of such data limits its ability to effectively capture cross-lingual semantics. In contrast, mult

2025

GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning

ACL 2025finding

Large Language Models (LLMs) fine-tuning technologies have achieved remarkable results. However, traditional LLM fine-tuning approaches face significant challenges: they require large Floating Point(FP) computation, raising privacy concerns when handling sensitive data, and are impractical for resou…

2025

Infer the Whole from a Glimpse of a Part: Keypoint-Based Knowledge Graph for Vehicle Re-Identification

AAAI 2025technical

Vehicle re-identification aims to match vehicles across non-overlapping camera views. Many existing methods extract features from one specific image, and these methods lack view-invariance when comparing vehicles of different orientations. As a result, discriminative parts obscured by viewpoint chan…

Cited by 0SourcePDFScholar
2025

KBAlign: Efficient Self Adaptation on Specific Textual Knowledge Bases

EMNLP 2025

Although retrieval-augmented generation (RAG) remains essential for knowledge-based question answering (KBQA), current paradigms face critical challenges under specific domains. Existing methods struggle with targeted adaptation on small-scale KBs: vanilla unsupervised training exhibits poor effecti

2025

LLM×MapReduce: Simplified Long-Sequence Processing using Large Language Models

ACL 2025long

We propose a training-free framework that enables large language models (LLMs) to effectively process long texts, using a divide-and-conquer strategy for comprehensive document understanding.The proposed LLM×MapReduce framework splits the entire document into several chunks for LLMs to read and then…

Cited by 0SourcePDFScholar
2025

Linguistics-Vision Monotonic Consistent Network for Sign Language Production

ICASSP 2025accepted

Sign Language Production (SLP) aims to generate sign videos corresponding to spoken language sentences, where the conversion of sign Glosses to Poses (G2P) is the key step. Due to the cross-modal semantic gap and the lack of word-action correspondence labels for strong supervision alignment, the SLP…

Cited by 0SourceScholar
2025

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information

ACL 2025finding

Recent advancements in large language models (LLMs) have markedly improved their capacity to handle long text inputs; however, current models, including GPT-4o, still exhibit unsatisfactory performance in long-form generation. Generating high-quality long-form content still remains a significant cha…

2025

MALoRA: Mixture of Asymmetric Low-Rank Adaptation for Enhanced Multi-Task Learning

NAACL 2025findings

Parameter-Efficient Fine-Tuning (PEFT) methods such as LoRA have significantly improved the adaptation of LLMs to downstream tasksin a resource-efficient manner. However, in multi-task scenarios, challenges such as training imbalance and the seesaw effect frequently emerge. Mixture-of-LoRA (MoLoRA),…

Cited by 1SourcePDFScholar
2025

MISCGrasp: Leveraging Multiple Integrated Scales and Contrastive Learning for Enhanced Volumetric Grasping

IROS 2025

Robotic grasping faces challenges in adapting to objects with varying shapes and sizes. In this paper, we introduce MISCGrasp, a volumetric grasping method that integrates multi-scale feature extraction with contrastive feature enhancement for self-adaptive grasping. We propose a query-based interac

Cited by 2SourcecodeScholar
2025

MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training Smoothing

CVPR 2025poster

Deep visual odometry has demonstrated great advancements by learning-to-optimize technology. This approach heavily relies on the visual matching across frames. However, ambiguous matching in challenging scenarios leads to significant errors in geometric modeling and bundle adjustment optimization,…

Cited by 0SourcePDFScholar
2025

MiLoRA: Harnessing Minor Singular Components for Parameter-Efficient LLM Finetuning

NAACL 2025long

Efficient finetuning of large language models (LLMs) aims to adapt the LLMs with reduced computational and memory costs. Previous LoRA-based approaches initialize the low-rank matrices with Gaussian distribution and zero values while keeping the original weight matrices frozen. However, the trainabl…

2025

Mixture of Multimodal Adapters for Sentiment Analysis

NAACL 2025long

Pre-trained language model (PLM) have achieved great success in text sentiment analysis. However, in practical applications, sentiment is not only conveyed through language but also hidden in other modalities. Therefore, multimodal sentiment analysis (MSA) has attracted increasing research interest.…

2025

MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic Segmentation

AAAI 2025technical

Weakly Supervised Semantic Segmentation (WSSS) with image-level labels typically uses Class Activation Maps (CAM) to achieve dense predictions. Recently, Vision Transformer (ViT) has provided an alternative to generate localization maps from class-patch attention. However, due to insufficient constr…

2025

Multi-Domain Graph Foundation Models: Robust Knowledge Transfer via Topology Alignment

ICML 2025poster

Recent advances in CV and NLP have inspired researchers to develop general-purpose graph foundation models through pre-training across diverse domains. However, a fundamental challenge arises from the substantial differences in graph topologies across domains. Additionally, real-world graphs are oft…

Cited by 2SourcePDFScholar
2025

NeuGrasp: Generalizable Neural Surface Reconstruction with Background Priors for Material-Agnostic Object Grasp Detection

ICRA 2025

Robotic grasping in scenes with transparent and specular objects presents great challenges for methods relying on accurate depth information. In this paper, we introduce NeuGrasp, a neural surface reconstruction method that leverages background priors for material-agnostic grasp detection. NeuGrasp

Cited by 3SourcecodeScholar
2025

OA-Stereo: Self-Supervised Opti-Acoustic Stereo for Robust 3D Perception of Underwater Vehicles

RA-L 2025

Accurate 3D perception is essential for underwater vehicles in tasks such as seabed mapping, structural reconstruction, and environmental monitoring. However, optical cameras struggle in underwater environments due to light attenuation, scattering, and blurring, while forward-looking sonar suffers f

Cited by 1SourcecodeScholar
2025

On LLM-Based Scientific Inductive Reasoning Beyond Equations

EMNLP 2025

As large language models (LLMs) increasingly exhibit human-like capabilities, a fundamental question emerges: How can we enable LLMs to learn the underlying patterns from limited examples in entirely novel environments and apply them effectively? This question is central to the ability of LLMs in in

2025

PGP-DOR: A Point-Grid-Point Scheme for Efficient Dynamic Object Removal

RA-L 2025

In the field of autonomous driving, constructing high-precision maps, typically represented as 3D point cloud maps or bird's-eye view (BEV) grid maps, is essential for both offline and online applications. However, the presence of dynamic objects within a scene can introduce artifacts and noise that

Cited by 0SourceScholar
2025

Pragmatic Inference Chain (PIC) Improving LLMs’ Reasoning of Authentic Implicit Toxic Language

EMNLP 2025

The rapid development of large language models (LLMs) gives rise to ethical concerns about their performance, while opening new avenues for developing toxic language detection techniques. However, LLMs’ unethical output and their capability of detecting toxicity have primarily been tested on languag

2025

QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation

NeurIPS 2025poster

The rise of GPU-based high-performance computing (HPC) has driven the widespread adoption of parallel programming models such as CUDA. Yet, the inherent complexity of parallel programming creates a demand for the automated sequential-to-parallel approaches. However, data scarcity poses a significant…

Cited by 0SourcecodeScholar
2025

RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards

ICLR 2025poster

Retrieval-Augmented Generation (RAG) has proven its effectiveness in mitigating hallucinations in Large Language Models (LLMs) by retrieving knowledge from external resources. To adapt LLMs for the RAG systems, current approaches use instruction tuning to optimize LLMs, improving their ability to ut…

2025

RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework

ACL 2025long

Retrieval-Augmented Generation (RAG) is a powerful approach that enables large language models (LLMs) to incorporate external knowledge. However, evaluating the effectiveness of RAG systems in specialized scenarios remains challenging due to the high costs of data construction and the lack of suitab…

2025

ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization

EMNLP 2025

Recent advances in Chain-of-Thought (CoT) prompting have substantially improved the reasoning capabilities of Large Language Models (LLMs). However, these methods often suffer from overthinking, leading to unnecessarily lengthy or redundant reasoning traces. Existing approaches attempt to mitigate t

2025

RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing

NeurIPS 2025poster

Recent advances in self-supervised learning for Vision Transformers (ViTs) have fueled breakthroughs in remote sensing (RS) foundation models. However, the quadratic complexity of self-attention poses a significant barrier to scalability, particularly for large models and high-resolution images. Whi…

Cited by 0SourcecodeScholar
2025

SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning

IROS 2025

Preference-based Reinforcement Learning (PbRL) methods provide a solution to avoid reward engineering by learning reward models based on human preferences. However, poor feedback- and sample- efficiency still remain the problems that hinder the application of PbRL. In this paper, we present a novel

Cited by 0SourcecodeScholar
2025

TIMotion: Temporal and Interactive Framework for Efficient Human-Human Motion Generation

CVPR 2025poster

Human-human motion generation is essential for understanding humans as social beings. Current methods fall into two main categories: single-person-based methods and separate modeling-based methods. To delve into this field, we abstract the overall generation process into a general framework MetaMoti…

Cited by 0SourcePDFScholar
2025

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

ICLR 2025poster

Retrieval-augmented generation (RAG) is an effective technique that enables large language models (LLMs) to utilize external knowledge sources for generation. However, current RAG systems are solely based on text, rendering it impossible to utilize vision information like layout and images that pla…

2025

WaveAR: Wavelet-Aware Continuous Autoregressive Diffusion for Accurate Human Motion Prediction

NeurIPS 2025poster

This work tackles a challenging problem: stochastic human motion prediction (SHMP), which aims to forecast diverse and physically plausible future pose sequences based on a short history of observed motion. While autoregressive sequence models have excelled in related generation tasks, their relianc…

Cited by 0SourceScholar
2025

Where Does This Data Come From? Enhanced Source Inference Attacks in Federated Learning

IJCAI 2025

Federated learning (FL) enables collaborative model training without exposing raw data, offering a privacy-aware alternative to centralized learning. However, FL remains vulnerable to various privacy attacks that exploit shared model updates, including membership inference, property inference, and g

Cited by 0SourcePDFScholar
2025

Why Stop at One Error? Benchmarking LLMs as Data Science Code Debuggers for Multi-Hop and Multi-Bug Errors

EMNLP 2025

LLMs are transforming software development, yet current code generation and code repair benchmarks mainly assess syntactic and functional correctness in simple, single-error cases. LLMs’ capabilities to autonomously find and fix runtime logical errors in complex data science code remain largely unex

2024

Beyond Redundancy: Information-aware Unsupervised Multiplex Graph Structure Learning

NeurIPS 2024poster

Unsupervised Multiplex Graph Learning (UMGL) aims to learn node representations on various edge types without manual labeling. However, existing research overlooks a key factor: the reliability of the graph structure. Real-world data often exhibit a complex nature and contain abundant task-irrelevan…

2024

Boosting Few-Shot Learning via Attentive Feature Regularization

AAAI 2024technical

Few-shot learning (FSL) based on manifold regularization aims to improve the recognition capacity of novel objects with limited training samples by mixing two samples from different categories with a blending factor. However, this mixing operation weakens the feature representation due to the linear…

Cited by 11SourcePDFScholar
2024

Breaking of Brightness Consistency in Optical Flow With a Lightweight CNN Network

RA-L 2024

The sparse optical flow method is a fundamental task in computer vision. However, its reliance on the assumption of constant environmental brightness constrains its applicability in high dynamic range (HDR) scenes. In this study, we propose a novel approach aimed at transcending image color informat

Cited by 13SourcecodeScholar
2024

Deep Neighbor Layer Aggregation for Lightweight Self-Supervised Monocular Depth Estimation

ICASSP 2024accepted

With the frequent use of self-supervised monocular depth estimation in robotics and autonomous driving, the model’s efficiency is becoming increasingly important. Most current approaches apply much larger and more complex networks to improve the precision of depth estimation. Some researchers incorp…

Cited by 0SourceScholar
2024

Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models

NeurIPS 2024poster

Fine-tuning is a crucial process for adapting large language models (LLMs) to diverse applications. In certain scenarios, such as multi-tenant serving, deploying multiple LLMs becomes necessary to meet complex demands. Recent studies suggest decomposing a fine-tuned LLM into a base model and corresp…

2024

Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages

ACL 2024long

While large language models (LLMs) have been pre-trained on multilingual corpora, their performance still lags behind in most languages compared to a few resource-rich languages. One common approach to mitigate this issue is to translate training data from resource-rich languages into other language…

2024

Enhancing Recipe Retrieval with Foundation Models: A Data Augmentation Perspective

ECCV 2024poster

"Learning recipe and food image representation in common embedding space is non-trivial but crucial for cross-modal recipe retrieval. In this paper, we propose a new perspective for this problem by utilizing foundation models for data augmentation. Leveraging on the remarkable capabilities of founda…

2024

Enhancing Zero-Shot Vision Models by Label-Free Prompt Distribution Learning and Bias Correcting

NeurIPS 2024spotlight

Vision-language models, such as CLIP, have shown impressive generalization capacities when using appropriate text descriptions. While optimizing prompts on downstream labeled data has proven effective in improving performance, these methods entail labor costs for annotations and are limited by their…

2024

Exploring the Limitations and Implications of the JIGSAWS Dataset for Robot-Assisted Surgery

RA-L 2024

The JHU-ISI Gesture and Skill Assessment Working Set (JIGSAWS) dataset has proven to be a foundational component of modern work on the skill analysis of robotic surgeons. In particular, methods using either the system's kinematics or video data have shown to be able to classify operators into distin

Cited by 3SourceScholar
2024

FAST: A Dual-tier Few-Shot Learning Paradigm for Whole Slide Image Classification

NeurIPS 2024poster

The expensive fine-grained annotation and data scarcity have become the primary obstacles for the widespread adoption of deep learning-based Whole Slide Images (WSI) classification algorithms in clinical practice. Unlike few-shot learning methods in natural images that can leverage the labels of…

2024

How to Learn Domain-Invariant Representations for Visual Reinforcement Learning: An Information-Theoretical Perspective

IJCAI 2024poster

Despite the impressive success in visual control challenges, Visual Reinforcement Learning (VRL) policies have struggled to generalize to other scenarios. Existing works attempt to empirically improve the generalization capability, lacking theoretical support. In this work, we explore how to learn d…

2024

INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair

ACL 2024findings

This paper introduces INTERVENOR (INTERactiVE chaiN Of Repair), a system designed to emulate the interactive code repair processes observed in humans, encompassing both code diagnosis and code repair. INTERVENOR prompts Large Language Models (LLMs) to play distinct roles during the code repair proce…

2024

Learning with Mixture of Prototypes for Out-of-Distribution Detection

ICLR 2024poster

Out-of-distribution (OOD) detection aims to detect testing samples far away from the in-distribution (ID) training data, which is crucial for the safe deployment of machine learning models in the real world. Distance-based OOD detection methods have emerged with enhanced deep representation learning…

2024

LoRA-Flow: Dynamic LoRA Fusion for Large Language Models in Generative Tasks

ACL 2024long

LoRA employs lightweight modules to customize large language models (LLMs) for each downstream task or domain, where different learned additional modules represent diverse skills. Combining existing LoRAs to address new tasks can enhance the reusability of learned LoRAs, particularly beneficial for…

2024

MCTS: A Multi-Reference Chinese Text Simplification Dataset

COLING 2024main

Text simplification aims to make the text easier to understand by applying rewriting transformations. There has been very little research on Chinese text simplification for a long time. The lack of generic evaluation data is an essential reason for this phenomenon. In this paper, we introduce MCTS,…

2024

MatPlotAgent: Method and Evaluation for LLM-Based Agentic Scientific Data Visualization

ACL 2024findings

Scientific data visualization plays a crucial role in research by enabling the direct display of complex information and assisting researchers in identifying implicit patterns. Despite its importance, the use of Large Language Models (LLMs) for scientific data visualization remains rather unexplored…

2024

OneBit: Towards Extremely Low-bit Large Language Models

NeurIPS 2024poster

Model quantification uses low bit-width values to represent the weight matrices of existing models to be quantized, which is a promising approach to reduce both storage and computational overheads of deploying highly anticipated LLMs. However, current quantization methods suffer severe performance d…

2024

PGN: The RNN's New Successor is Effective for Long-Range Time Series Forecasting

NeurIPS 2024poster

Due to the recurrent structure of RNN, the long information propagation path poses limitations in capturing long-term dependencies, gradient explosion/vanishing issues, and inefficient sequential execution. Based on this, we propose a novel paradigm called Parallel Gated Network (PGN) as the new suc…

2024

Pluggable Neural Machine Translation Models via Memory-augmented Adapters

COLING 2024main

Although neural machine translation (NMT) models perform well in the general domain, it remains rather challenging to control their generation behavior to satisfy the requirement of different users. Given the expensive training cost and the data scarcity challenge of learning a new model from scratc…

2024

Separate and Conquer: Decoupling Co-occurrence via Decomposition and Representation for Weakly Supervised Semantic Segmentation

CVPR 2024poster

Weakly supervised semantic segmentation (WSSS) with image-level labels aims to achieve segmentation tasks without dense annotations. However attributed to the frequent coupling of co-occurring objects and the limited supervision from image-level labels the challenging co-occurrence problem is widely…

2024

Shadow-Free Membership Inference Attacks: Recommender Systems Are More Vulnerable Than You Thought

IJCAI 2024poster

Recommender systems have been successfully applied in many applications. Nonetheless, recent studies demonstrate that recommender systems are vulnerable to membership inference attacks (MIAs), leading to the leakage of users’ membership privacy. However, existing MIAs relying on shadow training suff…

2024

Text2Reaction : Enabling Reactive Task Planning Using Large Language Models

RA-L 2024

To complete tasks in dynamic environments, robots need to timely update their plans to react to environment changes. Traditional stripe-like or learning-based planners struggle to achieve this due to their high reliance on meticulously predefined planning rules or labeled data. Fortunately, recent w

Cited by 24SourceScholar
2024

UltraLink: An Open-Source Knowledge-Enhanced Multilingual Supervised Fine-tuning Dataset

ACL 2024long

Open-source large language models (LLMs) have gained significant strength across diverse fields. Nevertheless, the majority of studies primarily concentrate on English, with only limited exploration into the realm of multilingual abilities.In this work, we therefore construct an open-source multilin…

2024

What Effects the Generalization in Visual Reinforcement Learning: Policy Consistency with Truncated Return Prediction

AAAI 2024technical

In visual Reinforcement Learning (RL), the challenge of generalization to new environments is paramount. This study pioneers a theoretical analysis of visual RL generalization, establishing an upper bound on the generalization objective, encompassing policy divergence and Bellman error components. M…

2024

∞Bench: Extending Long Context Evaluation Beyond 100K Tokens

ACL 2024long

Processing and reasoning over long contexts is crucial for many practical applications of Large Language Models (LLMs), such as document comprehension and agent construction. Despite recent strides in making LLMs process contexts with more than 100K tokens, there is currently a lack of a standardize…

2023

Bi-Directional Distribution Alignment for Transductive Zero-Shot Learning

CVPR 2023poster

It is well-known that zero-shot learning (ZSL) can suffer severely from the problem of domain shift, where the true and learned data distributions for the unseen classes do not match. Although transductive ZSL (TZSL) attempts to improve this by allowing the use of unlabelled examples from the unseen…

2023

Boosting Whole Slide Image Classification from the Perspectives of Distribution, Correlation and Magnification

ICCV 2023poster

Bag-based multiple instance learning (MIL) methods have become the mainstream for Whole Slide Image (WSI) classification. However, there are still three important issues that have not been fully addressed: (1) positive bags with a low positive instance ratio are prone to the influence of a large num…

Cited by 14PDFcodeScholar
2023

DETRDistill: A Universal Knowledge Distillation Framework for DETR-families

ICCV 2023poster

Transformer-based detectors (DETRs) are becoming popular for their simple framework, but the large model size and heavy time consumption hinder their deployment in the real world. While knowledge distillation (KD) can be an appealing technique to compress giant detectors into small ones for comparab…

Cited by 39PDFScholar
2023

Dasformer: Deep Alternating Spectrogram Transformer For Multi/Single-Channel Speech Separation

ICASSP 2023accepted

For the task of speech separation, previous study usually treats multi-channel and single-channel scenarios as two research tracks with specialized solutions developed respectively. Instead, we propose a simple and unified architecture - DasFormer (Deep alternating spectrogram transFormer) to handle…

Cited by 0SourceScholar
2023

Demystifying Uneven Vulnerability of Link Stealing Attacks against Graph Neural Networks

ICML 2023poster

While graph neural networks (GNNs) dominate the state-of-the-art for exploring graphs in real-world applications, they have been shown to be vulnerable to a growing number of privacy attacks. For instance, link stealing is a well-known membership inference attack (MIA) on edges that infers the prese…

Cited by 26SourcePDFScholar
2023

Explainable Recommendation with Personalized Review Retrieval and Aspect Learning

ACL 2023long

Explainable recommendation is a technique that combines prediction and generation tasks to produce more persuasive results. Among these tasks, textual generation demands large amounts of data to achieve satisfactory accuracy. However, historical user reviews of items are often insufficient, making i…

2023

Feature Shrinkage Pyramid for Camouflaged Object Detection With Transformers

CVPR 2023poster

Vision transformers have recently shown strong global context modeling capabilities in camouflaged object detection. However, they suffer from two major limitations: less effective locality modeling and insufficient feature aggregation in decoders, which are not conducive to camouflaged object detec…

2023

GPDAN: Grasp Pose Domain Adaptation Network for Sim-to-Real 6-DoF Object Grasping

RA-L 2023

In this letter, we propose a novel Grasp Pose Domain Adaptation Network (GPDAN) to achieve sim-to-real domain adaptation for 6-DoF grasp pose detection. The main task of GPDAN is to detect feasible 6-DoF grasp poses in cluttered scenes. A point-wise self-supervised domain classification module with

Cited by 16SourceScholar
2023

High-Resolution Iterative Feedback Network for Camouflaged Object Detection

AAAI 2023technical

Spotting camouflaged objects that are visually assimilated into the background is tricky for both object detection algorithms and humans who are usually confused or cheated by the perfectly intrinsic similarities between the foreground objects and the background surroundings. To tackle this challeng…

2023

Learning from Noisy Data for Semi-Supervised 3D Object Detection

ICCV 2023poster

Pseudo-Labeling (PL) is a critical approach in semi-supervised 3D object detection (SSOD). In PL, delicately selected pseudo-labels, generated by the teacher model, are provided for the student model to supervise the semi-supervised detection framework. However, such a paradigm may introduce misclas…

Cited by 15PDFcodeScholar
2023

Memory-Aided Contrastive Consensus Learning for Co-salient Object Detection

AAAI 2023technical

Co-salient object detection (CoSOD) aims at detecting common salient objects within a group of relevant source images. Most of the latest works employ the attention mechanism for finding common objects. To achieve accurate CoSOD results with high-quality maps and high efficiency, we propose a novel…

2023

Source-free Depth for Object Pop-out

ICCV 2023poster

Depth cues are known to be useful for visual perception. However, direct measurement of depth is often impracticable. Fortunately, though, modern learning-based methods offer promising depth maps by inference in the wild. In this work, we adapt such depth inference models for object segmentation usi…

Cited by 70PDFcodeScholar
2023

TemplateGEC: Improving Grammatical Error Correction with Detection Template

ACL 2023long

Grammatical error correction (GEC) can be divided into sequence-to-edit (Seq2Edit) and sequence-to-sequence (Seq2Seq) frameworks, both of which have their pros and cons. To utilize the strengths and make up for the shortcomings of these frameworks, this paper proposes a novel method, TemplateGEC, wh…

2023

Towards Domain Generalization for Multi-View 3D Object Detection in Bird-Eye-View

CVPR 2023poster

Multi-view 3D object detection (MV3D-Det) in Bird-Eye-View (BEV) has drawn extensive attention due to its low cost and high efficiency. Although new algorithms for camera-only 3D object detection have been continuously proposed, most of them may risk drastic performance degradation when the domain o…

Cited by 25SourcePDFScholar
2023

Transferable and Efficient: Unifying Dynamic Multi-Domain Product Categorization

ACL 2023industry

As e-commerce platforms develop different business lines, a special but challenging product categorization scenario emerges, where there are multiple domain-specific category taxonomies and each of them evolves dynamically over time. In order to unify the categorization process and ensure efficiency…

2022

A Bert Based Joint Learning Model with Feature Gated Mechanism for Spoken Language Understanding

ICASSP 2022accepted

Intent detection (ID) and slot filling (SF) are two major tasks for spoken language understanding (SLU). Recent joint learning approaches consider the relationship between intent detection and slot filling, which leverage the shared knowledge across two tasks to benefit each other. However, most exi…

Cited by 0SourceScholar
2022

A Template-based Method for Constrained Neural Machine Translation

EMNLP 2022main

Machine translation systems are expected to cope with various types of constraints in many practical scenarios. While neural machine translation (NMT) has achieved strong performance in unconstrained cases, it is non-trivial to impose pre-specified constraints into the translation process of NMT mod…

2022

An Efficient Training Approach for Very Large Scale Face Recognition

CVPR 2022poster

Face recognition has achieved significant progress in deep learning era due to the ultra-large-scale and welllabeled datasets. However, training on the outsize datasets is time-consuming and takes up a lot of hardware resource. Therefore, designing an efficient training approach is indispensable. Th…

Cited by 39PDFcodeScholar
2022

CAFE: Learning To Condense Dataset by Aligning Features

CVPR 2022poster

Dataset condensation aims at reducing the network training effort through condensing a cumbersome training set into a compact synthetic one. State-of-the-art approaches largely rely on learning the synthetic data by matching the gradients between the real and synthetic data batches. Despite the intu…

Cited by 277PDFcodeScholar
2022

Learning-based Six-axis Force/Torque Estimation Using GelStereo Fingertip Visuotactile Sensing

IROS 2022poster

Visuotactile sensors have recently attracted much attention in robot communities due to the benefit of high spatial resolution sensing. However, force/torque estimation by visuotactile sensors remains a challenging problem. In this paper, we propose a learning-based six-axis force/torque estimation…

Cited by 10SourceScholar
2022

Long-term Spatio-Temporal Forecasting via Dynamic Multiple-Graph Attention

IJCAI 2022poster

Many real-world ubiquitous applications, such as parking recommendations and air pollution monitoring, benefit significantly from accurate long-term spatio-temporal forecasting (LSTF). LSTF makes use of long-term dependency structure between the spatial and temporal domains, as well as the contextua…

2022

MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better Translators

ACL 2022long

Prompting has recently been shown as a promising approach for applying pre-trained language models to perform downstream tasks. We present Multi-Stage Prompting, a simple and automatic approach for leveraging pre-trained language models to translation tasks. To better mitigate the discrepancy betwee…

2022

Meta-Residual Policy Learning: Zero-Trial Robot Skill Adaptation via Knowledge Fusion

RA-L 2022

Adapting the mastered manipulation skill to novel objects is still challenging for robots. Recent works have attempted to endow the robot with the ability to adapt to unseen tasks by leveraging meta-learning. However, these methods are data-hungry in the training phase, which limits their applicatio

Cited by 23SourcecodeScholar
2022

Towards a Hybrid-ASP Planning Approach With Adjoint Observation for Incomplete Task-Relevant Information

RA-L 2022

In the real world, robot task plans may easily become invalid due to unexpected state dynamics, preventing the robot from accessing the complete task-relevant information. The possible occurrence of information incompleteness during robot plan execution expects the robot to sense the environment and

Cited by 1SourceScholar
2021

DIMSAN: Fast Exploration with the Synergy between Density-based Intrinsic Motivation and Self-adaptive Action Noise

ICRA 2021poster

Exploration in environments with sparse rewards remains a challenging problem in Deep Reinforcement Learning (DRL). For the off-policy method, it usually needs a large number of training samples. With the growing dimensions of state and action space, this method becomes more and more sample-ineffici…

Cited by 0SourceScholar
2021

Learning Motion-Appearance Co-Attention for Zero-Shot Video Object Segmentation

ICCV 2021poster

How to make the appearance and motion information interact effectively to accommodate complex scenarios is a fundamental issue in flow-based zero-shot video object segmentation. In this paper, we propose an Attentive Multi-Modality Collaboration Network (AMC-Net) to utilize appearance and motion inf…

Cited by 78PDFcodeScholar
2021

Towards Adjoint Sensing and Acting Schemes and Interleaving Task Planning for Robust Robot Plan

ICRA 2021poster

Robots operating in open environments expect to have robust plans to achieve tasks successfully under environment uncertainties. However, both partial observability and dynamics of environment states have significantly decreased the robustness of task achievement, making robot task planning much mor…

Cited by 2SourceScholar
2020

Dual Adversarial Network for Deep Active Learning

ECCV 2020poster

Active learning, reducing the cost and workload of annotations, attracts increasing attentions from the community. Current active learning approaches commonly adopted uncertainty-based acquisition functions for the data selection due to their effectiveness. However, data selection based on uncertain…

Cited by 40SourcePDFScholar
2020

Exclusivity-Consistency Regularized Knowledge Distillation for Face Recognition

ECCV 2020poster

Knowledge distillation is an effective tool to compress large pre-trained Convolutional Neural Networks (CNNs) or their ensembles into models applicable to mobile and embedded devices. The success of which mainly comes from two aspects: the designed student network and the exploited knowledge. Howev…

2020

Grasp State Assessment of Deformable Objects Using Visual-Tactile Fusion Perception

ICRA 2020poster

Humans can quickly determine the force required to grasp a deformable object to prevent its sliding or excessive deformation through vision and touch, which is still a challenging task for robots. To address this issue, we propose a novel 3D convolution-based visual-tactile fusion deep neural networ…

Cited by 62SourceScholar
2020

Large-Scale Few-Shot Learning via Multi-Modal Knowledge Discovery

ECCV 2020poster

Large-scale few-shot learning aims at identifying hundreds of novel object categories where each category has only a few samples. It is a challenging problem since (1) the identifying process is susceptible to over-fitting with limited samples of an object, and (2) the sample imbalance between a bas…

Cited by 43SourcePDFScholar
2020

Self-Attention Based Visual-Tactile Fusion Learning for Predicting Grasp Outcomes

RA-L 2020

Predicting whether a particular grasp will succeed is critical to performing stable grasping and manipulating tasks. Robots need to combine vision and touch as humans do to accomplish this prediction. The primary problem to be solved in this process is how to learn effective visual-tactile fusion fe

Cited by 67SourceScholar
2019

CityFlow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-Identification

CVPR 2019oral

Urban traffic optimization using traffic cameras as sensors is driving the need to advance state-of-the-art multi-target multi-camera (MTMC) tracking. This work introduces CityFlow, a city-scale traffic camera dataset consisting of more than 3 hours of synchronized HD videos from 40 cameras across 1…

Cited by 511PDFcodeScholar
2019

PAMTRI: Pose-Aware Multi-Task Learning for Vehicle Re-Identification Using Highly Randomized Synthetic Data

ICCV 2019poster

In comparison with person re-identification (ReID), which has been widely studied in the research community, vehicle ReID has received less attention. Vehicle ReID is challenging due to 1) high intra-class variability (caused by the dependency of shape and appearance on viewpoint), and 2) small inte…

Cited by 147PDFcodeScholar
2019

Self-modeling Tracking Control of Crawler Fire Fighting Robot Based on Causal Network

IROS 2019poster

In this paper, a self-modeling method based on a causal network is proposed for the tracking control of the Crawler Fire Fighting Robot (CFFR). The method mainly consists of two parts, one is a motion model, based on data driving, learning to establish the correspondence between control signal seque…

Cited by 0SourceScholar
2018

Detect Globally, Refine Locally: A Novel Approach to Saliency Detection

CVPR 2018poster

Effective integration of contextual information is crucial for salient object detection. To achieve this, most existing methods based on 'skip' architecture mainly focus on how to integrate hierarchical features of Convolutional Neural Networks (CNNs). They simply apply concatenation or element-wise…

Cited by 507SourcePDFScholar