← Search

dong wang

212 accepted papers

2026

A Data-Observation Hybrid Compensation Method for Precise Force Control of Cable-Driven Wrist Exoskeletons in Teleoperation

RA-L 2026

High-precision force control of wearable exoskeletons enables highly transparent force interaction operations, effectively improving the feasibility of teleoperation tasks. We adopted a cable-driven spherical parallel wrist exoskeleton (SPWE) which enables it to reduce the volume and enhance operati

Cited by 0SourceScholar
2026

CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT Tracking

AAAI 2026technical

RGB-Thermal (RGBT) tracking aims to exploit visible and thermal infrared modalities for robust all-weather object tracking. However, existing RGBT trackers struggle to resolve modality discrepancies, which poses great challenges for robust feature representation. This limitation hinders effective cr

Cited by 0SourcePDFScholar
2026

CHAIN OF CORRECTION FOR FULL-TEXT SPEECH RECOGNITION WITH LARGE LANGUAGE MODELS

ICASSP 2026poster

Full-text error correction with Large Language Models (LLMs) for Automatic Speech Recognition (ASR) is attracting increased attention for its ability to address a wide range of error types, such as punctuation restoration and inverse text normalization, across long context. However, challenges remai…

Cited by 0SourcePDFScholar
2026

Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion Policy

ICRA 2026poster

Diffusion-based policies have achieved remarkable results in robotic manipulation but often struggle to adapt rapidly in dynamic scenarios, leading to delayed responses or task failures. We present DCDP, a Dynamic Closed-Loop Diffusion Policy framework that integrates chunk-based action generation w…

2026

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

CVPR 2026

Vision-and-Language Navigation (VLN) requires the agent to navigate based on natural instructions. This task is challenging due to partial observability, which makes it difficult to align perception with language. Recent methods mitigate this by imagining future scenes, yet they rely on vision-based

Cited by 0SourcecodeScholar
2026

Cut Less, Fold More: Model Compression through the Lens of Projection Geometry

ICLR 2026poster

Compressing neural networks without retraining is vital for deployment at scale. We study calibration-free compression through the lens of projection geometry: structured pruning is an axis-aligned projection, whereas model folding performs a low-rank projection via weight clustering. We formalize b…

Cited by 0SourceScholar
2026

Doppler-SLAM: Doppler-Aided Radar-Inertial and LiDAR-Inertial Simultaneous Localization and Mapping

ICRA 2026poster

Simultaneous localization and mapping is a critical capability for autonomous systems. Traditional SLAM approaches often rely on visual or LiDAR sensors and face significant challenges in adverse conditions such as low light or featureless environments. To overcome these limitations, we propose a no…

2026

Dynamic-ICP: Doppler-Aware Iterative Closest Point Registration for Dynamic Scenes

RA-L 2026

Reliable odometry in highly dynamic environments remains challenging when it relies on ICP-based registration: ICP assumes near-static scenes and degrades in repetitive or low-texture geometry. We introduce Dynamic-ICP, a Doppler-aware registration framework. The method (i) estimates ego translation

Cited by 0SourcecodeScholar
2026

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy

CVPR 2026

Humans constantly reason about 3D proximity, the relations between their body and surrounding objects, to guide perception and action in daily life. Whether multimodal large language models (MLLMs) can perform such embodied 3D reasoning remains unclear. To this end, we introduce EgoProx, a benchmark

Cited by 0SourceScholar
2026

Eva-Tracker: ESDF-Update-Free, Visibility-Aware Planning with Target Reacquisition for Robust Aerial Tracking

ICRA 2026poster

The Euclidean Signed Distance Field (ESDF) is widely used in visibility evaluation to prevent occlusions and collisions during tracking. However, frequent ESDF updates introduce considerable computational overhead. To address this issue, we propose Eva-Tracker, a visibility-aware trajectory planning…

2026

Exploring the Potential of Encoder-free Architectures in 3D LMMs

ICLR 2026poster

Encoder-free architectures have been preliminarily explored in the 2D Large Multimodal Models (LMMs), yet it remains an open question whether they can be effectively applied to 3D understanding scenarios. In this paper, we present the first comprehensive investigation into the potential of encoder-f…

Cited by 0SourcecodeScholar
2026

FM-Steer: Enhance Generalist Policies with Value-Guided Cascaded Denoising

CVPR 2026

Humans naturally allocate more time before acting when handling complex tasks in the physical world. This paradigm has recently led to remarkable advances in boosting Large Language Models (LLMs) on complex tasks in digital domains. However, the potential of test-time computing remains largely unexp

Cited by 0SourcecodeScholar
2026

FreeGaussian: Annotation-free Control of Articulated Objects via 3D Gaussian Splats with Flow Derivatives

AAAI 2026technical

Reconstructing controllable Gaussian splats for articulated objects from monocular video is especially challenging due to its inherently insufficient constraints. Existing methods address this by relying on dense masks and manually defined control signals, limiting their real-world applications. In

Cited by 0SourcePDFScholar
2026

From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors

ICLR 2026poster

Existing vision-language-action (VLA) models act in 3D real-world but are typically built on 2D encoders, leaving a spatial reasoning gap that limits generalization and adaptability. Recent 3D integration techniques for VLAs either require specialized sensors and transfer poorly across modalities, o…

Cited by 0SourcecodeScholar
2026

Illumination-Consistent Human-Scene Reconstruction from Monocular Video

CVPR 2026

Reconstructing 3D humans and scenes from monocular videos is a challenging task, particularly due to human motion, varying illumination, and dynamic scene shadows. While recent works have explored scene disentanglement by jointly modeling humans and their surrounding scenes, they often overlook illu

Cited by 0SourceScholar
2026

In Pursuit of Pixel Supervision for Visual Pre-training

CVPR 2026

Data matters. In computer vision, data (or pixels) are the primary source of information containing signals that span from low-level attributes to high-level concepts. At scale, the success of modern vision systems has been closely tied to how data is curated for semantic understanding (e.g., ImageN

Cited by 0SourcecodeScholar
2026

MT-HUBERT: SELF-SUPERVISED MIX-TRAINING FOR FEW-SHOT KEYWORD SPOTTING IN MIXED SPEECH

ICASSP 2026oral

Few-shot keyword spotting aims to detect previously unseen keywords with very limited labeled samples. A pre-training and adaptation paradigm is typically adopted for this task. While effective in clean conditions, most existing approaches struggle with mixed keyword spotting--detecting multiple ove…

Cited by 0SourcePDFScholar
2026

Multi-Session Mapping and Long-Term Localization for Autonomous Vehicles Using Radar

RA-L 2026

Localization of autonomous vehicles in existing maps is crucial for reliable navigation. Using previously constructed maps allows vehicles to estimate their pose without the inherent odometry drift. Building such maps involves aligning data recorded at different times and maintaining the map over ti

Cited by 0SourceScholar
2026

OpenFly: A COMPREHENSIVE PLATFORM FOR AERIAL VISION-LANGUAGE NAVIGATION

ICLR 2026poster

Aerial Vision-Language Navigation (VLN) seeks to guide UAVs by leveraging language instructions and visual cues, establishing a new paradigm for human-UAV interaction. However, the collection of VLN data demands extensive human effort to construct trajectories and corresponding instructions, hinderi…

Cited by 0SourcecodeScholar
2026

Partitioning for Intrinsic Model Inversion Resistance in Collaborative Inference

ICML 2026poster

In collaborative inference (CI), transmitting intermediate representations $Z$ from edge devices enables model inversion attacks (MIA) that reconstruct the original inputs $X$, while existing defenses mainly perturb shallow-layer $Z$ at the cost of utility. We instead ask: *where should an edge–clou…

Cited by 0SourceScholar
2026

Prefix cache aware data reordering for LLM augmented database analytics

ICML 2026poster

LLM-augmented database analytics face a major bottleneck in the costly prefill phase. Although relational tables inherently contain repeated attribute values, standard row-by-row processing produces fragmented prompt layouts that obscure shared prefixes, thereby minimizing opportunities for prefix K…

Cited by 0SourceScholar
2026

PureProof: Diffusion-Resistant Black-box Targeted Attack on Large Vision-Language Models

CVPR 2026

Large Vision-Language Models (VLMs) are increasingly deployed across diverse applications, such as AI agents, yet remain vulnerable to targeted adversarial attacks. However, the practical robustness of such attacks often remains unclear with limited evaluation under defenses. Diffusion-based purific

Cited by 0SourceScholar
2026

RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation

CVPR 2026

RGB-Thermal (RGBT) tracking aims to achieve robust object localization across diverse environmental conditions by fusing visible and thermal infrared modalities. However, existing RGBT trackers rely solely on initial-frame visual information for target modeling, failing to adapt to appearance variat

Cited by 0SourcecodeScholar
2026

RELO: Reinforcement Learning to Localize for Visual Object Tracking

ICML 2026poster

Existing one-stream Transformer-based visual trackers localize targets by training a classification head with a handcrafted spatial prior encoded as a heatmap. However, this heuristic supervision merely serves as a surrogate objective, which misaligns with evaluation metrics such as IoU and AUC. To …

Cited by 0SourceScholar
2026

Rh-3DGS: Robust Open-Vocabulary Scene Understanding via Riemannian Huber Distillation and Manifold-Aware Sampling

ICML 2026poster

Open-vocabulary 3D scene understanding answers free-form text queries over reconstructed scenes. However, lifting dense 2D foundation-model embeddings into 3D Gaussian Splatting (3DGS) is still challenging. Existing 3DGS-based methods often average normalized embeddings in Euclidean space. This igno…

Cited by 0SourceScholar
2026

RoboOmni: Actions Are Just Another Modality for Your Vision-Language Models

ICML 2026poster

Integrating Vision-Language Models (VLMs) into robotics has facilitated the development of generalizable Vision-Language Action (VLA) policies. However, unified discrete frameworks lag behind decoupled continuous designs due to limitations in action chunking and temporal modeling. To address this, w…

Cited by 0SourceScholar
2026

SpecBridge: Spectral Structure Alignment and Transitive Bridging for 3D–2D–Text Pre-Training

IJCAI 2026

Open-vocabulary 3D understanding aims to align 3D representations with a unified vision-language semantic space. However, existing methods suffer from the challenge of structural asymmetry caused by sparse observations and holistic geometries. Additionally, the inherent semantic chasm between discre

Cited by 0Scholar
2026

TGTrack: Temporal Generative Learning for Unified Single Object Tracking

CVPR 2026

Existing single object trackers typically treat temporal modeling superficially by passing limited inter-frame information, such as propagated tokens or template updates, without intrinsic temporal supervision learning. To address this limitation, we propose TGTrack, a new unified tracking framework

Cited by 0SourcecodeScholar
2026

Trajectory Conditioned Cross-Embodiment Skill Transfer

ICRA 2026poster

Learning manipulation skills from human demonstration videos presents a promising yet challenging problem, primarily due to the significant embodiment gap between human body and robot manipulators. Existing methods rely on paired datasets or hand-crafted rewards, which limit scalability and generali…

2026

UETrack: A Unified and Efficient Framework for Single Object Tracking

CVPR 2026

With growing real-world demands, efficient tracking has received increasing attention. However, most existing methods are limited to RGB inputs and struggle in multi-modal scenarios. Moreover, current multi-modal tracking approaches typically use complex designs, making them too heavy and slow for r

Cited by 0SourcecodeScholar
2025

A Non-Contrastive Learning Framework for Sequential Recommendation with Preference-Preserving Profile Generation

ICLR 2025poster

Contrastive Learning (CL) proves to be effective for learning generalizable user representations in Sequential Recommendation (SR), but it suffers from high computational costs due to its reliance on negative samples. To overcome this limitation, we propose the first Non-Contrastive Learning (NCL) f…

Cited by 0SourcePDFScholar
2025

AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations

ICCV 2025poster

Visual grounding (VG) aims to localize target objects in an image based on natural language descriptions. In this paper, we propose AerialVG, a new task focusing on visual grounding from aerial views. Compared to traditional VG, AerialVG poses new challenges, e.g., appearance-based grounding is insu…

2025

AlignBot: Aligning VLM-Powered Customized Task Planning with User Reminders Through Fine-Tuning for Household Robots

ICRA 2025

This paper presents AlignBot, a novel framework designed to optimize VLM-powered customized task planning for household robots by effectively aligning with user reminders. In domestic settings, aligning task planning with user reminders poses significant challenges due to the limited quantity, diver

Cited by 9SourceScholar
2025

Anti-Tamper Protection for Unauthorized Individual Image Generation

ICCV 2025poster

With the advancement of personalized image generation technologies, concerns about forgery attacks that infringe on portrait rights and privacy are growing. To address these concerns, protection perturbation algorithms have been developed to disrupt forgery generation. However, the protection algori…

2025

Bidirectional Human–AI Collaboration for Equitable Student Performance Prediction via Deep Uncertainty Learning

IJCAI 2025

This paper studies a bidirectional human-AI collaborative student performance prediction problem to enhance equitable online education, aligning with the United Nations' Sustainable Development Goal (SDG) of ensuring inclusive and equitable quality education for all. The goal is to leverage collabor

Cited by 0SourcePDFScholar
2025

CAT: A Unified Click-and-Track Framework for Realistic Tracking

ICCV 2025poster

Modern visual trackers have achieved robust performance with precisely initialized target bounding boxes. However, providing high-precision initial annotations is a process both labor-intensive and error-prone in real-world scenarios. Interactive initialization (e.g., click-based, scribble-based) pr…

2025

CLIP-driven View-aware Prompt Learning for Unsupervised Vehicle Re-identification

AAAI 2025technical

With the emergence of vision-language pre-trained models, such as CLIP, some textual prompts have been gradually introduced recently into re-identification (Re-ID) tasks to obtain considerably robust multimodal information. However, most textual descriptions based on vehicle Re-ID tasks only contain…

Cited by 0SourcePDFScholar
2025

COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language Models

ICRA 2025

Leveraging the powerful reasoning capabilities of large language models (LLMs), recent LLM-based robot task planning methods yield promising results. However, they mainly focus on single or multiple homogeneous robots on simple tasks. Practically, complex long-horizon tasks always require collaborat

Cited by 41SourcecodeScholar
2025

Collaborative Reasoner: Self-Improving Social Agents with Synthetic Conversations

NeurIPS 2025poster

With increasingly powerful large language models (LLMs) and LLM-based agents tackling an ever-growing list of tasks, we envision a future where numerous LLM agents work seamlessly with other AI agents and humans to solve complex problems and enhance daily life. To achieve these goals, LLM agents mus…

Cited by 0SourceScholar
2025

ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning

EMNLP 2025

Large Reasoning Models (LRMs) perform strongly in complex reasoning tasks via Chain-of-Thought (CoT) prompting, but often suffer from verbose outputs, increasing computational overhead. Existing fine-tuning-based compression methods either operate post-hoc pruning, risking disruption to reasoning co

Cited by 0SourcePDFScholar
2025

DeformCL: Learning Deformable Centerline Representation for Vessel Extraction in 3D Medical Image

CVPR 2025poster

In the field of 3D medical imaging, accurately extracting and representing the blood vessels with curvilinear structures holds paramount importance for clinical diagnosis. Previous methods have commonly relied on discrete representation like mask, often resulting in local fractures or scattered frag…

2025

Detection for Harvesting with an Active Illumination Camera System and DUTU2-Net+

IROS 2025

Robots operating in agricultural environments require a robust, fast perception system to accurately identify picking points. This paper proposed a lightweight method for detecting sweet pepper peduncles, which uses an active illumination camera system and DUTU<sup xmlns:mml="http://www.w3.org/1998/

Cited by 0SourceScholar
2025

Doppler-SLAM: Doppler-Aided Radar-Inertial and LiDAR-Inertial Simultaneous Localization and Mapping

RA-L 2025

Simultaneous localization and mapping is a critical capability for autonomous systems. Traditional SLAM approaches often rely on visual or LiDAR sensors and face significant challenges in adverse conditions such as low light or featureless environments. To overcome these limitations, we propose a no

Cited by 9SourcecodeScholar
2025

Efficient Diffusion as Low Light Enhancer

CVPR 2025poster

The computational burden of the iterative sampling process remains a major challenge in diffusion-based Low-Light Image Enhancement (LLIE). Current acceleration methods, whether training-based or training-free, often lead to significant performance degradation, highlighting the trade-off between per…

Cited by 0SourcePDFScholar
2025

Efficient Motion Prompt Learning for Robust Visual Tracking

ICML 2025poster

Due to the challenges of processing temporal information, most trackers depend solely on visual discriminability and overlook the unique temporal coherence of video data. In this paper, we propose a lightweight and plug-and-play motion prompt tracking method. It can be easily integrated into existin…

2025

Equilibrium Policy Generalization: A Reinforcement Learning Framework for Cross-Graph Zero-Shot Generalization in Pursuit-Evasion Games

NeurIPS 2025poster

Equilibrium learning in adversarial games is an important topic widely examined in the fields of game theory and reinforcement learning (RL). Pursuit-evasion game (PEG), as an important class of real-world games from the fields of robotics and security, requires exponential time to be accurately sol…

Cited by 0SourceScholar
2025

Exploring Enhanced Contextual Information for Video-Level Object Tracking

AAAI 2025technical

Contextual information at the video level has become increasingly crucial for visual object tracking. However, existing methods typically use only a few tokens to convey this information, which can lead to information loss and limit their ability to fully capture the context. To address this issue,…

2025

FastUMI: A Scalable and Hardware-Independent Universal Manipulation Interface with Dataset

CoRL 2025poster

Real-world manipulation datasets for robotic arms remain scarce due to the high costs, rigid hardware dependencies, and complex setup procedures associated with existing data collection methods. We introduce, a redesigned Universal Manipulation Interface (UMI) that addresses these challenges, enabli…

Cited by 0SourceScholar
2025

Forget the Data and Fine-Tuning! Just Fold the Network to Compress

ICLR 2025poster

We introduce model folding, a novel data-free model compression technique that merges structurally similar neurons across layers, significantly reducing the model size without the need for fine-tuning or access to training data. Unlike existing methods, model folding preserves data statistics during…

2025

Full-text Error Correction for Chinese Speech Recognition with Large Language Model

ICASSP 2025accepted

Large Language Models (LLMs) have demonstrated substantial potential for error correction in Automatic Speech Recognition (ASR). However, most research focuses on utterances from short-duration speech recordings, which are the predominant form of speech data for supervised ASR training. This paper i…

Cited by 0SourceScholar
2025

GFM-Planner: Perception-Aware Trajectory Planning with Geometric Feature Metric

IROS 2025

Like humans who rely on landmarks for orientation, autonomous robots depend on feature-rich environments for accurate localization. In this paper, we propose the GFM-Planner, a perception-aware trajectory planning framework based on the geometric feature metric, which enhances LiDAR localization acc

Cited by 1SourceScholar
2025

Hybrid Latent Reasoning via Reinforcement Learning

NeurIPS 2025poster

Recent advances in large language models (LLMs) have introduced latent reasoning as a promising alternative to autoregressive reasoning. By performing internal computation with hidden states from previous steps, latent reasoning benefit from more informative features rather than sampling a discrete…

Cited by 0SourcecodeScholar
2025

HyperDiff: Masked Diffusion Model with High-efficient Transformer for Hyperspectral Image Cross-Scene Classification

ICASSP 2025accepted

Hyperspectral Image (HSI) cross-scene classification is a challenging task in remote sensing, particularly when real-time processing of Target Domain (TD) HSI is required, and data cannot be reused for training. While deep learning methods have shown promising results, the generalization ability of…

Cited by 0SourceScholar
2025

Improving Transferable Targeted Attacks with Feature Tuning Mixup

CVPR 2025poster

Deep neural networks (DNNs) exhibit vulnerability to adversarial examples that can transfer across different DNN models. A particularly challenging problem is developing transferable targeted attacks that can mislead DNN models into predicting specific target classes. While various methods have been…

2025

Inference Scaling for Long-Context Retrieval Augmented Generation

ICLR 2025oral

The scaling of inference computation has unlocked the potential of long-context large language models (LLMs) across diverse settings. For knowledge-intensive tasks, the increased compute is often allocated to incorporate more external knowledge. However, without effectively utilizing such knowledge…

Cited by 25SourcePDFScholar
2025

Knowledge Graph Completion with Relation-Aware Anchor Enhancement

AAAI 2025technical

Text-based knowledge graph completion methods take advantage of pre-trained language models (PLM) to enhance intrinsic semantic connections of raw triplets with detailed text descriptions. Typical methods in this branch map an input query (textual descriptions associated with an entity and a relatio…

2025

LOMIA: Label-Only Membership Inference Attacks against Pre-trained Large Vision-Language Models

NeurIPS 2025poster

Large vision-language models (VLLMs) have driven significant progress in multi-modal systems, enabling a wide range of applications across domains such as healthcare, education, and content generation. Despite the success, the large-scale datasets used to train these models often contain sensitive o…

Cited by 0SourceScholar
2025

Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding

AAAI 2025technical

3D Object Affordance Grounding aims to predict the functional regions on a 3D object and has laid the foundation for a wide range of applications in robotics. Recent advances tackle this problem via learning a mapping between 3D regions and a single human-object interaction image. However, the geome…

2025

Logical DA: Enhancing Data Augmentation for Logical Reasoning via a Multi-Agent System

ACL 2025finding

Recent advancements in large language models (LLMs) have highlighted the importance of improving their reasoning capabilities. A critical challenge lies in the scarcity of high-quality reasoning data—characterized by diversity and rich supervisory signals—necessary for robust model training. While d…

Cited by 0SourcePDFScholar
2025

MPPQ: Enhancing Post-Training Quantization for LLMs via Mixed Supervision, Proxy Rounding, and Pre-Searching

IJCAI 2025

Recently, post-training quantization (PTQ) methods for large language models (LLMs) primarily focus on tackling the challenges caused by outliers. Scaling transformation has proven to be effective while how to enhance the performance of extremely low-bitwidth (e.g., 2-bit) PTQ under it remains large

Cited by 0SourcePDFScholar
2025

Meta CLIP 2: A Worldwide Scaling Recipe

NeurIPS 2025spotlight

Contrastive Language-Image Pretraining (CLIP) is a popular foundation model, supporting from zero-shot classification, retrieval to encoders for multimodal large language models (MLLMs). Although CLIP is successfully trained on billion-scale image-text pairs from the English world, scaling CLIP's tr…

Cited by 0SourcecodeScholar
2025

MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation

ICCV 2025poster

In mobile manipulation, navigation and manipulation are often treated as separate problems, resulting in a significant gap between merely approaching an object and engaging with it effectively. Many navigation approaches primarily define success by proximity to the target, often overlooking the nece…

Cited by 0SourcePDFScholar
2025

Multi-Material 3D-Printed Magnetic Millirobot for Quadrupedal Locomotion in Endoluminal Spaces

IROS 2025

Quadrupedal locomotion have advantages of a low center of gravity, broad support base, and four-legged coordination, enabling outstanding stability in complex terrains. Drawing inspiration from this, researchers have developed robots emulate such locomotion through multiple actuations controlling an

Cited by 0SourceScholar
2025

NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions

NeurIPS 2025poster

Scaling reasoning capabilities beyond traditional domains such as math and coding is hindered by the lack of diverse and high-quality questions. To overcome this limitation, we introduce a scalable approach for generating diverse and challenging reasoning questions, accompanied by reference answers.…

Cited by 0SourceScholar
2025

Phoenix: A Motion-based Self-Reflection Framework for Fine-grained Robotic Action Correction

CVPR 2025poster

Building a generalizable self-correction system is crucial for robots to recover from failures. Despite advancements in Multimodal Large Language Models (MLLMs) that empower robots with semantic reflection ability for failure, translating semantic reflection into how to correct fine-grained robotic…

2025

ProcWorld: Benchmarking Large Model Planning in Reachability-Constrained Environments

EMNLP 2025

We introduce ProcWorld, a large-scale benchmark for partially observable embodied spatial reasoning and long-term planning with large language models (LLM) and vision language models (VLM). ProcWorld features a wide range of challenging embodied navigation and object manipulation tasks, covering 16

Cited by 0SourcePDFScholar
2025

RaI-SLAM: Radar-Inertial SLAM for Autonomous Vehicles

RA-L 2025

Simultaneous localization and mapping are essential components for the operation of autonomous vehicles in unknown environments. While localization focuses on estimating the vehicle's pose, mapping captures the surrounding environment to enhance future localization and decision-making. Localization

Cited by 24SourcecodeScholar
2025

SIDE: Socially Informed Drought Estimation Toward Understanding Societal Impact Dynamics of Environmental Crisis

AAAI 2025technical

Drought has become a critical global threat with significant societal impact. Existing drought monitoring solutions primarily focus on assessing drought severity using quantitative measurements, overlooking the diverse societal impact of drought from human-centric perspectives. Motivated by the coll…

Cited by 0SourcePDFScholar
2025

SUTrack: Towards Simple and Unified Single Object Tracking

AAAI 2025technical

In this paper, we propose a simple yet unified single object tracking (SOT) framework, dubbed SUTrack. It consolidates five SOT tasks (RGB-based, RGB-Depth, RGB-Thermal, RGB-Event, RGB-Language Tracking) into a unified model trained in a single session. Due to the distinct nature of the data, curren…

2025

Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language

ICASSP 2025accepted

We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoken language in Ethiopia and neighboring regions. The dataset was collected through a crowdsourcing initiative, encompassing a diverse range of speakers and phonetic variations. It consists of 100 hours…

Cited by 0SourceScholar
2025

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Models

RSS 2025poster

In this paper, we claim that spatial understanding is the keypoint in robot manipulation, and propose SpatialVLA to explore effective spatial representations for the robot foundation model. Specifically, we propose Ego3D Position Encoding to inject 3D information into VLA’s input observations, and i…

Cited by 18PDFScholar
2025

TALON: A Multi-Agent Framework for Long-Table Exploration and Question Answering

EMNLP 2025

Table question answering (TQA) requires accurate retrieval and reasoning over tabular data. Existing approaches attempt to retrieve query-relevant content before leveraging large language models (LLMs) to reason over long tables. However, these methods often fail to accurately retrieve contextually

2025

Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation

CVPR 2025poster

Learning a generalist robot that can effectively leverage prior knowledge for continuous skill acquisition remains significantly challenging. Despite the success of experience replay and parameter-efficient methods in maintaining knowledge across skills, naively applying these methods causes a failu…

Cited by 0SourcePDFScholar
2025

TimeRAG: Boosting LLM Time Series Forecasting via Retrieval-Augmented Generation

ICASSP 2025accepted

Although the rise of large language models (LLMs) has introduced new opportunities for time series forecasting, existing LLM-based solutions require excessive training and exhibit limited transferability. In view of these challenges, we propose TimeRAG, a framework that incorporates Retrieval-Augmen…

Cited by 0SourceScholar
2025

Two-stream Beats One-stream: Asymmetric Siamese Network for Efficient Visual Tracking

AAAI 2025technical

Efficient tracking has garnered attention for its ability to operate on resource-constrained platforms for real-world deployment beyond desktop GPUs. Current efficient trackers mainly follow precision-oriented trackers, adopting a one-stream framework with lightweight modules. However, blindly adher…

2025

VehicleMAE: View-asymmetry Mutual Learning for Vehicle Re-identification Pre-training via Masked AutoEncoders

ICCV 2025poster

Large-scale pre-training technology has achieved remarkable performance in diversified object re-identification (Re-ID) downstream tasks. Nevertheless, to our best knowledge, the pre-training model specifically for vehicle Re-ID, which focuses on tackling the challenge of multi-view variations, has…

Cited by 0SourcePDFScholar
2025

Zero-Shot Cross-Domain Aspect-Based Sentiment Analysis via Domain-Contextualized Chain-of-Thought Reasoning

EMNLP 2025

Cross-domain aspect-based sentiment analysis (ABSA) aims at learning specific knowledge from a source domain to perform various ABSA tasks on a target domain. Recent works mainly focus on how to use domain adaptation techniques to transfer the domain-agnostic features from the labeled source domain

Cited by 0SourcePDFScholar
2024

An Investigation of Distribution Alignment in Multi-Genre Speaker Recognition

ICASSP 2024accepted

Multi-genre speaker recognition is becoming increasingly popular due to its ability to better represent the complexities of real-world applications. However, a major challenge is the significant shift in the distribution of speaker vectors across different genres. While distribution alignment is a c…

Cited by 0SourceScholar
2024

Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters

CVPR 2024poster

Continual learning can empower vision-language models to continuously acquire new knowledge without the need for access to the entire historical dataset. However mitigating the performance degradation in large-scale models is non-trivial due to (i) parameter shifts throughout lifelong learning and (…

2024

Box2Poly: Memory-Efficient Polygon Prediction of Arbitrarily Shaped and Rotated Text

AAAI 2024technical

Recently, Transformer-based text detection techniques have sought to predict polygons by encoding the coordinates of individual boundary vertices using distinct query features. However, this approach incurs a significant memory overhead and struggles to effectively capture the intricate relationship…

2024

Color Event Enhanced Single-Exposure HDR Imaging

AAAI 2024technical

Single-exposure high dynamic range (HDR) imaging aims to reconstruct the wide-range intensities of a scene by using its single low dynamic range (LDR) image, thus providing significant efficiency. Existing methods pay high attention to restoring the luminance by inversing the tone-mapping process, w…

2024

Depth Helps: Improving Pre-trained RGB-based Policy with Depth Information Injection

IROS 2024poster

3D perception ability is crucial for generalizable robotic manipulation. While recent foundation models have made significant strides in perception and decision-making with RGB-based input, their lack of 3D perception limits their effectiveness in fine-grained robotic manipulation tasks. To address…

Cited by 2SourcecodeScholar
2024

Development of an Automatic Sweet Pepper Harvesting Robot and Experimental Evaluation

ICRA 2024poster

The aging population and diminishing working population in agriculture motivate the development of autonomous harvesting robots. Although autonomous harvesting is expanding rapidly, the commercial application of sweet pepper harvesting robots still faces challenges. This paper presents the developme…

Cited by 1SourceScholar
2024

EvSign: Sign Language Recognition and Translation with Streaming Events

ECCV 2024poster

"Sign language is one of the most effective communication tools for people with hearing difficulties. Most existing works focus on improving the performance of sign language tasks on RGB videos, which may suffer from degraded recording conditions, such as fast movement of hands with motion blur and…

2024

Evidence-Driven Retrieval Augmented Response Generation for Online Misinformation

NAACL 2024long

The proliferation of online misinformation has posed significant threats to public interest. While numerous online users actively participate in the combat against misinformation, many of such responses can be characterized by the lack of politeness and supporting facts. As a solution, text generati…

Cited by 27SourcePDFScholar
2024

Fair Federated Learning with Biased Vision-Language Models

ACL 2024findings

Existing literature that integrates CLIP into federated learning (FL) largely ignores the inherent group unfairness within CLIP and its ethical implications on FL applications. Furthermore, such CLIP bias may be amplified in FL, due to the unique issue of data heterogeneity across clients. However,…

Cited by 4SourcePDFScholar
2024

Flexible and Topological Consistent Local Replanning for Multirotors

IROS 2024poster

In many situations such as city delivery and wild inspection, quadrotors are often required to follow a predefined reference trajectory. However, these reference trajectories cannot be perfectly safe, resulting in conflicts between tracking the reference precisely, flying safely, and finishing the m…

Cited by 1SourceScholar
2024

GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting

CVPR 2024highlight

In this paper we introduce GS-SLAM that first utilizes 3D Gaussian representation in the Simultaneous Localization and Mapping (SLAM) system. It facilitates a better balance between efficiency and accuracy. Compared to recent SLAM methods employing neural implicit representations our method utilizes…

2024

GraphMorph: Tubular Structure Extraction by Morphing Predicted Graphs

NeurIPS 2024poster

Accurately restoring topology is both challenging and crucial in tubular structure extraction tasks, such as blood vessel segmentation and road network extraction. Diverging from traditional approaches based on pixel-level classification, our proposed method, named GraphMorph, focuses on branch-leve…

Cited by 0SourcePDFScholar
2024

HPL-ESS: Hybrid Pseudo-Labeling for Unsupervised Event-based Semantic Segmentation

CVPR 2024poster

Event-based semantic segmentation has gained popularity due to its capability to deal with scenarios under high-speed motion and extreme lighting conditions which cannot be addressed by conventional RGB cameras. Since it is hard to annotate event data previous approaches rely on event-to-image recon…

Cited by 5SourcePDFScholar
2024

Hybrid-SORT: Weak Cues Matter for Online Multi-Object Tracking

AAAI 2024technical

Multi-Object Tracking (MOT) aims to detect and associate all desired objects across frames. Most methods accomplish the task by explicitly or implicitly leveraging strong cues (i.e., spatial and appearance information), which exhibit powerful instance-level discrimination. However, when object occlu…

2024

KOI: Accelerating Online Imitation Learning via Hybrid Key-state Guidance

CoRL 2024poster

Online Imitation Learning methods struggle with the gap between extensive online exploration space and limited expert trajectories, which hinder efficient exploration due to inaccurate task-aware reward estimation. Inspired by the findings from cognitive neuroscience that task decomposition coul…

Cited by 0SourceScholar
2024

Kinematic-aware Prompting for Generalizable Articulated Object Manipulation with LLMs

ICRA 2024poster

Generalizable articulated object manipulation is essential for home-assistant robots. Recent efforts focus on imitation learning from demonstrations or reinforcement learning in simulation, however, due to the prohibitive costs of real-world data collection and precise object simulation, it still re…

Cited by 23SourcecodeScholar
2024

LLMs Can Evolve Continually on Modality for $\mathbb{X}$-Modal Reasoning

NeurIPS 2024poster

Multimodal Large Language Models (MLLMs) have gained significant attention due to their impressive capabilities in multimodal understanding. However, existing methods rely heavily on extensive modal-specific pretraining and joint-modal tuning, leading to significant computational burdens when expand…

2024

Learning Manipulation by Predicting Interaction

RSS 2024poster

Representation learning approaches for robotic manipulation have boomed in recent years. Due to the scarcity of in-domain robot data, prevailing methodologies tend to leverage large-scale human video datasets to extract generalizable features for visuomotor policy learning. Despite the progress achi…

2024

LiveScene: Language Embedding Interactive Radiance Fields for Physical Scene Control and Rendering

NeurIPS 2024poster

This paper scales object-level reconstruction to complex scenes, advancing interactive scene reconstruction. We introduce two datasets, OmniSim and InterReal, featuring 28 scenes with multiple interactive objects. To tackle the challenge of inaccurate interactive motion recovery in complex scenes, w…

2024

Multi-Fov-Constrained Trajectory Planning for Multirotor Safe Landing

IROS 2024poster

In recent years, multirotors have become more and more widely used, such as in aerial photography and delivery. Ensuring a safe landing in emergencies is the most basic requirement, and it is important to make full use of all the sensors of the multirotor. To improve the safety of UAV landing in unk…

Cited by 0SourceScholar
2024

Off-Policy Primal-Dual Safe Reinforcement Learning

ICLR 2024poster

Primal-dual safe RL methods commonly perform iterations between the primal update of the policy and the dual update of the Lagrange Multiplier. Such a training paradigm is highly susceptible to the error in cumulative cost estimation since this estimation serves as the key bond connecting the primal…

2024

Point-PEFT: Parameter-Efficient Fine-Tuning for 3D Pre-trained Models

AAAI 2024technical

The popularity of pre-trained large models has revolutionized downstream tasks across diverse fields, such as language, vision, and multi-modality. To minimize the adaption cost for downstream tasks, many Parameter-Efficient Fine-Tuning (PEFT) techniques are proposed for language and 2D image pre-tr…

2024

Retrieval Augmented Fact Verification by Synthesizing Contrastive Arguments

ACL 2024long

The rapid propagation of misinformation poses substantial risks to public interest. To combat misinformation, large language models (LLMs) are adapted to automatically verify claim credibility. Nevertheless, existing methods heavily rely on the embedded knowledge within LLMs and / or black-box APIs…

2024

Robust Quadrupedal Locomotion via Risk-Averse Policy Learning

ICRA 2024poster

The robustness of legged locomotion is crucial for quadrupedal robots in challenging terrains. Recently, Reinforcement Learning (RL) has shown promising results in legged locomotion and various methods try to integrate privileged distillation, scene modeling, and external sensors to improve the gene…

Cited by 13SourceScholar
2024

Safety-First Tracker: A Trajectory Planning Framework for Omnidirectional Robot Tracking

IROS 2024poster

This paper introduces a Safety-First Tracker (SF-Tracker) designed for omnidirectional autonomous tracking robots. The position and orientation of omnidirectional robots are decoupled for stepwise planning to ensure trajectory safety and maintain target visibility. SF-Tracker puts the trajectory saf…

Cited by 0SourcecodeScholar
2024

Train Once, Deploy Anywhere: Matryoshka Representation Learning for Multimodal Recommendation

EMNLP 2024finding

Despite recent advancements in language and vision modeling, integrating rich multimodal knowledge into recommender systems continues to pose significant challenges. This is primarily due to the need for efficient recommendation, which requires adaptive and interactive responses. In this study, we f…

2024

X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-Modal Knowledge Transfer

AAAI 2024technical

The field of 4D point cloud understanding is rapidly developing with the goal of analyzing dynamic 3D point cloud sequences. However, it remains a challenging task due to the sparsity and lack of texture in point clouds. Moreover, the irregularity of point cloud poses a difficulty in aligning tempo…

2023

A Crowd-AI Collaborative Duo Relational Graph Learning Framework towards Social Impact Aware Photo Classification

AAAI 2023technical

In artificial intelligence (AI), negative social impact (NSI) represents the negative effect on the society as a result of mistakes conducted by AI agents. While the photo classification problem has been widely studied in the AI community, the NSI made by photo misclassification is largely ignored d…

Cited by 0SourcePDFScholar
2023

Affordance-Driven Next-Best-View Planning for Robotic Grasping

CoRL 2023poster

Grasping occluded objects in cluttered environments is an essential component in complex robotic manipulation tasks. In this paper, we introduce an AffordanCE-driven Next-Best-View planning policy (ACE-NBV) that tries to find a feasible grasp for target object via continuously observing scenes from…

Cited by 14SourceScholar
2023

CN-CVS: A Mandarin Audio-Visual Dataset for Large Vocabulary Continuous Visual to Speech Synthesis

ICASSP 2023accepted

Research on Video to Speech Synthesis (VTS) surges recently and the focus is gradually shifting from small-vocabulary short-phrase VTS to large-vocabulary continuous VTS (LVC-VTS). A large-scale dataset with sufficient speakers and utterances is a prerequisite for such research, and the database is…

Cited by 0SourceScholar
2023

Cross-Domain Policy Adaptation via Value-Guided Data Filtering

NeurIPS 2023poster

Generalizing policies across different domains with dynamics mismatch poses a significant challenge in reinforcement learning. For example, a robot learns the policy in a simulator, but when it is deployed in the real world, the dynamics of the environment may be different. Given the source and targ…

Cited by 19SourcePDFScholar
2023

Decision-Making Context Interaction Network for Click-Through Rate Prediction

AAAI 2023technical

Click-through rate (CTR) prediction is crucial in recommendation and online advertising systems. Existing methods usually model user behaviors, while ignoring the informative context which influences the user to make a click decision, e.g., click pages and pre-ranking candidates that inform inferenc…

Cited by 13SourcePDFScholar
2023

Diffusion Model is an Effective Planner and Data Synthesizer for Multi-Task Reinforcement Learning

NeurIPS 2023poster

Diffusion models have demonstrated highly-expressive generative capabilities in vision and NLP. Recent studies in reinforcement learning (RL) have shown that diffusion models are also powerful in modeling complex policies or trajectories in offline datasets. However, these works have been limited to…

2023

Direct Heterogeneous Causal Learning for Resource Allocation Problems in Marketing

AAAI 2023technical

Marketing is an important mechanism to increase user engagement and improve platform revenue, and heterogeneous causal learning can help develop more effective strategies. Most decision-making problems in marketing can be formulated as resource allocation problems and have been studied for decades.…

2023

Dual Memory Aggregation Network for Event-Based Object Detection with Learnable Representation

AAAI 2023technical

Event-based cameras are bio-inspired sensors that capture brightness change of every pixel in an asynchronous manner. Compared with frame-based sensors, event cameras have microsecond-level latency and high dynamic range, hence showing great potential for object detection under high-speed motion and…

2023

Exploring Lightweight Hierarchical Vision Transformers for Efficient Visual Tracking

ICCV 2023poster

Transformer-based visual trackers have demonstrated significant progress owing to their superior modeling capabilities. However, existing trackers are hampered by low speed, limiting their applicability on devices with limited computational power. To alleviate this problem, we propose HiT, a new fam…

Cited by 70PDFcodeScholar
2023

Fully Self-Supervised Depth Estimation From Defocus Clue

CVPR 2023poster

Depth-from-defocus (DFD), modeling the relationship between depth and defocus pattern in images, has demonstrated promising performance in depth estimation. Recently, several self-supervised works try to overcome the difficulties in acquiring accurate depth ground-truth. However, they depend on the…

2023

HSR-Diff: Hyperspectral Image Super-Resolution via Conditional Diffusion Models

ICCV 2023poster

Despite the proven significance of hyperspectral images (HSIs) in performing various computer vision tasks, its potential is adversely affected by the low-resolution (LR) property in the spatial domain, resulting from multiple physical factors. Inspired by recent advancements in deep generative mode…

Cited by 47PDFScholar
2023

KEPL: Knowledge Enhanced Prompt Learning for Chinese Hypernym-Hyponym Extraction

EMNLP 2023long main

Modeling hypernym-hyponym ("is-a") relations is very important for many natural language processing (NLP) tasks, such as classification, natural language inference and relation extraction. Existing work on is-a relation extraction is mostly in the English language environment. Due to the flexibility…

Cited by 0SourceScholar
2023

MetaAdapt: Domain Adaptive Few-Shot Misinformation Detection via Meta Learning

ACL 2023long

With emerging topics (e.g., COVID-19) on social media as a source for the spreading misinformation, overcoming the distributional shifts between the original training domain (i.e., source domain) and such target domains remains a non-trivial task for misinformation detection. This presents an elusiv…

2023

Not All Features Matter: Enhancing Few-shot CLIP with Adaptive Prior Refinement

ICCV 2023poster

The popularity of Contrastive Language-Image Pre-training (CLIP) has propelled its application to diverse downstream vision tasks. To improve its capacity on downstream tasks, few-shot learning has become a widely-adopted technique. However, existing methods either exhibit limited performance or suf…

Cited by 89PDFcodeScholar
2023

On Adversarial Robustness of Demographic Fairness in Face Attribute Recognition

IJCAI 2023poster

Demographic fairness has become a critical objective when developing modern visual models for identity-sensitive applications, such as face attribute recognition (FAR). While great efforts have been made to improve the fairness of the models, the investigation on the adversarial robustness of the fa…

Cited by 5SourcePDFScholar
2023

On Optimizing Model Generality in AI-based Disaster Damage Assessment: A Subjective Logic-driven Crowd-AI Hybrid Learning Approach

IJCAI 2023poster

This paper focuses on the AI-based damage assessment (ADA) applications that leverage state-of-the-art AI techniques to automatically assess the disaster damage severity using online social media imagery data, which aligns well with the ''disaster risk reduction'' target under United Nations' Sustai…

Cited by 3SourcePDFScholar
2023

One-Shot High-Fidelity Talking-Head Synthesis With Deformable Neural Radiance Field

CVPR 2023poster

Talking head generation aims to generate faces that maintain the identity information of the source image and imitate the motion of the driving image. Most pioneering methods rely primarily on 2D representations and thus will inevitably suffer from face distortion when large head rotations are encou…

Cited by 54SourcePDFScholar
2023

Polynomial-Based Online Planning for Autonomous Drone Racing in Dynamic Environments

IROS 2023poster

In recent years, there is a noteworthy advance-ment in autonomous drone racing. However, the primary focus is on attaining execution times, while scant attention is given to the challenges of dynamic environments. The high-speed nature of racing scenarios, coupled with the potential for unforeseeabl…

Cited by 8SourceScholar
2023

Propagate and Calibrate: Real-Time Passive Non-Line-of-Sight Tracking

CVPR 2023poster

Non-line-of-sight (NLOS) tracking has drawn increasing attention in recent years, due to its ability to detect object motion out of sight. Most previous works on NLOS tracking rely on active illumination, e.g., laser, and suffer from high cost and elaborate experimental conditions. Besides, these te…

2023

Safe Offline Reinforcement Learning with Real-Time Budget Constraints

ICML 2023poster

Aiming at promoting the safe real-world deployment of Reinforcement Learning (RL), research on safe RL has made significant progress in recent years. However, most existing works in the literature still focus on the online setting where risky violations of the safety budget are likely to be incurred…

2023

SeqTrack: Sequence to Sequence Learning for Visual Object Tracking

CVPR 2023poster

In this paper, we present a new sequence-to-sequence learning framework for visual tracking, dubbed SeqTrack. It casts visual tracking as a sequence generation problem, which predicts object bounding boxes in an autoregressive fashion. This is different from prior Siamese trackers and transformer tr…

2023

Towards Benchmarking and Assessing Visual Naturalness of Physical World Adversarial Attacks

CVPR 2023poster

Physical world adversarial attack is a highly practical and threatening attack, which fools real world deep learning systems by generating conspicuous and maliciously crafted real world artifacts. In physical world attacks, evaluating naturalness is highly emphasized since human can easily detect an…

2023

Towards Nonlinear-Motion-Aware and Occlusion-Robust Rolling Shutter Correction

ICCV 2023poster

This paper addresses the problem of rolling shutter correction in complex nonlinear and dynamic scenes with extreme occlusion. Existing methods suffer from two main drawbacks. Firstly, they face challenges in estimating the accurate correction field due to the uniform velocity assumption, leading t…

Cited by 8PDFcodeScholar
2023

Universal Instance Perception As Object Discovery and Retrieval

CVPR 2023poster

All instance perception tasks aim at finding certain objects specified by some queries such as category names, language expressions, and target annotations, but this complete field has been split into multiple independent subtasks. In this work, we present a universal instance perception model of th…

2023

ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding

ICCV 2023poster

Understanding 3D scenes from multi-view inputs has been proven to alleviate the view discrepancy issue in 3D visual grounding. However, existing methods normally neglect the view cues embedded in the text modality and fail to weigh the relative importance of different views. In this paper, we propos…

Cited by 64PDFScholar
2023

Zero- and Few-Shot Event Detection via Prompt-Based Meta Learning

ACL 2023long

With emerging online topics as a source for numerous new events, detecting unseen / rare event types presents an elusive challenge for existing event detection methods, where only limited data access is provided for training. To address the data scarcity problem in event detection, we propose MetaEv…

2022

A Copy-Augmented Generative Model for Open-Domain Question Answering

ACL 2022short

Open-domain question answering is a challenging task with a wide variety of practical applications. Existing modern approaches mostly follow a standard two-stage paradigm: retriever then reader. In this article, we focus on improving the effectiveness of the reader module and propose a novel copy-au…

Cited by 3SourcePDFScholar
2022

Balanced Multimodal Learning via On-the-Fly Gradient Modulation

CVPR 2022oral

Audio-visual learning helps to comprehensively understand the world, by integrating different senses. Accordingly, multiple input modalities are expected to boost model performance, but we actually find that they are not fully exploited even when the multi-modal model outperforms its uni-modal count…

Cited by 247PDFcodeScholar
2022

Check and Link: Pairwise Lesion Correspondence Guides Mammogram Mass Detection

ECCV 2022poster

"Detecting mass in mammogram is significant due to the high occurrence and mortality of breast cancer. In mammogram mass detection, modeling pairwise lesion correspondence explicitly is particularly important. However, most of the existing methods build relatively coarse correspondence and have not…

Cited by 6SourcePDFScholar
2022

Crowd, Expert & AI: A Human-AI Interactive Approach Towards Natural Language Explanation Based COVID-19 Misinformation Detection

IJCAI 2022poster

In this paper, we study an explainable COVID-19 misinformation detection problem where the goal is to accurately identify COVID-19 misleading posts on social media and explain the posts with natural language explanations (NLEs). Our problem is motivated by the limitations of current explainable misi…

Cited by 19SourcePDFScholar
2022

D-DPCC: Deep Dynamic Point Cloud Compression via 3D Motion Prediction

IJCAI 2022poster

The non-uniformly distributed nature of the 3D Dynamic Point Cloud (DPC) brings significant challenges to its high-efficient inter-frame compression. This paper proposes a novel 3D sparse convolution-based Deep Dynamic Point Cloud Compression (D-DPCC) network to compensate and compress the DPC geome…

2022

Domain Adaptation for Question Answering via Question Classification

COLING 2022main

Question answering (QA) has demonstrated impressive progress in answering questions from customized domains. Nevertheless, domain adaptation remains one of the most elusive challenges for QA systems, especially when QA systems are trained in a source domain but deployed in a different target domain.…

2022

Feature Dense Relevance Network for Single Image Dehazing

IJCAI 2022poster

Existing learning-based dehazing methods do not fully use non-local information, which makes the restoration of seriously degraded region very tough. We propose a novel dehazing network by defining the Feature Dense Relevance module (FDR) and the Shallow Feature Mapping module (SFM). The FDR is defi…

Cited by 4SourcePDFScholar
2022

Gradient Importance Learning for Incomplete Observations

ICLR 2022poster

Though recent works have developed methods that can generate estimates (or imputations) of the missing entries in a dataset to facilitate downstream analysis, most depend on assumptions that may not align with real-world applications and could suffer from poor performance in subsequent tasks such as…

2022

MultiQuant: Training Once for Multi-bit Quantization of Neural Networks

IJCAI 2022poster

Quantization has become a popular technique to compress deep neural networks (DNNs) and reduce computational costs, but most prior work focuses on training DNNs at each individual fixed bit-width and accuracy trade-off point. How to produce a model with flexible precision is largely unexplored. This…

Cited by 9SourcePDFScholar
2022

On Attacking Out-Domain Uncertainty Estimation in Deep Neural Networks

IJCAI 2022poster

In many applications with real-world consequences, it is crucial to develop reliable uncertainty estimation for the predictions made by the AI decision systems. Targeting at the goal of estimating uncertainty, various deep neural network (DNN) based uncertainty estimation algorithms have been propos…

Cited by 13SourcePDFScholar
2022

Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-training

NeurIPS 2022accept

Masked Autoencoders (MAE) have shown great potentials in self-supervised pre-training for language and 2D image transformers. However, it still remains an open question on how to exploit masked autoencoding for learning 3D representations of irregular point clouds. In this paper, we propose Point-M2…

2022

PointScatter: Point Set Representation for Tubular Structure Extraction

ECCV 2022poster

"This paper explores the point set representation for tubular structure extraction tasks. Compared with the traditional mask representation, the point set representation enjoys its flexibility and representation ability, which would not be restricted by the fixed grid as the mask. Inspired by this,…

2022

QA Domain Adaptation using Hidden Space Augmentation and Self-Supervised Contrastive Adaptation

EMNLP 2022main

Question answering (QA) has recently shown impressive results for answering questions from customized domains. Yet, a common challenge is to adapt QA models to an unseen target domain. In this paper, we propose a novel self-supervised framework called QADA for QA domain adaptation. QADA introduces a…

2022

Show, Deconfound and Tell: Image Captioning With Causal Inference

CVPR 2022poster

The transformer-based encoder-decoder framework has shown remarkable performance in image captioning. However, most transformer-based captioning methods ever overlook two kinds of elusive confounders: the visual confounder and the linguistic confounder, which generally lead to harmful bias, induce t…

Cited by 67PDFcodeScholar
2022

Tight Mutual Information Estimation With Contrastive Fenchel-Legendre Optimization

NeurIPS 2022accept

Successful applications of InfoNCE (Information Noise-Contrastive Estimation) and its variants have popularized the use of contrastive variational mutual information (MI) estimators in machine learning . While featuring superior stability, these estimators crucially depend on costly large-batch trai…

2022

Towards Grand Unification of Object Tracking

ECCV 2022poster

"We present a unified method, termed Unicorn, that can simultaneously solve four tracking problems (SOT, MOT, VOS, MOTS) with a single network using the same model parameters. Due to the fragmented definitions of the object tracking problem itself, most existing trackers are developed to address a s…

2022

Visible-Thermal UAV Tracking: A Large-Scale Benchmark and New Baseline

CVPR 2022poster

With the popularity of multi-modal sensors, visible-thermal (RGB-T) object tracking is to achieve robust performance and wider application scenarios with the guidance of objects' temperature information. However, the lack of paired training samples is the main bottleneck for unlocking the power of R…

Cited by 208PDFcodeScholar
2021

A Meta-Learning Framework for Few-Shot Classification of Remote Sensing Scene

ICASSP 2021accepted

While achieving remarkable success in remote sensing (RS) scene classification for the past few years, convolutional neural network (CNN) based methods suffer from the demand for large amounts of training data. The bottleneck in prediction accuracy has shifted from data processing limits toward a la…

Cited by 0SourceScholar
2021

A Streaming End-to-End Framework For Spoken Language Understanding

IJCAI 2021poster

End-to-end spoken language understanding (SLU) recently attracted increasing interest. Compared to the conventional tandem-based approach that combines speech recognition and language understanding as separate modules, the new approach extracts users' intentions directly from the speech signals, res…

Cited by 12SourcePDFScholar
2021

Alpha-Refine: Boosting Tracking Performance by Precise Bounding Box Estimation

CVPR 2021poster

Visual object tracking aims to precisely estimate the bounding box for the given target, which is a challenging problem due to factors such as deformation and occlusion. Many recent trackers adopt the multiple-stage tracking strategy to improve the quality of bounding box estimation. These methods f…

Cited by 268PDFcodeScholar
2021

CLINE: Contrastive Learning with Semantic Negative Examples for Natural Language Understanding

ACL 2021long

Despite pre-trained language models have proven useful for learning high-quality semantic representations, these models are still vulnerable to simple perturbations. Recent works aimed to improve the robustness of pre-trained models mainly focus on adversarial training from perturbed examples with s…

2021

FAST-Dynamic-Vision: Detection and Tracking Dynamic Objects with Event and Depth Sensing

IROS 2021poster

The development of aerial autonomy has enabled aerial robots to fly agilely in complex environments. However, dodging fast-moving objects in flight remains a challenge, limiting the further application of unmanned aerial vehicles (UAVs). The bottleneck of solving this problem is the accurate percept…

Cited by 45SourcecodeScholar
2021

Heterogeneous two-Stream Network with Hierarchical Feature Prefusion for Multispectral Pan-Sharpening

ICASSP 2021accepted

Multispectral (MS) pan-sharpening aims at producing a high spatial resolution (HR) MS image by fusing a single-band HR panchromatic (PAN) image and a corresponding MS image with low spatial resolution. In this paper, we propose a heterogeneous two-stream network (HTSNet) with hierarchical feature pr…

Cited by 0SourceScholar
2021

KeSpeech: An Open Source Speech Dataset of Mandarin and Its Eight Subdialects

NeurIPS 2021poster

This paper introduces an open source speech dataset, KeSpeech, which involves 1,542 hours of speech signals recorded by 27,237 speakers in 34 cities in China, and the pronunciation includes standard Mandarin and its 8 subdialects. The new dataset possesses several properties. Firstly, the dataset pr…

Cited by 41SourceScholar
2021

Learning Spatio-Temporal Transformer for Visual Tracking

ICCV 2021poster

In this paper, we present a new tracking architecture with an encoder-decoder transformer as the key component. The encoder models the global spatio-temporal feature dependencies between target objects and search regions, while the decoder learns a query embedding to predict the spatial positions of…

Cited by 1080PDFcodeScholar
2021

LightTrack: Finding Lightweight Neural Networks for Object Tracking via One-Shot Architecture Search

CVPR 2021poster

Object tracking has achieved significant progress over the past few years. However, state-of-the-art trackers become increasingly heavy and expensive, which limits their deployments in resource-constrained applications. In this work, we present LightTrack, which uses neural architecture search (NAS)…

Cited by 241PDFcodeScholar
2021

Pyramid Spatial-Temporal Aggregation for Video-Based Person Re-Identification

ICCV 2021poster

Video-based person re-identification aims to associate the video clips of the same person across multiple non-overlapping cameras. Spatial-temporal representations can provide richer and complementary information between frames, which are crucial to distinguish the target person when occlusion occur…

Cited by 107PDFcodeScholar
2021

Squeezing Value of Cross-Domain Labels: A Decoupled Scoring Approach for Speaker Verification

ICASSP 2021accepted

Domain mismatch often occurs in real applications and causes serious performance reduction on speaker verification systems. The common wisdom is to collect cross-domain data and train a multi-domain PLDA model, with the hope to learn a domain-independent speaker subspace. In this paper, we firstly p…

Cited by 0SourceScholar
2021

Temporal Relational Modeling with Self-Supervision for Action Segmentation

AAAI 2021technical

Temporal relational modeling in video is essential for human action understanding, such as action recognition and action segmentation. Although Graph Convolution Networks (GCNs) have shown promising advantages in relation reasoning on many tasks, it is still a challenge to apply graph convolution ne…

2021

User Retention: A Causal Approach with Triple Task Modeling

IJCAI 2021poster

For many Internet companies, it has been an important focus to improve user retention rate. To achieve this goal, we need to recommend proper services in order to meet the demands of users. Unlike conventional click-through rate (CTR) estimation, there are lots of noise in the collected data when m…

Cited by 9SourcePDFScholar
2021

Video Annotation for Visual Tracking via Selection and Refinement

ICCV 2021poster

Deep learning based visual trackers entail offline pre-training on large volumes of video datasets with accurate bounding box annotations that are labor-expensive to achieve. We present a new framework to facilitate bounding box annotations for video sequences, which investigates a selection-and-ref…

Cited by 11PDFcodeScholar
2021

Vset: A Multimodal Transformer for Visual Speech Enhancement

ICASSP 2021accepted

The transformer architecture has shown great capability in learning long-term dependency and works well in multiple domains. However, transformer has been less considered in audio-visual speech enhancement (AVSE) research, partly due to the convention that treats speech enhancement as a short-time s…

Cited by 0SourceScholar
2021

Wasserstein Contrastive Representation Distillation

CVPR 2021poster

The primary goal of knowledge distillation (KD) is to encapsulate the information of a model learned from a teacher network into a student network, with the latter being more compact than the former. Existing work, e.g., using Kullback-Leibler divergence for distillation, may fail to capture importa…

Cited by 128PDFScholar
2021

eTREE: Learning Tree-structured Embeddings

AAAI 2021technical

Matrix factorization (MF) plays an important role in a wide range of machine learning and data mining models. MF is commonly used to obtain item embeddings and feature representations due to its ability to capture correlations and higher-order statistical dependencies across dimensions. In many appl…

2020

CN-Celeb: A Challenging Chinese Speaker Recognition Dataset

ICASSP 2020accepted

Recently, researchers set an ambitious goal of conducting speaker recognition in unconstrained conditions where the variations on ambient, channel and emotion could be arbitrary. However, most publicly available datasets are collected under constrained environments, i.e., with little noise and limit…

Cited by 0SourceScholar
2020

Cooling-Shrinking Attack: Blinding the Tracker With Imperceptible Noises

CVPR 2020poster

Adversarial attack of CNN aims at deceiving models to misbehave by adding imperceptible perturbations to images. This feature facilitates to understand neural networks deeply and to improve the robustness of deep learning models. Although several works have focused on attacking image classifiers and…

Cited by 105PDFcodeScholar
2020

High-Performance Long-Term Tracking With Meta-Updater

CVPR 2020oral

Long-term visual tracking has drawn increasing attention because it is much closer to practical applications than short-term tracking. Most top-ranked long-term trackers adopt the offline-trained Siamese architectures, thus,they cannot benefit from great progress of short-term trackers with online u…

Cited by 319PDFcodeScholar
2020

Integrating User History into Heterogeneous Graph for Dialogue Act Recognition

COLING 2020main

Dialogue Act Recognition (DAR) is a challenging problem in Natural Language Understanding, which aims to attach Dialogue Act (DA) labels to each utterance in a conversation. However, previous studies cannot fully recognize the specific expressions given by users due to the informality and diversity…

Cited by 8SourcePDFScholar
2020

Summarize before Aggregate: A Global-to-local Heterogeneous Graph Inference Network for Conversational Emotion Recognition

COLING 2020main

Conversational Emotion Recognition (CER) is a crucial task in Natural Language Processing (NLP) with wide applications. Prior works in CER generally focus on modeling emotion influences solely with utterance-level features, with little attention paid on phrase-level semantic connection between utter…

Cited by 44SourcePDFScholar
2019

'Skimming-Perusal' Tracking: A Framework for Real-Time and Robust Long-Term Tracking

ICCV 2019poster

Compared with traditional short-term tracking, long-term tracking poses more challenges and is much closer to realistic applications. However, few works have been done and their performance have also been limited. In this work, we present a novel robust and real-time long-term tracking framework bas…

Cited by 229PDFcodeScholar
2019

A Mutual Learning Method for Salient Object Detection With Intertwined Multi-Supervision

CVPR 2019poster

Though deep learning techniques have made great progress in salient object detection recently, the predicted saliency maps still suffer from incomplete predictions due to the internal complexity of objects and inaccurate boundaries caused by strides in convolution and pooling operations. To alleviat…

Cited by 289PDFcodeScholar
2019

GradNet: Gradient-Guided Network for Visual Object Tracking

ICCV 2019oral

The fully-convolutional siamese network based on template matching has shown great potentials in visual tracking. During testing, the template is fixed with the initial target feature and the performance totally relies on the general matching ability of the siamese network. However, this manner cann…

Cited by 400PDFcodeScholar
2019

Language Person Search with Mutually Connected Classification Loss

ICASSP 2019accepted

In this work, we develop an effective person search algorithm with natural language descriptions. The contributions of this work mainly include two aspects. First, we design a baseline language person search framework including three basic components: a deep CNN model to extract visual features, a b…

Cited by 0SourceScholar
2019

Online Single Person Tracking for Unmanned Aerial Vehicles: Benchmark and New Baseline

ICASSP 2019accepted

Online tracking a specific person from a low-altitude unmanned aerial vehicle (UAV) is a very interesting and challenging problem to be solved. However, there exists no large-scale aerial video dataset regarding this online single person tracking (OSPT) task. To promote the study of the OSPT problem…

Cited by 0SourceScholar
2019

Visual Tracking via Adaptive Spatially-Regularized Correlation Filters

CVPR 2019oral

In this work, we propose a novel adaptive spatially-regularized correlation filters (ASRCF) model to simultaneously optimize the filter coefficients and the spatial regularization weight. First, this adaptive spatial regularization scheme could learn an effective spatial weight for a specific object…

Cited by 498PDFcodeScholar
2018

BRITS: Bidirectional Recurrent Imputation for Time Series

NeurIPS 2018poster

Time series are widely used as signals in many classification/regression tasks. It is ubiquitous that time series contains many missing values. Given multiple correlated time series data, how to fill in missing values and to predict their class labels? Existing imputation methods often impose strong…

2018

Correlation Tracking via Joint Discrimination and Reliability Learning

CVPR 2018poster

For visual tracking, an ideal filter learned by the correlation filter (CF) method should take both discrimination and reliability information. However, existing attempts usually focus on the former one while pay less attention to reliability learning. This may make the learned filter be dominated b…

2018

Defocus Blur Detection via Multi-Stream Bottom-Top-Bottom Fully Convolutional Network

CVPR 2018poster

Defocus blur detection (DBD) is the separation of infocus and out-of-focus regions in an image. This process has been paid considerable attention because of its remarkable potential applications. Accurate differentiation of homogeneous regions and detection of low-contrast focal regions, as well as…

Cited by 98SourcePDFScholar
2018

Human and Machine Speaker Recognition Based on Short Trivial Events

ICASSP 2018accepted

Human speech often has events that we will call trivial events, e.g., cough, laugh and sniff. Compared to regular speech, these trivial events are usually short and variable, thus generally regarded as not speaker discriminative and so are largely ignored by present speaker recognition research. How…

Cited by 0SourceScholar
2018

Learning Spatial-Aware Regressions for Visual Tracking

CVPR 2018poster

In this paper, we analyze the spatial information of deep features, and propose two complementary regressions for robust visual tracking. First, we propose a kernelized ridge regression model wherein the kernel value is defined as the weighted sum of similarity scores of all pairs of patches between…