← Search

Jian Wang

176 accepted papers

2026

3DAlign-DAER: Dynamic Attention Policy and Efficient Retrieval Strategy for Fine-grained 3D-Text Alignment at Scale

AAAI 2026technical

Despite recent advancements in 3D-text cross-modal alignment, existing state-of-the-art methods still struggle to align fine-grained textual semantics with detailed geometric structures, and their alignment performance degrades significantly when scaling to large-scale 3D databases. To overcome this

Cited by 13SourcePDFScholar
2026

Beyond Text: Visual Description Assembly by Probabilistic Model for CLIP-based Weakly Supervised Semantic Segmentation

CVPR 2026

Contrastive Language-Image Pre-training (CLIP) offers a new paradigm for Weakly Supervised Semantic Segmentation (WSSS) by generating Class Activation Maps (CAMs) from text-image alignment. Existing methods primarily rely on hand-crafted templates or general attribute descriptions generated by a lar

Cited by 0SourceScholar
2026

Dynamics-Aware Preference Optimization for Vision-Language Models

CVPR 2026

Preference-based finetuning of vision-language models (VLMs) is notoriously unstable, as trivially wrong negatives inject uninformative gradients that distort optimization and degrade calibration. This work revisits this issue through the lens of learning dynamics and identifies a core pathology, th

Cited by 0SourcecodeScholar
2026

Efficient Reinforcement Learning for Zero-Shot Coordination in Evolving Games

AAAI 2026technical

Zero-shot coordination(ZSC), a key challenge in multi-agent game theory, has become a hot topic in reinforcement learning (RL) research recently, especially in complex evolving games. It focuses on the generalization ability of agents, requiring them to coordinate well with collaborators from a dive

Cited by 0SourcePDFScholar
2026

Eliminate Distance Differences Induced by Backdoor Attacks: Layer-Selective Training and Clipping to Mask Backdoor Models

CVPR 2026

Federated learning (FL) enables a central server to collaboratively train a global model with multiple clients while preserving data privacy. However, the distributed nature of FL makes the paradigm vulnerable to backdoor attacks, as proved by numerous recent studies. Although existing studies impro

Cited by 0SourceScholar
2026

Global Policy-Space Response Oracles for Two-Player Zero-Sum Games

ICML 2026poster

The Policy-Space Response Oracles (PSRO) framework scales equilibrium computation to large zero-sum games by iteratively expanding a restricted strategy set using deep reinforcement learning (DRL). A central challenge is to construct, under limited computational budgets, a small strategy population …

Cited by 0SourceScholar
2026

HandX: Scaling Bimanual Motion and Interaction Generation

CVPR 2026

Synthesizing human motion has advanced rapidly, yet realistic hand motion and bimanual interaction remain underexplored. Whole-body models often miss the fine-grained cues that drive dexterous behavior, finger articulation, contact timing, and inter-hand coordination, and existing resources lack hig

Cited by 0SourcecodeScholar
2026

Hybrid Token Compression for Vision-Language Models

CVPR 2026

Vision-language models (VLMs) have transformed multimodal reasoning, but feeding hundreds of visual patch tokens to LLMs incurs quadratic computational costs, straining memory and context windows. Traditional approaches face a trade-off: continuous compression dilutes high-level semantics like objec

Cited by 0SourcecodeScholar
2026

IF-Prune: Information-Flow Guided Token Pruning for Efficient Vision-Language Models

CVPR 2026

Vision-language models (VLMs) with dynamic resolution vision encoders achieve strong performance, but face significant efficiency challenges due to long input sequences. A common approach is to assess the importance of tokens and prune those that are less informative. Recent methods utilizing a smal

Cited by 0SourcecodeScholar
2026

Invisible Triggers, Visible Threats! Road-Style Adversarial Creation Attack for Visual 3D Detection in Autonomous Driving

AAAI 2026technical

Modern autonomous driving (AD) systems leverage 3D object detection to perceive foreground objects in 3D environments for subsequent prediction and planning. Visual 3D detection based on RGB cameras provides a cost-effective solution compared to the LiDAR paradigm. While achieving promising detectio

Cited by 0SourcePDFScholar
2026

LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction

AAAI 2026technical

Large language models (LLMs) often generate hallucinated content lacking factual or contextual grounding, hindering their reliability in critical applications. Traditional methods like supervised fine-tuning and reinforcement learning from human feedback are data-intensive and computationally expens

Cited by 0SourcePDFScholar
2026

Learning Locally, Revising Globally: Global Reviser for Federated Learning with Noisy Labels

ICML 2026poster

In pursuit of data privacy, federated learning (FL) collaboratively trains a global model by aggregating local models learned from decentralized data. However, FL heavily depends on high-quality labels, which are often impractical in the real world, leading to the federated label-noise (F-LN) proble…

Cited by 0SourceScholar
2026

LiveClin: A Live Clinical Benchmark without Leakage

ICLR 2026poster

The reliability of medical LLM evaluation is critically undermined by data contamination and knowledge obsolescence, leading to inflated scores on static benchmarks. To address these challenges, we introduce LiveClin, a live benchmark designed for the approximating real-world clinical practice. Buil…

Cited by 0SourcecodeScholar
2026

MHB: Medical Hallucination Benchmark for Large Language Models in Complex Clinical Tasks

AAAI 2026technical

The integration of Large Language Models (LLMs) into clinical applications presents transformative potential but is undermined by the critical risk of hallucination, the generation of plausible but factually incorrect information. Such failures pose a direct threat to patient safety and the integrit

Cited by 0SourcePDFScholar
2026

MetaToolAgent: Towards Generalizable Tool Usage in LLMs through Meta-Learning

ICASSP 2026oral

Tool learning is increasingly important for large language models (LLMs) to effectively coordinate and utilize a diverse set of tools in order to solve complex real-world tasks. By selecting and integrating appropriate tools, LLMs extend their capabilities beyond pure language understanding to perfo…

Cited by 0SourcePDFScholar
2026

OptScale: Probabilistic Optimality for Inference-time Scaling

AAAI 2026technical

Inference-time scaling has emerged as a powerful technique for enhancing the reasoning performance of Large Language Models (LLMs). However, existing approaches often rely on heuristic strategies for parallel sampling, lacking a principled foundation. To address this gap, we propose a probabilistic

Cited by 0SourcePDFScholar
2026

Probing Cultural Awareness in LLMs: A Case Study of Cross-Culture Aesthetic Stylistics

IJCAI 2026

Large Language Models (LLMs) are increasingly deployed in diverse cultural contexts, yet their ability to master aesthetic stylistics, i.e., the strategic use of language to evoke cultural resonance, remains underexplored. We curate C4Styli, a benchmark of highly stylized translated movie titles and

Cited by 0Scholar
2026

PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis

AAAI 2026technical

Recent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of real-world clinical diagnostics, which involve heterogeneous inputs and require ongoing contextual understanding during pa

Cited by 0SourcePDFScholar
2026

RigMo: Unifying Rig and Motion Learning for Generative Animation

CVPR 2026

Despite significant progress in 4D generation, rig and motion--the core structural and dynamic components of animation--are typically modeled as separate problems. Existing pipelines rely on ground-truth skeletons and skinning weights for motion generation and treat auto-rigging as an independent pr

Cited by 0SourceScholar
2026

SOLAR for Offline MARL: Plateau-Triggered Potential Shaping under World-Model Uncertainty

ICML 2026poster

Reward shaping can accelerate reinforcement learning, but in sparse-reward \emph{offline} multi-agent RL it is often brittle: dense intrinsic rewards may alter the underlying Markov game, while world-model guidance can amplify model bias. We find that shaping becomes reliable when it is (i) activate…

Cited by 0SourceScholar
2026

STFMamba: A Neurocognitive-Inspired Dual-Path State Space Model for Surgical Phase Recognition

RA-L 2026

Surgical phase recognition is critical in computer-assisted surgery. Clinically, surgeons discriminate surgical phases through visuospatial analysis of instrument-tissue interactions. However, existing methods fail to adequately account for the critical role of the visual-neural mechanisms of the su

Cited by 0SourceScholar
2026

Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction Generation

ICLR 2026poster

Generating realistic and diverse human-human interactions from text is a crucial yet challenging task in computer vision, graphics, and robotics. Despite recent advances, existing methods have two key limitations. First, two-person interaction synthesis is highly complex, simultaneously requiring in…

Cited by 0SourcecodeScholar
2026

The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation

ICRA 2026poster

While scaling laws for imitation learning have primarily focused on generalization in open-world settings, the relationship between data and precision in closed-world tasks like robotic assembly remains largely unexplored. This paper systematically investigates this relationship and introduces a nov…

Cited by 0Scholar
2026

Top-Down Semantic Refinement for Image Captioning

AAAI 2026technical

Large Vision-Language Models (VLMs) face an inherent contradiction in image captioning: their powerful single-step generation capabilities often lead to a myopic decision-making process. This makes it difficult to maintain global narrative coherence while capturing rich details, a limitation that is

Cited by 0SourcePDFScholar
2026

Towards Better Optimization For Listwise Preference in Diffusion Models

ICLR 2026poster

Reinforcement learning from human feedback (RLHF) has proven effectiveness for aligning text-to-image (T2I) diffusion models with human preferences. Although Direct Preference Optimization (DPO) is widely adopted for its computational efficiency and avoidance of explicit reward modeling, its applica…

Cited by 0SourceScholar
2026

Unleashing Guidance Without Classifiers for Human-Object Interaction Animation

ICLR 2026poster

Generating realistic human-object interaction (HOI) animations remains challenging because it requires jointly modeling dynamic human actions and diverse object geometries. Prior diffusion-based approaches often rely on handcrafted contact priors or human-imposed kinematic constraints to improve con…

Cited by 0SourceScholar
2026

Unleashing the Representational Power of Fourier Shapes for Attacking Infrared Object Detection

ICML 2026poster

Infrared object detection is crucial for perception in autonomous driving and surveillance but remains vulnerable to physical adversarial attacks. Unlike in the RGB domain, where attacks rely on color texture, infrared attacks must manipulate thermal signatures, making the geometry shape of heat-blo…

Cited by 0SourceScholar
2026

What Makes Effective Supervision in Latent Chain-of-Thought: An Information-Theoretic Analysis

ICML 2026poster

Latent Chain-of-Thought (CoT) aims to internalize reasoning into continuous hidden states, promising to transcend the computational bottlenecks of explicit tokens. However, the precise mechanisms ensuring its validity remain opaque. To bridge this gap, we establish an Information-Theoretic Framework…

Cited by 0SourceScholar
2026

Why Keep Your Doubts to Yourself? Trading Visual Uncertainties in Multi-Agent Bandit Systems

ICLR 2026poster

Vision-Language Models (VLMs) enable powerful multi-agent systems, but scaling them is economically unsustainable: coordinating heterogeneous agents under information asymmetry often spirals costs. Existing paradigms, such as Mixture-of-Agents and knowledge-based routers, rely on heuristic proxies t…

Cited by 0SourceScholar
2025

4KAgent: Agentic Any Image to 4K Super-Resolution

NeurIPS 2025poster

We present 4KAgent, a unified agentic super-resolution generalist system designed to universally upscale any image to 4K resolution (and even higher, if applied iteratively). Our system can transform images from extremely low resolutions with severe degradations, for example, highly distorted inputs…

Cited by 0SourcecodeScholar
2025

An Empirical Study of Federated Prompt Learning for Vision Language Model

IJCAI 2025

The Vision Language Model (VLM) excels in aligning vision and language representations, and prompt learning has emerged as a key technique for adapting such models to downstream tasks. However, the application of prompt learning with VLM in federated learning (FL) scenarios remains underexplored. Th

2025

Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond

EMNLP 2025

The increasing demand for domain-specific evaluation of large language models (LLMs) has led to the development of numerous benchmarks. These efforts often adhere to the principle of data scaling, relying on large corpora or extensive question-answer (QA) sets to ensure broad coverage. However, the

2025

Bioinspired Suction Cup Arrays Enabling Intelligent Adhesion in Underwater Environments

RA-L 2025

Biological adhesion represents a remarkable evolutionary adaptation, as exemplified by remoras using dorsal suction cups for hitchhiking and octopuses employing independently controllable suction cups to adhere to complex surfaces. Therefore, this study focuses on the design, quantitative analysis,

Cited by 0SourceScholar
2025

Bring Your Rear Cameras for Egocentric 3D Human Pose Estimation

ICCV 2025poster

Egocentric 3D human pose estimation has been actively studied using cameras installed in front of a head-mounted device (HMD). While frontal placement is the optimal and the only option for some tasks, such as hand tracking, it remains unclear if the same holds for full-body tracking due to self-occ…

2025

Discrete Curvature Graph Information Bottleneck

AAAI 2025technical

Graph neural networks(GNNs) have been demonstrated to depend on whether the node effective information is sufficiently passing. Discrete curvature (Ricci curvature) is used to study graph connectivity and information propagation efficiency with a geometric perspective, and has been raised in recent…

2025

DrDiff: Dynamic Routing Diffusion with Hierarchical Attention for Breaking the Efficiency-Quality Trade-off

EMNLP 2025

This paper introduces DrDiff, a novel framework for long-text generation that overcomes the efficiency-quality trade-off through three core technologies. First, we design a dynamic expert scheduling mechanism that intelligently allocates computational resources during the diffusion process based on

Cited by 0SourcePDFScholar
2025

Efficient and Hardware-Friendly Online Adaptation for Deep Stereo Depth Estimation on Embedded Robots

RA-L 2025

Accurate and real-time stereo depth estimation is important for autonomous robots, such as autonomous aerial vehicles (AAVs). Due to the computation constraints of these miniaturized robots, current state-of-the-art algorithms deploy light-weight neural networks while using self-supervised online ad

Cited by 3SourceScholar
2025

Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input

CVPR 2025poster

This work focuses on tracking and understanding human motion using consumer wearable devices, such as VR/AR headsets, smart glasses, cellphones, and smartwatches. These devices provide diverse, multi-modal sensor inputs, including egocentric images, and 1-3 sparse IMU sensors in varied combinations.…

Cited by 0SourcePDFScholar
2025

FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video

CVPR 2025highlight

Egocentric motion capture with a head-mounted body-facing stereo camera is crucial for VR and AR applications but presents significant challenges such as heavy occlusions and limited annotated real-world data. Existing methods rely on synthetic pretraining and struggle to generate smooth and accurat…

Cited by 0SourcePDFScholar
2025

Federated Recommendation with Explicitly Encoding Item Bias

AAAI 2025technical

With the development of federated learning techniques and the increased need for user privacy protection, the federated recommendation has become a new recommendation paradigm. However, most existing works focus on user-level federated recommendation, leaving platform-level federated recommendation…

Cited by 0SourcePDFScholar
2025

GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual Reasoning

NeurIPS 2025poster

We propose **GAM-Agent**, a game-theoretic multi-agent framework for enhancing vision-language reasoning. Unlike prior single-agent or monolithic models, GAM-Agent formulates the reasoning process as a non-zero-sum game between base agents—each specializing in visual perception subtasks—and a critic…

Cited by 0SourceScholar
2025

HIRAG: Hierarchical-Thought Instruction-Tuning Retrieval-Augmented Generation

EMNLP 2025

Retrieval-augmented generation (RAG) has become a fundamental paradigm for addressing the challenges faced by large language models in handling real-time information and domain-specific problems. Traditional RAG systems primarily rely on the in-context learning (ICL) capabilities of the large langua

Cited by 0SourcePDFScholar
2025

KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems

ICML 2025poster

As scaling large language models faces prohibitive costs, multi-agent systems emerge as a promising alternative, though challenged by static knowledge assumptions and coordination inefficiencies. We introduce Knowledge-Aware Bayesian Bandits (KABB), a novel framework that enhances multi-agent system…

Cited by 0SourcePDFScholar
2025

KVQ: Boosting Video Quality Assessment via Saliency-guided Local Perception

CVPR 2025poster

Video Quality Assessment (VQA), which intends to predict the perceptual quality of videos, has attracted increasing attention. Due to factors like motion blur or specific distortions, the quality of different regions in a video varies. Recognizing the region-wise local quality within a video is bene…

2025

Latent Chain-of-Thought for Visual Reasoning

NeurIPS 2025poster

Chain-of-thought (CoT) reasoning is critical for improving the interpretability and reliability of Large Vision-Language Models (LVLMs). However, existing training algorithms such as SFT, PPO, and GRPO may not generalize well across unseen reasoning tasks and heavily rely on a biased reward model. T…

Cited by 0SourceScholar
2025

MR-COGraphs: Communication-Efficient Multi-Robot Open-Vocabulary Mapping System via 3D Scene Graphs

RA-L 2025

Collaborative perception in unknown environments is crucial for multi-robot systems. With the emergence of foundation models, robots can now not only perceive geometric information but also achieve open-vocabulary scene understanding. However, existing map representations that support open-vocabular

Cited by 12SourcecodeScholar
2025

Ponimator: Unfolding Interactive Pose for Versatile Human-human Interaction Animation

ICCV 2025poster

Close-proximity human-human interactive poses convey rich contextual information about interaction dynamics. Given such poses, humans can intuitively infer the context and anticipate possible past and future dynamics, drawing on strong priors of human behavior. Inspired by this observation, we propo…

2025

Prototype Tuning: A Meta-Learning Approach for Few-Shot Document-Level Relation Extraction with Large Language Models

NAACL 2025findings

Few-Shot Document-Level Relation Extraction (FSDLRE) aims to develop models capable of generalizing to new categories with minimal support examples. Although Large Language Models (LLMs) demonstrate exceptional In-Context Learning (ICL) capabilities on many few-shot tasks, their performance on FSDLR…

2025

RAGDiffusion: Faithful Cloth Generation via External Knowledge Assimilation

ICCV 2025poster

Standard clothing asset generation involves restoring forward-facing flat-lay garment images displayed on a clear background by extracting clothing information from diverse real-world contexts, which presents significant challenges due to highly standardized structure sampling distributions and clot…

Cited by 0SourcePDFScholar
2025

STeCa: Step-level Trajectory Calibration for LLM Agent Learning

ACL 2025finding

Large language model (LLM)-based agents have shown promise in tackling complex tasks by interacting dynamically with the environment. Existing work primarily focuses on behavior cloning from expert demonstrations or preference learning through exploratory trajectory sampling. However, these methods…

2025

SceneMI: Motion In-betweening for Modeling Human-Scene Interaction

ICCV 2025poster

Modeling human-scene interactions (HSI) is essential for understanding and simulating everyday human behaviors. Recent approaches utilizing generative modeling have made progress in this domain; however, they are limited in controllability and flexibility for real-world applications. To address thes…

Cited by 0SourcePDFScholar
2025

Similar Modality Enhancement and Action Consistency Learning for Weakly Supervised Temporal Action Localization

AAAI 2025technical

Weakly-supervised temporal action localization (WTAL) aims to identify and localize action instances in untrimmed videos using only video-level labels. Existing methods typically rely on original features from frozen pre-trained encoders designed for trimmed action classification (TAC) tasks, which…

2025

SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modeling

CVPR 2025poster

Open-world interpretation aims to accurately localize and recognize all objects within images by vision-language models (VLMs). While substantial progress has been made in this task for natural images, the advancements for remote sensing (RS) images still remain limited, primarily due to these two c…

2025

T2Bs: Text-to-Character Blendshapes via Video Generation

ICCV 2025poster

We present T2Bs, a framework for generating high-quality, animatable character head morphable models from text by combining static text-to-3D generation with video diffusion. Text-to-3D models produce detailed static geometry but lack motion synthesis, while video diffusion models generate motion wi…

Cited by 0SourcePDFScholar
2025

TextMaster: A Unified Framework for Realistic Text Editing via Glyph-Style Dual-Control

ICCV 2025poster

In image editing tasks, high-quality text editing capabilities can significantly reduce both human and material resource costs. Existing methods, however, face significant limitations in terms of stroke accuracy for complex text and controllability of generated text styles. To address these challeng…

Cited by 0SourcePDFScholar
2025

Time-varying EEG Signal Reconstruction Based on Local Graph Signal Smoothness

ICASSP 2025accepted

Electroencephalography (EEG) signals record the electrical activity of the brain and have significant applications in neuroscience and medicine. However, accurately reconstructing EEG signals has been a challenge due to potential signal missing and noise interference during signal acquisition. The p…

Cited by 0SourceScholar
2025

Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding Tutors

ACL 2025finding

Intelligent tutoring agents powered by large language models (LLMs) have been increasingly explored to deliver personalized knowledge in areas such as language learning and science education. However, their capabilities in guiding users to solve complex real-world tasks remain underexplored. To addr…

2025

Training-Free Text-Guided Image Editing with Visual Autoregressive Model

ICCV 2025poster

Text-guided image editing is an essential task, enabling users to modify images through natural language descriptions. Recent advances in diffusion models and rectified flows have significantly improved editing quality, primarily relying on inversion techniques to extract structured noise from input…

2025

Tri-MARF: A Tri-Modal Multi-Agent Responsive Framework for Comprehensive 3D Object Annotation

NeurIPS 2025poster

Driven by the applications in autonomous driving, robotics, and augmented reality, 3D object annotation is a critical task compared to 2D annotation, such as spatial complexity, occlusion, and viewpoint inconsistency. The existing methods relying on single models often struggle with these issues. In…

Cited by 0SourceScholar
2025

WebRTC and 5G Based Remote Control System for a Vascular Intervention Robot

IROS 2025

Cardiovascular and cerebrovascular diseases are significant health issues that threaten human life. They typically develop insidiously and progress gradually, but when an event occurs, the consequences can be severe. These conditions often manifest suddenly and acutely, necessitating prompt treatmen

Cited by 0SourceScholar
2025

Why Safeguarded Ships Run Aground? Aligned Large Language Models’ Safety Mechanisms Tend to Be Anchored in The Template Region

ACL 2025long

The safety alignment of large language models (LLMs) remains vulnerable, as their initial behavior can be easily jailbroken by even relatively simple attacks. Since infilling a fixed template between the input instruction and initial model output is a common practice for existing LLMs, we hypothesiz…

2024

3D Human Pose Perception from Egocentric Stereo Videos

CVPR 2024highlight

While head-mounted devices are becoming more compact they provide egocentric views with significant self-occlusions of the device user. Hence existing methods often fail to accurately estimate complex 3D poses from egocentric views. In this work we propose a new transformer-based framework to improv…

Cited by 19SourcePDFScholar
2024

A Novel SEA-based Haptic Interface for Robot-Assisted Vascular Interventional Surgery

ICRA 2024poster

Robot-assisted vascular interventional surgery can isolate interventionists and X-ray radiation, and improve surgical accuracy. However, the leader side outside the operating room still has problems such as incomplete collection of operating information and unrealistic tactile feedback. The main obj…

Cited by 1SourceScholar
2024

Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight

NeurIPS 2024poster

This paper introduces Chain-of-Sight, a vision-language bridge module that accelerates the pre-training of Multimodal Large Language Models (MLLMs). Our approach employs a sequence of visual resamplers that capture visual details at various spacial scales. This architecture not only leverages globa…

Cited by 3SourcePDFScholar
2024

Autonomous Recovery Control of Biomimetic Robotic Fish Based on Multi-Sensory System

RA-L 2024

As a new-type underwater robot with outstanding motion performance and environmental friendliness, bionic robotic fish holds enormous potential in the field of ocean exploration. In practical application, it is a vital mission for underwater robots to accomplish efficient and stable autonomous recov

Cited by 10SourceScholar
2024

Cooper: Coordinating Specialized Agents towards a Complex Dialogue Goal

AAAI 2024technical

In recent years, there has been a growing interest in exploring dialogues with more complex goals, such as negotiation, persuasion, and emotional support, which go beyond traditional service-focused dialogue systems. Apart from the requirement for much more sophisticated strategic reasoning and comm…

2024

DSL-FIQA: Assessing Facial Image Quality via Dual-Set Degradation Learning and Landmark-Guided Transformer

CVPR 2024poster

Generic Face Image Quality Assessment (GFIQA) evaluates the perceptual quality of facial images which is crucial in improving image restoration algorithms and selecting high-quality face images for downstream tasks. We present a novel transformer-based method for GFIQA which is aided by two unique m…

Cited by 9SourcePDFScholar
2024

Dual Directional Complementary Gradient Fusion and Deep Refinement for Hyperspectral Image Super Resolution

ICASSP 2024accepted

The spatial and spectral resolution trade-off in the hyperspectral imaging is a fundamental and essential issue, and automatically generating high-resolution images in both spatial and spectral domains (HR-HS) by merging a low spatial resolution hyperspectral (LR-HS) image and a high spatial resolut…

Cited by 0SourceScholar
2024

E2CL: Exploration-based Error Correction Learning for Embodied Agents

EMNLP 2024finding

Language models are exhibiting increasing capability in knowledge utilization and reasoning. However, when applied as agents in embodied environments, they often suffer from misalignment between their intrinsic knowledge and environmental knowledge, leading to infeasible actions. Traditional environ…

2024

ESC-Eval: Evaluating Emotion Support Conversations in Large Language Models

EMNLP 2024main

Emotion Support Conversation (ESC) is a crucial application, which aims to reduce human stress, offer emotional guidance, and ultimately enhance human mental and physical well-being. With the advancement of Large Language Models (LLMs), many researchers have employed LLMs as the ESC models. However,…

2024

EcoMatcher: Efficient Clustering Oriented Matcher for Detector-free Image Matching

ECCV 2024poster

"Detector-free local feature matching methods have demonstrated significant performance improvements since leveraging the power of Transformer architecture. The global receptive field allows for simultaneous interaction among all elements, proving particularly beneficial in regions with low texture…

Cited by 1SourcePDFScholar
2024

Egocentric Whole-Body Motion Capture with FisheyeViT and Diffusion-Based Motion Refinement

CVPR 2024poster

In this work we explore egocentric whole-body motion capture using a single fisheye camera which simultaneously estimates human body and hand motion. This task presents significant challenges due to three factors: the lack of high-quality datasets fisheye camera distortion and human body self-occlus…

Cited by 21SourcePDFScholar
2024

EventEgo3D: 3D Human Motion Capture from Egocentric Event Streams

CVPR 2024poster

Monocular egocentric 3D human motion capture is a challenging and actively researched problem. Existing methods use synchronously operating visual sensors (e.g. RGB cameras) and often fail under low lighting and fast motions which can be restricting in many applications involving head-mounted device…

2024

Exponential Spectral Pursuit: An Effective Initialization Method for Sparse Phase Retrieval

ICML 2024poster

Sparse phase retrieval aims to reconstruct an $n$-dimensional $k$-sparse signal from its phaseless measurements. For most of the existing reconstruction algorithms, their sampling complexity is known to be dominated by the initialization stage. In this paper, in order to improve the sampling complex…

Cited by 3SourcePDFScholar
2024

Instruct Once, Chat Consistently in Multiple Rounds: An Efficient Tuning Framework for Dialogue

ACL 2024long

Tuning language models for dialogue generation has been a prevalent paradigm for building capable dialogue agents. Yet, traditional tuning narrowly views dialogue generation as resembling other language generation tasks, ignoring the role disparities between two speakers and the multi-round interact…

2024

Low Overhead DMG Sensing for Vital Signs Detection

ICASSP 2024accepted

Sensing biometric markers such as respiration rate (RR) and heart rate (HR) in non-medical contexts using the high resolution of Millimeter-Wave (mmWave) Wi-Fi networks has recently gathered considerable attention. A significant challenge in deploying a Wi-Fi system capable of performing sensing tas…

Cited by 0SourceScholar
2024

MS$^3$D: A RG Flow-Based Regularization for GAN Training with Limited Data

ICML 2024poster

Generative adversarial networks (GANs) have made impressive advances in image generation, but they often require large-scale training data to avoid degradation caused by discriminator overfitting. To tackle this issue, we investigate the challenge of training GANs with limited data, and propose a no…

Cited by 1SourcePDFScholar
2024

Mobile Attention: Mobile-Friendly Linear-Attention for Vision Transformers

ICML 2024poster

Vision Transformers (ViTs) excel in computer vision tasks due to their ability to capture global context among tokens. However, their quadratic complexity $\mathcal{O}(N^2D)$ in terms of token number $N$ and feature dimension $D$ limits practical use on mobile devices, necessitating more mobile-frie…

2024

POA: Pre-training Once for Models of All Sizes

ECCV 2024poster

"Large-scale self-supervised pre-training has paved the way for one foundation model to handle many different vision tasks. Most pre-training methodologies train a single model of a certain size at one time. Nevertheless, various computation or storage constraints in real-world scenarios require sub…

2024

Robust Communicative Multi-Agent Reinforcement Learning with Active Defense

AAAI 2024technical

Communication in multi-agent reinforcement learning (MARL) has been proven to effectively promote cooperation among agents recently. Since communication in real-world scenarios is vulnerable to noises and adversarial attacks, it is crucial to develop robust communicative MARL technique. However, exi…

Cited by 5SourcePDFScholar
2024

RobustSAM: Segment Anything Robustly on Degraded Images

CVPR 2024highlight

Segment Anything Model (SAM) has emerged as a transformative approach in image segmentation acclaimed for its robust zero-shot segmentation capabilities and flexible prompting system. Nonetheless its performance is challenged by images with degraded quality. Addressing this limitation we propose the…

2024

SAM: A Self-Adaptive Attention Module for Context-Aware Recommendation System

ICASSP 2024accepted

Recently, textual information has been proven to positively affect recommendation systems. However, most of the existing methods only focus on representation learning of textual information in ratings, while potential selection bias induced by the textual information is ignored. In this work, we pro…

Cited by 0SourceScholar
2024

SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery

CVPR 2024poster

Prior studies on Remote Sensing Foundation Model (RSFM) reveal immense potential towards a generic model for Earth Observation. Nevertheless these works primarily focus on a single modality without temporal and geo-context modeling hampering their capabilities for diverse tasks. In this study we pre…

Cited by 140SourcePDFScholar
2024

Towards Better Vision-Inspired Vision-Language Models

CVPR 2024poster

Vision-language (VL) models have achieved unprecedented success recently in which the connection module is the key to bridge the modality gap. Nevertheless the abundant visual clues are not sufficiently exploited in most existing methods. On the vision side most existing approaches only use the last…

Cited by 2SourcePDFScholar
2023

A Performance Optimization Strategy Based on Improved NSGA-II for a Flexible Robotic Fish

ICRA 2023poster

The high speed and low energy cost are two conflicting objectives in the motion optimization of bio-inspired underwater robots, but playing a very important role. To this end, this paper proposes an optimization strategy for swimming speed and power cost using an improved NSGA-II for a flexible robo…

Cited by 2SourceScholar
2023

A Unified Conditional Framework for Diffusion-based Image Restoration

NeurIPS 2023poster

Diffusion Probabilistic Models (DPMs) have recently shown remarkable performance in image generation tasks, which are capable of generating highly realistic images. When adopting DPMs for image restoration tasks, the crucial aspect lies in how to integrate the conditional information to guide the DP…

2023

COLA: Improving Conversational Recommender Systems by Collaborative Augmentation

AAAI 2023technical

Conversational recommender systems (CRS) aim to employ natural language conversations to suggest suitable products to users. Understanding user preferences for prospective items and learning efficient item representations are crucial for CRS. Despite various attempts, earlier studies mostly learned…

2023

Design and Modeling of a Sperm-Inspired Helical Propulsion Robot

RA-L 2023

The development of biomimetics and the demand for higher propulsion efficiency lead to more research in helical propulsion robots. This letter presents a novel sperm-inspired robot that utilizes flexible tail as propulsion and analyzes the motion performance. Firstly, the robot's propulsion system i

Cited by 1SourceScholar
2023

Dialogue Planning via Brownian Bridge Stochastic Process for Goal-directed Proactive Dialogue

ACL 2023findings

Goal-directed dialogue systems aim to proactively reach a pre-determined target through multi-turn conversations. The key to achieving this task lies in planning dialogue paths that smoothly and coherently direct conversations towards the target. However, this is a challenging and under-explored tas…

2023

Energy-Efficient Adaptive 3D Sensing

CVPR 2023poster

Active depth sensing achieves robust depth estimation but is usually limited by the sensing range. Naively increasing the optical power can improve sensing range but induces eye-safety concerns for many applications, including autonomous robots and augmented reality. In this paper, we propose an ada…

2023

Graph Contrastive Learning for Skeleton-based Action Recognition

ICLR 2023poster

In the field of skeleton-based action recognition, current top-performing graph convolutional networks (GCNs) exploit intra-sequence context to construct adaptive graphs for feature aggregation. However, we argue that such context is still $\textit{local}$ since the rich cross-sequence relations hav…

2023

Group DETR: Fast DETR Training with Group-Wise One-to-Many Assignment

ICCV 2023poster

Detection transformer (DETR) relies on one-to-one assignment, assigning one ground-truth object to one prediction, for end-to-end detection without NMS post-processing. It is known that one-to-many assignment, assigning one ground-truth object to multiple predictions, succeeds in detection methods s…

Cited by 160PDFcodeScholar
2023

Group Pose: A Simple Baseline for End-to-End Multi-Person Pose Estimation

ICCV 2023poster

In this paper, we study the problem of end-to-end multi-person pose estimation. State-of-the-art solutions adopt the DETR-like framework, and mainly develop the complex decoder, e.g., regarding pose estimation as keypoint box detection and combining with human detection in ED-Pose, hierarchically pr…

Cited by 41PDFcodeScholar
2023

HAP: Structure-Aware Masked Image Modeling for Human-Centric Perception

NeurIPS 2023poster

Model pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting the MIM training strategy, we reveal that human structure priors offer significant potential. Motivated by this insight…

2023

Medical Dialogue Generation via Dual Flow Modeling

ACL 2023findings

Medical dialogue systems (MDS) aim to provide patients with medical services, such as diagnosis and prescription. Since most patients cannot precisely describe their symptoms, dialogue understanding is challenging for MDS. Previous studies mainly addressed this by extracting the mentioned medical en…

2023

Model-Based Event-Triggered Dynamic Pursuing and Surrounding Control for a Multi-Robotic Fish System

RA-L 2023

This paper investigates the event-triggered-based pursuing and surrounding control of a multi-robotic fish system. A distributed surrounding control framework is put forward to deploy the whole system and enclose the dynamic evader into a convex hull. In particular, distributed kinematics-based moti

Cited by 9SourceScholar
2023

PSVT: End-to-End Multi-Person 3D Pose and Shape Estimation With Progressive Video Transformers

CVPR 2023poster

Existing methods of multi-person video 3D human Pose and Shape Estimation (PSE) typically adopt a two-stage strategy, which first detects human instances in each frame and then performs single-person PSE with temporal model. However, the global spatio-temporal context among spatial instances can not…

Cited by 35SourcePDFScholar
2023

Promoting Cooperation in Multi-Agent Reinforcement Learning via Mutual Help

ICASSP 2023accepted

Multi-agent reinforcement learning (MARL) has achieved great progress in cooperative tasks in recent years. However, in the local reward scheme, where only local rewards for each agent are given without global rewards shared by all the agents, traditional MARL algorithms lack sufficient consideratio…

Cited by 0SourceScholar
2023

Scene-Aware Egocentric 3D Human Pose Estimation

CVPR 2023poster

Egocentric 3D human pose estimation with a single head-mounted fisheye camera has recently attracted attention due to its numerous applications in virtual and augmented reality. Existing methods still struggle in challenging poses where the human body is highly occluded or is closely interacting wit…

2023

Self-Detoxifying Language Models via Toxification Reversal

EMNLP 2023long main

Language model detoxification aims to minimize the risk of generating offensive or harmful content in pretrained language models (PLMs) for safer deployment. Existing methods can be roughly categorized as finetuning-based and decoding-based. However, the former is often resource-intensive, while the…

Cited by 0SourcecodeScholar
2023

Simultaneously Short- and Long-Term Temporal Modeling for Semi-Supervised Video Semantic Segmentation

CVPR 2023poster

In order to tackle video semantic segmentation task at a lower cost, e.g., only one frame annotated per video, lots of efforts have been devoted to investigate the utilization of those unlabeled frames by either assigning pseudo labels or performing feature enhancement. In this work, we propose a no…

Cited by 12SourcePDFScholar
2023

Stroke Extraction of Chinese Character Based on Deep Structure Deformable Image Registration

AAAI 2023technical

Stroke extraction of Chinese characters plays an important role in the field of character recognition and generation. The most existing character stroke extraction methods focus on image morphological features. These methods usually lead to errors of cross strokes extraction and stroke matching due…

2023

Target-oriented Proactive Dialogue Systems with Personalization: Problem Formulation and Dataset Curation

EMNLP 2023short main

Target-oriented dialogue systems, designed to proactively steer conversations toward predefined targets or accomplish specific system-side goals, are an exciting area in conversational AI. In this work, by formulating a <dialogue act, topic> pair as the conversation target, we explore a novel proble…

Cited by 0SourcecodeScholar
2023

Uncertainty-guided Learning for Improving Image Manipulation Detection

ICCV 2023poster

Image manipulation detection (IMD) is of vital importance as faking images and spreading misinformation can be malicious and harm our daily life. IMD is the core technique to solve these issues and poses challenges in two main aspects: (1) Data Uncertainty, i.e., the manipulated artifacts are often…

Cited by 18PDFcodeScholar
2023

Unified Pre-Training with Pseudo Texts for Text-To-Image Person Re-Identification

ICCV 2023poster

The pre-training task is indispensable for the text-to-image person re-identification (T2I-ReID) task. However, there are two underlying inconsistencies between these two tasks that may impact the performance: i) Data inconsistency. A large domain gap exists between the generic images/texts used in…

Cited by 45PDFcodeScholar
2023

s-Adaptive Decoupled Prototype for Few-Shot Object Detection

ICCV 2023poster

Meta-learning-based few-shot detectors use one K-average-pooled prototype (averaging along K-shot dimension) in both Region Proposal Network (RPN) and Detection head (DH) for query detection. Such plain operation would harm the FSOD performance in two aspects: 1) the poor quality of the prototype, a…

Cited by 14PDFScholar
2022

3D Photo Stylization: Learning To Generate Stylized Novel Views From a Single Image

CVPR 2022oral

Visual content creation has spurred a soaring interest given its applications in mobile photography and AR / VR. Style transfer and single-image 3D photography as two representative tasks have so far evolved independently. In this paper, we make a connection between the two, and address the challeng…

Cited by 60PDFScholar
2022

Action Quality Assessment with Temporal Parsing Transformer

ECCV 2022poster

"Action Quality Assessment(AQA) is important for action understanding and resolving the task poses unique challenges due to subtle visual differences. Existing state-of-the-art methods typically rely on the holistic video representations for score regression or ranking, which limits the generalizati…

Cited by 60SourcePDFScholar
2022

Development and Stiffness Optimization for a Flexible-Tail Robotic Fish

RA-L 2022

The integral flexible tail has the potential advantage of lifelike undulating motion. However, due to the complex manufacturing process and difficult modification of structural parameters, its application in robotic fish encounters many challenges. Combining rigid structure and flexible material, th

Cited by 17SourceScholar
2022

Dynamic Modeling and Performance Analysis for a Wire-Driven Elastic Robotic Fish

RA-L 2022

The complex and continuous undulation of fishtail facilitates extraordinary underwater motion performance for natural fish. For the widely used Multi-Joint robotic fish, a lot of joints are used to simulate continuum fishtail, resulting in some challenges, e.g., the mechanism complexity, friction lo

Cited by 12SourceScholar
2022

Estimating Egocentric 3D Human Pose in the Wild With External Weak Supervision

CVPR 2022poster

Egocentric 3D human pose estimation with a single fisheye camera has drawn a significant amount of attention recently. However, existing methods struggle with pose estimation from in-the-wild images, because they can only be trained on synthetic data due to the unavailability of large-scale in-the-w…

Cited by 37PDFScholar
2022

Explore-Bench: Data Sets, Metrics and Evaluations for Frontier-based and Deep-reinforcement-learning-based Autonomous Exploration

ICRA 2022poster

Autonomous exploration and mapping of unknown terrains employing single or multiple robots is an essential task in mobile robotics and has therefore been widely investigated. Nevertheless, given the lack of unified data sets, metrics, and platforms to evaluate the exploration approaches, we develop…

Cited by 41SourcecodeScholar
2022

Hierarchical Memory Learning for Fine-Grained Scene Graph Generation

ECCV 2022poster

"Regarding Scene Graph Generation (SGG), coarse and fine predicates mix in the dataset due to the crowd-sourced labeling, and the long-tail problem is also pronounced. Given this tricky situation, many existing SGG methods treat the predicates equally and learn the model under the supervision of mix…

Cited by 31SourcePDFScholar
2022

Human-Object Interaction Detection via Disentangled Transformer

CVPR 2022poster

Human-Object Interaction Detection tackles the problem of joint localization and classification of human object interactions. Existing HOI transformers either adopt a single decoder for triplet prediction, or utilize two parallel decoders to detect individual objects and interactions separately, and…

Cited by 77PDFScholar
2022

Implicit Sample Extension for Unsupervised Person Re-Identification

CVPR 2022poster

Most existing unsupervised person re-identification (Re-ID) methods use clustering to generate pseudo labels for model training. Unfortunately, clustering sometimes mixes different true identities together or splits the same identity into two or more sub clusters. Training on these noisy clusters su…

Cited by 135PDFcodeScholar
2022

MixFormer: Mixing Features Across Windows and Dimensions

CVPR 2022oral

While local-window self-attention performs notably in vision tasks, it suffers from limited receptive field and weak modeling capability issues. This is mainly because it performs self-attention within non-overlapped windows and shares weights on the channel dimension. We propose MixFormer to find a…

Cited by 161PDFcodeScholar
2022

Point Cloud Change Detection With Stereo V-SLAM: Dataset, Metrics and Baseline

RA-L 2022

Localization and navigation are basic robotic tasks requiring an accurate and up-to-date map to finish these tasks, with crowdsourced data to detect map changes posing an appealing solution. Collecting and processing crowdsourced data requires low-cost sensors and algorithms, but existing methods re

Cited by 4SourcecodeScholar
2022

RTFormer: Efficient Design for Real-Time Semantic Segmentation with Transformer

NeurIPS 2022accept

Recently, transformer-based networks have shown impressive results in semantic segmentation. Yet for real-time semantic segmentation, pure CNN-based approaches still dominate in this field, due to the time-consuming computation mechanism of transformer. We propose RTFormer, an efficient dual-resolut…

2022

RealMedDial: A Real Telemedical Dialogue Dataset Collected from Online Chinese Short-Video Clips

COLING 2022main

Intelligent medical services have attracted great research interests for providing automated medical consultation. However, the lack of corpora becomes a main obstacle to related research, particularly data from real scenarios. In this paper, we construct RealMedDial, a Chinese medical dialogue data…

2022

Relative Distributed Formation and Obstacle Avoidance with Multi-agent Reinforcement Learning

ICRA 2022poster

Multi-agent formation as well as obstacle avoid-ance is one of the most actively studied topics in the field of multi-agent systems. Although some classic controllers like model predictive control (MPC) and fuzzy control achieve a certain measure of success, most of them require precise global infor…

Cited by 24SourceScholar
2022

Self-Guided Hard Negative Generation for Unsupervised Person Re-Identification

IJCAI 2022poster

Recent unsupervised person re-identification (reID) methods mostly apply pseudo labels from clustering algorithms as supervision signals. Despite great success, this fashion is very likely to aggregate different identities with similar appearances into the same cluster. In result, the hard negative…

Cited by 12SourcePDFScholar
2022

Singular Value Fine-tuning: Few-shot Segmentation requires Few-parameters Fine-tuning

NeurIPS 2022accept

Freezing the pre-trained backbone has become a standard paradigm to avoid overfitting in few-shot segmentation. In this paper, we rethink the paradigm and explore a new regime: {\em fine-tuning a small part of parameters in the backbone}. We present a solution to overcome the overfitting problem, le…

2022

Training Object Detectors From Scratch: An Empirical Study in the Era of Vision Transformer

CVPR 2022poster

Modeling in computer vision has long been dominated by convolutional neural networks (CNNs). Recently, in light of the excellent performances of self-attention mechanism in the language field, transformers tailored for visual data have drawn numerous attention and triumphed CNNs in various vision ta…

Cited by 16PDFScholar
2022

Two Languages Are Better than One: Bilingual Enhancement for Chinese Named Entity Recognition

COLING 2022main

Chinese Named Entity Recognition (NER) has continued to attract research attention. However, most existing studies only explore the internal features of the Chinese language but neglect other lingual modal features. Actually, as another modal knowledge of the Chinese language, English contains rich…

2022

Uncertainty Modeling in Generative Compressed Sensing

ICML 2022spotlight

Compressed sensing (CS) aims to recover a high-dimensional signal with structural priors from its low-dimensional linear measurements. Inspired by the huge success of deep neural networks in modeling the priors of natural signals, generative neural networks have been recently used to replace the han…

2022

UnrealEgo: A New Dataset for Robust Egocentric 3D Human Motion Capture

ECCV 2022poster

"We present UnrealEgo, a new large-scale naturalistic dataset for egocentric 3D human pose estimation. UnrealEgo is based on an advanced concept of eyeglasses equipped with two fisheye cameras that can be used in unconstrained environments. We design their virtual prototype and attach them to 3D hum…

Cited by 54SourcePDFScholar
2021

A Ranked Similarity Loss Function with pair Weighting for Deep Metric Learning

ICASSP 2021accepted

Metric learning is a widely-used method for image retrieval. The object of metric learning is to limit the distance between similar samples and increase the distance between samples of different classes through learning. Many studies tend to pay more attention to keep the distance between positive a…

Cited by 0SourceScholar
2021

An Open-Source, Fiducial-Based, Underwater Stereo Visual-Inertial Localization Method with Refraction Correction

IROS 2021poster

Underwater visual localization is an essential technique for the autonomous operation of underwater robots. However, the unique underwater image characteristics, including refraction, sparse features, and severe noise, pose an enormous challenge to it. For addressing these issues, this paper propose…

Cited by 10SourceScholar
2021

Autonomous Navigation of an Ultrasound Probe Towards Standard Scan Planes with Deep Reinforcement Learning

ICRA 2021poster

Autonomous ultrasound (US) acquisition is an important yet challenging task, as it involves interpretation of the highly complex and variable images and their spatial relationships. In this work, we propose a deep reinforcement learning framework to autonomously control the 6-D pose of a virtual US…

Cited by 66SourceScholar
2021

Estimating Egocentric 3D Human Pose in Global Space

ICCV 2021poster

Egocentric 3D human pose estimation using a single fisheye camera has become popular recently as it allows capturing a wide range of daily activities in unconstrained environments, which is difficult for traditional outside-in motion capture with external cameras. However, existing methods have seve…

Cited by 83PDFcodeScholar
2021

Focus on Interaction: A Novel Dynamic Graph Model for Joint Multiple Intent Detection and Slot Filling

IJCAI 2021poster

Intent detection and slot filling are two main tasks for building a spoken language understanding (SLU) system. Since the two tasks are closely related, the joint models for the two tasks always outperform the pipeline models in SLU. However, most joint models directly incorporate multiple intent in…

2021

MFNet: Multi-Filter Directive Network for Weakly Supervised Salient Object Detection

ICCV 2021poster

Weakly supervised salient object detection (WSOD) targets to train a CNNs-based saliency network using only low-cost annotations. Existing WSOD methods take various techniques to pursue single "high-quality" pseudo label from low-cost annotations and then develop their saliency networks. Though thes…

Cited by 82PDFcodeScholar
2021

Marine Autonomous Navigation for Biomimetic Underwater Robots Based on Deep Stereo Attention Network

IROS 2021poster

This paper proposes a multi-objective visionbased navigation network for biomimetic underwater robots to cope with scientific observation, target selection, and obstacle avoidance in marine missions. Structurally, a stereo block attention module is first constructed to serially extract the channel a…

Cited by 4SourceScholar
2021

Mining Contextual Information Beyond Image for Semantic Segmentation

ICCV 2021poster

This paper studies the context aggregation problem in semantic image segmentation. The existing researches focus on improving the pixel representations by aggregating the contextual information within individual images. Though impressive, these methods neglect the significance of the representations…

Cited by 105PDFcodeScholar
2021

RNNRepair: Automatic RNN Repair via Model-based Analysis

ICML 2021spotlight

Deep neural networks are vulnerable to adversarial attacks. Due to their black-box nature, it is rather challenging to interpret and properly repair these incorrect behaviors. This paper focuses on interpreting and repairing the incorrect behaviors of Recurrent Neural Networks (RNNs). We propose a l…

Cited by 25SourcePDFScholar
2021

Unsupervised Multi-Source Domain Adaptation for Person Re-Identification

CVPR 2021poster

Unsupervised domain adaptation (UDA) methods for person re-identification (re-ID) aim at transferring re-ID knowledge from labeled source data to unlabeled target data. Among these methods, the pseudo-label-based branch has achieved great success, whereas most of them only use limited data from a si…

Cited by 112PDFScholar
2020

Dual Dynamic Memory Network for End-to-End Multi-turn Task-oriented Dialog Systems

COLING 2020main

Existing end-to-end task-oriented dialog systems struggle to dynamically model long dialog context for interactions and effectively incorporate knowledge base (KB) information into dialog generation. To conquer these limitations, we propose a Dual Dynamic Memory Network (DDMN) for multi-turn dialog…

2020

FakeSpotter: A Simple yet Robust Baseline for Spotting AI-Synthesized Fake Faces

IJCAI 2020poster

In recent years, generative adversarial networks (GANs) and its variants have achieved unprecedented success in image synthesis. They are widely adopted in synthesizing facial images which brings potential security concerns to humans as the fakes spread and fuel the misinformation. However, robust d…

2020

Graph-PCNN: Two Stage Human Pose Estimation with Graph Pose Refinement

ECCV 2020poster

Recently, most of the state-of-the-art human pose estimation methods are based on heatmap regression. The final coordinates of keypoints are obtained by decoding heatmap directly. In this paper, we aim to find a better approach to get more accurate localization results. We mainly put forward two sug…

Cited by 117SourcePDFScholar
2020

Group Contextual Encoding for 3D Point Clouds

NeurIPS 2020poster

Global context is crucial for 3D point cloud scene understanding tasks. In this work, we extended the contextual encoding layer that was originally designed for 2D tasks to 3D Point Cloud scenarios. The encoding layer learns a set of code words in the feature space of the 3D point cloud to characte…

2020

Watch out! Motion is Blurring the Vision of Your Deep Neural Networks

NeurIPS 2020poster

The state-of-the-art deep neural networks (DNNs) are vulnerable against adversarial examples with additive random-like noise perturbations. While such examples are hardly found in the physical world, the image blurring effect caused by object motion, on the other hand, commonly occurs in practice, m…

2019

Agile Depth Sensing Using Triangulation Light Curtains

ICCV 2019oral

Depth sensors like LIDARs and Kinect use a fixed depth acquisition strategy that is independent of the scene of interest. Due to the low spatial and temporal resolution of these sensors, this strategy can undersample parts of the scene that are important (small or fast moving objects), or oversample…

Cited by 28PDFScholar
2018

Neural Network Language Modeling with Letter-Based Features and Importance Sampling

ICASSP 2018accepted

In this paper we describe an extension of the Kaldi software toolkit to support neural-based language modeling, intended for use in automatic speech recognition (ASR) and related tasks. We combine the use of subword features (letter n-grams) and one-hot encoding of frequent words so that the models…

Cited by 0SourceScholar
2018

Programmable Triangulation Light Curtains

ECCV 2018poster

A vehicle on a road or a robot in the field does not need a full-featured 3D depth sensor to detect potential collisions or monitor its blind spot. Instead, it needs to only monitor if any object comes within its near proximity which is an easier task than full depth scanning. We introduce a novel d…

Cited by 51SourcePDFScholar
2017

Automatic radar waveform recognition based on time-frequency analysis and convolutional neural network

ICASSP 2017accepted

In this paper, we apply the idea of deep learning to radar waveform recognition. Since the frequency variation with time is the most essential distinction among radar signals with different modulation types, we transform one-dimensional radar signals into time-frequency images (TFIs) using time-freq…

Cited by 0SourceScholar
2017

Premise Selection for Theorem Proving by Deep Graph Embedding

NeurIPS 2017spotlight

We propose a deep learning-based approach to the problem of premise selection: selecting mathematical statements relevant for proving a given conjecture. We represent a higher-order logic formula as a graph that is invariant to variable renaming but still fully preserves syntactic and semantic infor…

2017

Reflectance Capture Using Univariate Sampling of BRDFs

ICCV 2017poster

We propose the use of a light-weight setup consisting of a collocated camera and light source --- commonly found on mobile devices --- to reconstruct surface normals and spatially-varying BRDFs of near-planar material samples. A collocated setup provides only a 1-D "univariate" sampling of the 4-D B…

Cited by 73PDFScholar
2016

Single underwater image restoration by blue-green channels dehazing and red channel correction

ICASSP 2016accepted

Restoring underwater image from a single image is know to be ill-posed, and some assumptions made in previous methods are not suitable for many situations. In this paper, we propose a method based on blue-green channels dehazing and red channel correction for underwater image restoration. Firstly, b…

Cited by 0SourceScholar