← Search

Zhongyuan Wang

94 accepted papers

2026

A Sanity Check for Multi-In-Domain Face Forgery Detection in the Real World

CVPR 2026

Existing methods for deepfake detection aim to develop generalizable detectors. Although "generalizable" could be the ultimate target once and for all, with limited training forgeries and domains, it appears idealistic to expect generalization that covers entirely unseen variations, especially given

Cited by 1SourceScholar
2026

Action-Sketcher: From Reasoning to Action via Visual Sketches for Robotic Manipulation

CVPR 2026

Long-horizon robotic manipulation is increasingly important for real-world deployment, requiring spatial disambiguation in complex layouts and temporal resilience under dynamic interaction. However, existing end-to-end and hierarchical Vision-Language-Action (VLA) policies often rely on text-only cu

Cited by 0SourcecodeScholar
2026

Agentic Reinforced Policy Optimization

ICLR 2026poster

Large-scale reinforcement learning with verifiable rewards (RLVR) has proven effective in harnessing the potential of large language models (LLMs) for single-turn reasoning tasks. In realistic reasoning scenarios, LLMs often rely on external tools to assist in task-solving processes. However, curren…

Cited by 0SourcecodeScholar
2026

DeformTrace: A Deformable State Space Model with Relay Tokens for Temporal Forgery Localization

AAAI 2026technical

Temporal Forgery Localization (TFL) aims to precisely identify manipulated segments in video and audio, offering strong interpretability for security and forensics. While recent State Space Models (SSMs) show promise in precise temporal reasoning, their use in TFL is hindered by ambiguous boundaries

Cited by 0SourcePDFScholar
2026

Divide and Conquer: Reliable Multi-View Evidential Learning for Deepfake Detection

ICML 2026poster

With the evolution of generative models, deepfakes have achieved near-perfect semantic realism, leaving forensic traces only in subtle structural anomalies. However, existing single-view paradigms often fail to generalize, as dominant semantic features overwhelm subtle artifact cues within entangled…

Cited by 0SourceScholar
2026

Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control

CVPR 2026

Humans intuitively move to sound, but current humanoid robots lack expressive improvisational capabilities, confined to predefined motions or sparse commands. Generating motion from audio and then retargeting it to robots relies on explicit motion reconstruction, leading to cascaded errors, high lat

Cited by 0SourceScholar
2026

From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance

ICLR 2026poster

Natural language offers a natural interface for humanoid robots, but existing text-to-motion pipelines remain cumbersome and unreliable. They typically decode human motion, retarget it to robot morphology, and then track it with a physics-based controller. However, this multi-stage process is prone…

Cited by 0SourceScholar
2026

GEM-TFL: Bridging Weak and Full Supervision for Forgery Localization through EM-Guided Decomposition and Temporal Refinement

CVPR 2026

Temporal Forgery Localization (TFL) aims to precisely identify manipulated segments within videos or audio streams, providing interpretable evidence for multimedia forensics and security. While most existing TFL methods rely on dense frame-level labels in a fully supervised manner, Weakly Supervised

Cited by 0SourceScholar
2026

GLoMOT: Efficient Online GNN-based Low-Frame-Rate Multi-Object Tracker

AAAI 2026technical

Low-frame-rate (LFR) Multi-Object Tracking (MOT) is crucial for efficient tracking on edge devices, as it significantly reduces computational and storage demands. However, existing trackers struggle in LFR settings due to large temporal gaps, extreme appearance changes, and motion non-linearity. Whi

Cited by 0SourcePDFScholar
2026

General Process Reward Modeling for Robotic Reinforcement Learning

CVPR 2026

The primary obstacle for applying reinforcement learning (RL) to real-world robotics is the design of effective reward functions. While recently learning-based Process Reward Models (PRMs) are a promising direction, they are often hindered by two fundamental limitations: their reward models lack ste

Cited by 0SourcecodeScholar
2026

Knowledge Reasoning Language Model: Unifying Knowledge and Language for Inductive Knowledge Graph Reasoning

ICLR 2026poster

Inductive Knowledge Graph Reasoning (KGR) aims to discover facts in open-domain KGs containing unknown entities and relations, which poses a challenge for KGR models in comprehending uncertain KG components. Existing studies have proposed Knowledge Graph Foundation Models (KGFMs) that learn structur…

Cited by 0SourceScholar
2026

Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models

ICML 2026poster

Vision-Language-Action (VLA) models benefit from Chain-of-Thought (CoT) reasoning, but existing approaches incur high inference overhead and rely on discrete reasoning representations that mismatch continuous perception and control. We propose Latent Reasoning VLA (LaRA-VLA), a unified VLA framework…

Cited by 0SourceScholar
2026

MS^2Gait: A Multi-Scale Spatio-Temporal Fusion Network for LiDAR-based Gait Recognition

CVPR 2026

3D LiDAR-based gait recognition has gained increasing attention due to its robustness to illumination, privacy preservation, and capability for long-range and non-contact identity verification. However, existing point cloud-based methods suffer from two critical limitations: they fail to model seman

Cited by 0SourceScholar
2026

OmniGen2: Towards Instruction-Aligned Multimodal Generation

CVPR 2026

Multimodal generative models can process instructions in various modalities and demonstrate outstanding performance across a wide range of image generation tasks. However, their robustness in complex real-world scenarios remains limited due to insufficient generalized instruction alignment. We intro

Cited by 0SourcecodeScholar
2026

Rethinking Surgical Smoke: A Smoke-Type-Aware Laparoscopic Video Desmoking Method and Dataset

AAAI 2026technical

Electrocautery or lasers will inevitably generate surgical smoke, which hinders the visual guidance of laparoscopic videos for surgical procedures. The surgical smoke can be classified into different types based on its motion patterns, leading to distinctive spatio-temporal characteristics across sm

Cited by 0SourcePDFScholar
2026

SaPaVe: Towards Active Perception and Manipulation in Vision-Language Action Models for Robotics

CVPR 2026

Active perception and manipulation are crucial for robots to interact with complex scenes. Existing methods struggle to unify semantic-driven perception actively with robust, viewpoint-invariant execution accordingly. To this end, we propose SaPaVe, an end-to-end framework that jointly learns these

Cited by 0SourceScholar
2026

TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics

ICRA 2026poster

Vision-Language Models (VLMs) have shown remarkable capabilities in spatial reasoning, yet they remain fundamentally limited to qualitative assessments and lack the computational precision required for real-world robotics. Current approaches fail to leverage metric information from depth sensors and…

2026

Towards Effective Code-Integrated Reasoning

AAAI 2026technical

In this paper, we investigate code-integrated reasoning (CIR), where models generate code when necessary and integrate feedback by executing it through a code interpreter. To acquire this capability, models must learn when and how to use external code tools effectively, which is supported by tool-au

Cited by 0SourcePDFScholar
2026

Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection

CVPR 2026

Standard supervised training for deepfake detection treats all samples with uniform importance, which can be suboptimal for learning robust and generalizable features. In this work, we propose a novel Tutor-Student Reinforcement Learning (TSRL) framework to dynamically optimize the training curricul

Cited by 0SourcecodeScholar
2025

AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation

NeurIPS 2025poster

Recently, mobile manipulation has attracted increasing attention for enabling language-conditioned robotic control in household tasks. However, existing methods still face challenges in coordinating mobile base and manipulator, primarily due to two limitations. On the one hand, they fail to explicit…

Cited by 0SourceScholar
2025

ACRL-10K: A Dataset for Air Conditioner Refrigerant Leak Smoke Detection

ICASSP 2025accepted

In this paper, we introduce a new dataset for air conditioner refrigerant leak smoke detection, called ACRL-10K. The dataset is designed to develop algorithms for detecting refrigerant leak smoke faults during air conditioner recycling. It contains a total of 10,724 images covering three common scen…

Cited by 0SourceScholar
2025

AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter

IROS 2025

Inferring the affordance of an object and grasping it in a task-oriented manner is crucial for robots to successfully complete manipulation tasks. Affordance indicates where and how to grasp an object by taking its functionality into account, serving as the foundation for effective task-oriented gra

Cited by 27SourcecodeScholar
2025

Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection

CVPR 2025poster

Automatic detection and prevention of open-set failures are crucial in closed-loop robotic systems. Recent studies often struggle to simultaneously identify unexpected failures reactively after they occur and prevent foreseeable ones proactively. To this end, we propose Code-as-Monitor (CaM), a nove…

Cited by 7SourcePDFScholar
2025

FriendsQA: A New Large-Scale Deep Video Understanding Dataset with Fine-grained Topic Categorization for Story Videos

AAAI 2025technical

Video question answering (VideoQA) aims to answer natural language questions according to the given videos. Although existing models perform well in the factoid VideoQA task, they still face challenges in deep video understanding (DVU) task, which focuses on story videos. Compared to factoid videos,…

2025

Lift3D Policy: Lifting 2D Foundation Models for Robust 3D Robotic Manipulation

CVPR 2025poster

3D geometric information is essential for manipulation tasks, as robots need to perceive the 3D environment, reason about spatial relationships, and interact with intricate spatial configurations. Recent research has increasingly focused on the explicit extraction of 3D features, while still facing…

Cited by 0SourcePDFScholar
2025

Link-based Contrastive Learning for One-Shot Unsupervised Domain Adaptation

CVPR 2025poster

Unsupervised domain adaptation (UDA) aims to learn discriminative features from a labeled source domain by supervised learning and to transfer the knowledge to an unlabeled target domain via distribution alignment. However, in some real-world scenarios, e.g., public safety or access control, it's di…

Cited by 0SourcePDFScholar
2025

MapNav: A Novel Memory Representation via Annotated Semantic Maps for VLM-based Vision-and-Language Navigation

ACL 2025long

Vision-language navigation (VLN) is a key task in Embodied AI, requiring agents to navigate diverse and unseen environments while following natural language instructions. Traditional approaches rely heavily on historical observations as spatio-temporal contexts for decision making, leading to signif…

2025

Multi-Shape Matching with Cycle Consistency Basis via Functional Maps

AAAI 2025technical

Multi-shape matching is a central problem in various applications of computer vision and graphics, where cycle consistency constraints play a pivotal role. For this issue, we propose a novel and efficient approach that models multi-shapes as directed graphs for two-stage optimization, i.e., optimizi…

2025

Na Vid-4D: Unleashing Spatial Intelligence in Egocentric RGB-D Videos for Vision-and-Language Navigation

ICRA 2025

Understanding and reasoning about the 4D space-time is crucial for Vision-and-Language Navigation (VLN). However, previous works lack in-depth exploration in this aspect, resulting in bottlenecked spatial perception and action precision of VLN agents. In this work, we introduce NaVid-4D, a Vision La

Cited by 4SourceScholar
2025

OODML: Whole Slide Image Classification Meets Online Pseudo-Supervision and Dynamic Mutual Learning

AAAI 2025technical

Bag-label-based multi-instance learning (MIL) has demonstrated significant performance in whole slide image (WSI) analysis, particularly in pseudo-label-based learning schemes. However, due to inaccurate feature representation and interference, existing MIL methods often yield unreliable pseudo-labe…

Cited by 0SourcePDFScholar
2025

ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges

ICCV 2025poster

As multi-modal large language models (MLLMs) frequently exhibit errors when solving scientific problems, evaluating the validity of their reasoning processes is critical for ensuring reliability and uncovering fine-grained model weaknesses. Since human evaluation is laborious and costly, prompting M…

2025

Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models

NeurIPS 2025poster

Visual reasoning abilities play a crucial role in understanding complex multimodal data, advancing both domain-specific applications and artificial general intelligence (AGI). Existing methods enhance Vision-Language Models (VLMs) through Chain-of-Thought (CoT) supervised fine-tuning using meticulou…

Cited by 0SourceScholar
2025

Rethinking the Adversarial Robustness of Multi-Exit Neural Networks in an Attack-Defense Game

CVPR 2025poster

Multi-exit neural networks represent a promising approach to enhancing model inference efficiency, yet like common neural networks, they suffer from significantly reduced robustness against adversarial attacks. While some defense methods have been raised to strengthen the adversarial robustness of m…

Cited by 0SourcePDFScholar
2025

RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

CVPR 2025poster

Recent advancements in Multimodal Large Language Models (MLLMs) have shown remarkable capabilities across various multimodal contexts. However, their application in robotic scenarios, particularly for long-horizon manipulation tasks, reveals significant limitations. These limitations arise from the…

Cited by 9SourcePDFScholar
2025

RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics

NeurIPS 2025poster

Spatial referring is a fundamental capability of embodied robots to interact with the 3D physical world. However, even with the powerful pretrained VLMs, recent approaches are still not qualified to accurately understand the complex 3D scenes and dynamically reason about the instruction-indicated lo…

Cited by 0SourceScholar
2025

SimpleDeepSearcher: Deep Information Seeking via Web-Powered Reasoning Trajectory Synthesis

EMNLP 2025

Retrieval-augmented generation (RAG) systems have advanced large language models (LLMs) in complex deep search scenarios requiring multi-step reasoning and iterative information retrieval. However, existing approaches face critical limitations that lack high-quality training trajectories or suffer f

2025

Stacking Brick by Brick: Aligned Feature Isolation for Incremental Face Forgery Detection

CVPR 2025poster

The rapid advancement of face forgery techniques has introduced a growing variety of forgeries.Incremental Face Forgery Detection (IFFD), involvinggradually adding new forgery data to fine-tune the previously trained model, has been introduced as a promising strategy to deal with evolving forgery me…

2025

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

RSS 2025poster

Embodied Navigation is a fundamental capability for intelligent robots, requiring robots to follow human commands and move autonomously within physical environments. Despite significant advancements, most existing navigation approaches are tailored to specific navigation tasks, such as instruction f…

Cited by 12PDFScholar
2024

Can We Leave Deepfake Data Behind in Training Deepfake Detector?

NeurIPS 2024poster

The generalization ability of deepfake detectors is vital for their applications in real-world scenarios. One effective solution to enhance this ability is to train the models with manually-blended data, which we termed ''blendfake'', encouraging models to learn generic forgery artifacts like blendi…

2024

Code-Style In-Context Learning for Knowledge-Based Question Answering

AAAI 2024technical

Current methods for Knowledge-Based Question Answering (KBQA) usually rely on complex training techniques and model frameworks, leading to many limitations in practical applications. Recently, the emergence of In-Context Learning (ICL) capabilities in Large Language Models (LLMs) provides a simple a…

2024

CogGPT: Unleashing the Power of Cognitive Dynamics on Large Language Models

EMNLP 2024finding

Cognitive dynamics, which refer to the evolution in human cognitive processes, are pivotal to advance human understanding of the world. Recent advancements in large language models (LLMs) highlight their potential for cognitive simulation. However, these LLM-based cognitive studies primarily focus o…

2024

Decoding at the Speed of Thought: Harnessing Parallel Decoding of Lexical Units for LLMs

COLING 2024main

Large language models have demonstrated exceptional capability in natural language understanding and generation. However, their generation speed is limited by the inherently sequential nature of their decoding process, posing challenges for real-time applications. This paper introduces Lexical Unit…

2024

Decompose, Prioritize, and Eliminate: Dynamically Integrating Diverse Representations for Multimodal Named Entity Recognition

COLING 2024main

Multi-modal Named Entity Recognition, a fundamental task for multi-modal knowledge graph construction, requires integrating multi-modal information to extract named entities from text. Previous research has explored the integration of multi-modal representations at different granularities. However,…

Cited by 1SourcePDFScholar
2024

GFMAE: Self-Supervised GNN-Free Masked Autoencoders

ICASSP 2024accepted

Generative self-supervised learning, represented by graph autoencoders (GAEs), has begun to exhibit significant potential in addressing graph tasks. However, GAEs often rely on Graph Neural Networks (GNNs) for encoding and decoding, this can pose a computation challenge due to the inherent complexit…

Cited by 0SourceScholar
2024

GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension

IJCAI 2024poster

There are substantial instructional videos on the Internet, which provide us tutorials for completing various tasks. Existing instructional video datasets only focus on specific steps at the video level, lacking experiential guidelines at the task level, which can lead to beginners struggling to lea…

Cited by 1SourcePDFScholar
2024

Just Ask One More Time! Self-Agreement Improves Reasoning of Language Models in (Almost) All Scenarios

ACL 2024findings

Although chain-of-thought (CoT) prompting combined with language models has achieved encouraging results on complex reasoning tasks, the naive greedy decoding used in CoT prompting usually causes the repetitiveness and local optimality. To address this shortcoming, ensemble-optimization tries to obt…

2024

Learning Multi-Dimensional Human Preference for Text-to-Image Generation

CVPR 2024poster

Current metrics for text-to-image models typically rely on statistical metrics which inadequately represent the real preference of humans. Although recent work attempts to learn these preferences via human annotated images they reduce the rich tapestry of human preference to a single overall score.…

Cited by 24SourcePDFScholar
2024

Relational Graph-Bridged Image-Text Interaction: A Novel Method for Multi-Modal Relation Extraction

ICASSP 2024accepted

Multi-modal relation extraction (MRE) requires the integration of multi-modal information to identify relationships between entities. Although fine-grained correlations between visual objects and textual words have the potential to improve cross-modal interaction, they are typically modeled implicit…

Cited by 0SourceScholar
2023

Augmentation-Aware Self-Supervision for Data-Efficient GAN Training

NeurIPS 2023poster

Training generative adversarial networks (GANs) with limited data is challenging because the discriminator is prone to overfitting. Previously proposed differentiable augmentation demonstrates improved data efficiency of training GANs. However, the augmentation implicitly introduces undesired invari…

2023

ConTextual Masked Auto-Encoder for Dense Passage Retrieval

AAAI 2023technical

Dense passage retrieval aims to retrieve the relevant passages of a query from a large corpus based on dense representations (i.e., vectors) of the query and the passages. Recent studies have explored improving pre-trained language models to boost dense retrieval performance. This paper proposes CoT…

2023

Continuous Learning for Blind Image Quality Assessment with Contrastive Transformer

ICASSP 2023accepted

Most existing blind image quality assessment (BIQA) models focus on improving performance on existing datasets and are weak in adapting to unknown distortion or degradation types. In this paper, we propose a Transformer-based BIQA contrastive continual learning approach to improve model transfer per…

Cited by 0SourceScholar
2023

FEditNet: Few-Shot Editing of Latent Semantics in GAN Spaces

AAAI 2023technical

Generative Adversarial networks (GANs) have demonstrated their powerful capability of synthesizing high-resolution images, and great efforts have been made to interpret the semantics in the latent spaces of GANs. However, existing works still have the following limitations: (1) the majority of works…

2023

GTR: A Grafting-Then-Reassembling Framework for Dynamic Scene Graph Generation

IJCAI 2023poster

Dynamic scene graph generation aims to identify visual relationships (subject-predicate-object) in frames based on spatio-temporal contextual information in the video. Previous work implicitly models the spatio-temporal interaction simultaneously, which leads to entanglement of spatio-temporal conte…

Cited by 2SourcePDFScholar
2023

Implicit Identity Driven Deepfake Face Swapping Detection

CVPR 2023poster

In this paper, we consider the face swapping detection from the perspective of face identity. Face swapping aims to replace the target face with the source face and generate the fake face that the human cannot distinguish between real and fake. We argue that the fake face contains the explicit ident…

Cited by 135SourcePDFScholar
2023

Improving Prosody for Cross-Speaker Style Transfer by Semi-Supervised Style Extractor and Hierarchical Modeling in Speech Synthesis

ICASSP 2023accepted

Cross-speaker style transfer in speech synthesis aims at transferring a style from source speaker to synthesized speech of a target speaker’s timbre. In most previous methods, the synthesized fine-grained prosody features often represent the source speaker’s average style, similar to the one-to-many…

Cited by 0SourceScholar
2023

LSTFE-Net:Long Short-Term Feature Enhancement Network for Video Small Object Detection

CVPR 2023poster

Video small object detection is a difficult task due to the lack of object information. Recent methods focus on adding more temporal information to obtain more potent high-level features, which often fail to specify the most vital information for small objects, resulting in insufficient or inappropr…

2023

Structure-Aware Multi-Feature Co-Learning for Dual Branch Face Super Resolution

ICASSP 2023accepted

Recently, face super-resolution has achieved pleasing performance. Numerous works have shown that texture features and structural information play a crucial role for super-resolution reconstruction. However, effective co-learning of both has been limiting the performance improvement of existing stat…

Cited by 0SourceScholar
2022

Adaptive Unsupervised Self-training for Disfluency Detection

COLING 2022main

Supervised methods have achieved remarkable results in disfluency detection. However, in real-world scenarios, human-annotated data is difficult to obtain. Recent works try to handle disfluency detection with unsupervised self-training, which can exploit existing large-scale unlabeled data efficient…

2022

DANet: Image Deraining via Dynamic Association Learning

IJCAI 2022poster

Rain streaks and background components in a rainy input are highly correlated, making the deraining task a composition of the rain streak removal and background restoration. However, the correlation of these two components is barely considered, leading to unsatisfied deraining results. To this end,…

Cited by 21SourcePDFScholar
2022

Degrade Is Upgrade: Learning Degradation for Low-Light Image Enhancement

AAAI 2022technical

Low-light image enhancement aims to improve an image's visibility while keeping its visual naturalness. Different from existing methods, which tend to accomplish the relighting task directly, we investigate the intrinsic degradation and relight the low-light image while refining the details and colo…

2022

Domain Generalization via Shuffled Style Assembly for Face Anti-Spoofing

CVPR 2022poster

With diverse presentation attacks emerging continually, generalizable face anti-spoofing (FAS) has drawn growing attention. Most existing methods implement domain generalization (DG) on the complete representations. However, different image statistics may have unique properties for the FAS tasks. In…

Cited by 203PDFcodeScholar
2022

EAD-Conformer: a Conformer-Based Encoder-Attention-Decoder-Network for Multi-Task Audio Source Separation

ICASSP 2022accepted

In this paper, we propose a Conformer-based network to improve the performance of multi-task audio source separation. This network, named EAD-Conformer, employs Conformer blocks to capture both local and global information, and an encoder-attention-decoder manner encourages the network to perform at…

Cited by 8SourceScholar
2022

ESimCSE: Enhanced Sample Building Method for Contrastive Learning of Unsupervised Sentence Embedding

COLING 2022main

Contrastive learning has been attracting much attention for learning unsupervised sentence embeddings. The current state-of-the-art unsupervised method is the unsupervised SimCSE (unsup-SimCSE). Unsup-SimCSE takes dropout as a minimal data augmentation method, and passes the same input sentence to a…

2022

InfoCSE: Information-aggregated Contrastive Learning of Sentence Embeddings

EMNLP 2022finding

Contrastive learning has been extensively studied in sentence embedding learning, which assumes that the embeddings of different views of the same sentence are closer. The constraint brought by this assumption is weak, and a good sentence representation should also be able to reconstruct the origina…

2022

RaP: Redundancy-aware Video-language Pre-training for Text-Video Retrieval

EMNLP 2022finding

Video language pre-training methods have mainly adopted sparse sampling techniques to alleviate the temporal redundancy of videos. Though effective, sparse sampling still suffers inter-modal redundancy: visual redundancy and textual redundancy. Compared with highly generalized text, sparsely sampled…

2022

Smoothed Contrastive Learning for Unsupervised Sentence Embedding

COLING 2022main

Unsupervised contrastive sentence embedding models, e.g., unsupervised SimCSE, use the InfoNCE loss function in training. Theoretically, we expect to use larger batches to get more adequate comparisons among samples and avoid overfitting. However, increasing batch size leads to performance degradati…

2021

A Triplet Appearance Parsing Network for Person Re-Identification

ICASSP 2021accepted

As one of the specific vision tasks, person re-identification has become a prevalent research topic in the field of multimedia and computer vision. However, existing feature extraction methods, originating from the quality of the bounding boxes which could cause the inhomogeneity and incoherence of…

Cited by 0SourceScholar
2021

Converse, Focus and Guess – Towards Multi-Document Driven Dialogue

AAAI 2021technical

We propose a novel task, Multi-Document Driven Dialogue (MD3), in which an agent can guess the target document that the user is interested in by leading a dialogue. To benchmark progress, we introduce a new dataset of GuessMovie, which contains 16,881 documents, each describing a movie, and associat…

2021

Dynamic Inconsistency-aware DeepFake Video Detection

IJCAI 2021poster

The spread of DeepFake videos causes a serious threat to information security, calling for effective detection methods to distinguish them. However, the performance of recent frame-based detection methods become limited due to their ignorance of the inter-frame inconsistency of fake videos. In this…

Cited by 0SourcePDFScholar
2021

Frequency-Aware Discriminative Feature Learning Supervised by Single-Center Loss for Face Forgery Detection

CVPR 2021poster

Face forgery detection is raising ever-increasing interest in computer vision since facial manipulation technologies cause serious worries. Though recent works have reached sound achievements, there are still unignorable problems: a) learned features supervised by softmax loss are separable but not…

Cited by 333PDFScholar
2021

HiT: Hierarchical Transformer With Momentum Contrast for Video-Text Retrieval

ICCV 2021poster

Video-Text Retrieval has been a hot research topic with the growth of multimedia data on the internet. Transformer for video-text learning has attracted increasing attention due to its promising performance. However, existing cross-modal transformer approaches typically suffer from two major limitat…

Cited by 191PDFScholar
2021

One-Shot Voice Conversion Based on Speaker Aware Module

ICASSP 2021accepted

Voice conversion (VC) is a task to convert the voice of speech while preserving its linguistic content. Although several methods have been proposed to enable VC with non-parallel data, it is still difficult to model the voice without a great number of data or an adaptive process. In this paper, we p…

Cited by 0SourceScholar
2021

When Face Recognition Meets Occlusion: A New Benchmark

ICASSP 2021accepted

The existing face recognition datasets usually lack occlusion samples, which hinders the development of face recognition. Especially during the COVID-19 coronavirus epidemic, wearing a mask has become an effective means of preventing the virus spread. Traditional CNN-based face recognition models tr…

Cited by 0SourceScholar
2020

Attention-Guided Deraining Network Via Stage-Wise Learning

ICASSP 2020accepted

Due to diverse rain shapes, directions, densities as well as different distances to cameras, rain streaks in the air are interweaved and overlapped. However, most existing deraining methods are inherently oblivious this phenomenon and tend to learn a single rain streak layer to simulate this complex…

Cited by 0SourceScholar
2020

Cartoon-Texture Decomposition-Based Variational Pansharpening

ICASSP 2020accepted

Pansharpening is widely used to increase the spatial resolution of a multispectral (MS) image by fusing with a panchromatic (PAN) image that has high-spatial resolution and the same scene. In this paper, the similarities of MS and PAN images in cartoon-texture space are exploited. The cartoon and te…

Cited by 0SourceScholar
2020

Fusionndvi: A Novel Fusion Method for NDVI in Remote Sensing

ICASSP 2020accepted

Normalized difference vegetation index (NDVI) is widely utilized to examine vegetation coverage and estimate crop yield. To obtain a high-resolution (HR) NDVI, fusion techniques, which first generates a HR multispectral (MS) image by fusing a low-resolution (LR) MS image and a HR panchromatic image,…

Cited by 0SourceScholar
2020

Image Super-Resolution Using Residual Global Context Network

ICASSP 2020accepted

Recent studies have showed that convolutional neural networks (CNN) can effectively improve the performance of single image super-resolution (SR). However, previous methods rarely considered long-range dependencies between pixels and channel-wise interdependencies at the same time. They ignores the…

Cited by 0SourceScholar
2020

Learn with Noisy Data via Unsupervised Loss Correction for Weakly Supervised Reading Comprehension

COLING 2020main

Weakly supervised machine reading comprehension (MRC) task is practical and promising for its easily available and massive training data, but inevitablely introduces noise. Existing related methods usually incorporate extra submodels to help filter noise before the noisy data is input to main models…

Cited by 5SourcePDFScholar
2020

Multi-Scale Progressive Fusion Network for Single Image Deraining

CVPR 2020poster

Rain streaks in the air appear in various blurring degrees and resolutions due to different distances from their positions to the camera. Similar rain patterns are visible in a rain image as well as its multi-scale (or multi-resolution) versions, which makes it possible to exploit such complementary…

Cited by 844PDFcodeScholar
2020

Syntactic Graph Convolutional Network for Spoken Language Understanding

COLING 2020main

Slot filling and intent detection are two major tasks for spoken language understanding. In most existing work, these two tasks are built as joint models with multi-task learning with no consideration of prior linguistic knowledge. In this paper, we propose a novel joint model that applies a graph c…

Cited by 10SourcePDFScholar
2019

Blind Quality Assessment for 3D-synthesized Images by Measuring Geometric Distortions and Image Complexity

ICASSP 2019accepted

Free viewpoint video (FVV), owing to its comprehensive applications in immersive entertainment, remote surveillance and distanced education, has received extensive attention and been regarded as a new important direction of video technology development. Depth image-based rendering (DIBR) technologie…

Cited by 0SourceScholar
2019

Long Term Background Reference Based Satellite Video Coding

ICASSP 2019accepted

Video transmission from satellites to terrestrial devices usually requires a large amount of channel resources due to the huge amount of satellite video data. Subject to limited transmission bandwidth in space environment, the video encoder for video satellite calls for higher coding efficiency. In…

Cited by 0SourceScholar
2019

Multisource Surveillance Video Coding by Exploiting 3D and 2D Knolwedge

ICASSP 2019accepted

The rapidly increasing surveillance video data has challenged the existing video coding standards. Even though knowledge based video coding scheme proposed for moving objects so far has achieved high efficiency, it does not take full advantages of local information and highly relies on the accuracy…

Cited by 0SourceScholar
2019

Progressive Fusion Video Super-Resolution Network via Exploiting Non-Local Spatio-Temporal Correlations

ICCV 2019oral

Most previous fusion strategies either fail to fully utilize temporal information or cost too much time, and how to effectively fuse temporal information from consecutive frames plays an important role in video super-resolution (SR). In this study, we propose a novel progressive fusion network for v…

Cited by 337PDFcodeScholar
2019

Rain Streak Removal via Multi-scale Mixture Exponential Power Model

ICASSP 2019accepted

Rain streaks severely hamper the visible performance of the outdoor surveillance videos, which becomes an attractive issue in recent computer vision research. Existing methods usually encode rain streaks into Gaussian Mixture Model (GM-M). However, the limited number of Gaussian components in the GM…

Cited by 0SourceScholar
2018

Improving Convolutional Neural Networks Via Compacting Features

ICASSP 2018accepted

Convolutional neural networks (CNNs) have shown great advantages in computer vision fields, and loss functions are of great significance to their gradient descent algorithms. Softmax loss, a combination of cross-entropy loss and Softmax function, is the most commonly used one for CNNs. Hence, it can…

Cited by 0SourceScholar
2017

A joint learning based Face Super Resolution approach via contextual topological structure

ICASSP 2017accepted

Face Super Resolution(FSR) is to infer High Resolution(HR) facial images from given Low Resolution(LR) ones with the assistance of LR and HR training pairs. Among existing methods, local patch based methods are superior in visual and objective quality than global based methods. These local patch bas…

Cited by 0SourceScholar
2016

L1-L1 norms for face super-resolution with mixed Gaussian-impulse noise

ICASSP 2016accepted

In real world surveillance application, the captured faces are often low resolution (LR) and corrupted by mixed Gaussian-impulse noise during the acquisition and transmission processes. In this paper, we propose an effective patch-based face super-resolution method to reconstruct a high resolution (…

Cited by 0SourceScholar
2015

Face hallucination via Cauchy regularized sparse representation

ICASSP 2015accepted

In dictionary-learning-based face hallucination, the testing image is represented as a linear combination of the training samples, and how to obtain the optimal coefficients is the primary issue. Sparse representation (SR) has ever been widely used in face hallucination, however, due to the fact tha…

Cited by 0SourceScholar