← Search

Bowen Zhang

77 accepted papers

2026

GauSem-SLAM: Gaussian Semantic Submaps with Loop Closure for Globally Consistent SLAM

ICRA 2026poster

3DGS has shown outstanding performance in multi-view geometry, driving its adoption in visual SLAM. However, real-time semantic 3DGS mapping faces challenges. Current methods typically treat semantics as external priors, making it hard to integrate them into SLAM tracking or loop closure correction.…

Cited by 0Scholar
2026

Improving Day-Ahead Grid Carbon Intensity Forecasting by Joint Modeling of Local-Temporal and Cross-Variable Dependencies Across Different Frequencies

AAAI 2026technical

Accurate forecasting of the grid carbon intensity factor (CIF) is critical for enabling demand-side management and reducing emissions in modern electricity systems. Leveraging multiple interrelated time series, CIF prediction is typically formulated as a multivariate time series forecasting problem.

Cited by 0SourcePDFScholar
2026

Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning

AAAI 2026technical

Zero-shot stance detection (ZSSD) seeks to determine the stance of text toward previously unseen targets, a task critical for analyzing dynamic and polarized online discourse with limited labeled data. While large language models (LLMs) offer zero-shot capabilities, prompting-based approaches often

Cited by 0SourcePDFScholar
2026

MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer

ICLR 2026poster

Unified multimodal Large Language Models (LLMs) that can both understand and generate visual content hold immense potential. However, existing open-source models often suffer from a performance trade-off between these capabilities. We present Manzano, a simple and scalable unified framework that sub…

Cited by 0SourceScholar
2026

MoETTA: Test-Time Adaptation Under Mixed Distribution Shifts with MoE-LayerNorm

AAAI 2026technical

Test-time adaptation (TTA) has proven effective in mitigating performance drops under single-domain distribution shifts by updating model parameters during inference. However, real-world deployments often involve mixed distribution shifts---where test samples are affected by diverse and potentially

Cited by 0SourcePDFScholar
2026

OccDriver: Future Occupancy Guided Dual-branch Trajectory Planner in Autonomous Driving

ICLR 2026poster

Trajectory planning for autonomous driving is challenging due to agents' behavioral uncertainty and intricate multi-agent interaction modeling. Most existing studies generate trajectories without explicitly exploiting possible scene evolution, while world models predict consequences from ego behavio…

Cited by 0SourceScholar
2026

PBR3DGen: A VLM-Guided Mesh Generation with High-Quality PBR Texture

AAAI 2026technical

Generating high-quality physically based rendering (PBR) materials is important to achieve realistic rendering in the downstream tasks, yet it remains challenging due to the intertwined effects of materials and lighting. While existing methods have made breakthroughs by incorporating material decomp

Cited by 0SourcePDFScholar
2026

PoLi-RL: A Point-to-List Reinforcement Learning Framework for Conditional Semantic Textual Similarity

ICLR 2026poster

Conditional Semantic Textual Similarity (C-STS) measures the semantic proximity between text segments under a specific condition, thereby overcoming the ambiguity inherent in traditional STS. However, existing methods are largely confined to discriminative models, failing to fully integrate recent b…

Cited by 0SourcecodeScholar
2026

Resisting Label Drift: Real-Time Multi-View Clustering with Semantic Consistency

IJCAI 2026

Real-time clustering of dynamic multi-view data streams is a critical yet challenging task in open-world applications. While several methods have been proposed to address this task, most of them extract features incrementally but fail to output instant clustering results for the current batch. In ad

Cited by 0Scholar
2026

Sat2Flow: A Structure-Aware Diffusion Framework for Human Flow Generation from Satellite Imagery

AAAI 2026technical

Origin-Destination (OD) flow matrices are critical for urban mobility analysis, supporting traffic forecasting, infrastructure planning, and policy design. Existing methods face two key limitations: (1) reliance on costly auxiliary features (e.g., Points of Interest, socioeconomic statistics) with l

Cited by 0SourcePDFScholar
2026

Spiking-Aided Neural Architecture for Efficient and Robust WiFi Sensing

AAAI 2026technical

This paper introduces a spiking-aided wifi sensing network (SWS-Net), a novel hybrid neural architecture that integrates Spiking Neural Networks (SNNs) with conventional Artificial Neural Networks (ANNs) for robust WiFi-based indoor sensing. WiFi signals offer a low-cost and device-free solution for

Cited by 0SourcePDFScholar
2026

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

CVPR 2026

The video reasoning ability of multimodal large language models (MLLMs) is crucial for downstream tasks like video question answering and temporal grounding. While recent approaches have explored text-based chain-of-thought (CoT) reasoning for MLLMs, these methods often suffer from limited cross-mod

Cited by 0SourcecodeScholar
2026

WiFi-GEN: High-Resolution Indoor Imaging from WiFi Signals Using Generative AI

ICASSP 2026poster

Indoor imaging is a critical task for robotics and internet-ofthings. WiFi as an omnipresent signal is a promising candidate for carrying out passive imaging and synchronizing the up-to-date information to all connected devices. This is the first research work to consider WiFi indoor imaging as a mu…

Cited by 0SourcePDFScholar
2025

Adapting to Observation Length of Trajectory Prediction via Contrastive Learning

CVPR 2025poster

The ability to adapt to varying observation lengths is crucial for human trajectory prediction tasks, particularly in scenarios with limited observation lengths or missing data. Existing approaches mainly focus on introducing novel architectures or additional structural components, which substantial…

2025

CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling

EMNLP 2025

Mixture-of-Experts (MoE) models are crucial for scaling model capacity while controlling inference costs. While integrating MoE into multimodal models like CLIP improves performance, training these models is notoriously challenging and expensive. We propose CLIP-Upcycling (CLIP-UP), an efficient alt

Cited by 0SourcePDFScholar
2025

Contrastive Localized Language-Image Pre-Training

ICML 2025poster

CLIP has been a celebrated method for training vision encoders to generate image/text representations facilitating various applications. Recently, it has been widely adopted as the vision backbone of multimodal large language models (MLLMs). The success of CLIP relies on aligning web-crawled noisy t…

Cited by 10SourcePDFScholar
2025

Core Knowledge Learning Framework for Graph

AAAI 2025technical

Graph classification is a pivotal challenge in machine learning, especially within the realm of graph-based data, given its importance in numerous real-world applications such as social network analysis, recommendation systems, and bioinformatics. Despite its significance, graph classification faces…

Cited by 0SourcePDFScholar
2025

EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice Routing

ICLR 2025poster

Diffusion transformers have been widely adopted for text-to-image synthesis. While scaling these models up to billions of parameters shows promise, the effectiveness of scaling beyond current sizes remains underexplored and challenging. By explicitly exploiting the computational heterogeneity of ima…

Cited by 0SourcePDFScholar
2025

GaussR-SLAM: Gaussian Robust SLAM in Data Loss and Interference Environments

RA-L 2025

Recent advancements in 3DGS-based explicit mapping have significantly improved SLAM performance, achieving more realistic environment reconstruction and faster processing. However, issues such as data loss caused by unstable data transmission, textureless and repetitive-texture often occur in real-w

Cited by 0SourceScholar
2025

Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis

ICCV 2025poster

In this paper, we present a novel framework for video-to-4D generation that creates high-quality dynamic 3D content from single video inputs. Direct 4D diffusion modeling is extremely challenging due to costly data construction and the high-dimensional nature of jointly representing 3D shape, appear…

2025

Hi-Patch: Hierarchical Patch GNN for Irregular Multivariate Time Series

ICML 2025poster

Multi-scale information is crucial for multivariate time series modeling. However, most existing time series multi-scale analysis methods treat all variables in the same manner, making them unsuitable for Irregular Multivariate Time Series (IMTS), where variables have distinct origin scales/sampling…

Cited by 0SourcePDFScholar
2025

Improve Vision Language Model Chain-of-thought Reasoning

ACL 2025long

Chain-of-thought (CoT) reasoning in vision language models (VLMs) is crucial for improving interpretability and trustworthiness. However, current training recipes often relying on datasets dominated by short annotations with minimal rationales. In this work, we show that training VLM on short answer…

2025

Kangaroo Tail-Inspired Variable Stiffness Executive Mechanism for Rescue Robots

RA-L 2025

In the field of casualty extraction and rescue, a key challenge is avoiding secondary injuries to the human body during rescue operations. Currently, most rescue robot executive mechanisms are rigid, which increases the risk of contact-related injuries. To address this issue, a rescue executive mech

Cited by 3SourceScholar
2025

MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

ICLR 2025poster

We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture, MM1.5 adopts a data-centric approach to model training, systema…

Cited by 29SourcePDFScholar
2025

MMEgo: Towards Building Egocentric Multimodal LLMs for Video QA

ICLR 2025poster

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as there is a lack of QA data for egocentric video understanding, we automatically generate 7M high-quality QA samples for e…

Cited by 0SourcePDFScholar
2025

MuSCLe-Reg: Multi-Scale Contextual Embedding and Local Correspondence Rectification for Robust Two-Stage Point Cloud Registration

RA-L 2025

Algorithm of outlier removal for learning-based 3D point cloud registration is usually regarded as a classification problem. The core for this to be successful is to learn the discriminative inlier/outlier feature representations. This letter proposes a two-stage efficient network (MuSCLe-Reg) with

Cited by 0SourceScholar
2025

Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

ICLR 2025poster

Recent advancements in multimodal models highlight the value of rewritten captions for improving performance, yet key challenges remain. For example, while synthetic captions often provide superior quality and image-text alignment, it is not clear whether they can fully replace AltTexts: the role of…

Cited by 4SourcePDFScholar
2025

SPARK: Simulating the Co-evolution of Stance and Topic Dynamics in Online Discourse with LLM-based Agents

EMNLP 2025

Topic evolution and stance dynamics are deeply intertwined in online social media, shaping the fragmentation and polarization of public discourse. Yet existing dynamic topic models and stance analysis approaches usually consider these processes in isolation, relying on abstractions that lack interpr

2025

STIV: Scalable Text and Image Conditioned Video Generation

ICCV 2025poster

We present a simple and scalable text and image conditioned video generation method. Our approach, named STIV, integrates a variable number of image conditions into a Diffusion Transformer (DiT) through frame replacement. This design enables STIV to perform both text-to-video (T2V) and text-image-to…

2025

Semantics-Guided Dynamic Hypergraph Network for Human Mobility Nowcasting in Disaster

ICASSP 2025accepted

Human mobility nowcasting is crucial for public safety, especially during disasters when human mobility significantly differs from normal patterns, posing unique challenges. Recent studies have shown a correlation between disaster-related social media information and abnormal patterns in human mobil…

Cited by 0SourceScholar
2025

SemiGPS: GraphGPS-based Semi-supervised Graph Learning for Sector-Specific GDP Mapping

ICASSP 2025accepted

Accurate forecasting of granular socioeconomic indicators, such as GDP, is essential for informed economic decision-making. Despite pioneering efforts to harness multi-modal data using traditional supervised or self-supervised learning methods for economic prediction, effectively integrating their f…

Cited by 0SourceScholar
2025

Structured 3D Latents for Scalable and Versatile 3D Generation

CVPR 2025highlight

We introduce a novel 3D generation method for versatile and high-quality 3D asset creation.The cornerstone is a unified Structured LATent (SLAT) representation which allows decoding to different output formats, such as Radiance Fields, 3D Gaussians, and meshes. This is achieved by integrating a spar…

2025

Zero-shot Stance Detection with Logically Consistent Data Augmentation

ICASSP 2025accepted

Zero-shot stance detection (ZSSD) is a challenging task that requires classifying stances towards unseen targets without large, well-curated training datasets. Existing data augmentation methods for ZSSD often suffer from semantic inconsistencies, hindering their effectiveness. To address these limi…

Cited by 0SourceScholar
2024

"MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training"

ECCV 2024poster

"In this work, we discuss building performant Multimodal Large Language Models (MLLMs). In particular, we study the importance of various architecture components and data choices. Through careful and comprehensive ablations of the image encoder, the vision language connector, and various pre-trainin…

2024

A Challenge Dataset and Effective Models for Conversational Stance Detection

COLING 2024main

Previous stance detection studies typically concentrate on evaluating stances within individual instances, thereby exhibiting limitations in effectively modeling multi-party discussions concerning the same specific topic, as naturally transpire in authentic social media interactions. This constraint…

2024

Advancing Semantic Textual Similarity Modeling: A Regression Framework with Translated ReLU and Smooth K2 Loss

EMNLP 2024main

Since the introduction of BERT and RoBERTa, research on Semantic Textual Similarity (STS) has made groundbreaking progress. Particularly, the adoption of contrastive learning has substantially elevated state-of-the-art performance across various STS benchmarks. However, contrastive learning categori…

2024

Amodal Scene Analysis via Holistic Occlusion Relation Inference and Generative Mask Completion

AAAI 2024technical

Amodal scene analysis entails interpreting the occlusion relationship among scene elements and inferring the possible shapes of the invisible parts. Existing methods typically frame this task as an extended instance segmentation or a pair-wise object de-occlusion problem. In this work, we propose a…

2024

Compressing LLMs: The Truth is Rarely Pure and Never Simple

ICLR 2024poster

Despite their remarkable achievements, modern Large Language Models (LLMs) encounter exorbitant computational and memory footprints. Recently, several works have shown significant success in *training-free* and *data-free* compression (pruning and quantization) of LLMs achieving 50-60\% sparsity an…

2024

Cross-Target Stance Detection by Exploiting Target Analytical Perspectives

ICASSP 2024accepted

Cross-target stance detection (CTSD) is an important task, which infers the attitude of the destination target by utilizing annotated data derived from the source target. One important approach in CTSD is to extract domain-invariant features to bridge the knowledge gap between multiple targets. Howe…

Cited by 0SourceScholar
2024

Decomposing Disease Descriptions for Enhanced Pathology Detection: A Multi-Aspect Vision-Language Pre-training Framework

CVPR 2024poster

Medical vision language pre-training (VLP) has emerged as a frontier of research enabling zero-shot pathological recognition by comparing the query image with the textual descriptions for each disease. Due to the complex semantics of biomedical texts current methods struggle to align medical images…

2024

EDDA: An Encoder-Decoder Data Augmentation Framework for Zero-Shot Stance Detection

COLING 2024main

Stance detection aims to determine the attitude expressed in text towards a given target. Zero-shot stance detection (ZSSD) has emerged to classify stances towards unseen targets during inference. Recent data augmentation techniques for ZSSD increase transferable knowledge between targets through te…

2024

Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction

EMNLP 2024main

In this work, we are interested in automated methods for knowledge graph creation (KGC) from input text. Progress on large language models (LLMs) has prompted a series of recent works applying them to KGC, e.g., via zero/few-shot prompting. Despite successes on small domain-specific datasets, these…

2024

Ferret: Refer and Ground Anything Anywhere at Any Granularity

ICLR 2024spotlight

We introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding open-vocabulary descriptions. To unify referring and grounding in the LLM paradigm, Ferret employs a novel and powerful hy…

2024

GaussianCube: A Structured and Explicit Radiance Representation for 3D Generative Modeling

NeurIPS 2024poster

We introduce a radiance representation that is both structured and fully explicit and thus greatly facilitates 3D generative modeling. Existing radiance representations either require an implicit feature decoder, which significantly degrades the modeling power of the representation, or are spatially…

Cited by 9SourcePDFScholar
2024

MOFI: Learning Image Representations from Noisy Entity Annotated Images

ICLR 2024poster

We present MOFI, Manifold OF Images, a new vision foundation model designed to learn image representations from noisy entity annotated images. MOFI differs from previous work in two key aspects: 1. pre-training data, and 2. training recipe. Regarding data, we introduce a new approach to automaticall…

2024

MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval

CVPR 2024poster

State-of-the-art video-text retrieval (VTR) methods typically involve fully fine-tuning a pre-trained model (e.g. CLIP) on specific datasets. However this can result in significant storage costs in practical applications as a separate model per task must be stored. To address this issue we present o…

2024

MetaDiff: Meta-Learning with Conditional Diffusion for Few-Shot Learning

AAAI 2024technical

Equipping a deep model the ability of few-shot learning (FSL) is a core challenge for artificial intelligence. Gradient-based meta-learning effectively addresses the challenge by learning how to learn novel tasks. Its key idea is learning a deep model in a bi-level optimization manner, where the out…

Cited by 53SourcePDFScholar
2024

MoZIP: A Multilingual Benchmark to Evaluate Large Language Models in Intellectual Property

COLING 2024main

Large language models (LLMs) have demonstrated impressive performance in various natural language processing (NLP) tasks. However, there is limited understanding of how well LLMs perform in specific domains (e.g, the intellectual property (IP) domain). In this paper, we contribute a new benchmark, t…

2024

Pcc-tuning: Breaking the Contrastive Learning Ceiling in Semantic Textual Similarity

EMNLP 2024main

Semantic Textual Similarity (STS) constitutes a critical research direction in computational linguistics and serves as a key indicator of the encoding capabilities of embedding models. Driven by advances in pre-trained language models and contrastive learning, leading sentence representation methods…

2024

RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models

ECCV 2024poster

"We present RodinHD, which can generate high-fidelity 3D avatars from a portrait image. Existing methods fail to capture intricate details such as hairstyles which we tackle in this paper. We first identify an overlooked problem of catastrophic forgetting that arises when fitting triplanes sequentia…

2024

VeCLIP: Improving CLIP Training via Visual-enriched Captions

ECCV 2024poster

"Large-scale web-crawled datasets are fundamental for the success of pre-training vision-language models, such as CLIP. However, the inherent noise and potential irrelevance of web-crawled AltTexts pose challenges in achieving precise image-text alignment. Existing methods utilizing large language m…

2024

Weak-eval-Strong: Evaluating and Eliciting Lateral Thinking of LLMs with Situation Puzzles

NeurIPS 2024poster

While advancements in NLP have significantly improved the performance of Large Language Models (LLMs) on tasks requiring vertical thinking, their lateral thinking capabilities remain under-explored and challenging to measure due to the complexity of assessing creative thought processes and the scarc…

2024

WildlifeMapper: Aerial Image Analysis for Multi-Species Detection and Identification

CVPR 2024poster

We introduce WildlifeMapper (WM) a flexible model designed to detect locate and identify multiple species in aerial imagery. It addresses the limitations of traditional labor-intensive wildlife population assessments that are central to advancing environmental conservation efforts worldwide. While a…

2024

iTrendRNN: An Interpretable Trend-Aware RNN for Meteorological Spatiotemporal Prediction

AAAI 2024technical

Accurate prediction of meteorological elements, such as temperature and relative humidity, is important to human livelihood, early warning of extreme weather, and urban governance. Recently, neural network-based methods have shown impressive performance in this field. However, most of them are overc…

2023

Dynamic Token Pruning in Plain Vision Transformers for Semantic Segmentation

ICCV 2023poster

Vision transformers have achieved leading performance on various visual tasks yet still suffer from high computational complexity. The situation deteriorates in dense prediction tasks like semantic segmentation, as high-resolution inputs and outputs usually imply more tokens involved in computations…

Cited by 28PDFcodeScholar
2023

Int-GNN: A User Intention Aware Graph Neural Network for Session-Based Recommendation

ICASSP 2023accepted

Session-Based Recommendation (SBR) is a spotlight research problem. Although many efforts have been made, challenges still exist. The key to unlocking this shackle is the user intention, an intuitive but hard-to-model concept in the anonymous session. Unlike previous research, we suggest mining pote…

Cited by 0SourceScholar
2023

Knowledge-Aware Few Shot Learning for Event Detection from Short Texts

ICASSP 2023accepted

Event detection in a city is crucial for the government to listen to the voice of the citizens, be aware of the real occurrences in a city, and then make wiser policies. However, in reality some important events with few samples are easily to be overwhelmed by the massive information and hard to be…

Cited by 0SourceScholar
2023

MetaPortrait: Identity-Preserving Talking Head Generation With Fast Personalized Adaptation

CVPR 2023poster

In this work, we propose an ID-preserving talking head generation framework, which advances previous methods in two aspects. First, as opposed to interpolating from sparse flow, we claim that dense landmarks are crucial to achieving accurate geometry-aware flow fields. Second, inspired by face-swapp…

2023

STAIR: Learning Sparse Text and Image Representation in Grounded Tokens

EMNLP 2023long main

Image and text retrieval is one of the foundational tasks in the vision and language domain with multiple real-world applications. State-of-the-art contrastive approaches, e.g. CLIP, ALIGN, represent images and texts as dense embeddings and calculate the similarity in the dense embedding space as th…

Cited by 0SourceScholar
2023

Stance Detection on Social Media with Background Knowledge

EMNLP 2023long main

Identifying users' stances regarding specific targets/topics is a significant route to learning public opinion from social media platforms. Most existing studies of stance detection strive to learn stance information about specific targets from the context, in order to determine the user's stance on…

Cited by 0SourceScholar
2023

Twitter Stance Detection via Neural Production Systems

ICASSP 2023accepted

Stance detection is an important task, which aims to classify the attitude of an opinionated text toward a given target. In this paper, we develop an interpretable neural production system for stance detection (NPS4SD). NPS4SD is an end-to-end deep learning model, which consists of a set of knowledg…

Cited by 0SourceScholar
2023

ZegCLIP: Towards Adapting CLIP for Zero-Shot Semantic Segmentation

CVPR 2023poster

Recently, CLIP has been applied to pixel-level zero-shot learning tasks via a wo-stage scheme. The general idea is to first generate class-agnostic region proposals and then feed the cropped proposal regions to CLIP to utilize its image-level zero-shot classification capability. While effective, suc…

2022

Noise Learning for Text Classification: A Benchmark

COLING 2022main

Noise Learning is important in the task of text classification which depends on massive labeled data that could be error-prone. However, we find that noise learning in text classification is relatively underdeveloped: 1. many methods that have been proven effective in the image domain are not explor…

Cited by 11SourcePDFScholar
2022

SegViT: Semantic Segmentation with Plain Vision Transformers

NeurIPS 2022accept

We explore the capability of plain Vision Transformers (ViTs) for semantic segmentation and propose the SegViT. Previous ViT-based segmentation networks usually learn a pixel-level representation from the output of the ViT. Differently, we make use of the fundamental component—attention mechanism, t…

2022

Sentiment Interpretable Logic Tensor Network for Aspect-Term Sentiment Analysis

COLING 2022main

Aspect-term sentiment analysis (ATSA) is an important task that aims to infer the sentiment towards the given aspect-terms. It is often required in the industry that ATSA should be performed with interpretability, computational efficiency and high accuracy. However, such an ATSA method has not yet b…

Cited by 18SourcePDFScholar
2022

StyleSwin: Transformer-Based GAN for High-Resolution Image Generation

CVPR 2022poster

Despite the tantalizing success in a broad of vision tasks, transformers have not yet demonstrated on-par ability as ConvNets in high-resolution image generative modeling. In this paper, we seek to explore using pure transformers to build a generative adversarial network for high-resolution image sy…

Cited by 318PDFcodeScholar
2021

Dynamic Neural Representational Decoders for High-Resolution Semantic Segmentation

NeurIPS 2021poster

Semantic segmentation requires per-pixel prediction for a given image. Typically, the output resolution of a segmentation network is severely reduced due to the downsampling operations in the CNN backbone. Most previous methods employ upsampling decoders to recover the spatial resolution. Various de…

Cited by 16SourcePDFScholar
2021

FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling

NeurIPS 2021poster

The recently proposed FixMatch achieved state-of-the-art results on most semi-supervised learning (SSL) benchmarks. However, like other modern SSL algorithms, FixMatch uses a pre-defined constant threshold for all classes to select unlabeled data that contribute to the training, thus failing to cons…

2021

Systematic Generalization on gSCAN: What is Nearly Solved and What is Next?

EMNLP 2021main

We analyze the grounded SCAN (gSCAN) benchmark, which was recently proposed to study systematic generalization for grounded language understanding. First, we study which aspects of the original benchmark can be solved by commonly used methods in multi-modal research. We find that a general-purpose T…

2020

Effectiveness of Random Deep Feature Selection for Securing Image Manipulation Detectors Against Adversarial Examples

ICASSP 2020accepted

We investigate if the random feature selection approach proposed in [1] to improve the robustness of forensic detectors to targeted attacks, can be extended to detectors based on deep learning features. In particular, we study the transferability of adversarial examples targeting an original CNN ima…

Cited by 0SourceScholar
2016

Real-Time Action Recognition With Enhanced Motion Vector CNNs

CVPR 2016poster

The deep two-stream architecture exhibited excellent performance on video based action recognition. The most computationally expensive step in this approach comes from the calculation of optical flow which prevents it to be real-time. This paper accelerates this architecture by replacing optical flo…

Cited by 546PDFcodeScholar