← Search

Zhen Yang

79 accepted papers

2026

A Local-Rotation-Driven Global Consistency Framework with Dual-View Decoding for Semi-Supervised Medical Image Segmentation

IJCAI 2026

Medical image segmentation is still challenging, especially in semi-supervised scenarios where limited annotations are expected to support both accurate boundary delineation and coherent anatomical structures. We propose LR-GCF, a Local-Rotation-Driven Global Consistency Framework that couples stron

Cited by 0Scholar
2026

CCAHCL: Multi-Level Hypergraph Contrastive Learning for Connected Component Awareness

AAAI 2026technical

Hypergraph contrastive learning has emerged as a powerful unsupervised paradigm for hypergraph representation learning. Traditional hypergraph contrastive learning methods typically leverage neighbor aggregation strategy to obtain entity (node and hyperedge) representations within each connected com

Cited by 0SourcePDFScholar
2026

Cooperative Multi-View Graph Learning via High-Rank Tensor Specificity

IJCAI 2026

Graph-based multi-view clustering, with its ability to mine potential associations between samples, has attracted extensive attention. To capture high-order correlations, tensor-based frameworks have been introduced to model multiple graphs jointly. Although these methods have achieved promising per

Cited by 0Scholar
2026

CorrectManip: A Data-Driven Closed-Loop Framework for Autonomous Skill Learning with Failure Recovery

ICRA 2026poster

Simulation-based training offers an efficient paradigm for robotic skill learning, providing scalable data generation while reducing reliance on costly hardware trials and manual data collection. However, existing methods that rely on handcrafted scenarios fail to fully cover the complexity of open-…

Cited by 0Scholar
2026

DF^2-VB: Dual-level Fuzzy Fusion with View-specific Boosting for Multi-view Multi-label Classification

CVPR 2026

Multi-view multi-label classification (MVMLC) aims to utilize both consensus and complementarity information to predict potentially relevant labels for samples. Existing MVMLC approaches typically focus on either feature-level fusion, which integrates complementary features for more expressive repre

Cited by 0SourceScholar
2026

Event-Based Motion Deblurring Using Task-Oriented 3D Gaussian Event Representations

CVPR 2026

Event-based motion deblurring has attracted increasing attention, as the high temporal resolution of event cameras provides motion cues unavailable to conventional RGB sensors, thereby enabling more effective deblurring. In real-world scenes, motion blur is often complex and nonlinear, with differen

Cited by 0SourceScholar
2026

Hypergraph-Based Multi-View Multi-Label Classification via Adaptive High-Order Semantic Fusion

AAAI 2026technical

In multi-view multi-label (MVML) classification, each sample is represented by multiple heterogeneous views and annotated with multiple labels. Existing methods typically exploit pairwise semantic relationships to mine intra-view correlations and align inter-view features for generating structural r

Cited by 0SourcePDFScholar
2026

Learning Structured Reasoning via Tractable Trajectory Control

ICML 2026spotlight

Large language models can exhibit emergent reasoning behaviors, often manifested as recurring lexical patterns (e.g., “wait,” indicating verification). However, complex reasoning trajectories remain sparse in unconstrained sampling, and standard RL often fails to guarantee the acquisition of diverse…

Cited by 0SourceScholar
2026

MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning

AAAI 2026technical

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in vision-language answering tasks. Despite their strengths, these models often encounter challenges in achieving complex reasoning tasks such as mathematical problem-solving. Previous works have focused on fine-tunin

Cited by 0SourcePDFScholar
2026

One-Shot Flow, Any-Time Frame: A Bidirectional Warping Framework for Event-Based Video Frame Interpolation

CVPR 2026

Video Frame Interpolation (VFI) is a crucial task in video processing. Flow-based methods, despite their success, are constrained by a fundamental dilemma: forward warping is efficient but prone to artifacts, while backward warping yields higher quality at a significant computational cost, especiall

Cited by 0SourcecodeScholar
2026

Selective Actuation for Microrobots Based on Distributed Magnetic Field Design

ICRA 2026poster

Mechanical stimulation is essential for regulating cellular processes such as proliferation, differentiation, and apoptosis. Magnetic microrobot swarms offer a promising platform for delivering targeted mechanical stimulation to cells via remote actuation under rotating magnetic fields. However, mag…

Cited by 0Scholar
2026

TAMMs: Change Understanding and Forecasting in Satellite Image Time Series with Temporal-Aware Multimodal Models

ICLR 2026poster

Temporal Change Description (TCD) and Future Satellite Image Forecasting (FSIF) are critical, yet historically disjointed tasks in Satellite Image Time Series (SITS) analysis. Both are fundamentally limited by the common challenge of modeling long-range temporal dynamics. To explore how to improve t…

Cited by 0SourceScholar
2026

TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language Model

AAAI 2026technical

Transformers are the cornerstone of modern large language models, but their quadratic computational complexity limits efficiency in long-sequence processing. Recent advancements in Mamba, a state space model (SSM) with linear complexity, offer promising efficiency gains but suffer from unstable cont

Cited by 0SourcePDFScholar
2026

UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization

ICML 2026poster

UI-to-code aims to translate UI screenshots into executable front-end code. Despite progress with vision-language models (VLMs), most existing methods formulate UI-to-code as a single-pass generation, which mismatches real-world UI development that is inherently iterative and feedback-driven. We ref…

Cited by 0SourceScholar
2026

Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability

CVPR 2026

Large language models (LLMs) often generate self-contradictory outputs, which severely impacts their reliability and hinders their adoption in practical applications. In video-language models (Video-LLMs), this phenomenon recently draws the attention of researchers. Specifically, these models fail t

Cited by 0SourceScholar
2026

VisionWebDev: A Hierarchical Benchmark for Visual Website Development with Agent Verification

ICML 2026spotlight

Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, end-to-end website development remains limited. To address this gap, we introduce \benchname{}, a hierarchical benchmark for visual website development, spanning from stati…

Cited by 0SourceScholar
2026

WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection

ICLR 2026poster

Search agents have achieved significant advancements in enabling intelligent information retrieval and decision-making within interactive environments. Although reinforcement learning has been employed to train agentic models capable of more dynamic interactive retrieval, existing methods are limite…

Cited by 0SourcecodeScholar
2025

AF-UMC: An Alignment-Free Fusion Framework for Unaligned Multi-View Clustering

NeurIPS 2025poster

The Unaligned Multi-view Clustering (UMC) aims to learn a discriminative cluster structure from unaligned multi-view data, where the features of samples are not completely aligned across multiple views. Most existing methods usually prioritize employing various alignment strategies to align sample r…

Cited by 0SourceScholar
2025

CFDM: Contrastive Fusion and Disambiguation for Multi-View Partial-Label Learning

AAAI 2025technical

When dealing with multi-view data, the heterogeneity of data attributes across different views often leads to label ambiguity. To effectively address this challenge, this paper designs a Multi-View Partial-Label Learning (MVPLL) framework, where each training instance is described by multiple view f…

Cited by 0SourcePDFScholar
2025

CaliGCL: Calibrated Graph Contrastive Learning via Partitioned Similarity and Consistency Discrimination

NeurIPS 2025poster

Graph contrastive learning (GCL) aims to learn self-supervised representations by distinguishing positive and negative sample pairs generated from multiple augmented graph views. Despite showing promising performance, GCL still suffers from two critical biases: (1) ***Similarity estimation bias*** a…

Cited by 0SourceScholar
2025

Causality Meets the Table: Debiasing LLMs for Faithful TableQA via Front-Door Intervention

NeurIPS 2025poster

Table Question Answering (TableQA) combines natural language understanding and structured data reasoning, posing challenges in semantic interpretation and logical inference. Recent advances in Large Language Models (LLMs) have improved TableQA performance through Direct Prompting and Agent paradigms…

Cited by 0SourceScholar
2025

Constructing Your Model’s Value Distinction: Towards LLM Alignment with Anchor Words Tuning

EMNLP 2025

With the widespread applications of large language models (LLMs), aligning LLMs with human values has emerged as a critical challenge. For alignment, we always expect LLMs to be honest, positive, harmless, etc. And LLMs appear to be capable of generating the desired outputs after the alignment tunin

2025

Critical Node-aware Augmentation for Hypergraph Contrastive Learning

IJCAI 2025

Hypergraph contrastive learning enables effective representation learning for hypergraphs without requiring labels. However, existing methods typically rely on randomly deleting or replacing nodes during hypergraph augmentation, which may lead to the absence of critical nodes and further disrupt the

Cited by 0SourcePDFScholar
2025

EPA: Boosting Event-based Video Frame Interpolation with Perceptually Aligned Learning

NeurIPS 2025poster

Event cameras, with their capacity to provide high temporal resolution information between frames, are increasingly utilized for video frame interpolation (VFI) in challenging scenarios characterized by high-speed motion and significant occlusion. However, prevalent issues of blur and distortion wit…

Cited by 0SourceScholar
2025

ESEG: Event-Based Segmentation Boosted by Explicit Edge-Semantic Guidance

AAAI 2025technical

Event-based semantic segmentation (ESS) has attracted researchers' attention recently, as event cameras can solve problems such as under/over-exposure or motion blur that are difficult for RGB cameras to handle. However, event data are noisy and sparse, resulting in difficulties for the model to loc…

2025

Enhance Multi-View Classification Through Multi-Scale Alignment and Expanded Boundary

ICLR 2025poster

Multi-view classification aims at unifying the data from multiple views to complementarily enhance the classification performance. Unfortunately, two major problems in multi-view data are damaging model performance. The first is feature heterogeneity, which makes it hard to fuse features from differ…

Cited by 0SourcePDFScholar
2025

Forget the Unneeded: Backdooring Large Language Models via Contrastive-enhanced Machine Unlearning

EMNLP 2025

Prompt tuning for Large Language Models (LLMs) is vulnerable to backdoor attacks. Existing methods find backdoor attacks to be a significant threat in data-rich scenarios. However, in data-limited scenarios, these methods have difficulty capturing precise backdoor patterns, leading to weakened backd

2025

GLoCIM: Global-view Long Chain Interest Modeling for news recommendation

COLING 2025main

Accurately recommending candidate news articles to users has always been the core challenge of news recommendation system. News recommendations often require modeling of user interest to match candidate news. Recent efforts have primarily focused on extracting local subgraph information in a global…

Cited by 1SourcePDFScholar
2025

GNCL: A Graph Neural Network with Consistency Loss for Segment-Level Spoofed Speech Detection

ICASSP 2025accepted

Segment-level spoofed speech detection focuses on recognizing fake or synthetic segments within identifying partially spoofed speech. Nevertheless, existing models for this segment-level task usually overlook latent local relationships between fake and bona fide segments, and further, a lack of inte…

Cited by 0SourceScholar
2025

Graph Consistency and Diversity Measurement for Federated Multi-View Clustering

AAAI 2025technical

Federated Multi-View Clustering (FMVC) aims to learn a global clustering model from heterogeneous data distributed across different devices, where each device only stores one view of all clustering samples. The key to deal with such problem lies in how to effectively fuse these heterogeneous samples…

Cited by 0SourcePDFScholar
2025

HMoE: Heterogeneous Mixture of Experts for Language Modeling

EMNLP 2025

Mixture of Experts (MoE) offers remarkable performance and computational efficiency by selectively activating subsets of model parameters. Traditionally, MoE models use homogeneous experts, each with identical capacity. However, varying complexity in input data necessitates experts with diverse capa

2025

Know Where You Are From: Event-Based Segmentation via Spatio-Temporal Propagation

AAAI 2025technical

Event cameras have gained attention in segmentation due to their higher temporal resolution and dynamic range compared to traditional cameras. However, they struggle with issues like lack of color perception and triggering only at motion edges, making it hard to distinguish objects with similar cont…

2025

LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image Generation

AAAI 2025technical

Diffusion models have exhibited substantial success in text-to-image generation. However, they often encounter challenges when dealing with complex and dense prompts involving multiple objects, attribute binding, and long descriptions. In this paper, we propose a novel framework called LLM4GEN, whic…

Cited by 20SourcePDFScholar
2025

MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

ICLR 2025poster

We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture, MM1.5 adopts a data-centric approach to model training, systema…

Cited by 29SourcePDFScholar
2025

MSV-PCT: Multi-Sparse-View Enhanced Transformer Framework for Salient Object Detection in Point Clouds

AAAI 2025technical

Salient object detection (SOD) methods for 2D images have great significance in the field of human-computer interaction (HCI). However, as a common data format in HCI, the SOD research in the form of 3D point cloud data remains limited. Previous works commonly treat this task as point cloud segmenta…

Cited by 0SourcePDFScholar
2025

Memory or Reasoning? Explore How LLMs Compute Mixed Arithmetic Expressions

ACL 2025finding

Large language models (LLMs) can solve complex multi-step math reasoning problems, but little is known about how these computations are implemented internally. Many recent studies have investigated the mechanisms of LLMs on simple arithmetic tasks (e.g., a+b, a× b), but how LLMs solve mixed arithmet…

Cited by 0SourcePDFScholar
2025

Mitigating Local Cohesion and Global Sparseness in Graph Contrastive Learning with Fuzzy Boundaries

ICML 2025poster

Graph contrastive learning (GCL) aims at narrowing positives while dispersing negatives, often causing a minority of samples with great similarities to gather as a small group. It results in two latent shortcomings in GCL: 1) **local cohesion** that a class cluster contains numerous independent smal…

Cited by 0SourcePDFScholar
2025

Multi-View Multi-Label Classification via View-Label Matching Selection

AAAI 2025technical

In multi-view multi-label classification (MVML), each object is described by several heterogeneous views while annotated with multiple related labels. The key to learn from such complicate data lies in how to fuse cross-view features and explore multi-label correlations, while accordingly obtain cor…

Cited by 0SourcePDFScholar
2025

Option Symbol Matters: Investigating and Mitigating Multiple-Choice Option Symbol Bias of Large Language Models

NAACL 2025long

Multiple-Choice Question Answering (MCQA) is a widely used task in the evaluation of Large Language Models (LLMs). In this work, we reveal that current LLMs’ performance in MCQA could be heavily influenced by the choice of option symbol sets, due to the option symbol bias. That is, when altering onl…

Cited by 0SourcePDFScholar
2025

Scaling Laws for Floating–Point Quantization Training

ICML 2025poster

Low-precision training is considered an effective strategy for reducing both training and downstream inference costs. Previous scaling laws for precision mainly focus on integer quantization, which pay less attention to the constituents in floating-point (FP) quantization, and thus cannot well fit t…

Cited by 1SourcePDFScholar
2025

Separation of Powers: On Segregating Knowledge from Observation in LLM-enabled Knowledge-based Visual Question Answering

CVPR 2025poster

Knowledge-based visual question answering (KBVQA) separates image interpretation and knowledge retrieval into separate processes, motivated in part by the fact that they are very different tasks. In this paper, we transform the KBVQA into linguistic question-answering tasks so that we can leverage t…

Cited by 0SourcePDFScholar
2025

Tensorized Multi-View Multi-Label Classification via Laplace Tensor Rank

ICML 2025poster

In multi-view multi-label classification (MVML), each object has multiple heterogeneous views and is annotated with multiple labels. The key to deal with such problem lies in how to capture cross-view consistent correlations while excavate multi-label semantic relationships. Existing MVML methods us…

Cited by 0SourcePDFScholar
2025

Thought-Path Contrastive Learning via Premise-Oriented Data Augmentation for Logical Reading Comprehension

AAAI 2025technical

Logical reading comprehension is a challenging task that entails grasping the underlying semantics of text and applying reasoning to deduce the correct answer. Prior researches have primarily focused on enhancing logical reasoning capabilities through Chain-of-Thought (CoT) or data augmentation. How…

2025

Time-Graph Frequency Representation with Singular Value Decomposition for Neural Speech Enhancement

ICASSP 2025accepted

Time-frequency (T-F) domain methods for monaural speech enhancement have benefited from the success of deep learning. Recently, focus has been put on designing two-stream network models to predict amplitude mask and phase separately, or, coupling the amplitude and phase into Cartesian coordinates an…

Cited by 0SourceScholar
2025

Triples as the Key: Structuring Makes Decomposition and Verification Easier in LLM-based TableQA

ICLR 2025poster

As the mainstream approach, LLMs have been widely applied and researched in TableQA tasks. Currently, the core of LLM-based TableQA methods typically include three phases: question decomposition, sub-question TableQA reasoning, and answer verification. However, several challenges remain in this proc…

Cited by 0SourcePDFScholar
2024

Common-Individual Semantic Fusion for Multi-View Multi-Label Learning

IJCAI 2024poster

In Multi-View Multi-Label Learning, each instance is described by several heterogeneous features and associated with multiple valid labels simultaneously. Existing methods mainly focus on leveraging feature-level view fusion to capture a common representation for multi-label classifier induction. In…

Cited by 5SourcePDFScholar
2024

FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior

ECCV 2024poster

"[width=0.985]assets/teaser.pdf Figure 1: harnesses the generative prior of pre-trained diffusion models to achieve versatile image composition, such as appearance editing (image harmonization) and semantic editing (semantic image composition). Furthermore, it can be extended to various downstream a…

2024

FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition

CVPR 2024poster

Benefiting from large-scale pre-trained text-to-image (T2I) generative models impressive progress has been achieved in customized image generation which aims to generate user-specified concepts. Existing approaches have extensively focused on single-concept customization and still encounter challeng…

2024

Improving Implicit Discourse Relation Recognition with Semantics Confrontation

COLING 2024main

Implicit Discourse Relation Recognition (IDRR), which infers discourse logical relations without explicit connectives, is one of the most challenging tasks in natural language processing (NLP). Recently, pre-trained language models (PLMs) have yielded impressive results across numerous NLP tasks, bu…

Cited by 0SourcePDFScholar
2024

LightVLP: A Lightweight Vision-Language Pre-training via Gated Interactive Masked AutoEncoders

COLING 2024main

This paper studies vision-language (V&L) pre-training for deep cross-modal representations. Recently, pre-trained V&L models have shown great success in V&L tasks. However, most existing models apply multi-modal encoders to encode the image and text, at the cost of high training complexity because o…

Cited by 1SourcePDFScholar
2024

LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning

ACL 2024findings

Large Language Models (LLMs), such as LLaMA and T5, have shown exceptional performance across various tasks through fine-tuning. Although low-rank adaption (LoRA) has emerged to cheaply fine-tune these LLMs on downstream tasks, their deployment is still hindered by the vast model scale and computati…

2024

Multi-Level Cross-Modal Alignment for Speech Relation Extraction

EMNLP 2024main

Speech Relation Extraction (SpeechRE) aims to extract relation triplets from speech data. However, existing studies usually use synthetic speech to train and evaluate SpeechRE models, hindering the further development of SpeechRE due to the disparity between synthetic and real speech. Meanwhile, the…

Cited by 0SourcePDFScholar
2024

Object-Aware Inversion and Reassembly for Image Editing

ICLR 2024poster

Diffusion-based image editing methods have achieved remarkable advances in text-driven image editing. The editing task aims to convert an input image with the original text prompt into the desired image that is well-aligned with the target text prompt. By comparing the original and target prompts, w…

2024

SAM-Event-Adapter: Adapting Segment Anything Model for Event-RGB Semantic Segmentation

ICRA 2024poster

Semantic segmentation, a fundamental visual task ubiquitously employed in sectors ranging from transportation and robotics to healthcare, has always captivated the research community. In the wake of rapid advancements in large model research, the foundation model for semantic segmentation tasks, ter…

Cited by 11SourceScholar
2024

SDformer: Transformer with Spectral Filter and Dynamic Attention for Multivariate Time Series Long-term Forecasting

IJCAI 2024poster

Transformer has gained widespread adoption in modeling time series due to the exceptional ability of its self-attention mechanism in capturing long-range dependencies. However, when processing time series data with numerous variates, the vanilla self-attention mechanism tends to distribute attention…

2024

TriSampler: A Better Negative Sampling Principle for Dense Retrieval

AAAI 2024technical

Negative sampling stands as a pivotal technique in dense retrieval, essential for training effective retrieval models and significantly impacting retrieval performance. While existing negative sampling methods have made commendable progress by leveraging hard negatives, a comprehensive guiding princ…

Cited by 3SourcePDFScholar
2024

Video Frame Interpolation via Direct Synthesis with the Event-based Reference

CVPR 2024poster

Video Frame Interpolation (VFI) has witnessed a surge in popularity due to its abundant downstream applications. Event-based VFI (E-VFI) has recently propelled the advancement of VFI. Thanks to the high temporal resolution benefits event cameras can bridge the informational void present between succ…

Cited by 6SourcePDFScholar
2024

XAL: EXplainable Active Learning Makes Classifiers Better Low-resource Learners

NAACL 2024long

Active learning (AL), which aims to construct an effective training set by iteratively curating the most formative unlabeled data for annotation, has been widely used in low-resource tasks. Most active learning techniques in classification rely on the model’s uncertainty or disagreement to choose un…

2023

CLIP2: Contrastive Language-Image-Point Pretraining From Real-World Point Cloud Data

CVPR 2023poster

Contrastive Language-Image Pre-training, benefiting from large-scale unlabeled text-image pairs, has demonstrated great performance in open-world vision understanding tasks. However, due to the limited Text-3D data pairs, adapting the success of 2D Vision-Language Models (VLM) to the 3D space remain…

Cited by 107SourcePDFScholar
2023

Enhancing Argument Structure Extraction with Efficient Leverage of Contextual Information

EMNLP 2023short findings

Argument structure extraction (ASE) aims to identify the discourse structure of arguments within documents. Previous research has demonstrated that contextual information is crucial for developing an effective ASE model. However, we observe that merely concatenating sentences in a contextual window…

Cited by 0SourcecodeScholar
2023

Rethinking the Word-level Quality Estimation for Machine Translation from Human Judgement

ACL 2023findings

Word-level Quality Estimation (QE) of Machine Translation (MT) aims to detect potential translation errors in the translated sentence without reference. Typically, conventional works on word-level QE are usually designed to predict the quality of translated words in terms of the post-editing effort,…

2023

Towards Domain Generalization for Multi-View 3D Object Detection in Bird-Eye-View

CVPR 2023poster

Multi-view 3D object detection (MV3D-Det) in Bird-Eye-View (BEV) has drawn extensive attention due to its low cost and high efficiency. Although new algorithms for camera-only 3D object detection have been continuously proposed, most of them may risk drastic performance degradation when the domain o…

Cited by 25SourcePDFScholar
2023

ViLTA: Enhancing Vision-Language Pre-training through Textual Augmentation

ICCV 2023poster

Vision-language pre-training (VLP) methods are blossoming recently, and its crucial goal is to jointly learn visual and textual features via a transformer-based architecture, demonstrating promising improvements on a variety of vision-language tasks. Prior arts usually focus on how to align visual a…

Cited by 13PDFScholar
2023

Zero-Shot Speech Emotion Recognition Using Generative Learning with Reconstructed Prototypes

ICASSP 2023accepted

Zero-shot Speech Emotion Recognition (SER) enables machines to perceive unseen-emotional speech without knowing any samples from these emotional states, which is helpful in audio-based autonomous affective computing. However, existing works on zero-shot SER directly employ original prototypes and on…

Cited by 7SourceScholar
2022

Cross-Align: Modeling Deep Cross-lingual Interactions for Word Alignment

EMNLP 2022main

Word alignment which aims to extract lexicon translation equivalents between source and target sentences, serves as a fundamental tool for natural language processing. Recent studies in this area have yielded substantial improvements by generating alignments from contextualized embeddings of the pre…

2022

EAG: Extract and Generate Multi-way Aligned Corpus for Complete Multi-lingual Neural Machine Translation

ACL 2022long

Complete Multi-lingual Neural Machine Translation (C-MNMT) achieves superior performance against the conventional MNMT by constructing multi-way aligned corpus, i.e., aligning bilingual training examples from different language pairs when either their source or target sides are identical. However, s…

Cited by 4SourcePDFScholar
2022

Generating Authentic Adversarial Examples beyond Meaning-preserving with Doubly Round-trip Translation

NAACL 2022long

Generating adversarial examples for Neural Machine Translation (NMT) with single Round-Trip Translation (RTT) has achieved promising results by releasing the meaning-preserving restriction. However, a potential pitfall for this approach is that we cannot decide whether the generated examples are adv…

2022

Laneformer: Object-Aware Row-Column Transformers for Lane Detection

AAAI 2022technical

We present Laneformer, a conceptually simple yet powerful transformer-based architecture tailored for lane detection that is a long-standing research topic for visual perception in autonomous driving. The dominant paradigms rely on purely CNN-based architectures which often fail in incorporating rel…

Cited by 60SourcePDFScholar
2022

Maximized Hydrodynamic Stimulation Strategy for Placement of Differential Pressure and Velocity Sensors in Artificial Lateral Line Systems

RA-L 2022

Fish can perceive the surrounding flow field using their lateral line systems, consisting of canal neuromasts (CNs) for flow pressure gradient perception and superficial neuromasts (SNs) for flow velocity detection. Although various artificial lateral line (ALL) systems have been developed inspired

Cited by 21SourceScholar
2022

ONCE-3DLanes: Building Monocular 3D Lane Detection

CVPR 2022poster

We present ONCE-3DLanes, a real-world autonomous driving dataset with lane layout annotation in 3D space. Conventional 2D lane detection from a monocular image yields poor performance of following planning and control tasks in autonomous driving due to the case of uneven road. Predicting the 3D lane…

Cited by 76PDFcodeScholar
2022

Semi-supervised Semantic Segmentation with Prototype-based Consistency Regularization

NeurIPS 2022accept

Semi-supervised semantic segmentation requires the model to effectively propagate the label information from limited annotated images to unlabeled ones. A challenge for such a per-pixel prediction task is the large intra-class variation, i.e., regions belonging to the same class may exhibit a very d…

2022

TSAM: A Two-Stream Attention Model for Causal Emotion Entailment

COLING 2022main

Causal Emotion Entailment (CEE) aims to discover the potential causes behind an emotion in a conversational utterance. Previous works formalize CEE as independent utterance pair classification problems, with emotion and speaker information neglected. From a new perspective, this paper considers CEE…

2021

Learning Transferable Features for Point Cloud Detection via 3D Contrastive Co-training

NeurIPS 2021poster

Most existing point cloud detection models require large-scale, densely annotated datasets. They typically underperform in domain adaptation settings, due to geometry shifts caused by different physical environments or LiDAR sensor configurations. Therefore, it is challenging but valuable to learn t…

Cited by 34SourcePDFScholar