← Search

Yu Zhou

108 accepted papers

2026

ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding

CVPR 2026

The rapid advancement of educational applications, artistic creation, and AI-generated content (AIGC) technologies has substantially increased practical requirements for comprehensive Image Aesthetics Assessment (IAA), particularly demanding methods capable of delivering both quantitative scoring an

Cited by 32SourcecodeScholar
2026

Beyond Detection: A Structure-Aware Framework for Scene Text Tracking

ICML 2026poster

Modern visual object trackers show impressive results on general targets, yet their performance drops substantially when dealing with scene text. Although currently underexplored, tracking text in videos is essential for dynamic text manipulations such as segmentation, removal, and editing. To fill …

Cited by 0SourceScholar
2026

Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach

ICLR 2026poster

Recently, Multimodal Large Language Models (MLLMs) have achieved exceptional performance across diverse tasks, continually surpassing previous expectations regarding their capabilities. Nevertheless, their proficiency in perceiving emotions from images remains debated, with studies yielding divergen…

Cited by 0SourcecodeScholar
2026

ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction

ICLR 2026poster

Embodied cognition argues that intelligence arises from continuous sensorimotor interaction with the world. It raises an intriguing question: do modern vision-language models (VLMs), trained largely in a disembodied manner, exhibit signs of embodied cognition? To investigate this, we introduce **ENA…

Cited by 0SourcecodeScholar
2026

Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning

ICLR 2026poster

The maturation of Large Audio Language Models (LALMs) has raised growing expectations for them to comprehend complex audio much like humans. Current efforts primarily replicate text-based reasoning by contextualizing audio content through a one-time encoding, which introduces a critical information…

Cited by 0SourcecodeScholar
2026

FDSPC: Fast and Direct Smooth Motion Planning Via Continuous Curvature Integration

ICRA 2026poster

In recent decades, mobile robot motion planning has seen significant advancements. Both search-based and sampling-based methods have demonstrated capabilities to find feasible solutions in complex scenarios. Mainstream path planning algorithms divide the map into occupied and free spaces, considerin…

2026

Focus, Align, and Sustain: Counteracting Gradient Dilution in Incremental Object Detection

ICML 2026poster

Adapting Detection Transformers to Incremental Object Detection (IOD) poses a systemic challenge, as set-based optimization is inherently destabilized by sequential learning. In this work, we identify Gradient Dilution as the root cause of performance degradation, wherein optimization signals requir…

Cited by 0SourceScholar
2026

MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation

CVPR 2026

End-to-end text-image machine translation (TIMT), which directly translates textual content in images across languages, is crucial for real-world multilingual scene understanding. Despite advances in vision-language large models (VLLMs), robustness across diverse visual scenes and low-resource langu

Cited by 0SourceScholar
2026

Position: Quantum Program Generation Must Prioritize Validity Over Probabilistic Scaling

ICML 2026poster

The scaling hypothesis assumes that increasing model parameters yields emergent reasoning capabilities. This position paper argues that applying this probabilistic paradigm to generic quantum circuit synthesis is a category error. Unlike natural languages, quantum circuits require strict adherence t…

Cited by 0SourceScholar
2026

SAM 3: Segment Anything with Concepts

ICLR 2026poster

We present Segment Anything Model (SAM) 3, a unified model that detects, segments, and tracks objects in images and videos based on concept prompts, which we define as either short noun phrases (e.g., “yellow school bus”), image exemplars, or a combination of both. Promptable Concept Segmentation (P…

Cited by 687SourcecodeScholar
2026

ST-SAM: Multimodal Scene Text Segmentation with Dense Visual and Sparse Textual Prompts via SAM

AAAI 2026technical

Scene text segmentation is a critical preprocessing step in various text-based applications. Specialist text segmentation methods, often relying on a detect-then-segment paradigm, tend to exhibit reduced robustness and can lead to cascading errors. The introduction of the Segment Anything Model (SAM

Cited by 0SourcePDFScholar
2026

SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition

AAAI 2026technical

Large Language Models (LLMs) hold rich implicit knowledge and powerful transferability. In this paper, we explore the combination of LLMs with the human skeleton to perform action classification and description. However, when treating LLM as a recognizer, two questions arise: 1) How can LLMs underst

Cited by 0SourcePDFScholar
2026

Semantic Audio-Visual Navigation in Continuous Environments

CVPR 2026

Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on precomputed room impulse responses (RIRs) for binaural audio rendering, restricting agents to discrete grid positions and l

Cited by 0SourcecodeScholar
2026

Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric Refinement

AAAI 2026technical

Understanding 3D scene-level affordances from natural language instructions is essential for enabling embodied agents to interact meaningfully in complex environments. However, this task remains challenging due to the need for semantic reasoning and spatial grounding. Existing methods mainly focus o

Cited by 0SourcePDFScholar
2026

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training

CVPR 2026

Document parsing has recently advanced with multimodal large language models (MLLMs) that directly map document images to structured outputs. Traditional cascaded pipelines depend on precise layout analysis and often fail under casually captured or non-standard conditions. Although end-to-end approa

Cited by 0SourceScholar
2026

URS: A Unified Neural Routing Solver for Cross-Problem Zero-Shot Generalization

ICML 2026poster

Multi-task neural routing solvers have emerged as a promising paradigm for their ability to solve multiple vehicle routing problems (VRPs) using a single model. However, existing neural solvers typically rely on predefined problem constraints or require per-problem fine-tuning, which substantially l…

Cited by 0SourceScholar
2026

UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture

ICML 2026spotlight

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks such as visual grounding, segmentation, and captioning. However, their ability to perceive perceptual-level image features remains limited. In this work, we present UniPercept-Bench, a unified fr…

Cited by 0SourceScholar
2026

When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?

AAAI 2026technical

Can Multimodal Large Language Models (MLLMs) discern confused objects that are visually present but audio-absent? To study this, we introduce a new benchmark, AV-ConfuseBench, which simulates an “Audio-Visual Confusion” scene by modifying the corresponding sound of an object in the video, e.g., mute

Cited by 0SourcePDFScholar
2025

A Query-Response Framework for Whole-Page Complex-Layout Document Image Translation with Relevant Regional Concentration

ACL 2025finding

Document Image Translation (DIT), which aims at translating documents in images from source language to the target, plays an important role in Document Intelligence. It requires a comprehensive understanding of document multi-modalities and a focused concentration on relevant textual regions during…

Cited by 0SourcePDFScholar
2025

A Robust Quality Evaluator for Panoramic Videos

ICASSP 2025accepted

Most of the existing methods to evaluate the quality of panoramic content mainly focus on studying the quality evaluation of static panoramic images, rather than the more widely used dynamic panoramic videos. Also, the few panoramic video quality metrics that are available have obvious weaknesses in…

Cited by 0SourceScholar
2025

An Adaptive Quantum Circuit of Dempster's Rule of Combination for Uncertain Pattern Classification

NeurIPS 2025poster

In pattern classification, efficient uncertainty reasoning plays a critical role, particularly in real-time applications involving noisy data, ambiguous class boundaries, or overlapping categories. Leveraging the advanced computational power of quantum computing, an Adaptive Quantum Circuit for Demp…

Cited by 0SourceScholar
2025

An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability

ICML 2025poster

The advancements in Multimodal Large Language Models (MLLMs) have enabled various multimodal tasks to be addressed under a zero-shot paradigm. This paradigm sidesteps the cost of model fine-tuning, emerging as a dominant trend in practical application. Nevertheless, Multimodal Sentiment Analysis (MS…

Cited by 0SourcePDFScholar
2025

AnomalyNCD: Towards Novel Anomaly Class Discovery in Industrial Scenarios

CVPR 2025poster

Recently, multi-class anomaly classification has garnered increasing attention. Previous methods directly cluster anomalies but often struggle due to the lack of anomaly-prior knowledge. Acquiring this knowledge faces two issues: the non-prominent and weak-semantics anomalies. In this paper, we prop…

2025

Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance

AAAI 2025technical

Scene text spotting has attracted the enthusiasm of relative researchers in recent years. Most existing scene text spotters follow the detection-then-recognition paradigm, where the vanilla detection module hardly determines the reading order and leads to failure recognition. After rethinking the au…

Cited by 2SourcePDFScholar
2025

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts

ICML 2025poster

Chinese scene text retrieval is a practical task that aims to search for images containing visual instances of a Chinese query text. This task is extremely challenging because Chinese text often features complex and diverse layouts in real-world scenes. Current efforts tend to inherit the solution f…

Cited by 0SourcePDFScholar
2025

Beyond Sequences: Two-dimensional Representation and Dependency Encoding for Code Generation

ACL 2025long

The advent of large language models has significantly advanced automatic code generation, transforming the way programmers writing code. Inspired by natural language processing, mainstream code generation approaches represent code as a linear sequence of tokens. In this paper, we propose to represen…

Cited by 0SourcePDFScholar
2025

Bridging the Reality Gap: Communication-Aware Task Allocation with Multi-Objective Asynchronous Policy Learning

IROS 2025

Distributed task allocation in the UAV swarm is sensitive to excessive communication overhead and frequent transmissions. Combining reinforcement learning and task allocation demonstrates great potential in enhancing algorithm performance and optimizing communication. However, existing studies rely

Cited by 0SourceScholar
2025

Char-SAM: Turning Segment Anything Model into Scene Text Segmentation Annotator with Character-level Visual Prompts

ICASSP 2025accepted

The recent emergence of the Segment Anything Model (SAM) enables various domain-specific segmentation tasks to be tackled cost-effectively by using bounding boxes as prompts. However, in scene text segmentation, SAM can not achieve desirable performance. The word-level bounding box as prompts is too…

Cited by 0SourceScholar
2025

Contrastive Visual Data Augmentation

ICML 2025poster

Large multimodal models (LMMs) often struggle to recognize novel concepts, as they rely on pre-trained knowledge and have limited ability to capture subtle visual details. Domain-specific knowledge gaps in training also make them prone to confusing visually similar, commonly misrepresented, or low-r…

Cited by 0SourcePDFScholar
2025

DCA: Dividing and Conquering Amnesia in Incremental Object Detection

AAAI 2025technical

Incremental object detection (IOD) aims to cultivate an object detector that can continuously localize and recognize novel classes while preserving its performance on previous classes. Existing methods achieve certain success by improving knowledge distillation and exemplar replay for transformer-ba…

2025

Decoupled Distillation to Erase: A General Unlearning Method for Any Class-centric Tasks

CVPR 2025highlight

In this work, we present DEcoupLEd Distillation To Erase (DELETE), a general and strong unlearning method for any class-centric tasks. To derive this, we first propose a theoretical framework to analyze the general form of unlearning loss and decompose it into forgetting and retention terms. Through…

Cited by 2SourcePDFScholar
2025

Distilling LLM Prior to Flow Model for Generalizable Agent’s Imagination in Object Goal Navigation

NeurIPS 2025poster

The Object Goal Navigation (ObjectNav) task challenges agents to locate a specified object in an unseen environment by imagining unobserved regions of the scene. Prior approaches rely on deterministic and discriminative models to complete semantic maps, overlooking the inherent uncertainty in indoor…

Cited by 0SourceScholar
2025

Dual-S3D: Hierarchical Dual-Path Selective SSM-CNN for High-Fidelity Implicit Reconstruction

ICCV 2025poster

Single-view 3D reconstruction aims to recover the complete 3D geometry and appearance of objects from a single RGB image. Due to incomplete image information and ambiguity, this task remains challenging. Existing methods struggle with the trade-off between local detail and global topology, and with…

Cited by 0SourcePDFScholar
2025

From Chaotic OCR Words to Coherent Document: A Fine-to-Coarse Zoom-Out Network for Complex-Layout Document Image Translation

COLING 2025main

Document Image Translation (DIT) aims to translate documents in images from one language to another. It requires visual layouts and textual contents understanding, as well as document coherence capturing. However, current methods often rely on the quality of OCR output, which, particularly in comple…

2025

HM3: Hierarchical Multi-Objective Model Merging for Pretrained Models

NeurIPS 2025spotlight

Model merging is a technique that combines multiple large pretrained models into a single model, enhancing performance and broadening task adaptability without original data or additional training. However, most existing model merging methods focus primarily on exploring the parameter space, merging…

Cited by 0SourceScholar
2025

Improving MLLM’s Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency

ACL 2025finding

Multimodal Large Language Models (MLLMs) have shown strong performance in document image tasks, especially Optical Character Recognition (OCR). However, they struggle with Document Image Machine Translation (DIMT), which requires handling both cross-modal and cross-lingual challenges. Previous effor…

Cited by 0SourcePDFScholar
2025

Inverse Kinematics on Guiding Vector Fields for Robot Path Following

ICRA 2025

Inverse kinematics is a fundamental technique for motion and positioning control in robotics, typically applied to end-effectors. In this paper, we extend the concept of inverse kinematics to guiding vector fields for path following in autonomous mobile robots. The desired path is defined by its imp

Cited by 4SourcecodeScholar
2025

Investigating Hallucinations in Simultaneous Machine Translation: Knowledge Distillation Solution and Components Analysis

NAACL 2025long

Simultaneous Machine Translation (SiMT) generates target translation before receiving the whole source sentence and faces a serious hallucination problem. In contrast, traditional offline machine translation (OMT) models exhibit significantly fewer hallucinations. Motivated by this disparity, we pro…

Cited by 0SourcePDFScholar
2025

LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining

AAAI 2025technical

Visual Information Extraction (VIE) plays a crucial role in the comprehension of semi-structured documents, and several pre-trained models have been developed to enhance performance. However, most of these works are monolingual (usually English). Due to the extremely unbalanced quantity and quality…

Cited by 2SourcePDFScholar
2025

Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition

CVPR 2025poster

Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while the linguistic dimension incorporates contextual and semantic elements. In scenarios with degraded visual quality, lingui…

2025

MTL-KD: Multi-Task Learning Via Knowledge Distillation for Generalizable Neural Vehicle Routing Solver

NeurIPS 2025poster

Multi-Task Learning (MTL) in Neural Combinatorial Optimization (NCO) is a promising approach for training a unified model capable of solving multiple Vehicle Routing Problem (VRP) variants. However, existing Reinforcement Learning (RL)-based multi-task methods can only train light decoder models on…

Cited by 0SourcecodeScholar
2025

Orochi: Versatile Biomedical Image Processor

NeurIPS 2025spotlight

Deep learning has emerged as a pivotal tool for accelerating research in the life sciences, with the low-level processing of biomedical images (e.g., registration, fusion, restoration, super-resolution) being one of its most critical applications. Platforms such as ImageJ (Fiji) and napari have enab…

Cited by 0SourceScholar
2025

Pay More Attention to Images: Numerous Images-Oriented Multimodal Summarization

NAACL 2025long

Existing multimodal summarization approaches struggle with scenarios involving numerous images as input, leading to a heavy load for readers. Summarizing both the input text and numerous images helps readers quickly grasp the key points of multimodal input. This paper introduces a novel task, Numero…

2025

SeaS: Few-shot Industrial Anomaly Image Generation with Separation and Sharing Fine-tuning

ICCV 2025poster

We introduce SeaS, a unified industrial generative model for automatically creating diverse anomalies, authentic normal products, and precise anomaly masks. While extensive research exists, most efforts either focus on specific tasks, i.e., anomalies or normal products only, or require separate mode…

2025

SimulPL: Aligning Human Preferences in Simultaneous Machine Translation

ICLR 2025poster

Simultaneous Machine Translation (SiMT) generates translations while receiving streaming source inputs. This requires the SiMT model to learn a read/write policy, deciding when to translate and when to wait for more source input. Numerous linguistic studies indicate that audiences in SiMT scenarios…

2025

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation

ACL 2025long

Document Image Machine Translation (DIMT) aims to translate text within document images, facing generalization challenges due to limited training data and the complex interplay between visual and textual information. To address these challenges, we introduce M4Doc, a novel single-to-mix Modality ali…

Cited by 0SourcePDFScholar
2025

Specifying What You Know or Not for Multi-Label Class-Incremental Learning

AAAI 2025technical

Existing class incremental learning is mainly designed for single-label classification task, which is ill-equipped for multi-label scenarios due to the inherent contradiction of learning objectives for samples with incomplete labels. We argue that the main challenge to overcome this contradiction in…

2025

TROVE: A Challenge for Fine-Grained Text Provenance via Source Sentence Tracing and Relationship Classification

ACL 2025long

LLMs have achieved remarkable fluency and coherence in text generation, yet their widespread adoption has raised concerns about content reliability and accountability. In high-stakes domains, it is crucial to understand where and how the content is created. To address this, we introduce the Text pRO…

2025

The Devil is in Fine-tuning and Long-tailed Problems: A New Benchmark for Scene Text Detection

IJCAI 2025

Scene text detection has seen the emergence of high-performing methods that excel on academic benchmarks. However, these detectors often fail to replicate such success in real-world scenarios. We uncover two key factors contributing to this discrepancy through extensive experiments. First, a Fine-tu

2025

The Four Color Theorem for Cell Instance Segmentation

ICML 2025poster

Cell instance segmentation is critical to analyzing biomedical images, yet accurately distinguishing tightly touching cells remains a persistent challenge. Existing instance segmentation frameworks, including detection-based, contour-based, and distance mapping-based approaches, have made significan…

2025

The Role of Video Generation in Enhancing Data-Limited Action Understanding

IJCAI 2025

Video action understanding tasks in real-world scenarios often suffer from data limitations. In this paper, we address the data-limited action understanding problem by bridging data scarcity. We propose a novel method that leverages a text-to-video diffusion transformer to generate annotated data fo

Cited by 0SourcePDFScholar
2025

Towards Robustness and Explainability of Automatic Algorithm Selection

ICML 2025spotlight

Algorithm selection aims to identify the optimal performing algorithm before execution. Existing techniques typically focus on the observed correlations between algorithm performance and meta-features. However, little research has explored the underlying mechanisms of algorithm selection, specifical…

Cited by 0SourcePDFScholar
2025

Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues

AAAI 2025technical

Video text-based visual question answering (TextVQA) is a practical task that aims to answer questions by jointly reasoning textual and visual information in a given video. Inspired by the development of TextVQA in image domain, existing Video TextVQA approaches leverage a language model (e.g. T5) t…

2025

When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding

NeurIPS 2025poster

Large Multimodal Models (LMMs) have achieved impressive progress in visual perception and reasoning. However, when confronted with visually ambiguous or non-semantic scene text, they often struggle to accurately spot and understand the content, frequently generating semantically plausible yet visual…

Cited by 0SourceScholar
2025

monoVLN: Bridging the Observation Gap between Monocular and Panoramic Vision and Language Navigation

ICCV 2025poster

Vision and Language Navigation(VLN) requires agents to navigate 3D environments by following natural language instructions. While existing methods predominantly assume access to panoramic observations, many practical robotics are equipped with monocular RGBD cameras, creating a significant configura…

Cited by 0SourcePDFScholar
2024

A Complete Landscape of EFX Allocations on Graphs: Goods, Chores and Mixed Manna

IJCAI 2024poster

We study envy-free up to any item (EFX) allocations on graphs where vertices and edges represent agents and items respectively. An agent is only interested in items that are incident to her and all other items have zero marginal values to her. Christodoulou et al. first proposed this setting and stu…

Cited by 8SourcePDFScholar
2024

Accurate and Robust Scene Text Recognition via Adversarial Training

ICASSP 2024accepted

Adversarial training (AT) is a methodology that utilizes adversarial examples in the training process to enhance a model’s resistance to adversarial attacks and improve generalization. Despite its efficacy in several non-sequential computer vision tasks such as classification and object detection, i…

Cited by 0SourceScholar
2024

Born a BabyNet with Hierarchical Parental Supervision for End-to-End Text Image Machine Translation

COLING 2024main

Text image machine translation (TIMT) aims at translating source language texts in images into another target language, which has been proven successful by bridging text image recognition encoder and text translation decoder. However, it is still an open question of how to incorporate fine-grained k…

2024

DIUSum: Dynamic Image Utilization for Multimodal Summarization

AAAI 2024technical

Existing multimodal summarization approaches focus on fusing image features in the encoding process, ignoring the individualized needs for images when generating different summaries. However, whether intuitively or empirically, not all images can improve summary quality. Therefore, we propose a nove…

Cited by 5SourcePDFScholar
2024

Document Image Machine Translation with Dynamic Multi-pre-trained Models Assembling

NAACL 2024long

Text image machine translation (TIMT) is a task that translates source texts embedded in the image to target translations. The existing TIMT task mainly focuses on text-line-level images. In this paper, we extend the current TIMT task and propose a novel task, **D**ocument **I**mage **M**achine **T*…

2024

Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making

NeurIPS 2024oral

We aim to evaluate Large Language Models (LLMs) for embodied decision making. While a significant body of work has been leveraging LLMs for decision making in embodied environments, we still lack a systematic understanding of their performance because they are usually applied in different domains, f…

Cited by 33SourcePDFScholar
2024

Generalized Taxonomy-Guided Graph Neural Networks

IJCAI 2024poster

Graph neural networks have been demonstrated to be effective analytic apparatus for mining network data. Most real-world networks are inherently hierarchical, offering unique opportunities to acquire latent, intrinsic network organizational properties by utilizing network taxonomies. The existing ap…

Cited by 0SourcePDFScholar
2024

MuSc: Zero-Shot Industrial Anomaly Classification and Segmentation with Mutual Scoring of the Unlabeled Images

ICLR 2024poster

This paper studies zero-shot anomaly classification (AC) and segmentation (AS) in industrial vision. We reveal that the abundant normal and abnormal cues implicit in unlabeled test images can be exploited for anomaly determination, which is ignored by prior methods. Our key observation is that for t…

2024

Self-Modifying State Modeling for Simultaneous Machine Translation

ACL 2024long

Simultaneous Machine Translation (SiMT) generates target outputs while receiving stream source inputs and requires a read/write policy to decide whether to wait for the next source token or generate a new target token, whose decisions form a decision path. Existing SiMT methods, which learn the poli…

2024

TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control

NeurIPS 2024spotlight

Centred on content modification and style preservation, Scene Text Editing (STE) remains a challenging task despite considerable progress in text-to-image synthesis and text-driven image manipulation recently. GAN-based STE methods generally encounter a common issue of model generalization, while Di…

2024

Towards More Accurate Diffusion Model Acceleration with A Timestep Tuner

CVPR 2024poster

A diffusion model which is formulated to produce an image using thousands of denoising steps usually suffers from a slow inference speed. Existing acceleration algorithms simplify the sampling by skipping most steps yet exhibit considerable performance degradation. By viewing the generation of diffu…

2024

Vector Quantization Knowledge Transfer for End-to-End Text Image Machine Translation

ICASSP 2024accepted

End-to-end text image machine translation (TIMT) aims at translating source language embedded in images into target language without recognizing intermediate texts in images. However, the data scarcity of end-to-end TIMT task limits the translation performance. Existing research explores aligning co…

Cited by 0SourceScholar
2023

A Novel Heart Rate Estimation Method Exploiting Heartbeat Second Harmonic Reconstruction Via Millimeter Wave Radar

ICASSP 2023accepted

Millimeter wave radar has been extensively exploited in heart rate estimation tasks, but there is still potential for improvement in estimation accuracy. At present, the interference of the second and third harmonics of respiration has become a significant problem that hinders further improvement of…

Cited by 0SourceScholar
2023

An Effective Anomalous Sound Detection Method Based on Representation Learning with Simulated Anomalies

ICASSP 2023accepted

In this paper, we propose an effective anomalous sound detection (ASD) method based on representation learning with simulated anomalies. Recently, ASD systems have used Outlier Exposure (OE) strategy to achieve promising performance in DCASE challenges. These exploit deep Convolutional Neural Networ…

Cited by 0SourceScholar
2023

CCIM: Cross-modal Cross-lingual Interactive Image Translation

EMNLP 2023short findings

Text image machine translation (TIMT) which translates source language text images into target language texts has attracted intensive attention in recent years. Although the end-to-end TIMT model directly generates target translation from encoded text image features with an efficient architecture, i…

Cited by 0SourceScholar
2023

CFSum Coarse-to-Fine Contribution Network for Multimodal Summarization

ACL 2023long

Multimodal summarization usually suffers from the problem that the contribution of the visual modality is unclear. Existing multimodal summarization approaches focus on designing the fusion methods of different modalities, while ignoring the adaptive conditions under which visual modalities are usef…

2023

Divide Rows and Conquer Cells: Towards Structure Recognition for Large Tables

IJCAI 2023poster

Recent advanced Table Structure Recognition (TSR) models adopt image-to-text solutions to parse table structure. These methods can be formulated as image caption problem, i.e., input a single-table image and output table structure description in a specific text format, e.g., HTML. With the impressiv…

Cited by 20SourcePDFScholar
2023

EI2SR: Learning an Enhanced Intra-Instance Semantic Relationship for Arbitrary-Shaped Scene Text Detection

ICASSP 2023accepted

Text detection in natural scenarios, has made significant progress with the deep learning architecture. Towards arbitrary-shaped text detection, fracture detection is the major concern due to the lack of semantic relationship within an instance in existing methods. To circumvent this dilemma, we pro…

Cited by 0SourceScholar
2023

Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection

ICASSP 2023accepted

In this paper, we propose a joint generative and contrastive representation learning method (GeCo) for anomalous sound detection (ASD). GeCo exploits a Predictive AutoEncoder (PAE) equipped with self-attention as a generative model to perform frame-level prediction. The output of the PAE together wi…

Cited by 28SourceScholar
2023

LayoutDIT: Layout-Aware End-to-End Document Image Translation with Multi-Step Conductive Decoder

EMNLP 2023long findings

Document image translation (DIT) aims to translate text embedded in images from one language to another. It is a challenging task that needs to understand visual layout with text semantics simultaneously. However, existing methods struggle to capture the crucial visual layout in real-world complex d…

Cited by 0SourceScholar
2023

Localizing Active Objects from Egocentric Vision with Symbolic World Knowledge

EMNLP 2023long main

The ability to actively ground task instructions from an egocentric view is crucial for AI agents to accomplish tasks or assist humans virtually. One important step towards this goal is to localize and track key active objects that undergo major state change as a consequence of human actions/interac…

Cited by 0SourcecodeScholar
2023

Multilingual Knowledge Graph Completion with Language-Sensitive Multi-Graph Attention

ACL 2023long

Multilingual Knowledge Graph Completion (KGC) aims to predict missing links with multilingual knowledge graphs. However, existing approaches suffer from two main drawbacks: (a) alignment dependency: the multilingual KGC is always realized with joint entity or relation alignment, which introduces add…

Cited by 5SourcePDFScholar
2023

Non-Sequential Graph Script Induction via Multimedia Grounding

ACL 2023long

Online resources such as WikiHow compile a wide range of scripts for performing everyday tasks, which can assist models in learning to reason about procedures. However, the scripts are always presented in a linear manner, which does not reflect the flexibility displayed by people executing tasks in…

2023

One-Shot Replay: Boosting Incremental Object Detection via Retrospecting One Object

AAAI 2023technical

Modern object detectors are ill-equipped to incrementally learn new emerging object classes over time due to the well-known phenomenon of catastrophic forgetting. Due to data privacy or limited storage, few or no images of the old data can be stored for replay. In this paper, we design a novel One-S…

Cited by 7SourcePDFScholar
2023

Stargan-vc Based Cross-Domain Data Augmentation for Speaker Verification

ICASSP 2023accepted

Automatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors, such as recording device and speaking style, in real-world applications, which leads to severe performance degradation. Since single-speaker multi-condition (SSMC) data is difficult to…

Cited by 0SourceScholar
2023

UATVR: Uncertainty-Adaptive Text-Video Retrieval

ICCV 2023poster

With the explosive growth of web videos and emerging large-scale vision-language pre-training models, e.g., CLIP, retrieving videos of interest with text instructions has attracted increasing attention. A common practice is to transfer text-video pairs to the same embedding space and craft cross-mod…

Cited by 64PDFcodeScholar
2022

GitNet: Geometric Prior-Based Transformation for Birds-Eye-View Segmentation

ECCV 2022poster

"Birds-eye-view (BEV) semantic segmentation is critical for autonomous driving for its powerful spatial representation ability. It is challenging to estimate the BEV semantic maps from monocular images due to the spatial gap, since it is implicitly required to realize both the perspective-to-BEV tra…

Cited by 35SourcePDFScholar
2022

Imagine by Reasoning: A Reasoning-Based Implicit Semantic Data Augmentation for Long-Tailed Classification

AAAI 2022technical

Real-world data often follows a long-tailed distribution, which makes the performance of existing classification algorithms degrade heavily. A key issue is that the samples in tail categories fail to depict their intra-class diversity. Humans can imagine a sample in new poses, scenes and view angles…

2022

Other Roles Matter! Enhancing Role-Oriented Dialogue Summarization via Role Interactions

ACL 2022long

Role-oriented dialogue summarization is to generate summaries for different roles in the dialogue, e.g., merchants and consumers. Existing methods handle this task by summarizing each role’s content separately and thus are prone to ignore the information from other roles. However, we believe that ot…

2021

CSDS: A Fine-Grained Chinese Dataset for Customer Service Dialogue Summarization

EMNLP 2021main

Dialogue summarization has drawn much attention recently. Especially in the customer service domain, agents could use dialogue summaries to help boost their works by quickly knowing customer’s issues and service progress. These applications require summaries to contain the perspective of a single sp…

2021

Deep Adversarial Quantization Network for Cross-Modal Retrieval

ICASSP 2021accepted

In this paper, we propose a seamless multimodal binary learning method for cross-modal retrieval. First, we utilize adversarial learning to learn modality-independent representations of different modalities. Second, we formulate loss function through the Bayesian approach, which aims to jointly maxi…

Cited by 0SourceScholar
2021

FC2RN: A Fully Convolutional Corner Refinement Network for Accurate Multi-Oriented Scene Text Detection

ICASSP 2021accepted

Accurate detection of multi-oriented text that accounts for a large proportion in real practice is of great significance. The performance has improved rapidly on common benchmarks in recent years. However, dense long text case and the quality of detection are easy to be overlooked. Direct regression…

Cited by 0SourceScholar
2020

Knowledge Graph Enhanced Neural Machine Translation via Multi-task Learning on Sub-entity Granularity

COLING 2020main

Previous studies combining knowledge graph (KG) with neural machine translation (NMT) have two problems: i) Knowledge under-utilization: they only focus on the entities that appear in both KG and training sentence pairs, making much knowledge in KG unable to be fully utilized. ii) Granularity mismat…

2020

SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition

CVPR 2020poster

Scene text recognition is a hot research topic in computer vision. Recently, many recognition methods based on the encoder-decoder framework have been proposed, and they can handle scene texts of perspective distortion and curve shape. Nevertheless, they still face lots of challenges like image blur…

Cited by 340PDFcodeScholar
2020

Video Playback Rate Perception for Self-Supervised Spatio-Temporal Representation Learning

CVPR 2020poster

In self-supervised spatio-temporal representation learning, the temporal resolution and long-short term characteristics are not yet fully explored, which limits representation capabilities of learned models. In this paper, we propose a novel self-supervised method, referred to as video Playback Rate…

Cited by 210PDFcodeScholar
2019

Occlusion-Shared and Feature-Separated Network for Occlusion Relationship Reasoning

ICCV 2019poster

Occlusion relationship reasoning demands closed contour to express the object, and orientation of each contour pixel to describe the order relationship between objects. Current CNN-based methods neglect two critical issues of the task: (1) simultaneous existence of the relevance and distinction for…

Cited by 33PDFcodeScholar
2018

Historical Data is Useful for Navigation Planning: Data Driven Route Generation for Autonomous Ship

ICRA 2018poster

This work presents a method for automated generation of navigation plan for autonomous or robotic surface vessel. Historical Automatic Identification System (AIS) data is of significant value to this problem. The method joins AIS locations of a same vessel at different time and locations in a region…

Cited by 18SourceScholar
2018

Visual Homing via Guided Locality Preserving Matching

ICRA 2018poster

This study proposes a simple yet surprisingly effective feature matching approach, termed as guided locality preserving matching (GLPM), for visual homing of panoramic images. The key idea of our approach is merely to preserve the neighborhood structures of potential true matches between two panoram…

Cited by 16SourceScholar
2015

A parallel distributed strategy for arraying a scattered robot swarm

IROS 2015poster

We consider the problem of organizing a scattered group of n robots in two-dimensional space. The communication graph of the swarm is connected, but there is no central authority for organizing it. We want to arrange them into a sorted and equally-spaced array between the robots with lowest and high…

Cited by 11SourceScholar