← Search

Yang Zhao

128 accepted papers

2026

CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering

ICLR 2026poster

Knowledge graphs provide structured context for multi‑hop question answering, but deployed systems must balance answer accuracy with strict latency and cost targets while preserving provenance. Static $k$‑hop expansions and ``think‑longer'' prompting often over‑retrieve, inflate context, and yield u…

Cited by 0SourceScholar
2026

Captain Cinema: Towards Short Movie Generation

ICLR 2026poster

We present **Captain Cinema**, a generation framework for short movie generation. Given a detailed textual description of a movie storyline, our approach firstly generates a sequence of keyframes that outline the entire narrative, which ensures long-range coherence in both the storyline and visual a…

Cited by 0SourceScholar
2026

Deep Clustering Based on Sparse Kolmogorov-Arnold Network and Spectral Constraint

AAAI 2026technical

At present, spectral clustering is an important branch of unsupervised learning, and its application in deep learning has been widely concerned. However, for high-dimensional sparse datasets, the complexity of network scale leads to parameter explosion, and static Gaussian kernel often has wrong pre

Cited by 0SourcePDFScholar
2026

Depth Anything 3: Recovering the Visual Space from Any Views

ICLR 2026oral

We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single plain transformer (e.g., vanilla DINOv2 encoder) is sufficient…

Cited by 0SourcecodeScholar
2026

Design of an Active Haptic Interface Using Proprioception Feedback for Continuous Endovascular Teleoperation

RA-L 2026

Force feedback is essential for safe endovascular teleoperation, yet typically constrained by complex sensor integration. This article presents a compact active haptic interface system designed for robotic catheterization. Leveraging the intrinsic proprioception of a Permanent Magnet Synchronous Mot

Cited by 0SourceScholar
2026

Diagnosing and Remedying Knowledge Deficiencies in LLMs via Label-free Curricular Meaningful Learning

ICLR 2026poster

Large Language Models (LLMs) have demonstrated impressive generalization ability by learning from extensive unlabeled text. However, they still exhibit reasoning mistakes, which can affect their trustworthiness and reliability. Although users can interact with LLMs and provide diverse and comprehens…

Cited by 0SourcecodeScholar
2026

Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs

ICLR 2026poster

Model merging plays a crucial role in consolidating multiple specialized models into a single, unified model, especially in the era of large language models (LLMs). Recent research has primarily focused on developing strategies to enhance merging performance with the trained models, while the impact…

Cited by 0SourceScholar
2026

ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning

ICLR 2026poster

The rapid advancement of text-to-image (T2I) models has increased the need for reliable human preference modeling, a demand further amplified by recent progress in reinforcement learning for preference alignment. However, existing approaches typically quantify the quality of a generated image using…

Cited by 0SourceScholar
2026

LSAP-PV: High-Fidelity Palm Vein Image Synthesis via Layered Spectral Absorption Projection-Guided Diffusion Model

AAAI 2026technical

Palm vein recognition has emerged as a promising biometric technology, yet its development remains constrained by the scarcity of large-scale publicly available datasets. Several methods of palm vein image generation have been proposed to address this issue. These methods usually focus on the anatom

Cited by 0SourcePDFScholar
2026

MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer

ICLR 2026poster

Unified multimodal Large Language Models (LLMs) that can both understand and generate visual content hold immense potential. However, existing open-source models often suffer from a performance trade-off between these capabilities. We present Manzano, a simple and scalable unified framework that sub…

Cited by 0SourceScholar
2026

Multi-Dimensional Perturbation Strategies for Adversarial Attacks in Multi-Agent Deep Reinforcement Learning

ICRA 2026poster

Research indicates that single-agent reinforcement learning is vulnerable to adversarial attacks, which can lead to decision-making errors. Similarly, multi-agent deep reinforcement learning (MADRL) systems face analogous adversarial threats. However, existing attack methods require substantial inve…

Cited by 0Scholar
2026

Orchestrating Spatial Semantics via a Zone-Graph Paradigm for Intricate Indoor Scene Generation

ICML 2026poster

Autonomous 3D indoor scene synthesis breaks down in non-convex rooms with tightly coupled spatial constraints. Data-driven generators lack topological priors for long-horizon planning, while iterative agents fragment semantics and become geometrically brittle. We present \textbf{ZoneMaestro}, a unif…

Cited by 0SourceScholar
2026

Robust Vision-Language Models via Manifold-Adversarial Adapters

ICML 2026poster

Vision-language models (VLMs) have progressed rapidly with large-scale high-quality data and adaptation strategies, yet remain brittle under real-world corruptions, where both visual recognition and language-grounded reasoning degrade. Beyond cascaded image restoration, a natural alternative is para…

Cited by 0SourceScholar
2026

SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training

ICLR 2026poster

Recent advances in diffusion-based video restoration (VR) demonstrate significant improvement in visual quality, yet yield a prohibitive computational cost during inference. While several distillation-based approaches have exhibited the potential of one-step image restoration, extending existing app…

Cited by 0SourcecodeScholar
2026

Unified Latent Space for Understanding and Generation via Semantic Auto-encoder

CVPR 2026

Latent generative modeling has emerged as the dominant paradigm for Diffusion Transformers (DiT), where a pretrained autoencoder compresses image pixels into a latent space to facilitate the diffusion process. Recently, the use of semantic encoders within autoencoders (AEs) has gained attention, yet

Cited by 0SourceScholar
2026

VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery

ICLR 2026poster

Vision-Language Models (VLMs) have achieved significant progress in multimodal understanding tasks, demonstrating strong capabilities particularly in general tasks such as image captioning and visual reasoning. However, when dealing with specialized cultural heritage domains like 3D vase artifacts,…

Cited by 0SourcecodeScholar
2025

3DMolFormer: A Dual-channel Framework for Structure-based Drug Discovery

ICLR 2025poster

Structure-based drug discovery, encompassing the tasks of protein-ligand docking and pocket-aware 3D drug design, represents a core challenge in drug discovery. However, no existing work can deal with both tasks to effectively leverage the duality between them, and current methods for each task are…

2025

A Query-Response Framework for Whole-Page Complex-Layout Document Image Translation with Relevant Regional Concentration

ACL 2025finding

Document Image Translation (DIT), which aims at translating documents in images from source language to the target, plays an important role in Document Intelligence. It requires a comprehensive understanding of document multi-modalities and a focused concentration on relevant textual regions during…

Cited by 0SourcePDFScholar
2025

A Simple-Yet-Efficient Instruction Augmentation Method for Zero-Shot Sentiment Classification

COLING 2025main

Instruction tuning significantly enhances the performance of large language models in tasks such as sentiment classification. Previous studies have leveraged labeled instances from sentiment benchmark datasets to instruction-tune LLMs, improving zero-shot sentiment classification performance. In thi…

2025

Analyzing the Rapid Generalization of SFT via the Perspective of Attention Head Activation Patterns

ACL 2025long

LLMs’ performance on complex tasks is still unsatisfactory. A key issue is that presently LLMs learn in a data-driven schema, while the instructions about these complex tasks are both scarce and hard to collect or construct. On the contrary, a prominent phenomenon is that LLMs can learn rather fast…

2025

Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation

NeurIPS 2025poster

Existing large-scale video generation models are computationally intensive, preventing adoption in real-time and interactive applications. In this work, we propose autoregressive adversarial post-training (AAPT) to turn a pre-trained latent video diffusion model into a real-time, interactive, stream…

Cited by 0SourceScholar
2025

Beyond Similarity: A Gradient-based Graph Method for Instruction Tuning Data Selection

ACL 2025long

Large language models (LLMs) have shown great potential across various industries due to their remarkable ability to generalize through instruction tuning. However, the limited availability of domain-specific data significantly hampers their performance on specialized tasks. While existing methods p…

2025

Bias Analysis and Mitigation through Protected Attribute Detection and Regard Classification

EMNLP 2025

Large language models (LLMs) acquire general linguistic knowledge from massive-scale pretraining. However, pretraining data mainly comprised of web-crawled texts contain undesirable social biases which can be perpetuated or even amplified by LLMs. In this study, we propose an efficient yet effective

Cited by 0SourcePDFScholar
2025

Boosting LLM Translation Skills without General Ability Loss via Rationale Distillation

ACL 2025finding

Large Language Models (LLMs) have achieved impressive results across numerous NLP tasks, and fine-tuning them for Machine Translation (MT) has improved their performance. However, vanilla fine-tuning often leads to catastrophic forgetting, compromising the broad general abilities of LLMs and introdu…

2025

Data-Efficiently Learn Large Language Model for Universal 3D Scene Perception

NAACL 2025findings

3D scene understanding has gained significant attention due to its wide range of applications. However, existing methods for 3D scene understanding are limited to specific downstream tasks, which hinders their practicality in real-world applications. This paper presents Chat-3D, which combines the 3…

Cited by 0SourcePDFScholar
2025

Diff-Palm: Realistic Palmprint Generation with Polynomial Creases and Intra-Class Variation Controllable Diffusion Models

CVPR 2025poster

Palmprint recognition is significantly limited by the lack of large-scale publicly available datasets. Previous methods have adopted Bezier curves to simulate the palm creases, which then serve as input for conditional GANs to generate realistic palmprints.However, without employing real data fine-t…

2025

E4: Energy-Efficient DNN Inference for Edge Video Analytics via Early Exiting and DVFS

AAAI 2025technical

Deep neural network (DNN) models are increasingly popular in edge video analytic applications. However, the computeintensive nature of DNN models pose challenges for energyefficient inference on resource-constrained edge devices. Most existing solutions focus on optimizing DNN inference latency and…

Cited by 0SourcePDFScholar
2025

From Chaotic OCR Words to Coherent Document: A Fine-to-Coarse Zoom-Out Network for Complex-Layout Document Image Translation

COLING 2025main

Document Image Translation (DIT) aims to translate documents in images from one language to another. It requires visual layouts and textual contents understanding, as well as document coherence capturing. However, current methods often rely on the quality of OCR output, which, particularly in comple…

2025

Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMs

EMNLP 2025

Multimodal large language models (MLLMs) are increasingly deployed in open-ended, real-world environments where inputs are messy, underspecified, and not always trustworthy. Unlike curated benchmarks, these settings frequently involve instructions that reference missing objects or contradictory fact

Cited by 0SourcePDFScholar
2025

High-Fidelity Stereoscopic Image Rain Removal with Texture Integrity and Disparity Consistency

ICASSP 2025accepted

This paper tackles the challenge of stereoscopic image rain removal by focusing on enhancing texture integrity and disparity consistency. Existing stereoscopic rain removal techniques often fall short due to 1) disruptions in texture coherence caused by complex rain streaks, and 2) inaccuracies in d…

Cited by 0SourceScholar
2025

How Far Is Video Generation from World Model: A Physical Law Perspective

ICML 2025poster

Scaling video generation models is believed to be promising in building world models that adhere to fundamental physical laws. However, whether these models can discover physical laws purely from vision can be questioned. A world model learning the true law should give predictions robust to nuances…

Cited by 35SourcePDFScholar
2025

How to Auto-optimize Prompts for Domain Tasks? Adaptive Prompting and Reasoning through Evolutionary Domain Knowledge Adaptation

NeurIPS 2025poster

Designing optimal prompts and reasoning processes for large language models (LLMs) on domain-specific tasks is both necessary and challenging in real-world applications. Determining how to integrate domain knowledge, enhance reasoning efficiency, and even provide domain experts with refined knowledg…

Cited by 0SourceScholar
2025

Improving MLLM’s Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency

ACL 2025finding

Multimodal Large Language Models (MLLMs) have shown strong performance in document image tasks, especially Optical Character Recognition (OCR). However, they struggle with Document Image Machine Translation (DIMT), which requires handling both cross-modal and cross-lingual challenges. Previous effor…

Cited by 0SourcePDFScholar
2025

MExD: An Expert-Infused Diffusion Model for Whole-Slide Image Classification

CVPR 2025poster

Whole Slide Image (WSI) classification poses unique challenges due to the vast image size and numerous non-informative regions, which introduce noise and cause data imbalance during feature aggregation. To address these issues, we propose MExD, an Expert-Infused Diffusion Model that combines the str…

Cited by 0SourcePDFScholar
2025

MSMAR-RL: Multi-Step Masked-Attention Recovery Reinforcement Learning for Safe Maneuver Decision in High-Speed Pursuit-Evasion Game

IJCAI 2025

Ensuring the safety of high-speed agent in dynamic adversarial environments, such as pursuit-evasion games with target-purchase and obstacle-avoidance, is a significant challenge. Existing reinforcement learning methods often fail to balance safety and reward under strict safety constraints and dive

2025

Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing

AAAI 2025technical

The Audio-Visual Video Parsing task aims to recognize and temporally localize all events occurring in either the audio or visual stream, or both. Capturing accurate event semantics for each audio/visual segment is vital. Prior works directly utilize the extracted holistic audio and visual features f…

2025

Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models

ACL 2025finding

Existing Multimodal Large Language Models (MLLMs) are predominantly trained and tested on consistent visual-textual inputs, leaving open the question of whether they can handle inconsistencies in real-world, layout-rich content. To bridge this gap, we propose the Multimodal Inconsistency Reasoning (…

2025

Occult: Optimizing Collaborative Communications across Experts for Accelerated Parallel MoE Training and Inference

ICML 2025poster

Mixture-of-experts (MoE) architectures could achieve impressive computational efficiency with expert parallelism, which relies heavily on all-to-all communication across devices. Unfortunately, such communication overhead typically constitutes a significant portion of the total runtime, hampering th…

Cited by 0SourcePDFScholar
2025

PVTree: Realistic and Controllable Palm Vein Generation for Recognition Tasks

AAAI 2025technical

Palm vein recognition is an emerging biometric technology that offers enhanced security and privacy. However, acquiring sufficient palm vein data for training deep learning-based recognition models is challenging due to the high costs of data collection and privacy protection constraints. This has l…

2025

Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning

NeurIPS 2025spotlight

Decision-making models for individuals, particularly in high-stakes scenarios like vaccine uptake, often diverge from population optimal predictions. This gap arises from the uniqueness of the individual decision-making process, shaped by numerical attributes (e.g., cost, time) and linguistic influe…

Cited by 0SourcecodeScholar
2025

Plug-and-Play Multi-Domain Fusion Adaptation for Cross-Subject EEG-Based Motor Imagery Classification

ICRA 2025

Motor imagery (MI) classification in rehabilitation brain-computer interfaces (RBCIs) faces significant challenges due to the variability of electroencephalography (EEG) signals across subjects. Existing methods typically require extensive EEG data collection from each new subject, which is time-con

Cited by 1SourceScholar
2025

QiMeng-CodeV-R1: Reasoning-Enhanced Verilog Generation

NeurIPS 2025poster

Large language models (LLMs) trained via reinforcement learning with verifiable reward (RLVR) have achieved breakthroughs on tasks with explicit, automatable verification, such as software programming and mathematical problems. Extending RLVR to electronic design automation (EDA), especially automat…

Cited by 0SourceScholar
2025

RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models

ICML 2025poster

Supervised fine-tuning is a standard method for adapting pre-trained large language models (LLMs) to downstream tasks. Quantization has been recently studied as a post-training technique for efficient LLM deployment. To obtain quantized fine-tuned LLMs, conventional pipelines would first fine-tune t…

2025

SHIFT: Selected Helpful Informative Frame for Video-guided Machine Translation

EMNLP 2025

Video-guided Machine Translation (VMT) aims to improve translation quality by integrating contextual information from paired short video clips. Mainstream VMT approaches typically incorporate multimodal information by uniformly sampling frames from the input videos. However, this paradigm frequently

2025

SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration

CVPR 2025highlight

Video restoration poses non-trivial challenges in maintaining fidelity while recovering temporally consistent details from unknown degradations in the wild. Despite recent advances in diffusion-based restoration, these methods often face limitations in generation capability and sampling efficiency.…

2025

SimulPL: Aligning Human Preferences in Simultaneous Machine Translation

ICLR 2025poster

Simultaneous Machine Translation (SiMT) generates translations while receiving streaming source inputs. This requires the SiMT model to learn a read/write policy, deciding when to translate and when to wait for more source input. Numerous linguistic studies indicate that audiences in SiMT scenarios…

2025

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation

ACL 2025long

Document Image Machine Translation (DIMT) aims to translate text within document images, facing generalization challenges due to limited training data and the complex interplay between visual and textual information. To address these challenges, we introduce M4Doc, a novel single-to-mix Modality ali…

Cited by 0SourcePDFScholar
2025

Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study

NeurIPS 2025spotlight

How to integrate and verify spatial intelligence in foundation models remains an open challenge. Current practice often proxies Visual-Spatial Intelligence (VSI) with purely textual prompts and VQA-style scoring, which obscures geometry, invites linguistic shortcuts, and weakens attribution to genui…

Cited by 0SourceScholar
2025

TriFine: A Large-Scale Dataset of Vision-Audio-Subtitle for Tri-Modal Machine Translation and Benchmark with Fine-Grained Annotated Tags

COLING 2025main

Current video-guided machine translation (VMT) approaches primarily use coarse-grained visual information, resulting in information redundancy, high computational overhead, and neglect of audio content. Our research demonstrates the significance of fine-grained visual and audio information in VMT fr…

2025

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection

NeurIPS 2025poster

A primary impediment to scaling reinforcement learning (RL) for large language model (LLM) training is the substantial computational cost, predominantly arising from the necessity of multi-sampling for policy optimization and evaluation. This underscores the critical yet challenging nature of effici…

Cited by 0SourceScholar
2025

Unified Adversarial Augmentation for Improving Palmprint Recognition

ICCV 2025poster

Current palmprint recognition models achieve strong performance on constrained datasets, yet exhibit significant limitations in handling challenging palmprint samples with geometric distortions and textural degradations. Data augmentation is widely adopted to improve model generalization. However, e…

2025

VideoAuteur: Towards Long Narrative Video Generation

ICCV 2025poster

Recent video generation models have shown promising results in producing high-quality video clips lasting several seconds. However, these models face challenges in generating long sequences that convey clear and informative events, limiting their ability to support coherent narrations. In this paper…

Cited by 0SourcePDFScholar
2025

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

NeurIPS 2025poster

This paper presents a multimodal framework that attempts to unify visual understanding and generation within a shared discrete semantic representation. At its core is the Text-Aligned Tokenizer (TA-Tok), which converts images into discrete tokens using a text-aligned codebook projected from a large…

Cited by 0SourceScholar
2024

An Optimization-Based Planner with B-spline Parameterized Continuous-Time Reference Signals

IROS 2024poster

For the cascaded planning and control modules implemented for robot navigation, the frequency gap between the planner and controller has received limited attention. In this study, we introduce a novel B-spline parameterized optimization-based planner (BSPOP) designed to address the frequency gap cha…

Cited by 1SourceScholar
2024

Born a BabyNet with Hierarchical Parental Supervision for End-to-End Text Image Machine Translation

COLING 2024main

Text image machine translation (TIMT) aims at translating source language texts in images into another target language, which has been proven successful by bridging text image recognition encoder and text translation decoder. However, it is still an open question of how to incorporate fine-grained k…

2024

Causal-Guided Active Learning for Debiasing Large Language Models

ACL 2024long

Although achieving promising performance, recent analyses show that current generative large language models (LLMs) may still capture dataset biases and utilize them for generation, leading to poor generalizability and harmfulness of LLMs. However, due to the diversity of dataset biases and the over…

2024

Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

NeurIPS 2024poster

Recent advancements in 3D Large Language Models (LLMs) have demonstrated promising capabilities for 3D scene understanding. However, previous methods exhibit deficiencies in general referencing and grounding capabilities for intricate scene comprehension. In this paper, we introduce the use of objec…

2024

Deciphering the Impact of Pretraining Data on Large Language Models through Machine Unlearning

ACL 2024findings

Through pretraining on a corpus with various sources, Large Language Models (LLMs) have gained impressive performance. However, the impact of each component of the pretraining corpus remains opaque. As a result, the organization of the pretraining corpus is still empirical and may deviate from the o…

2024

Deep Video Inverse Tone Mapping Based on Temporal Clues

CVPR 2024poster

Inverse tone mapping (ITM) aims to reconstruct high dynamic range (HDR) radiance from low dynamic range (LDR) content. Although many deep image ITM methods can generate impressive results the field of video ITM is still to be explored. Processing video sequences by image ITM methods may cause tempor…

2024

Document Image Machine Translation with Dynamic Multi-pre-trained Models Assembling

NAACL 2024long

Text image machine translation (TIMT) is a task that translates source texts embedded in the image to target translations. The existing TIMT task mainly focuses on text-line-level images. In this paper, we extend the current TIMT task and propose a novel task, **D**ocument **I**mage **M**achine **T*…

2024

Extending Multi-modal Contrastive Representations

NeurIPS 2024poster

Multi-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high dependence on large-scale, high-quality paired data and the expensive training costs limit their further development. Ins…

2024

FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion

ICML 2024poster

Unified multi-model representation spaces are the foundation of multimodal understanding and generation. However, the billions of model parameters and catastrophic forgetting problems make it challenging to further enhance pre-trained unified spaces. In this work, we propose FreeBind, an idea that t…

2024

Image Understanding Makes for A Good Tokenizer for Image Generation

NeurIPS 2024poster

Modern image generation (IG) models have been shown to capture rich semantics valuable for image understanding (IU) tasks. However, the potential of IU models to improve IG performance remains uncharted. We address this issue using a token-based IG framework, which relies on effective tokenizers to…

2024

Improving Subject-Driven Image Synthesis with Subject-Agnostic Guidance

CVPR 2024poster

In subject-driven text-to-image synthesis the synthesis process tends to be heavily influenced by the reference images provided by users often overlooking crucial attributes detailed in the text prompt. In this work we propose Subject-Agnostic Guidance (SAG) a simple yet effective solution to remedy…

Cited by 2SourcePDFScholar
2024

Incorporating Syntax and Lexical Knowledge to Multilingual Sentiment Classification on Large Language Models

ACL 2024findings

This paper exploits a sentiment extractor supported by syntactic and lexical resources to enhance multilingual sentiment classification solved through the generative approach, without retraining LLMs. By adding external information of words and phrases that have positive/negative polarities, the mul…

Cited by 4SourcePDFScholar
2024

Instruct-Imagen: Image Generation with Multi-modal Instruction

CVPR 2024poster

This paper presents Instruct-Imagen a model that tackles heterogeneous image generation tasks and generalizes across unseen tasks. We introduce multi-modal instruction for image generation a task representation articulating a range of generation intents with precision. It uses natural language to am…

Cited by 42SourcePDFScholar
2024

Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA

ACL 2024long

Multipanel images, commonly seen as web screenshots, posters, etc., pervade our daily lives. These images, characterized by their composition of multiple subfigures in distinct layouts, effectively convey information to people. Toward building advanced multimodal AI applications, such as agents that…

Cited by 20SourcePDFScholar
2024

OHTA: One-shot Hand Avatar via Data-driven Implicit Priors

CVPR 2024poster

In this paper we delve into the creation of one-shot hand avatars attaining high-fidelity and drivable hand representations swiftly from a single image. With the burgeoning domains of the digital human the need for quick and personalized hand avatar creation has become increasingly critical. Existin…

Cited by 9SourcePDFScholar
2024

PCE-Palm: Palm Crease Energy Based Two-Stage Realistic Pseudo-Palmprint Generation

AAAI 2024technical

The lack of large-scale data seriously hinders the development of palmprint recognition. Recent approaches address this issue by generating large-scale realistic pseudo palmprints from Bézier curves. However, the significant difference between Bézier curves and real palmprints limits their effective…

Cited by 8SourcePDFScholar
2024

Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding

EMNLP 2024main

Graphical User Interfaces (GUIs) are central to our interaction with digital devices and growing efforts have been made to build models for various GUI understanding tasks. However, these efforts largely overlook an important GUI-referring task: screen reading based on user-indicated points, which w…

2024

Stereo Vision Conversion from Planar Videos Based on Temporal Multiplane Images

AAAI 2024technical

With the rapid development of 3D movie and light-field displays, there is a growing demand for stereo videos. However, generating high-quality stereo videos from planar videos remains a challenging task. Traditional depth-image-based rendering techniques struggle to effectively handle the problem of…

2024

UFOGen: You Forward Once Large Scale Text-to-Image Generation via Diffusion GANs

CVPR 2024highlight

Text-to-image diffusion models have demonstrated remarkable capabilities in transforming text prompts into coherent images yet the computational cost of the multi-step inference remains a persistent challenge. To address this issue we present UFOGen a novel generative model designed for ultra-fast o…

2024

Vector Quantization Knowledge Transfer for End-to-End Text Image Machine Translation

ICASSP 2024accepted

End-to-end text image machine translation (TIMT) aims at translating source language embedded in images into target language without recognizing intermediate texts in images. However, the data scarcity of end-to-end TIMT task limits the translation performance. Existing research explores aligning co…

Cited by 0SourceScholar
2023

3DRP-Net: 3D Relative Position-aware Network for 3D Visual Grounding

EMNLP 2023long main

3D visual grounding aims to localize the target object in a 3D point cloud by a free-form language description. Typically, the sentences describing the target object tend to provide information about its relative relation between other objects and its position within the whole scene. In this work, w…

Cited by 0SourceScholar
2023

A Simple Yet Strong Domain-Agnostic De-bias Method for Zero-Shot Sentiment Classification

ACL 2023findings

Zero-shot prompt-based learning has made much progress in sentiment analysis, and considerable effort has been dedicated to designing high-performing prompt templates. However, two problems exist; First, large language models are often biased to their pre-training data, leading to poor performance i…

Cited by 7SourcePDFScholar
2023

CCIM: Cross-modal Cross-lingual Interactive Image Translation

EMNLP 2023short findings

Text image machine translation (TIMT) which translates source language text images into target language texts has attracted intensive attention in recent years. Although the end-to-end TIMT model directly generates target translation from encoded text image features with an efficient architecture, i…

Cited by 0SourceScholar
2023

Connecting Multi-modal Contrastive Representations

NeurIPS 2023poster

Multi-modal Contrastive Representation (MCR) learning aims to encode different modalities into a semantically aligned shared space. This paradigm shows remarkable generalization ability on numerous downstream tasks across various modalities. However, the reliance on massive high-quality data pairs l…

2023

CoopInit: Initializing Generative Adversarial Networks via Cooperative Learning

AAAI 2023technical

Numerous research efforts have been made to stabilize the training of the Generative Adversarial Networks (GANs), such as through regularization and architecture design. However, we identify the instability can also arise from the fragile balance at the early stage of adversarial learning. This pape…

Cited by 4SourcePDFScholar
2023

DATE: Domain Adaptive Product Seeker for E-Commerce

CVPR 2023poster

Product Retrieval (PR) and Grounding (PG), aiming to seek image and object-level products respectively according to a textual query, have attracted great interest recently for better shopping experience. Owing to the lack of relevant datasets, we collect two large-scale benchmark datasets from Taoba…

2023

De novo Drug Design using Reinforcement Learning with Multiple GPT Agents

NeurIPS 2023poster

*De novo* drug design is a pivotal issue in pharmacology and a new area of focus in AI for science research. A central challenge in this field is to generate molecules with specific properties while also producing a wide range of diverse candidates. Although advanced technologies such as transformer…

2023

Distilling Coarse-to-Fine Semantic Matching Knowledge for Weakly Supervised 3D Visual Grounding

ICCV 2023poster

3D visual grounding involves finding a target object in a 3D scene that corresponds to a given sentence query. Although many approaches have been proposed and achieved impressive performance, they all require dense object-sentence pair annotations in 3D point clouds, which are both time-consuming an…

Cited by 19PDFcodeScholar
2023

LayoutDIT: Layout-Aware End-to-End Document Image Translation with Multi-Step Conductive Decoder

EMNLP 2023long findings

Document image translation (DIT) aims to translate text embedded in images from one language to another. It is a challenging task that needs to understand visual layout with text semantics simultaneously. However, existing methods struggle to capture the crucial visual layout in real-world complex d…

Cited by 0SourceScholar
2023

Multilingual Knowledge Graph Completion with Language-Sensitive Multi-Graph Attention

ACL 2023long

Multilingual Knowledge Graph Completion (KGC) aims to predict missing links with multilingual knowledge graphs. However, existing approaches suffer from two main drawbacks: (a) alignment dependency: the multilingual KGC is always realized with joint entity or relation alignment, which introduces add…

Cited by 5SourcePDFScholar
2023

RPG-Palm: Realistic Pseudo-data Generation for Palmprint Recognition

ICCV 2023poster

Palmprint recently shows great potential in recognition applications as it is a privacy-friendly and stable biometric. However, the lack of large-scale public palmprint datasets limits further research and development of palmprint recognition. In this paper, we propose a novel realistic pseudo-palmp…

Cited by 12PDFScholar
2023

Scene-robust Natural Language Video Localization via Learning Domain-invariant Representations

ACL 2023findings

Natural language video localization(NLVL) task involves the semantic matching of a text query with a moment from an untrimmed video. Previous methods primarily focus on improving performance with the assumption of independently identical data distribution while ignoring the out-of-distribution data.…

Cited by 6SourcePDFScholar
2023

Towards Authentic Face Restoration with Iterative Diffusion Models and Beyond

ICCV 2023poster

An authentic face restoration system is becoming increasingly demanding in many computer vision applications, e.g., image enhancement, video communication, and taking portrait. Most of the advanced face restoration models can recover high-quality faces from low-quality ones but usually fail to faith…

Cited by 17PDFcodeScholar
2023

Towards Informative Open-ended Text Generation with Dynamic Knowledge Triples

EMNLP 2023long findings

Pretrained language models (PLMs), especially large language models (LLMs) demonstrate impressive capabilities in open-ended text generation. While our statistical results show that LLMs often suffer from over-concentrated information, where the generated texts overly focus on the given prompt and f…

Cited by 0SourceScholar
2022

A Versatile Adaptive Curriculum Learning Framework for Task-oriented Dialogue Policy Learning

NAACL 2022findings

Training a deep reinforcement learning-based dialogue policy with brute-force random sampling is costly. A new training paradigm was proposed to improve learning performance and efficiency by combining curriculum learning. However, attempts in the field of dialogue policy are very limited due to the…

Cited by 4SourcePDFScholar
2022

Attention-Based Deep Driving Model for Autonomous Vehicles with Surround-View Cameras

IROS 2022poster

Experienced human drivers always make safe driving decisions by selectively observing the front, rear and side- view mirrors. Several end - to-end methods have been pro-posed to learn driving models with multi-view visual infor-mation. However, these benchmark methods lack semantic understanding of…

Cited by 0SourceScholar
2022

MEJIGCLU: More Effective Jigsaw Clustering For Unsupervised Visual Representation Learning

ICASSP 2022accepted

Unsupervised visual representation learning aims to learn general features from unlabelled data. Early methods design intra-image pretext tasks as learning targets and can be achieved with low computational overhead but unsatisfactory performance. Recent methods introduce contrastive learning and ac…

Cited by 0SourceScholar
2022

Penalizing Gradient Norm for Efficiently Improving Generalization in Deep Learning

ICML 2022spotlight

How to train deep neural networks (DNNs) to generalize well is a central concern in deep learning, especially for severely overparameterized networks nowadays. In this paper, we propose an effective method to improve the model generalization by additionally penalizing the gradient norm of loss funct…

2022

Towards Effective Multi-Modal Interchanges in Zero-Resource Sounding Object Localization

NeurIPS 2022accept

Aiming to locate the object that emits a specified sound in complex scenes, the task of sounding object localization bridges two perception-oriented modalities of vision and acoustics, and brings enormous research value to the comprehensive perceptual understanding of machine intelligence. Although…

Cited by 8SourcePDFScholar
2021

Benchmark Platform for Ultra-Fine-Grained Visual Categorization Beyond Human Performance

ICCV 2021poster

Deep learning methods have achieved remarkable success in fine-grained visual categorization. Such successful categorization at sub-ordinate level, e.g., different animal or plant species, however relies heavily on the visual differences that human can observe and the ground-truths are labelled on t…

Cited by 37PDFcodeScholar
2021

HW-NAS-Bench: Hardware-Aware Neural Architecture Search Benchmark

ICLR 2021spotlight

HardWare-aware Neural Architecture Search (HW-NAS) has recently gained tremendous attention by automating the design of deep neural networks deployed in more resource-constrained daily life devices. Despite its promising performance, developing optimal HW-NAS solutions can be prohibitively challengi…

2021

Learning Energy-Based Generative Models via Coarse-to-Fine Expanding and Sampling

ICLR 2021poster

Energy-based models (EBMs) parameterized by neural networks can be trained by the Markov chain Monte Carlo (MCMC) sampling-based maximum likelihood estimation. Despite the recent significant success of EBMs in image generation, the current approaches to train EBMs are unstable and have difficulty sy…

Cited by 50SourcePDFScholar
2021

Synchronous Interactive Decoding for Multilingual Neural Machine Translation

AAAI 2021technical

To simultaneously translate a source language into multiple different target languages is one of the most common scenarios of multilingual translation. However, existing methods cannot make full use of translation model information during decoding, such as intra-lingual and inter-lingual future info…

2020

A Bottom-up Framework for Construction of Structured Semantic 3D Scene Graph

IROS 2020poster

For high-level human-robot interaction tasks, 3D scene understanding is important and non-trivial for autonomous robots. However, parsing and utilizing effective environment information of the 3D scene is not trivial due to the complexity of the 3D environment and the limited ability for reasoning a…

Cited by 7SourceScholar
2020

A Flexible Recurrent Residual Pyramid Network for Video Frame Interpolation

ECCV 2020poster

Video frame interpolation (VFI) aims at synthesizing new video frames in-between existing frames to generate smoother high frame rate videos. Current methods usually use the fixed pre-trained networks to generate interpolated-frames for different resolutions and scenes. However, the fixed pre-traine…

Cited by 46SourcePDFScholar
2020

Bayesian Meta Sampling for Fast Uncertainty Adaptation

ICLR 2020poster

Meta learning has been making impressive progress for fast model adaptation. However, limited work has been done on learning fast uncertainty adaption for Bayesian modeling. In this paper, we propose to achieve the goal by placing meta learning on the space of probability measures, inducing the conc…

Cited by 25SourcecodeScholar
2020

DNN-Chip Predictor: An Analytical Performance Predictor for DNN Accelerators with Various Dataflows and Hardware Architectures

ICASSP 2020accepted

The recent breakthroughs in deep neural networks (DNNs) have spurred a tremendously increased demand for DNN accelerators. However, designing DNN accelerators is non-trivial as it often takes months/years and requires cross-disciplinary knowledge. To enable fast and effective DNN accelerator develop…

Cited by 0SourceScholar
2020

Deconstruct to Reconstruct a Configurable Evaluation Metric for Open-Domain Dialogue Systems

COLING 2020main

Many automatic evaluation metrics have been proposed to score the overall quality of a response in open-domain dialogue. Generally, the overall quality is comprised of various aspects, such as relevancy, specificity, and empathy, and the importance of each aspect differs according to the task. For i…

2020

Feature Quantization Improves GAN Training

ICML 2020poster

The instability in GANs’ training has been a long-standing problem despite remarkable research efforts. We identify that instability issues stem from difficulties of performing feature matching with mini-batch statistics, due to a fragile balance between the fixed target distribution and the progres…

2020

FracTrain: Fractionally Squeezing Bit Savings Both Temporally and Spatially for Efficient DNN Training

NeurIPS 2020poster

Recent breakthroughs in deep neural networks (DNNs) have fueled a tremendous demand for intelligent edge devices featuring on-site learning, while the practical realization of such systems remains a challenge due to the limited resources available at the edge and the required massive training costs…

2020

Knowledge Graph Enhanced Neural Machine Translation via Multi-task Learning on Sub-entity Granularity

COLING 2020main

Previous studies combining knowledge graph (KG) with neural machine translation (NMT) have two problems: i) Knowledge under-utilization: they only focus on the entities that appear in both KG and training sentence pairs, making much knowledge in KG unable to be fully utilized. ii) Granularity mismat…

2020

Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences

CVPR 2020poster

In this paper, we consider a novel task, Spatio-Temporal Video Grounding for Multi-Form Sentences (STVG). Given an untrimmed video and a declarative/interrogative sentence depicting an object, STVG aims to localize the spatio-temporal tube of the queried object. STVG has two challenging settings: (1…

Cited by 134PDFcodeScholar
2019

E2-Train: Training State-of-the-art CNNs with Over 80% Energy Savings

NeurIPS 2019poster

Convolutional neural networks (CNNs) have been increasingly deployed to edge devices. Hence, many efforts have been made towards efficient CNN inference on resource-constrained platforms. This paper attempts to explore an orthogonal direction: how to conduct more energy-efficient training of CNNs, s…

Cited by 106SourcePDFScholar
2019

End-to-End Driving Model for Steering Control of Autonomous Vehicles with Future Spatiotemporal Features

IROS 2019poster

End-to-end deep learning has gained considerable interests in autonomous driving vehicles in both academic and industrial fields, especially in decision making process. One critical issue in decision making process of autonomous driving vehicles is steering control. Researchers has already trained d…

Cited by 43SourceScholar
2018

Multispectral Image Intrinsic Decomposition via Subspace Constraint

CVPR 2018poster

Multispectral images contain many clues of surface characteristics of the objects, thus can be used in many computer vision tasks, e.g., recolorization and segmentation. However, due to the complex geometry structure of natural scenes, the spectra curves of the same surface can look very different u…

Cited by 13SourcePDFScholar
2018

Speaker-Invariant Training Via Adversarial Learning

ICASSP 2018accepted

We propose a novel adversarial multi-task learning scheme, aiming at actively curtailing the inter-talker feature variability while maximizing its senone discriminability so as to enhance the performance of a deep neural network (DNN) based ASR system. We call the scheme speaker-invariant training (…

Cited by 0SourceScholar
2017

Automatic Spatially-Aware Fashion Concept Discovery

ICCV 2017poster

This paper proposes an automatic spatially-aware concept discovery approach using weakly labeled image-text data from shopping websites. We first fine-tune GoogleNet by jointly modeling clothing images and their corresponding descriptions in a visual-semantic embedding space. Then, for each attribut…

Cited by 310PDFScholar