← Search

Chao Yang

82 accepted papers

2026

AR-Nav Benchmark: Augmented Reality Navigation with Vision and Language

AAAI 2026technical

Augmented Reality (AR) navigation has emerged as a transformative tool for spatial intelligence, enabling users to interactively explore complex environments through wearable and mobile AR devices. However, current AR navigation systems struggle with low indoor localization accuracy, weak semantic u

Cited by 0SourcePDFScholar
2026

Advancing LLM Reasoning with Natural Language and Numerical Feedback

ICML 2026spotlight

Recent advances in reinforcement learning (RL) using numerical rewards have significantly enhanced the complex reasoning capabilities of large language models (LLMs). However, we identify three fundamental limitations of purely numerical feedback: performance plateaus, ineffective spontaneous self-r…

Cited by 0SourceScholar
2026

HDW-SR: High-Frequency Guided Diffusion Model based on Wavelet Decomposition for Image Super-Resolution

CVPR 2026

Diffusion-based methods have shown great promise in single image super-resolution (SISR); however, existing approaches often produce blurred fine details due to insufficient guidance in the high-frequency domain. To address this issue, we propose a High-Frequency Guided Diffusion Network based on Wa

Cited by 0SourcecodeScholar
2026

MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers

ICLR 2026poster

Large language models (LLMs) are evolving into agentic systems that reason, plan, and operate external tools. The Model Context Protocol (MCP) is a key enabler of this transition, offering a standardized interface for connecting LLMs with heterogeneous tools and services. Yet MCP's openness and mult…

Cited by 0SourcecodeScholar
2026

Native Reasoning Models: Training Language Models to Reason on Unverifiable Data

ICLR 2026poster

The dominant paradigm for training large reasoning models—combining Supervised Fine-Tuning (SFT) with Reinforcement Learning with Verifiable Rewards (RLVR)—is fundamentally constrained by its reliance on high-quality, human-annotated reasoning data and external verifiers. This dependency incurs sign…

Cited by 0SourceScholar
2026

Reflector: Internalizing Step-wise Reflection against Indirect Jailbreaks

ICML 2026poster

While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sophisticated, multi-step jailbreak attacks that circumvent conventional surface-level safety alignment by exploiting the internal generation process. To address these vulnerabilities, we propose Refle…

Cited by 0SourceScholar
2025

Adversarial Preference Learning for Robust LLM Alignment

ACL 2025finding

Modern language models often rely on Reinforcement Learning from Human Feedback (RLHF) to encourage safe behaviors. However, they remain vulnerable to adversarial attacks due to three key limitations: (1) the inefficiency and high cost of human annotation, (2) the vast diversity of potential adversa…

2025

BESTAnP: Bi-Step Efficient and Statistically Optimal Estimator for Acoustic-n-Point Problem

RA-L 2025

We consider the acoustic-n-point (AnP) problem, which estimates the pose of a 2D forward-looking sonar (FLS) according to <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$n$</tex-math></inline-formula> 3D-2D point c

Cited by 1SourcecodeScholar
2025

Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition

ICASSP 2025accepted

Recently, end-to-end automatic speech recognition has become the mainstream approach in both industry and academia. To optimize system performance in specific scenarios, the Weighted Finite-State Transducer (WFST) is extensively used to integrate acoustic and language models, leveraging its capacity…

Cited by 0SourceScholar
2025

C-3PO: Compact Plug-and-Play Proxy Optimization to Achieve Human-like Retrieval-Augmented Generation

ICML 2025poster

Retrieval-augmented generation (RAG) systems face a fundamental challenge in aligning independently developed retrievers and large language models (LLMs). Existing approaches typically involve modifying either component or introducing simple intermediate modules, resulting in practical limitations a…

Cited by 1SourcePDFScholar
2025

CMFS: CLIP-Guided Modality Interaction for Mitigating Noise in Multi-Modal Image Fusion and Segmentation

IJCAI 2025

Infrared-visible image fusion and semantic segmentation are pivotal tasks for robust scene understanding under challenging conditions such as low light. However, existing methods often struggle with high noise, modality inconsistencies, and inefficient cross-modal interactions, limiting fusion quali

Cited by 0SourcePDFScholar
2025

Dynamic Spectral Graph Anomaly Detection

AAAI 2025technical

Graph anomaly detection is crucial for identifying anomalous nodes within graphs and addressing applications like financial fraud detection and social spam detection. Recent spectral graph neural network methods advance graph anomaly detection by focusing on anomalies that notably affect the distrib…

2025

Evolving Minds: Logic-Informed Inference from Temporal Action Patterns

ICML 2025poster

Understanding human mental states—such as intentions and desires—is crucial for natural AI-human collaboration. However, this is challenging because human actions occur irregularly over time, and the underlying mental states that drive these actions are unobserved. To tackle this, we propose a novel…

Cited by 0SourcePDFScholar
2025

Fuz-RL: A Fuzzy-Guided Robust Framework for Safe Reinforcement Learning under Uncertainty

NeurIPS 2025poster

Safe Reinforcement Learning (RL) is crucial for achieving high performance while ensuring safety in real-world applications. However, the complex interplay of multiple uncertainty sources in real environments poses significant challenges for interpretable risk assessment and robust decision-making.…

Cited by 0SourceScholar
2025

Improving Open-Ended Referring Expression Comprehension via Dual-Language Constraints

ICASSP 2025accepted

Open-ended referring expression comprehension focuses on locating the text query within an image via scene knowledge, requiring complex reasoning across the triplet of the image, scene knowledge, and the text query. However, most existing methods struggle to integrate scene knowledge while performin…

Cited by 0SourceScholar
2025

Instant GaussianImage: A Generalizable and Self-Adaptive Image Representation via 2D Gaussian Splatting

ICCV 2025poster

Implicit Neural Representation (INR) has demonstrated remarkable advances in the field of image representation but demands substantial GPU resources. GaussianImage recently pioneered the use of Gaussian Splatting to mitigate this cost, however, the slow training process limits its practicality, and…

2025

Point Cloud Upsampling Using Conditional Diffusion Module with Adaptive Noise Suppression

CVPR 2025poster

Point cloud upsampling can improve the quality of the initial point cloud, significantly enhancing the performance of downstream tasks such as classification and segmentation. Existing methods mostly focus on generating the geometric details of point clouds, neglecting noise suppression. To address…

2025

SolidGeo: Measuring Multimodal Spatial Math Reasoning in Solid Geometry

NeurIPS 2025poster

Geometry is a fundamental branch of mathematics and plays a crucial role in evaluating the reasoning capabilities of multimodal large language models (MLLMs). However, existing multimodal mathematics benchmarks mainly focus on plane geometry and largely ignore solid geometry, which requires spatial…

Cited by 0SourceScholar
2025

SrSv: Integrating Sequential Rollouts with Sequential Value Estimation for Multi-agent Reinforcement Learning

AAAI 2025technical

Although multi-agent reinforcement learning (MARL) has shown its success across diverse domains, extending its application to large-scale real-world systems still faces significant challenges. Primarily, the high complexity of real-world environments exacerbates the credit assignment problem, substa…

Cited by 0SourcePDFScholar
2025

Subdomain Uncertainty Optimization for Cross-Speed Fault Diagnosis

ICASSP 2025accepted

Cross-speed bearing fault diagnosis based on unsupervised domain adaptation can handle data distribution differences across various operating speeds, supporting intelligent maintenance of equipment like wind turbines with variable operating speeds. Existing methods focus on aligning sample distribut…

Cited by 0SourceScholar
2025

Think Twice, Act Once: A Co-Evolution Framework of LLM and RL for Large-Scale Decision Making

ICML 2025poster

Recent advancements in Large Language Models (LLMs) and Reinforcement Learning (RL) have shown significant promise in decision-making tasks. Nevertheless, for large-scale industrial decision problems, both approaches face distinct challenges: LLMs lack real-time long-sequence decision-making capabil…

Cited by 0SourcePDFScholar
2025

TreeEval: Benchmark-Free Evaluation of Large Language Models through Tree Planning

AAAI 2025technical

Recently, numerous new benchmarks have been established to evaluate the performance of large language models (LLMs) via either computing a holistic score or employing another LLM as a judge. However, these approaches suffer from data leakage due to the open access of the benchmark and inflexible ev…

2024

Adversarial Adaptive Sampling: Unify PINN and Optimal Transport for the Approximation of PDEs

ICLR 2024poster

Solving partial differential equations (PDEs) is a central task in scientific computing. Recently, neural network approximation of PDEs has received increasing attention due to its flexible meshless discretization and its potential for high-dimensional problems. One fundamental numerical difficulty…

Cited by 15SourcePDFScholar
2024

Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

NAACL 2024long

Large Language Models (LLMs) are now commonplace in conversation applications. However, their risks of misuse for generating harmful responses have raised serious societal concerns and spurred recent research on LLM conversation safety. Therefore, in this survey, we provide a comprehensive overview…

2024

Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization

ACL 2024findings

A single language model, even when aligned with labelers through reinforcement learning from human feedback (RLHF), may not suit all human preferences. Recent approaches therefore prefer customization, gathering multi-dimensional feedback, and creating distinct reward models for each dimension.Diffe…

2024

Critic-Guided Decision Transformer for Offline Reinforcement Learning

AAAI 2024technical

Recent advancements in offline reinforcement learning (RL) have underscored the capabilities of Return-Conditioned Supervised Learning (RCSL), a paradigm that learns the action distribution based on target returns for each state in a supervised manner. However, prevailing RCSL methods largely focus…

2024

Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!

ACL 2024long

Large language models (LLMs) undergo safety alignment to ensure safe conversations with humans. However, this paper introduces a training-free attack method capable of reversing safety alignment, converting the outcomes of stronger alignment into greater potential for harm by accessing only LLM outp…

2024

GLinSAT: The General Linear Satisfiability Neural Network Layer By Accelerated Gradient Descent

NeurIPS 2024poster

Ensuring that the outputs of neural networks satisfy specific constraints is crucial for applying neural networks to real-life decision-making problems. In this paper, we consider making a batch of neural network outputs satisfy bounded and general linear constraints. We first reformulate the neural…

2024

GaussianGrasper: 3D Language Gaussian Splatting for Open-Vocabulary Robotic Grasping

RA-L 2024

Constructing a 3D scene capable of accommodating open-ended language queries, is a pivotal pursuit in the domain of robotics, which facilitates robots in executing object manipulations based on human language directives. To achieve this, some research efforts have been dedicated to the development o

Cited by 102SourcecodeScholar
2024

Inference-Time Language Model Alignment via Integrated Value Guidance

EMNLP 2024finding

Large language models are typically fine-tuned to align with human preferences, but tuning large models is computationally intensive and complex. In this work, we introduce **Integrated Value Guidance (IVG)**, a method that uses implicit and explicit value functions to guide language model decoding…

Cited by 6SourcePDFScholar
2024

Joint-Semantics Multi-Similarity Hashing for Cross-Modal Retrieval

ICASSP 2024accepted

Recently, cross-modal hashing has attracted much attention in large-scale image retrieval scenarios. However, most existing methods ignore the potential higher-order relationships and label semantic information between heterogeneous modality data. Besides, the imbalanced training samples could bias…

Cited by 0SourceScholar
2024

LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction

CVPR 2024poster

Existing methods to fine-tune LLMs like Adapter Prefix-tuning and LoRA which introduce extra modules or additional input sequences to inject new skills or knowledge may compromise the innate abilities of LLMs. In this paper we propose LLaMA-Excitor a lightweight method that stimulates the LLMs' pote…

Cited by 3SourcePDFScholar
2024

Language-aware Visual Semantic Distillation for Video Question Answering

CVPR 2024poster

Significant advancements in video question answering (VideoQA) have been made thanks to thriving large image-language pretraining frameworks. Although these image-language models can efficiently represent both video and language branches they typically employ a goal-free vision perception process an…

Cited by 3SourcePDFScholar
2024

Latent Logic Tree Extraction for Event Sequence Explanation from LLMs

ICML 2024poster

Modern high-stakes systems, such as healthcare or robotics, often generate vast streaming event sequences. Our goal is to design an efficient, plug-and-play tool to elicit logic tree-based explanations from Large Language Models (LLMs) to provide customized insights into each observed event sequence…

Cited by 5SourcePDFScholar
2024

MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models

ECCV 2024poster

"redWarning: This paper contains examples of harmful language and images, and reader discretion is recommended. The security concerns surrounding Large Language Models (LLMs) have been extensively explored, yet the safety of Multimodal Large Language Models (MLLMs) remains understudied. In this pape…

2024

RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis

ICML 2024poster

Robotic behavior synthesis, the problem of understanding multimodal inputs and generating precise physical control for robots, is an important part of Embodied AI. Despite successes in applying multimodal large language models for high-level understanding, it remains challenging to translate these c…

Cited by 18SourcePDFScholar
2024

SEER: Facilitating Structured Reasoning and Explanation via Reinforcement Learning

ACL 2024long

Elucidating the reasoning process with structured explanations from question to answer is crucial, as it significantly enhances the interpretability, traceability, and trustworthiness of question-answering (QA) systems. However, structured explanations demand models to perform intricately structured…

2024

Safety of Multimodal Large Language Models on Images and Text

IJCAI 2024poster

Attracted by the impressive power of Multimodal Large Language Models (MLLMs), the public is increasingly utilizing them to improve the efficiency of daily work. Nonetheless, the vulnerabilities of MLLMs to unsafe instructions bring huge safety risks when these models are deployed in real-world scen…

2024

Unveiling Latent Causal Rules: A Temporal Point Process Approach for Abnormal Event Explanation

AISTATS 2024poster

In high-stakes systems such as healthcare, it is critical to understand the causal reasons behind unusual events, such as sudden changes in patient’s health. Unveiling the causal reasons helps with quick diagnoses and precise treatment planning. In this paper, we propose an automated method for unco…

2024

Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models

NeurIPS 2024poster

Large language models are usually fine-tuned to align with human preferences. However, fine-tuning a large language model can be challenging. In this work, we introduce $\textit{weak-to-strong search}$, framing the alignment of a large language model as a test-time greedy search to maximize the log-…

2023

3D Implicit Transporter for Temporally Consistent Keypoint Discovery

ICCV 2023oral

Keypoint-based representation has proven advantageous in various visual and robotic tasks. However, the existing 2D and 3D methods for detecting keypoints mainly rely on geometric consistency to achieve spatial alignment, neglecting temporal consistency. To address this issue, the Transporter method…

Cited by 16PDFcodeScholar
2023

Discovering Intrinsic Spatial-Temporal Logic Rules to Explain Human Actions

NeurIPS 2023poster

We propose an interpretable model to uncover the behavioral patterns of human movements by analyzing their trajectories. Our approach is based on the belief that human actions are driven by intentions and are influenced by environmental factors such as spatial relationships with surrounding objects.…

Cited by 7SourcePDFScholar
2023

GoBigger: A Scalable Platform for Cooperative-Competitive Multi-Agent Interactive Simulation

ICLR 2023poster

The emergence of various multi-agent environments has motivated powerful algorithms to explore agents' cooperation or competition. Even though this has greatly promoted the development of multi-agent reinforcement learning (MARL), it is still not enough to support further exploration on the behavio…

2023

Learning Interaction Regions and Motion Trajectories Simultaneously From Egocentric Demonstration Videos

RA-L 2023

Learning to interact with objects is significant for robots to integrate into human environments. When the interaction semantic is definite, manually guiding the manipulator is a commonly used method to teach robots how to interact with objects. However, the learning results are robot-dependent beca

Cited by 9SourceScholar
2023

Tensor-Based Sketching Method for the Low-Rank Approximation of Data Streams.

ICLR 2023poster

Low-rank approximation in data streams is a fundamental and significant task in computing science, machine learning and statistics. Multiple streaming algorithms have emerged over years and most of them are inspired by randomized algorithms, more specifically, sketching methods. However, many algori…

Cited by 4SourcePDFScholar
2022

Attentional Gated Res2net for Multivariate Time Series Classification

ICASSP 2022accepted

Multivariate time series classification is a critical problem in data mining with broad applications. We design a novel convolutional neural network architecture, Attentional Gated Res2Net, for robust multivariate time series classification. AGRes2Net uses hierarchical residual-like connections to a…

Cited by 0SourceScholar
2022

Recommending Fine-Grained Tool Consistent With Common Sense Knowledge for Robot

RA-L 2022

When robots carry out task, selecting an appropriate tool is necessary. The current research ignores the fine-grained characteristic of tasks, and mainly focuses on whether the task can be completed. Little consideration is paid for the object being manipulated, which affects the task completion qua

Cited by 2SourceScholar
2022

SAS: Self-Augmentation Strategy for Language Model Pre-training

AAAI 2022technical

The core of self-supervised learning for pre-training language models includes pre-training task design as well as appropriate data augmentation. Most data augmentations in language model pre-training are context-independent. A seminal contextualized augmentation was recently proposed in ELECTRA and…

2022

Sim2Real Object-Centric Keypoint Detection and Description

AAAI 2022technical

Keypoint detection and description play a central role in computer vision. Most existing methods are in the form of scene-level prediction, without returning the object classes of different keypoints. In this paper, we propose the object-centric formulation, which, beyond the conventional setting, r…

Cited by 9SourcePDFScholar
2022

Stacked Multi-Scale Attention Network for Image Colorization

ICASSP 2022accepted

Deep convolutional networks (CNNs) show their potential in image colorization for producing plausible results. Recently, the attention mechanism further boosts the performances of CNNs by constructing channel and spatial interactions. However, existing attention methods are performed in a single-sca…

Cited by 0SourceScholar
2022

TA-MoE: Topology-Aware Large Scale Mixture-of-Expert Training

NeurIPS 2022accept

Sparsely gated Mixture-of-Expert (MoE) has demonstrated its effectiveness in scaling up deep neural networks to an extreme scale. Despite that numerous efforts have been made to improve the performance of MoE from the model design or system optimization perspective, existing MoE dispatch patterns ar…

2022

WENETSPEECH: A 10000+ Hours Multi-Domain Mandarin Corpus for Speech Recognition

ICASSP 2022accepted

In this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours in total. We collect the data from YouTube and Podcast, which covers a variety of…

Cited by 0SourceScholar
2021

Deep Convolutional Gaussian Processes for Mmwave Outdoor Localization

ICASSP 2021accepted

Millimeter Wave (mmWave) communications, as a core technique of 5G, can be leveraged for outdoor localization because of its large bandwidth and massive antenna array. Fingerprinting based mmWave outdoor localization methods using deep learning are highly suitable for non-line-of-sight (NLOS) enviro…

Cited by 0SourceScholar
2021

Evaluations of the Gap between Supervised and Reinforcement Lifelong Learning on Robotic Manipulation Tasks

CoRL 2021poster

Overcoming catastrophic forgetting is of great importance for deep learning and robotics. Recent lifelong learning research has great advances in supervised learning. However, little work focuses on reinforcement learning(RL). We focus on evaluating the performances of state-of-the-art lifelong lear…

Cited by 13SourceScholar
2021

Multi-Models Fusion for Light Field Angular Super-Resolution

ICASSP 2021accepted

Light field (LF) imaging has received increasing attention due to its richer interpretation of the scene. However, an inherent spatial-angular trade-off exists in LF that prevents LF from practical applications. Consequently, how to break such a trade-off has become one of the main challenges in spa…

Cited by 0SourceScholar
2020

AUTO3D: Novel view synthesis through unsupervisely learned variational viewpoint and global 3D representation

ECCV 2020poster

This paper targets on learning-based novel view synthesis from a single or limited 2D images without the pose supervision. In the viewer-centered coordinates, we construct an end-to-end trainable conditional variational framework to disentangle the unsupervisely learned relative-pose/rotation and im…

Cited by 26SourcePDFScholar
2020

METNet: A Mutual Enhanced Transformation Network for Aspect-based Sentiment Analysis

COLING 2020main

Aspect-based sentiment analysis (ABSA) aims to determine the sentiment polarity of each specific aspect in a given sentence. Existing researches have realized the importance of the aspect for the ABSA task and have derived many interactive learning methods that model context based on specific aspect…

Cited by 18SourcePDFScholar
2020

PEDNet: A Persona Enhanced Dual Alternating Learning Network for Conversational Response Generation

COLING 2020main

Endowing a chatbot with a personality is essential to deliver more realistic conversations. Various persona-based dialogue models have been proposed to generate personalized and diverse responses by utilizing predefined persona information. However, generating personalized responses is still a chall…

2020

Rethinking Image Inpainting via a Mutual Encoder-Decoder with Feature Equalizations

ECCV 2020poster

Deep encoder-decoder based CNNs have advanced image inpainting methods for hole filling. While existing methods recover structures and textures step-by-step in the hole regions, they typically use two encoder-decoders for separate recovery. The CNN features of each encoder are learned to capture eit…

2019

Imitation Learning from Observations by Minimizing Inverse Dynamics Disagreement

NeurIPS 2019spotlight

This paper studies Learning from Observations (LfO) for imitation learning with access to state-only demonstrations. In contrast to Learning from Demonstration (LfD) that involves both action and state supervisions, LfO is more practical in leveraging previously inapplicable resources (e.g., videos)…

Cited by 90SourcePDFScholar
2018

A Dual-Modal Vision-Based Tactile Sensor for Robotic Hand Grasping

ICRA 2018poster

Humans' fingertips can perceive not only the magnitude and the direction of force but also the texture of object. When we grasp an object, the surface texture sensing of the fingertip helps us recognize the object and the force feeling that is parallel to the skin helps us grasp stably. Focusing on…

Cited by 64SourceScholar
2018

Contextual-based Image Inpainting: Infer, Match, and Translate

ECCV 2018poster

We study the task of image inpainting, which is to fill in the missing region of an incomplete image with plausible contents. To this end, we propose a learning-based approach to generate visually coherent completion given a high-resolution image with missing components. In order to overcome the dif…

Cited by 346SourcePDFScholar
2018

Dependency-aware Attention Control for Unconstrained Face Recognition with Image Sets

ECCV 2018poster

This paper targets the problem of image set-based face verification and identification. Unlike traditional single media (an image or video) setting, we encounter a set of heterogeneous contents containing orderless images and videos. The importance of each image is usually considered either equal or…

Cited by 53SourcePDFScholar
2018

Quality Enhancement for Intra Frame Coding Via Cnns: An Adversarial Approach

ICASSP 2018accepted

Lossy compression is an indispensable technique in image/video processing, due to its highly desirable ability of reducing the huge data volume. However, lossy compression introduces complex compression artifacts. To reduce these artifacts, post-processing techniques have been extensively studied. I…

Cited by 0SourceScholar
2017

Coding of 3D holoscopic image by using spatial correlation of rendered view images

ICASSP 2017accepted

Holoscopic imaging is a prospective acquisition and display solution for providing natural and fatigue-free 3D visualization. However, large amount of data is required to represent the 3D holoscopic content. Therefore, efficient coding schemes for this particular type of image are needed. In this pa…

Cited by 0SourceScholar
2017

High-Resolution Image Inpainting Using Multi-Scale Neural Patch Synthesis

CVPR 2017poster

Recent advances in deep learning have shown exciting promise in filling large holes in natural images with semantically plausible and context aware details, impacting fundamental image manipulation tasks such as object removal. While these learning-based methods are significantly more effective in c…

Cited by 1114PDFScholar
2017

Realistic Dynamic Facial Textures From a Single Image Using GANs

ICCV 2017poster

We present a novel method to realistically puppeteer and animate a face from a single RGB image using a source video sequence. We begin by fitting a multilinear PCA model to obtain the 3D geometry and a single texture of the target face. In order for the animation to be realistic, however, we need d…

Cited by 116PDFScholar
2017

Shape Inpainting Using 3D Generative Adversarial Network and Recurrent Convolutional Networks

ICCV 2017poster

Recent advances in convolutional neural networks have shown promising results in 3D shape completion. But due to GPU memory limitations, these methods can only produce low-resolution outputs. To inpaint 3D models with semantic plausibility and contextual details, we introduce a hybrid framework that…

Cited by 214PDFScholar