← Search

Xin Tong

56 accepted papers

2026

CoFact: Dynamic Coordination of Attention Heads for Improving Factual Consistency in LLMs

AAAI 2026technical

Large language models (LLMs) frequently generate fluent yet factually inaccurate content, a phenomenon known as hallucination. Recent inference-time approaches aim to improve truthfulness by steering model activations toward semantically meaningful directions. While effective to some extent, these m

Cited by 0SourcePDFScholar
2026

Expert Divergence Learning for MoE-based Language Models

ICLR 2026poster

The Mixture-of-Experts (MoE) architecture is a powerful technique for scaling language models, yet it often suffers from expert homogenization, where experts learn redundant functionalities, thereby limiting MoE's full potential. To address this, we introduce Expert Divergence Learning, a novel pre-…

Cited by 0SourceScholar
2025

IDE: A Multi-Agent-Driven Iterative Framework for Dynamic Evaluation of LLMs

ICASSP 2025accepted

With the widespread use of large language models (LLMs) in natural language processing, traditional evaluation methods based on static datasets have become inadequate to fully capture their performance and generalization capabilities. To address this challenge, we propose an Iterative Dynamic Evalua…

Cited by 0SourceScholar
2025

Improving Transformer Based Line Segment Detection with Matched Predicting and Re-ranking

AAAI 2025technical

Classical Transformer-based line segment detection methods have delivered impressive results. However, we observe that some accurately detected line segments are assigned low confidence scores during prediction, causing them to be ranked lower and potentially suppressed. Additionally, these models o…

Cited by 0SourcePDFScholar
2025

MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details

NeurIPS 2025poster

We propose MoGe-2, an advanced open-domain geometry estimation model that recovers a metric-scale 3D point map of a scene from a single image. Our method builds upon the recent monocular geometry estimation approach, MoGe, which predicts affine-invariant point maps with unknown scales. We explore ef…

Cited by 0SourceScholar
2025

MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision

CVPR 2025poster

We present MoGe, a powerful model for recovering 3D geometry from monocular open-domain images. Given a single image, our model directly predicts a 3D point map of the captured scene with an affine-invariant representation, which is agnostic to true global scale and shift. This new representation pr…

Cited by 25SourcePDFScholar
2025

NCDI-Diffusion: Neural Contextual and Directional Inversion for Novel View Synthesis through Diffusion Models

ICASSP 2025accepted

Novel view synthesis typically requires a comprehensive set of multi-view images for either image-based rendering or scene representation-based optimization. However, achieving high-fidelity novel view rendering often demands a large number of images. To address this limitation, we propose NCDI-Diff…

Cited by 0SourceScholar
2025

ReverB-SNN: Reversing Bit of the Weight and Activation for Spiking Neural Networks

ICML 2025poster

The Spiking Neural Network (SNN), a biologically inspired neural network infrastructure, has garnered significant attention recently. SNNs utilize binary spike activations for efficient information transmission, replacing multiplications with additions, thereby enhancing energy efficiency. However,…

Cited by 0SourcePDFScholar
2025

Robust Optical Transceiver Manipulation in Cluttered Cable Environments Using 3D Scene Understanding and Planning

ICRA 2025

Robotic manipulation in cluttered environments presents significant challenges, particularly when the clutter includes thin, deformable objects like cables, which complicate perception and decision-making processes. In the context of datacenters, the automation of networking tasks often involves the

Cited by 0SourceScholar
2025

Structured 3D Latents for Scalable and Versatile 3D Generation

CVPR 2025highlight

We introduce a novel 3D generation method for versatile and high-quality 3D asset creation.The cornerstone is a unified Structured LATent (SLAT) representation which allows decoding to different output formats, such as Radiance Fields, 3D Gaussians, and meshes. This is achieved by integrating a spar…

2025

Unlocking the Effectiveness of LoRA-FP for Seamless Transfer Implantation of Fingerprints in Downstream Models

EMNLP 2025

With the rapid development of large language models (LLMs), protecting intellectual property (IP) has become increasingly crucial. To tackle high costs and potential contamination in fingerprint integration, we propose LoRA-FP, a lightweight plug-and-play framework that encodes backdoor fingerprints

2024

"Plan, Posture and Go: Towards Open-vocabulary Text-to-Motion Generation"

ECCV 2024poster

"Conventional text-to-motion generation methods are usually trained on limited text-motion pairs, making them hard to generalize to open-vocabulary scenarios. Some works use the CLIP model to align the motion space and the text space, aiming to enable motion generation from natural language motion d…

Cited by 1SourcePDFScholar
2024

3D Feature Prediction for Masked-AutoEncoder-Based Point Cloud Pretraining

ICLR 2024poster

Masked autoencoders (MAE) have recently been introduced to 3D self-supervised pretraining for point clouds due to their great success in NLP and computer vision. Unlike MAEs used in the image domain, where the pretext task is to restore features at the masked pixels, such as colors, the existing 3D…

2024

A Simple Baseline for Spoken Language to Sign Language Translation with 3D Avatars

ECCV 2024oral

"The objective of this paper is to develop a functional system for translating spoken languages into sign languages, referred to as Spoken2Sign translation. The Spoken2Sign task is orthogonal and complementary to traditional sign language to spoken language (Sign2Spoken) translation. To enable Spoke…

2024

Diffusion Models are Geometry Critics: Single Image 3D Editing Using Pre-Trained Diffusion Priors

ECCV 2024poster

"We propose a novel image editing technique that enables 3D manipulations on single images, such as object rotation and translation. Existing 3D-aware image editing approaches typically rely on synthetic multi-view datasets for training specialized models, thus constraining their effectiveness on op…

Cited by 3SourcePDFScholar
2024

EnOF-SNN: Training Accurate Spiking Neural Networks via Enhancing the Output Feature

NeurIPS 2024poster

Spiking neural networks (SNNs) have gained more and more interest as one of the energy-efficient alternatives of conventional artificial neural networks (ANNs). They exchange 0/1 spikes for processing information, thus most of the multiplications in networks can be replaced by additions. However, bi…

Cited by 3SourcePDFScholar
2024

Finding Educationally Supportive Contexts for Vocabulary Learning with Attention-Based Models

COLING 2024main

When learning new vocabulary, both humans and machines acquire critical information about the meaning of an unfamiliar word through contextual information in a sentence or passage. However, not all contexts are equally helpful for learning an unfamiliar ‘target’ word. Some contexts provide a rich se…

Cited by 0SourcePDFScholar
2024

Ultrafast capturing in-flight objects with reprogrammable working speed ranges

ICRA 2024poster

In-flight high-speed object capturing is crucial in nature to improve survival and adaptation to the environment, such as the predation of frogs, leopards, and eagles. Despite its ubiquitousness in nature, capturing fast-moving objects is extremely challenging in engineering implementations. In this…

Cited by 0SourceScholar
2024

VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

NeurIPS 2024oral

We introduce VASA, a framework for generating lifelike talking faces with appealing visual affective skills (VAS) given a single static image and a speech audio clip. Our premiere model, VASA-1, is capable of not only generating lip movements that are exquisitely synchronized with the audio, but als…

Cited by 92SourcePDFScholar
2023

GRAM-HD: 3D-Consistent Image Generation at High Resolution with Generative Radiance Manifolds

ICCV 2023poster

Recent works have shown that 3D-aware GANs trained on unstructured single image collections can generate multiview images of novel instances. The key underpinnings to achieve this are a 3D radiance field generator and a volume rendering process. However, existing methods either cannot generate high-…

Cited by 81PDFScholar
2023

NeRFInvertor: High Fidelity NeRF-GAN Inversion for Single-Shot Real Image Animation

CVPR 2023poster

Nerf-based Generative models have shown impressive capacity in generating high-quality images with consistent 3D geometry. Despite successful synthesis of fake identity images randomly sampled from latent space, adopting these models for generating face images of real subjects is still a challenging…

Cited by 30SourcePDFScholar
2022

AniFaceGAN: Animatable 3D-Aware Face Image Generation for Video Avatars

NeurIPS 2022accept

Although 2D generative models have made great progress in face image generation and animation, they often suffer from undesirable artifacts such as 3D inconsistency when rendering images from different camera viewpoints. This prevents them from synthesizing video animations indistinguishable from re…

2022

Knowledge Graph Embedding by Adaptive Limit Scoring Loss Using Dynamic Weighting Strategy

ACL 2022findings

Knowledge graph embedding aims to represent entities and relations as low-dimensional vectors, which is an effective way for predicting missing links in knowledge graphs. Designing a strong and effective loss framework is essential for knowledge graph embedding models to distinguish between correct…

Cited by 6SourcePDFScholar
2022

Learning Hierarchy-Aware Quaternion Knowledge Graph Embeddings with Representing Relations as 3D Rotations

COLING 2022main

Knowledge graph embedding aims to represent entities and relations as low-dimensional vectors, which is an effective way for predicting missing links. It is crucial for knowledge graph embedding models to model and infer various relation patterns, such as symmetry/antisymmetry. However, many existin…

2022

Transformer Based Line Segment Classifier With Image Context for Real-Time Vanishing Point Detection in Manhattan World

CVPR 2022poster

Previous works on vanishing point detection usually use geometric prior for line segment clustering. We find that image context can also contribute to accurate line classification. Based on this observation, we propose to classify line segments into three groups according to three unknown-but-sought…

Cited by 21PDFcodeScholar
2021

Deep Implicit Moving Least-Squares Functions for 3D Reconstruction

CVPR 2021poster

Point set is a flexible and lightweight representation widely used for 3D deep learning. However, their discrete nature prevents them from representing continuous and fine geometry, posing a major issue for learning-based shape generation. In this work, we turn the discrete point sets into smooth su…

Cited by 132PDFcodeScholar
2021

High-Resolution Optical Flow From 1D Attention and Correlation

ICCV 2021poster

Optical flow is inherently a 2D search problem, and thus the computational complexity grows quadratically with respect to the search window, making large displacements matching infeasible for high-resolution images. In this paper, we take inspiration from Transformers and propose a new method for hi…

Cited by 98PDFcodeScholar
2021

Improving Knowledge Graph Embedding Using Affine Transformations of Entities Corresponding to Each Relation

EMNLP 2021finding

To find a suitable embedding for a knowledge graph remains a big challenge nowadays. By using previous knowledge graph embedding methods, every entity in a knowledge graph is usually represented as a k-dimensional vector. As we know, an affine transformation can be expressed in the form of a matrix…

Cited by 10SourcePDFScholar
2021

Indoor Scene Generation From a Collection of Semantic-Segmented Depth Images

ICCV 2021poster

We present a method for creating 3D indoor scenes with a generative model learned from a collection of semantic-segmented depth images captured from different unknown scenes. Given a room with a specified size, our method automatically generates 3D objects in a room from a randomly sampled latent co…

Cited by 34PDFcodeScholar
2021

Profiling Pareto Front With Multi-Objective Stein Variational Gradient Descent

NeurIPS 2021spotlight

Finding diverse and representative Pareto solutions from the Pareto front is a key challenge in multi-objective optimization (MOO). In this work, we propose a novel gradient-based algorithm for profiling Pareto front by using Stein variational gradient descent (SVGD). We also provide a counterpart o…

2021

Sampling with Trusthworthy Constraints: A Variational Gradient Framework

NeurIPS 2021poster

Sampling-based inference and learning techniques, especially Bayesian inference, provide an essential approach to handling uncertainty in machine learning (ML). As these techniques are increasingly used in daily life, it becomes essential to safeguard the ML systems with various trustworthy-related…

2021

Spline Positional Encoding for Learning 3D Implicit Signed Distance Fields

IJCAI 2021poster

Multilayer perceptrons (MLPs) have been successfully used to represent 3D shapes implicitly and compactly, by mapping 3D coordinates to the corresponding signed distance values or occupancy values. In this paper, we propose a novel positional encoding scheme, called Spline Positional Encoding, t…

2021

Towards Cross-View Consistency in Semantic Segmentation While Varying View Direction

IJCAI 2021poster

Several images are taken for the same scene with many view directions. Given a pixel in any one image of them, its correspondences may appear in the other images. However, by using existing semantic segmentation methods, we find that the pixel and its correspondences do not always have the same infe…

Cited by 3SourcePDFScholar
2021

Unsupervised 3D Learning for Shape Analysis via Multiresolution Instance Discrimination

AAAI 2021technical

We propose an unsupervised method for learning a generic and efficient shape encoding network for different shape analysis tasks. Our key idea is to jointly encode and learn shape and point features from unlabeled 3D point clouds. For this purpose, we adapt HRNet to octree-based convolutional neural…

Cited by 49SourcePDFScholar
2020

3D Orientation Estimation and Vanishing Point Extraction from Single Panoramas Using Convolutional Neural Network

ICRA 2020poster

3D orientation estimation is a key component of many important computer vision tasks such as autonomous navigation and 3D scene understanding. This paper presents a new CNN architecture to estimate the 3D orientation of an omnidirectional camera with respect to the world coordinate system from a sin…

Cited by 1SourceScholar
2020

A Closer Look at Local Aggregation Operators in Point Cloud Analysis

ECCV 2020poster

Recent advances of network architecture for point cloud processing are mainly driven by new designs of local aggregation operators. However, the impact of these operators to network performance is not carefully investigated due to different overall network architecture and implementation details in…

2020

Disentangled and Controllable Face Image Generation via 3D Imitative-Contrastive Learning

CVPR 2020oral

We propose an approach for face image generation of virtual people with disentangled, precisely-controllable latent representations for identity of non-existing people, expression, pose, and illumination. We embed 3D priors into adversarial learning and train the network to imitate the image formati…

Cited by 404PDFcodeScholar
2020

Object-based Illumination Estimation with Rendering-aware Neural Networks

ECCV 2020poster

We present a scheme for fast environment light estimation from the RGBD appearance of individual objects and their local image areas. Conventional inverse rendering is too computationally demanding for real-time applications, and the performance of purely learning-based techniques may be limited by…

Cited by 29SourcePDFScholar
2020

PFCNN: Convolutional Neural Networks on 3D Surfaces Using Parallel Frames

CVPR 2020poster

Surface meshes are widely used shape representations and capture finer geometry data than point clouds or volumetric grids, but are challenging to apply CNNs directly due to their non-Euclidean structure. We use parallel frames on surface to define PFCNNs that enable effective feature learning on su…

Cited by 56PDFcodeScholar
2020

RDCFace: Radial Distortion Correction for Face Recognition

CVPR 2020poster

The effects of radial lens distortion often appear in wide-angle cameras of surveillance and safeguard systems, which may severely degrade performances of previous face recognition algorithms. Traditional methods for radial lens distortion correction usually employ line features in scenarios that ar…

Cited by 24PDFScholar
2020

TextureFusion: High-Quality Texture Acquisition for Real-Time RGB-D Scanning

CVPR 2020oral

Real-time RGB-D scanning technique has become widely used to progressively scan objects with a hand-held sensor. Existing online methods restore color information per voxel, and thus their quality is often limited by the tradeoff between spatial resolution and time performance. Also, such methods of…

Cited by 27PDFScholar
2019

A Skeleton-Bridged Deep Learning Approach for Generating Meshes of Complex Topologies From Single RGB Images

CVPR 2019oral

This paper focuses on the challenging task of learning 3D object surface reconstructions from single RGB images. Existing methods achieve varying degrees of success by using different geometric representations. However, they all have their own drawbacks, and cannot well reconstruct those surfaces of…

Cited by 104PDFScholar
2019

Synthesizing 3D Shapes From Silhouette Image Collections Using Multi-Projection Generative Adversarial Networks

CVPR 2019poster

We present a new weakly supervised learning-based method for generating novel category-specific 3D shapes from unoccluded image collections. Our method is weakly supervised and only requires silhouette annotations from unoccluded, category-specific objects. Our method does not require access to the…

Cited by 36PDFScholar
2018

HairNet: Single-View Hair Reconstruction using Convolutional Neural Networks

ECCV 2018poster

We introduce a deep learning-based method to generate full 3D hair geometry from an unconstrained image. Our method can recover local strand details and has real-time performance. State-of-the-art hair modeling techniques rely on large hairstyle collections for nearest neighbor retrieval and then pe…

Cited by 86SourcePDFScholar