← Search

Tong Wang

51 accepted papers

2026

MiVE: Multiscale Vision-language features for reference-guided video Editing

ICML 2026poster

Reference-guided video editing takes a source video, a text instruction, and a reference image as inputs, requiring the model to faithfully apply the instructed edits while preserving original motion and unedited content. Existing methods fall into two paradigms, each with inherent limitations: deco…

Cited by 0SourceScholar
2026

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection

CVPR 2026

Recent advances in Vision-Language Models (VLMs) have benefited from Reinforcement Learning (RL) for enhanced reasoning. However, existing methods still face critical limitations, including the lack of low-level visual information and effective visual feedback. To address these problems, this paper

Cited by 0SourceScholar
2026

Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Edit via In-Context Learning

ICML 2026poster

Scene text editing aims to modify text in a target region of an image while preserving its background style and texture. Existing methods rely solely on image background information while neglecting the visual details of target regions, which discards stylistic features in the original text and esse…

Cited by 0SourceScholar
2026

Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation

AAAI 2026technical

Referring Audio-Visual Segmentation (Ref-AVS) aims to segment target objects in audible videos based on given reference expressions. Prior works typically rely on learning latent embeddings via multimodal fusion to prompt a tunable SAM/SAM2 decoder for segmentation, which requires strong pixel-level

Cited by 0SourcePDFScholar
2026

pFedSAM: Personalized Federated Learning of Segment Anything Model for Medical Image Segmentation

ICASSP 2026poster

Medical image segmentation is crucial for computer-aided diagnosis, yet privacy constraints hinder data sharing across institutions. Federated learning addresses this limitation, but existing approaches often rely on lightweight architectures that struggle with complex, heterogeneous data. Recently,…

Cited by 0SourcePDFScholar
2025

Clients Collaborate: Flexible Differentially Private Federated Learning with Guaranteed Improvement of Utility-Privacy Trade-off

ICML 2025poster

To defend against privacy leakage of user data, differential privacy is widely used in federated learning, but it is not free. The addition of noise randomly disrupts the semantic integrity of the model and this disturbance accumulates with increased communication rounds. In this paper, we introduce…

2025

Efficient Cross-Boundary Grasping in Stacked Clutter with Single-Visual Mapping Multi-Step

ICRA 2025

In logistics applications, the vision-based technology for grasping target objects in the air is relatively mature. However, when operating across the air and water such as grasping marine products from the water, the visual information collected by the camera will be disturbed by ripples and bubble

Cited by 0SourceScholar
2025

Explore the LiDAR-Camera Dynamic Adjustment Fusion for 3D Object Detection

ICRA 2025

Camera and LiDAR serve as informative sensors for accurate and robust autonomous driving systems. However, these sensors often exhibit heterogeneous natures, resulting in distributional modality gaps that present significant challenges for fusion. To address this, a robust fusion technique is crucia

Cited by 0SourcecodeScholar
2025

GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editing

CVPR 2025poster

Scene text editing, a subfield of image editing, requires modifying texts in images while preserving style consistency and visual coherence with the surrounding environment. While diffusion-based methods have shown promise in text generation, they still struggle to produce high-quality results. Thes…

2025

LIRA: Reasoning Reconstruction via Multimodal Large Language Models

ICCV 2025poster

Existing language instruction-guided online 3D reconstruction systems mainly rely on explicit instructions or queryable maps, showing inadequate capability to handle implicit and complex instructions. In this paper, we first introduce a reasoning reconstruction task. This task inputs an implicit ins…

2025

Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs

ACL 2025long

We observe a novel phenomenon, *contextual entrainment*, across a wide range of language models (LMs) and prompt settings, providing a new mechanistic perspective on how LMs become distracted by “irrelevant” contextual information in the input prompt. Specifically, LMs assign significantly higher lo…

2025

ODA-GAN: Orthogonal Decoupling Alignment GAN Assisted by Weakly-supervised Learning for Virtual Immunohistochemistry Staining

CVPR 2025poster

Recently, virtual staining has emerged as a promising alternative to revolutionize histological staining by digitally generating stains. However, most existing methods suffer from the curse of staining unreality and unreliability. In this paper, we propose the Orthogonal Decoupling Alignment Generat…

2025

ProtoPairNet: Interpretable Regression through Prototypical Pair Reasoning

NeurIPS 2025poster

We present Prototypical Pair Network (ProtoPairNet), a novel interpretable architecture that combines deep learning with case-based reasoning to predict continuous targets. While prototype-based models have primarily addressed image classification with discrete outputs, extending these methods to co…

Cited by 0SourceScholar
2025

RayFusion: Ray Fusion Enhanced Collaborative Visual Perception

NeurIPS 2025poster

Collaborative visual perception methods have gained widespread attention in the autonomous driving community in recent years due to their ability to address sensor limitation problems. However, the absence of explicit depth information often makes it difficult for camera-based perception systems, e.…

Cited by 0SourceScholar
2025

THESAURUS: Contrastive Graph Clustering by Swapping Fused Gromov-Wasserstein Couplings

AAAI 2025technical

Graph node clustering is a fundamental unsupervised task. Existing methods typically train an encoder through self-supervised learning and then apply K-means to the encoder output. Some methods use this clustering result directly as the final assignment, while others initialize centroids based on th…

Cited by 0SourcePDFScholar
2025

Will Annotators Disagree? Identifying Subjectivity in Value-Laden Arguments

EMNLP 2025

Aggregating multiple annotations into a single ground truth label may hide valuable insights into annotator disagreement, particularly in tasks where subjectivity plays a crucial role. In this work, we explore methods for identifying subjectivity in recognizing the human values that motivate argumen

Cited by 0SourcePDFScholar
2024

Energy-induced Explicit quantification for Multi-modality MRI fusion

ECCV 2024poster

"Multi-modality magnetic resonance imaging (MRI) is crucial for accurate disease diagnosis and surgical planning by comprehensively analyzing multi-modality information fusion. This fusion is characterized by unique patterns of information aggregation for each disease across modalities, influenced b…

2024

Inspecting Prediction Confidence for Detecting Black-Box Backdoor Attacks

AAAI 2024technical

Backdoor attacks have been shown to be a serious security threat against deep learning models, and various defenses have been proposed to detect whether a model is backdoored or not. However, as indicated by a recent black-box attack, existing defenses can be easily bypassed by implanting the backdo…

Cited by 10SourcePDFScholar
2024

Long-Short-Range Message-Passing: A Physics-Informed Framework to Capture Non-Local Interaction for Scalable Molecular Dynamics Simulation

ICLR 2024poster

Computational simulation of chemical and biological systems using *ab initio* molecular dynamics has been a challenge over decades. Researchers have attempted to address the problem with machine learning and fragmentation-based methods. However, the two approaches fail to give a satisfactory descrip…

2024

OPEN: Object-wise Position Embedding for Multi-view 3D Object Detection

ECCV 2024poster

"Accurate depth information is crucial for enhancing the performance of multi-view 3D object detection. Despite the success of some existing multi-view 3D detectors utilizing pixel-wise depth supervision, they overlook two significant phenomena: 1) the depth supervision obtained from LiDAR points is…

2024

SEED: A Simple and Effective 3D DETR in Point Clouds

ECCV 2024poster

"Recently, detection transformers (DETRs) have gradually taken a dominant position in 2D detection thanks to their elegant framework. However, DETR-based detectors for 3D point clouds are still difficult to achieve satisfactory performance. We argue that the main challenges are twofold: 1) How to ob…

2024

Semantic-Aware Autoregressive Image Modeling for Visual Representation Learning

AAAI 2024technical

The development of autoregressive modeling (AM) in computer vision lags behind natural language processing (NLP) in self-supervised pre-training. This is mainly caused by the challenge that images are not sequential signals and lack a natural order when applying autoregressive modeling. In this stud…

2024

Sparse and Faithful Explanations Without Sparse Models

AISTATS 2024poster

Even if a model is not globally sparse, it is possible for decisions made from that model to be accurately and faithfully described by a small number of features. For instance, an application for a large loan might be denied to someone because they have no credit history, which overwhelms any eviden…

2023

An Empirical Study of Instruction-tuning Large Language Models in Chinese

EMNLP 2023long findings

The success of ChatGPT validates the potential of large language models (LLMs) in artificial general intelligence (AGI). Subsequently, the release of LLMs has sparked the open-source community's interest in instruction-tuning, which is deemed to accelerate ChatGPT's replication process. However, r…

Cited by 0SourcecodeScholar
2023

DropPos: Pre-Training Vision Transformers by Reconstructing Dropped Positions

NeurIPS 2023poster

As it is empirically observed that Vision Transformers (ViTs) are quite insensitive to the order of input tokens, the need for an appropriate self-supervised pretext task that enhances the location awareness of ViTs is becoming evident. To address this, we present DropPos, a novel pretext task desig…

2023

Efficiently incorporating quintuple interactions into geometric deep learning force fields

NeurIPS 2023poster

Machine learning force fields (MLFFs) have instigated a groundbreaking shift in molecular dynamics (MD) simulations across a wide range of fields, such as physics, chemistry, biology, and materials science. Incorporating higher order many-body interactions can enhance the expressiveness and accuracy…

2023

General-to-Specific Transfer Labeling for Domain Adaptable Keyphrase Generation

ACL 2023findings

Training keyphrase generation (KPG) models require a large amount of annotated data, which can be prohibitively expensive and often limited to specific domains. In this study, we first demonstrate that large distribution shifts among different domains severely hinder the transferability of KPG model…

2023

Geometric Transformer with Interatomic Positional Encoding

NeurIPS 2023poster

The widespread adoption of Transformer architectures in various data modalities has opened new avenues for the applications in molecular modeling. Nevertheless, it remains elusive that whether the Transformer-based architecture can do molecular modeling as good as equivariant GNNs. In this paper,…

2023

Selecting Better Samples from Pre-trained LLMs: A Case Study on Question Generation

ACL 2023findings

Large Language Models (LLMs) have in recent years demonstrated impressive prowess in natural language generation. A common practice to improve generation diversity is to sample multiple outputs from the model. However, partly due to the inaccessibility of LLMs, there lacks a simple and robust way of…

Cited by 30SourcePDFScholar
2023

Semantics-Consistent Feature Search for Self-Supervised Visual Representation Learning

ICCV 2023poster

In contrastive self-supervised learning, the common way to learn discriminative representation is to pull different augmented "views" of the same image closer while pushing all other images further apart, which has been proven to be effective. However, it is unavoidable to construct undesirable view…

Cited by 7PDFcodeScholar
2022

An Invisible Black-Box Backdoor Attack through Frequency Domain

ECCV 2022poster

"Backdoor attacks have been shown to be a serious threat against deep learning systems such as biometric authentication and autonomous driving. An effective backdoor attack could enforce the model misbehave under certain predefined conditions, i.e., triggers, but behave normally otherwise. The trigg…

2022

C2AM Loss: Chasing a Better Decision Boundary for Long-Tail Object Detection

CVPR 2022poster

Long-tail object detection suffers from poor performance on tail categories. We reveal that the real culprit lies in the extremely imbalanced distribution of the classifier's weight norm. For conventional softmax cross-entropy loss, such imbalanced weight norm distribution yields ill conditioned dec…

Cited by 28PDFScholar
2022

ProtoX: Explaining a Reinforcement Learning Agent via Prototyping

NeurIPS 2022accept

While deep reinforcement learning has proven to be successful in solving control tasks, the ``black-box'' nature of an agent has received increasing concerns. We propose a prototype-based post-hoc \emph{policy explainer}, ProtoX, that explains a black-box agent by prototyping the agent's behaviors i…

2022

Towards Lifelong Learning of Multilingual Text-to-Speech Synthesis

ICASSP 2022accepted

This work presents a lifelong learning approach to train a multilingual Text-To-Speech (TTS) system, where each language was seen as an individual task and was learned sequentially and continually. It does not require pooled data from all languages altogether, and thus alleviates the storage and com…

Cited by 0SourceScholar
2021

Adaptive Class Suppression Loss for Long-Tail Object Detection

CVPR 2021poster

To address the problem of long-tail distribution for the large vocabulary object detection task, existing methods usually divide the whole categories into several groups and treat each group with different strategies. These methods bring the following two problems. One is the training inconsistency…

Cited by 128PDFcodeScholar
2021

An Empirical Study on Neural Keyphrase Generation

NAACL 2021long

Recent years have seen a flourishing of neural keyphrase generation (KPG) works, including the release of several large-scale datasets and a host of new models to tackle them. Model performance on KPG tasks has increased significantly with evolving deep learning research. However, there lacks a comp…

2021

Bringing Structure into Summaries: a Faceted Summarization Dataset for Long Scientific Documents

ACL 2021short

Faceted summarization provides briefings of a document from different perspectives. Readers can quickly comprehend the main points of a long document with the help of a structured outline. However, little research has been conducted on this subject, partially due to the lack of large-scale faceted s…

2021

Diverse Distributions of Self-Supervised Tasks for Meta-Learning in NLP

EMNLP 2021main

Meta-learning considers the problem of learning an efficient learning process that can leverage its past experience to accurately solve new tasks. However, the efficacy of meta-learning crucially depends on the distribution of tasks available for training, and this is often assumed to be known a pri…

2021

Entity Resolution in Open-domain Conversations

NAACL 2021industry

In recent years, incorporating external knowledge for response generation in open-domain conversation systems has attracted great interest. To improve the relevancy of retrieved knowledge, we propose a neural entity linking (NEL) approach. Different from formal documents, such as news, conversationa…

Cited by 11SourcePDFScholar
2021

Optimizing NLU Reranking Using Entity Resolution Signals in Multi-domain Dialog Systems

NAACL 2021industry

In dialog systems, the Natural Language Understanding (NLU) component typically makes the interpretation decision (including domain, intent and slots) for an utterance before the mentioned entities are resolved. This may result in intent classification and slot tagging errors. In this work, we propo…

Cited by 2SourcePDFScholar
2021

Pheromone-Diffusion-based Conscientious Reactive Path Planning for Road Network Persistent Surveillance

ICRA 2021poster

Road Network Persistent Surveillance Problem (RPSP) involves path planning for an unmanned ground vehicle (UGV) with detection ability to timely detect the events randomly occurred. The road network is formed by edges and weighted viewpoints, where the UGV must move along the edges. The existing met…

Cited by 6SourceScholar
2020

Large Batch Optimization for Object Detection: Training COCO in 12 Minutes

ECCV 2020poster

Most of existing object detectors usually adopt a small training batch size ( ~16), which severely hinders the whole community from exploring large-scale datasets due to the extremely long training procedure. In this paper, we propose a versatile large batch optimization framework for object detecti…

2020

Transparency Promotion with Model-Agnostic Linear Competitors

ICML 2020poster

We propose a novel type of hybrid model for multi-class classification, which utilizes competing linear models to collaborate with an existing black-box model, promoting transparency in the decision-making process. Our proposed hybrid model, Model-Agnostic Linear Competitors (MALC), brings together…

Cited by 7SourcePDFScholar
2018

Multi-value Rule Sets for Interpretable Classification with Feature-Efficient Representations

NeurIPS 2018poster

We present the Multi-value Rule Set (MRS) for interpretable classification with feature efficient presentations. Compared to rule sets built from single-value rules, MRS adopts a more generalized form of association rules that allows multiple values in a condition. Rules of this form are more concis…