← Search

Tong Zhang

291 accepted papers

2026

An Emotion-Preserving Conditional Information Bottleneck for Domain-Generalizable Speech Emotion Recognition

IJCAI 2026

Domain-generalizable speech emotion recognition (DG-SER) aims to ensure the robustness of SER models across unknown domains, which is essential for real-world human-machine interaction systems. Most DG-SER approaches employ alignment or adversarial strategies with domain labels to promote generaliza

Cited by 0Scholar
2026

Cross-Domain Few-Shot Learning via Multi-View Collaborative Optimization with Vision-Language Models

AAAI 2026technical

Vision-language models (VLMs) pre-trained on natural image and language data, such as CLIP, have exhibited significant potential in few-shot image recognition tasks, leading to development of various efficient transfer learning methods. These methods exploit inherent pre-learned knowledge in VLMs an

Cited by 0SourcePDFScholar
2026

Cure-SFT: Diagnostic-Guided Data Curation for Instruction Tuning

ICML 2026poster

Instruction data curation is central to improving the instruction-following ability of large language models. However, existing approaches often struggle to simultaneously maintain data quality, diversity, and distributional consistency, largely because they do not explicitly distinguish semantic re…

Cited by 0SourceScholar
2026

EDGE COLLABORATIVE GAUSSIAN SPLATTING WITH INTEGRATED RENDERING AND COMMUNICATION

ICASSP 2026oral

Gaussian splatting (GS) struggles with degraded rendering quality on low-cost devices. To address this issue, we present edge collaborative GS (ECO-GS), where each user can switch between a local small GS model to guarantee timeliness and a remote large GS model to guarantee fidelity. However, decid…

Cited by 0SourcePDFScholar
2026

From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs

CVPR 2026

While Multimodal Large Language Models (MLLMs) have achieved impressive performance on semantic tasks, their spatial intelligence--crucial for robust and grounded AI systems--remains underdeveloped. Existing benchmarks fall short of diagnosing this limitation: they either focus on overly simplified

Cited by 0SourcecodeScholar
2026

GAR: Generative Adversarial Reinforcement Learning for Formal Theorem Proving

ICLR 2026poster

Solving math problems through verifiable languages such as Lean has significantly impacted both the mathematics and computer science communities. Current state-of-the-art models are often trained with expensive online Reinforcement Learning (RL) or expert iteration. However, these approaches rely on…

Cited by 0SourcecodeScholar
2026

GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling

CVPR 2026

Aligning video generative models with human preferences remains challenging: current approaches rely on Vision-Language Models (VLMs) for reward modeling, but these models struggle to capture subtle temporal dynamics. We propose a fundamentally different approach: repurposing video generative models

Cited by 0SourceScholar
2026

GVI-Switch: A High-Precision GNSS-Visual-Inertial State Estimator for Indoor-Outdoor Environments

RA-L 2026

We propose a high-precision GNSS-visual-inertial state estimator named GVI-Switch, augmented with a loosely coupled Error State Kalman Filter (ESKF), to achieve stable six-degree-of-freedom (6-DoF) pose estimation under varying GNSS signal conditions. First, we developed a GNSS degradation detection

Cited by 0SourceScholar
2026

HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models

CVPR 2026

Achieving human-like spatial intelligence for vision-language models (VLMs) requires inferring 3D structures from 2D observations, recognizing object properties and relations in 3D space, and performing high-level spatial reasoning. In this paper, we propose a principled hierarchical framework that

Cited by 0SourcecodeScholar
2026

Homophily-Heterogeneity Gradient Surgery for Federated Graph Learning

ICML 2026poster

Federated Graph Learning (FGL) facilitates privacy-preserving collaborative training of graph neural networks, yet homophily heterogeneity across subgraphs triggers optimization conflicts that degrade model generalization. Most existing solutions rely on multi-channel architectures to mitigate such …

Cited by 0SourceScholar
2026

LeanForPhysics: Comprehensive Reasoning Framework for University-level Physics in Lean4

ICLR 2026poster

We present **Lean4PHYS**, a comprehensive reasoning framework for college-level physics problems in Lean4. **Lean4PHYS** includes *LeanPhysBench*, a college-level benchmark for Lean4 formal physics reasoning, which contains 200 hand-crafted and peer-reviewed statements formalized from university tex…

Cited by 0SourcecodeScholar
2026

Learning to Label: A Reinforced Self-Evolving Framework for Semi-supervised Referring Expression Segmentation

ICML 2026poster

Semi-supervised referring expression segmentation (SS-RES) aims to achieve precise pixel-level language grounding under limited annotation, yet suffers from limited supervision and unreliable pseudo-labels when exploiting unlabeled image–text pairs. In this work, we propose Learning to Label, a rein…

Cited by 0SourceScholar
2026

Localize-and-Stitch: Efficient Model Merging via Sparse Task Arithmetic

ICML 2026poster

Model merging offers an effective strategy to combine the strengths of multiple finetuned models into a unified model that preserves the specialized capabilities of each. Existing methods merge models in a global manner, performing arithmetic operations across all model parameters. However, such glo…

Cited by 0SourceScholar
2026

Mixture Prototype Flow Matching for Open-Set Supervised Anomaly Detection

ICML 2026poster

Open-set supervised anomaly detection (OSAD) aims to identify unseen anomalies using limited anomalous supervision. However, existing prototype-based methods typically model normal data via a unimodal Gaussian prior, failing to capture inherent multi-modality and resulting in blurred decision bounda…

Cited by 0SourceScholar
2026

P$^2$-DPO:Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization

ICLR 2026poster

Hallucination has recently garnered significant research attention in Large Vision-Language Models (LVLMs). Direct Preference Optimization (DPO) aims to learn directly from the corrected preferences provided by humans, thereby addressing the hallucination issue. Despite its success, this paradigm ha…

Cited by 0SourceScholar
2026

Parameter-Masked Decoupled Optimization for Cross-Domain Class-Incremental Learning

ICML 2026poster

Cross-domain class-incremental learning (CD-CIL) requires models to continuously acquire new classes across shifting domains while retaining previously learned knowledge. Existing approaches often entangle what to update with how to update, resulting in unstable adaptation and severe forgetting unde…

Cited by 0SourceScholar
2026

Preserve and Sculpt: Manifold-Aligned Fine-tuning of Vision-Language Models for Few-Shot Learning

ICLR 2026poster

Pretrained vision-language models (VLMs), such as CLIP, have shown remarkable potential in few-shot image classification and led to numerous effective transfer learning strategies. These methods leverage the pretrained knowledge of VLMs to enable effective domain adaptation while mitigating overfitt…

Cited by 0SourceScholar
2026

Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection

IJCAI 2026

The rapid advancement of generative AI has made audio deepfakes increasingly indistinguishable from authentic human vocals, posing significant threats to persons-of-interest (POI) such as public figures. Current detection systems primarily rely on generic, black-box models that fail to capture speak

Cited by 0Scholar
2026

PsyPARSE: Retrieval-Augmented Slow Thinking for Personalized Empathetic Counseling

AAAI 2026technical

The escalating global demand for mental health services highlights the potential of Large Language Models (LLMs) in psychological counseling. However, current LLM-based approaches, particularly fine-tuned models, are constrained by data distribution biases, leading to limited therapeutic diversity a

Cited by 0SourcePDFScholar
2026

Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation

ICML 2026poster

Accurate and efficient Video Quality Assessment (VQA) has long been a key research challenge. Current mainstream VQA methods typically improve performance by pretraining on large-scale classification datasets, followed by fine-tuning on VQA datasets. However, this strategy presents two significant c…

Cited by 0SourceScholar
2026

Region-Aware Instance Consistency Learning for Micro-Expression Recognition

CVPR 2026

Micro-expression Recognition (MER) is challenging due to the subtle motion. Existing methods heavily rely on the onset/apex pair to capture the most discriminative motion clues. This paradigm struggles with labor-intensive apex annotation and effective utilization of data. In this paper, we propose

Cited by 0SourceScholar
2026

Steer Where It Matters: Token-Level Visual-Sensitivity Steering for LVLMs Hallucination Mitigation

ICML 2026poster

Large vision language models (LVLMs) have made rapid advancements and are deployed across various applications, yet hallucinations remain a major challenge. Activation steering is appealing due to its minimal training overhead and controllability at inference time. However we found that during autor…

Cited by 0SourceScholar
2026

Towards a Sharp Analysis of Learning Offline $f$-Divergence-Regularized Contextual Bandits

ICLR 2026poster

Many offline reinforcement learning algorithms are underpinned by $f$-divergence regularization, but their sample complexity *defined with respect to regularized objectives* still lacks tight analyses, especially in terms of concrete data coverage conditions. In this paper, we study the exact concen…

Cited by 0SourceScholar
2026

Transferable Attacks on Open-Vocabulary Video Instance Segmentation via Dual-Objective Triggers

IJCAI 2026

Open‑vocabulary video instance segmentation (OV‑VIS) couples spatial‑temporal reasoning with language grounding, yet its adversarial robustness has remained unexplored. We present the Dual-Objective Triggers (DOT), the first transferable attack on OV-VIS that simultaneously exploits the vision–langu

Cited by 0Scholar
2026

UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register

CVPR 2026

Representation learning with Vision Transformers (ViTs) has advanced rapidly, yet the utility of large-scale models in spatially sensitive tasks is hindered by spurious tokens. Prior efforts to mitigate this have been limited, often defining these artifacts narrowly, for example, as simple high-norm

Cited by 0SourceScholar
2026

VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved impressive progress in vision–language alignment, yet they remain limited in visual–spatial reasoning. We first identify that this limitation arises from the attention mechanism: visual tokens are overshadowed by language tokens, preventing the…

Cited by 0SourceScholar
2026

WebGym: Scaling Training Environments for Long-Horizon Visual Web Agents with Realistic Tasks

CVPR 2026

We present WebGym, the largest-to-date open-source environment for training realistic visual web agents. Real websites are non-stationary and diverse, making artificial or small-scale task sets insufficient for robust policy learning. WebGym contains nearly 300,000 tasks with rubric-based evaluation

Cited by 0SourcecodeScholar
2026

When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse

CVPR 2026

Audio-Visual Speech Recognition (AVSR) has achieved remarkable progress in offline conditions, yet its robustness in real-world video conferencing (VC) remains largely unexplored. This paper presents the first systematic evaluation of state-of-the-art AVSR models across mainstream VC platforms, reve

Cited by 0SourceScholar
2025

A Parameter-Efficient and Fine-Grained Prompt Learning for Vision-Language Models

ACL 2025long

Current vision-language models (VLMs) understand complex vision-text tasks by extracting overall semantic information from large-scale cross-modal associations. However, extracting from large-scale cross-modal associations often smooths out semantic details and requires large computations, limiting…

Cited by 0SourcePDFScholar
2025

ALRPHFS: Adversarially Learned Risk Patterns with Hierarchical Fast & Slow Reasoning for Robust Agent Defense

EMNLP 2025

LLM Agents are becoming central to intelligent systems. However, their deployment raises serious safety concerns. Existing defenses largely rely on “Safety Checks”, which struggle to capture the complex semantic risks posed by harmful user inputs or unsafe agent behaviors—creating a significant sema

2025

AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation

ACL 2025long

In modern large language models (LLMs), LLM alignment is of crucial importance and is typically achieved through methods such as reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO). However, in most existing methods for LLM alignment, all tokens in the response…

2025

An Orthogonal High-Rank Adaptation for Large Language Models

EMNLP 2025

Low-rank adaptation (LoRA) efficiently adapts LLMs to downstream tasks by decomposing LLMs’ weight update into trainable low-rank matrices for fine-tuning. However, the random low-rank matrices may introduce massive task-irrelevant information, while their recomposed form suffer from limited represe

Cited by 0SourcePDFScholar
2025

Bridge-Coder: Transferring Model Capabilities from High-Resource to Low-Resource Programming Language

ACL 2025finding

Most LLMs universally excel at generating code for high-resource programming languages (HRPLs) like Python, a capability that has become standard due to the abundance of training data. However, they struggle significantly with low-resource programming languages (LRPLs) such as D, exacerbating the di…

2025

Building Math Agents with Multi-Turn Iterative Preference Learning

ICLR 2025poster

Recent studies have shown that large language models' (LLMs) mathematical problem-solving capabilities can be enhanced by integrating external tools, such as code interpreters, and employing multi-turn Chain-of-Thought (CoT) reasoning. While current methods focus on synthetic data generation and Sup…

Cited by 24SourcePDFScholar
2025

CANDY: Benchmarking LLMs’ Limitations and Assistive Potential in Chinese Misinformation Fact-Checking

EMNLP 2025

The effectiveness of large language models (LLMs) to fact-check misinformation remains uncertain, despite their growing use. To this end, we present CANDY, a benchmark designed to systematically evaluate the capabilities and limitations of LLMs in fact-checking Chinese misinformation. Specifically,

2025

DFMU: Distribution-based Framework for Modeling Aleatoric Uncertainty in Multimodal Sentiment Analysis

IJCAI 2025

In Multimodal Sentiment Analysis (MSA), data noise arising from various sources can lead to uncertainty in Aleatoric Uncertainty (AU), significantly impacting model performance. Current efforts to address AU have insufficiently explored its sources. They primarily focus on modeling noise rather than

Cited by 0SourcePDFScholar
2025

Distribution Prototype Diffusion Learning for Open-set Supervised Anomaly Detection

CVPR 2025poster

In Open-set Supervised Anomaly Detection (OSAD), the existing methods typically generate pseudo anomalies to compensate for the scarcity of observed anomaly samples, while overlooking critical priors of normal samples, leading to less effective discriminative boundaries. To address this issue,…

Cited by 0SourcePDFScholar
2025

ELABORATION: A Comprehensive Benchmark on Human-LLM Competitive Programming

ACL 2025long

While recent research increasingly emphasizes the value of human-LLM collaboration in competitive programming and proposes numerous empirical methods, a comprehensive understanding remains elusive due to the fragmented nature of existing studies and their use of diverse, application-specific human f…

2025

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

ICML 2025oral

Leveraging Multi-modal Large Language Models (MLLMs) to create embodied agents offers a promising avenue for tackling real-world tasks. While language-centric embodied agents have garnered substantial attention, MLLM-based embodied agents remain underexplored due to the lack of comprehensive evaluat…

2025

Enhancing Generalized EEG Classification with Decomposed Statistics-diverse Feature Augmentation

ICASSP 2025accepted

Learning a generalized EEG representation under limited data and subject variability is a long-standing challenge. Most studies utilized data augmentation to extend the distribution of training data, which may hinder the diversity of augmented samples to cover more subject variability. In this paper…

Cited by 0SourceScholar
2025

FANS: Formal Answer Selection for LLM Natural Language Math Reasoning Using Lean4

EMNLP 2025

Large Language Models (LLMs) have displayed astonishing abilities in various tasks, especially in text generation, classification, question answering, etc. However, the reasoning ability of LLMs still faces many debates, especially in math reasoning. The inherent ambiguity of Natural Language (NL) l

2025

FDS: Frequency-Aware Denoising Score for Text-Guided Latent Diffusion Image Editing

CVPR 2025poster

Text-guided image editing using Text-to-Image (T2I) models often fails to yield satisfactory results, frequently introducing unintended modifications, such as the loss of local detail and color changes. In this paper, we analyze these failure cases and attribute them to the indiscriminate optimizati…

Cited by 0SourcePDFScholar
2025

FPE2M2: Approaching Lossless and Efficient Quantization with Native Floating Point

ACL 2025finding

Auto-regressive decoding is a memory-bound job, meaning decoding inference performance is limited by the bandwidth rather than the computational capabilities of the GPU. Weight-only quantization is a promising method to address the memory-bound limitations. Previous studies have followed one of two…

Cited by 0SourcePDFScholar
2025

From Lists to Emojis: How Format Bias Affects Model Alignment

ACL 2025long

In this paper, we study format biases in reinforcement learning from human feedback (RLHF). We observe that many widely-used preference models—including human evaluators, GPT-4, and top-ranking models on the RewardBench benchmark—exhibit strong biases towards specific format patterns, such as lists,…

2025

GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents

NeurIPS 2025poster

One of the principal challenges in building VLM-powered GUI agents is visual grounding—localizing the appropriate screen region for action execution based on both the visual content and the textual plans. Most existing work formulates this as a text-based coordinate generation task. However, these a…

Cited by 0SourceScholar
2025

Generating Multimodal Driving Scenes via Next-Scene Prediction

CVPR 2025poster

Generative models in Autonomous Driving (AD) enable diverse scenario creation, yet existing methods fall short by only capturing a limited range of modalities, restricting the capability of generating controllable scenes for comprehensive evaluation of AD systems. In this paper, we introduce a multi…

2025

Going Beyond Consistency: Target-oriented Multi-view Graph Neural Network

IJCAI 2025

Multi‐view learning has emerged as a pivotal research area driven by the growing heterogeneity of real‐world data, and graph neural network-based models, modeling multi-view data as multi-view graphs, have achieved remarkable performance by revealing its deep semantics. However, by assuming cross‐vi

2025

Incongruity-aware Tension Field Network for Multi-modal Sarcasm Detection

ACL 2025long

Multi-modal sarcasm detection (MSD) identifies sarcasm and accurately understands users’ real attitudes from text-image pairs. Most MSD researches explore the incongruity of text-image pairs as sarcasm information through consistency preference methods. However, these methods prioritize consistency…

Cited by 0SourcePDFScholar
2025

Let’s Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM’s Math Capability

EMNLP 2025

Enhancing the mathematical reasoning capabilities of LLMs has garnered significant attention in both the mathematical and computer science communities. Recent works have made substantial progress in both Natural Language (NL) reasoning and Formal Language (FL) reasoning by leveraging the potential o

Cited by 0SourcePDFScholar
2025

Logarithmic Regret for Online KL-Regularized Reinforcement Learning

ICML 2025poster

Recent advances in Reinforcement Learning from Human Feedback (RLHF) have shown that KL-regularization plays a pivotal role in improving the efficiency of RL fine-tuning for large language models (LLMs). Despite its empirical advantage, the theoretical difference between KL-regularized RL and standa…

Cited by 1SourcePDFScholar
2025

M3Rec: Selective State Space Models with Mixture-of-Modality Experts for Multi-Modal Sequential Recommendation

ICASSP 2025accepted

The rapid growth of multimedia-sharing platforms drives the development of recommender systems. While traditional ID-based methods for mining user behavior signals are well-studied, research into multimodal sequential recommendation remains nascent. Current approaches face three critical challenges:…

Cited by 0SourceScholar
2025

MA-LoT: Model-Collaboration Lean-based Long Chain-of-Thought Reasoning enhances Formal Theorem Proving

ICML 2025poster

Solving mathematical problems using computer-verifiable languages like Lean has significantly impacted the mathematical and computer science communities. State-of-the-art methods utilize a single Large Language Model (LLM) to generate complete proof or perform tree search, but they fail to balance t…

Cited by 0SourcePDFScholar
2025

MatchDiffusion: Training-free Generation of Match-Cuts

ICCV 2025poster

Match-cuts are powerful cinematic tools that create seamless transitions between scenes, delivering strong visual and metaphorical connections. However, crafting impactful match-cuts is a challenging and resource-intensive process that requires deliberate artistic planning throughout the production…

2025

MergeBench: A Benchmark for Merging Domain-Specialized LLMs

NeurIPS 2025poster

Model merging provides a scalable alternative to multi-task training by combining specialized finetuned models through parameter arithmetic, enabling efficient deployment without the need for joint training or access to all task data. While recent methods have shown promise, existing evaluations are…

Cited by 0SourcecodeScholar
2025

MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning

EMNLP 2025

Reward modeling is a key step in building safe foundation models when applying reinforcement learning from human feedback (RLHF) to align Large Language Models (LLMs). However, reward modeling based on the Bradley-Terry (BT) model assumes a global reward function, failing to capture the inherently d

2025

One QuantLLM for ALL: Fine-tuning Quantized LLMs Once for Efficient Deployments

ACL 2025long

Large Language Models (LLMs) have advanced rapidly but face significant memory demands. While quantization has shown promise for LLMs, current methods typically require lengthy training to alleviate the performance degradation from quantization loss. However, deploying LLMs across diverse scenarios…

2025

One for All: Universal Topological Primitive Transfer for Graph Structure Learning

NeurIPS 2025poster

The non-Euclidean geometry inherent in graph structures fundamentally impedes cross-graph knowledge transfer. Drawing inspiration from texture transfer in computer vision, we pioneer topological primitives as transferable semantic units for graph structural knowledge. To address three critical barri…

Cited by 0SourceScholar
2025

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

NeurIPS 2025poster

Chain-of-thought (CoT) reasoning in large language models (LLMs) can be formalized as a latent variable problem, where the model needs to generate intermediate reasoning steps. While prior approaches such as iterative reward-ranked fine-tuning (RAFT) have relied on such formulations, they typically…

Cited by 0SourcecodeScholar
2025

Personalized Visual Instruction Tuning

ICLR 2025poster

Recent advancements in multimodal large language models (MLLMs) have demonstrated significant progress; however, these models exhibit a notable limitation, which we refer to as "face blindness." Specifically, they can engage in general conversations but fail to conduct personalized dialogues targeti…

2025

PivotMesh: Generic 3D Mesh Generation via Pivot Vertices Guidance

ICLR 2025poster

Generating compact and sharply detailed 3D meshes poses a significant challenge for current 3D generative models. Different from extracting dense meshes from neural representation, some recent works try to model the native mesh distribution (i.e., a set of triangles), which generates more compact re…

Cited by 11SourcePDFScholar
2025

Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and Alignment

EMNLP 2025

Recent studies have shown that Contrastive Language-Image Pre-training (CLIP) models are threatened by targeted data poisoning and backdoor attacks due to massive training image-caption pairs crawled from the Internet. Previous defense methods correct poisoned image-caption pairs by matching a new c

Cited by 0SourcePDFScholar
2025

ROAR: A Robust Autonomous Aerial Tracking System for Challenging Scenarios

RA-L 2025

Autonomous tracking represents a significant advancement in the evolution of unmanned aerial vehicles (UAVs), offering applications in areas such as aerial photography and infrastructure inspection. Despite its potential, many autonomous tracking systems encounter challenges in maintaining consisten

Cited by 3SourceScholar
2025

Refining CLIP's Spatial Awareness: A Visual-Centric Perspective

ICLR 2025poster

Contrastive Language-Image Pre-training (CLIP) excels in global alignment with language but exhibits limited sensitivity to spatial information, leading to strong performance in zero-shot classification tasks but underperformance in tasks requiring precise spatial understanding. Recent approaches ha…

Cited by 0SourcePDFScholar
2025

Rotated Runtime Smooth: Training-Free Activation Smoother for accurate INT4 inference

ICLR 2025poster

Large language models have demonstrated promising capabilities upon scaling up parameters. However, serving large language models incurs substantial computation and memory movement costs due to their large scale. Quantization methods have been employed to reduce service costs and latency. Neverthele…

Cited by 0SourcePDFScholar
2025

SEOE: A Scalable and Reliable Semantic Evaluation Framework for Open Domain Event Detection

ACL 2025long

Automatic evaluation for Open Domain Event Detection (ODED) is a highly challenging task, because ODED is characterized by a vast diversity of un-constrained output labels from various domains. Nearly all existing evaluation methods for ODED usually first construct evaluation benchmarks with limited…

2025

ScaleBiO: Scalable Bilevel Optimization for LLM Data Reweighting

ACL 2025long

Bilevel optimization has shown its utility across various machine learning settings, yet most algorithms in practice require second-order information, making it challenging to scale them up. Only recently, a paradigm of first-order algorithms has emerged in the theoretical literature, capable of eff…

2025

Scaling Mesh Generation via Compressive Tokenization

CVPR 2025poster

We propose a compressive yet effective mesh tokenization, Blocked and Patchified Tokenization (BPT), facilitating the generation of meshes exceeding 8k faces. BPT compresses mesh sequences by employing block-wise indexing and patch aggregation, reducing their length by approximately 75% compared to…

2025

Scene Graph-Grounded Image Generation

AAAI 2025technical

With the beneft of explicit object-oriented reasoning capabilities of scene graphs, scene graph-to-image generation has made remarkable advancements in comprehending object coherence and interactive relations. Recent state-of-the-arts typically predict the scene layouts as an intermediate represent…

2025

Self-Ensembling Gaussian Splatting for Few-Shot Novel View Synthesis

ICCV 2025poster

3D Gaussian Splatting (3DGS) has demonstrated remarkable effectiveness in novel view synthesis (NVS). However, 3DGS tends to overfit when trained with sparse views, limiting its generalization to novel viewpoints. In this paper, we address this overfitting issue by introducing Self-Ensembling Gaussi…

2025

SigmoidGS: To Guide Depth More Effectively

ICASSP 2025accepted

The recent success of 3D Gaussian Splatting (3DGS) on the task of novel view synthesis has amazed every one with its photorealistic results with high training and rendering speed. This paper aims to increase the interpretability of the model in both geometric attribute and appearances by a simple ye…

Cited by 0SourceScholar
2025

TAGCOS: Task-agnostic Gradient Clustered Coreset Selection for Instruction Tuning Data

NAACL 2025findings

Instruction tuning has achieved unprecedented success in NLP, turning large language models into versatile chatbots. However, the increasing variety and volume of instruction datasets demand significant computational resources. To address this, it is essential to extract a small and highly informati…

2025

TWIST: Text-encoder Weight-editing for Inserting Secret Trojans in Text-to-Image Models

ACL 2025long

Text-to-image (T2I) models excel at generating high-quality images from text via powerful text encoders but training these encoders demands substantial computational resources. Consequently, many users seek pre-trained text encoders from model plugin-sharing platforms like Civitai and Hugging Face,…

Cited by 0SourcePDFScholar
2025

Thinking vs. Doing: Improving Agent Reasoning by Scaling Test-Time Interaction

NeurIPS 2025poster

Test-time scaling in agentic tasks often relies on generating long reasoning traces ("think" more) before acting, but this does not allow agents to acquire new information from the environment or adapt behavior over time. In this work, we propose scaling test-time interaction, an untapped dimension…

Cited by 0SourceScholar
2025

TimeBooth: Disentangled Facial Invariant Representation for Diverse and Personalized Face Aging

ICCV 2025poster

Face aging is a typical ill-posed problem influenced by various factors such as environment and genetics, leading to highly diverse outcomes. However, existing methods primarily rely on numerical age representations, making it difficult to accurately capture individual or group-level aging patterns.…

2025

Understanding Overadaptation in Supervised Fine-Tuning: The Role of Ensemble Methods

ICML 2025poster

Supervised fine-tuning (SFT) on domain-specific data is the dominant approach for adapting foundation models to specialized tasks. However, it has been observed that SFT models tend to forget knowledge acquired during pretraining. In vision models, ensembling a pretrained model with its fine-tuned c…

Cited by 0SourcePDFScholar
2025

UniHG: A Large-scale Universal Heterogeneous Graph Dataset and Benchmark for Representation Learning and Cross-Domain Transferring

NeurIPS 2025poster

Irregular data in the real world are usually organized as heterogeneous graphs consisting of multiple types of nodes and edges. However, current heterogeneous graph research confronts three fundamental challenges: i) Benchmark Deficiency, ii) Semantic Disalignment, and iii) Propagation Degradation.…

Cited by 0SourceScholar
2024

3D-Aware Hypothesis & Verification for Generalizable Relative Object Pose Estimation

ICLR 2024poster

Prior methods that tackle the problem of generalizable object pose estimation highly rely on having dense views of the unseen object. By contrast, we address the scenario where only a single reference view of the object is available. Our goal then is to estimate the relative object pose between this…

Cited by 10SourcePDFScholar
2024

A Sober Look at the Robustness of CLIPs to Spurious Features

NeurIPS 2024poster

Large vision language models, such as CLIP, demonstrate impressive robustness to spurious features than single-modal models trained on ImageNet. However, existing test datasets are typically curated based on ImageNet-trained models, which aim to capture the spurious features inherited in ImageNet. B…

Cited by 8SourcePDFScholar
2024

Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient Noise

ICLR 2024poster

Heavy-ball momentum with decaying learning rates is widely used with SGD for optimizing deep learning models. In contrast to its empirical popularity, the understanding of its theoretical property is still quite limited, especially under the standard anisotropic gradient noise condition for quadrati…

Cited by 3SourcePDFScholar
2024

Active Prompting with Chain-of-Thought for Large Language Models

ACL 2024long

The increasing scale of large language models (LLMs) brings emergent abilities to various complex tasks requiring reasoning, such as arithmetic and commonsense reasoning. It is known that the effective design of task-specific prompts is critical for LLMs’ ability to produce high-quality answers. In…

Cited by 212SourcePDFScholar
2024

AdanCA: Neural Cellular Automata As Adaptors For More Robust Vision Transformer

NeurIPS 2024poster

Vision Transformers (ViTs) demonstrate remarkable performance in image classification through visual-token interaction learning, particularly when equipped with local information via region attention or convolutions. Although such architectures improve the feature aggregation from different granular…

Cited by 0SourcePDFScholar
2024

An Incremental Unified Framework for Small Defect Inspection

ECCV 2024poster

"Artificial Intelligence (AI)-driven defect inspection is pivotal in industrial manufacturing. However, existing inspection systems are typically designed for specific industrial products and struggle with diverse product portfolios and evolving processes. Although some previous studies attempt to a…

2024

Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

ACL 2024long

Fine-grained control over large language models (LLMs) remains a significant challenge, hindering their adaptability to diverse user needs. While Reinforcement Learning from Human Feedback (RLHF) shows promise in aligning LLMs, its reliance on scalar rewards often limits its ability to capture diver…

2024

CLAMBER: A Benchmark of Identifying and Clarifying Ambiguous Information Needs in Large Language Models

ACL 2024long

Large language models (LLMs) are increasingly used to meet user information needs, but their effectiveness in dealing with user queries that contain various types of ambiguity remains unknown, ultimately risking user trust and satisfaction. To this end, we introduce CLAMBER, a benchmark for evaluati…

2024

DVMNet: Computing Relative Pose for Unseen Objects Beyond Hypotheses

CVPR 2024poster

Determining the relative pose of an object between two images is pivotal to the success of generalizable object pose estimation. Existing approaches typically approximate the continuous pose representation with a large number of discrete pose hypotheses which incurs a computationally expensive proce…

2024

Data Augmentation via Latent Diffusion for Saliency Prediction

ECCV 2024poster

"Saliency prediction models are constrained by the limited diversity and quantity of labeled data. Standard data augmentation techniques such as rotating and cropping alter scene composition, affecting saliency. We propose a novel data augmentation method for deep saliency prediction that edits natu…

2024

Desigen: A Pipeline for Controllable Design Template Generation

CVPR 2024poster

Templates serve as a good starting point to implement a design (e.g. banner slide) but it takes great effort from designers to manually create. In this paper we present Desigen an automatic template creation pipeline which generates background images as well as harmonious layout elements over the ba…

2024

Disentanglement Network: Disentangle the Emotional Features from Acoustic Features for Speech Emotion Recognition

ICASSP 2024accepted

Speech emotion recognition plays a crucial role in human-computer interaction. However, data distribution of speech signals varies among individuals for emotion recognition. It may guide models to focus more on identity information rather than emotional information, which impairs the generalization…

Cited by 0SourceScholar
2024

Distributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal Algorithms

NeurIPS 2024poster

The sim-to-real gap, which represents the disparity between training and testing environments, poses a significant challenge in reinforcement learning (RL). A promising approach to addressing this challenge is distributionally robust RL, often framed as a robust Markov decision process (RMDP). In th…

Cited by 7SourcePDFScholar
2024

Enhancing Dialogue State Tracking Models through LLM-backed User-Agents Simulation

ACL 2024long

Dialogue State Tracking (DST) is designed to monitor the evolving dialogue state in the conversations and plays a pivotal role in developing task-oriented dialogue systems. However, obtaining the annotated data for the DST task is usually a costly endeavor. In this paper, we focus on employing LLMs…

2024

Faster Sampling via Stochastic Gradient Proximal Sampler

ICML 2024poster

Stochastic gradients have been widely integrated into Langevin-based methods to improve their scalability and efficiency in solving large-scale sampling problems. However, the proximal sampler, which exhibits much faster convergence than Langevin-based algorithms in the deterministic setting (Lee et…

Cited by 7SourcePDFScholar
2024

Image Textualization: An Automatic Framework for Generating Rich and Detailed Image Descriptions

NeurIPS 2024poster

Image description datasets play a crucial role in the advancement of various applications such as image understanding, text-to-image generation, and text-image retrieval. Currently, image description datasets primarily originate from two sources. One source is the scraping of image-text pairs from t…

Cited by 0SourcePDFScholar
2024

InNeRF360: Text-Guided 3D-Consistent Object Inpainting on 360-degree Neural Radiance Fields

CVPR 2024poster

We propose InNeRF360 an automatic system that accurately removes text-specified objects from 360-degree Neural Radiance Fields (NeRF). The challenge is to effectively remove objects while inpainting perceptually consistent content for the missing regions which is particularly demanding for existing…

2024

Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

EMNLP 2024finding

Reinforcement learning from human feedback (RLHF) has emerged as the primary method for aligning large language models (LLMs) with human preferences. The RLHF process typically starts by training a reward model (RM) using human preference data. Conventional RMs are trained on pairwise responses to t…

2024

Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraint

ICML 2024poster

This paper studies the theoretical framework of the alignment process of generative models with Reinforcement Learning from Human Feedback (RLHF). We consider a standard mathematical formulation, the reverse-KL regularized contextual bandit for RLHF. Despite its widespread practical application, a r…

Cited by 139SourcePDFScholar
2024

LECES: A Low-Bandwidth and Efficient Collaborative Exploration System With Distributed Multi-UAV

RA-L 2024

Collaborative exploration is a prevailing trend of autonomous exploration by unmanned aerial vehicles (UAVs). However, most collaborative exploration systems rely on excessively high communication bandwidth for precise map maintenance and efficient task allocation. This letter proposes a low-bandwid

Cited by 10SourceScholar
2024

LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning

NeurIPS 2024poster

The machine learning community has witnessed impressive advancements since large language models (LLMs) first appeared. Yet, their massive memory consumption has become a significant roadblock to large-scale training. For instance, a 7B model typically requires at least 60 GB of GPU memory with full…

2024

LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models

NAACL 2024system demonstrations

Foundation models have demonstrated a great ability to achieve general human-level intelligence far beyond traditional approaches. As the technique keeps attracting attention from the AI community, more and more foundation models have become publicly available.However, most of those models exhibit a…

2024

Leveraging Locality to Boost Sample Efficiency in Robotic Manipulation

CoRL 2024poster

Given the high cost of collecting robotic data in the real world, sample efficiency is a consistently compelling pursuit in robotics. In this paper, we introduce SGRv2, an imitation learning framework that enhances sample efficiency through improved visual and action representations. Central to the…

Cited by 8SourcecodeScholar
2024

MEDL-U: Uncertainty-aware 3D Automatic Annotation based on Evidential Deep Learning

ICRA 2024poster

Advancements in deep learning-based 3D object detection necessitate the availability of large-scale datasets. However, this requirement introduces the challenge of manual annotation, which is often both burdensome and time-consuming. To tackle this issue, the literature has seen the emergence of sev…

Cited by 1SourcecodeScholar
2024

MLLM-Protector: Ensuring MLLM’s Safety without Hurting Performance

EMNLP 2024main

The deployment of multimodal large language models (MLLMs) has brought forth a unique vulnerability: susceptibility to malicious attacks through visual inputs. This paper investigates the novel challenge of defending MLLMs against such attacks. Compared to large language models (LLMs), MLLMs include…

2024

Mind Your Augmentation: The Key to Decoupling Dense Self-Supervised Learning

ICLR 2024poster

Dense Self-Supervised Learning (SSL) creates positive pairs by building positive paired regions or points, thereby aiming to preserve local features, for example of individual objects. However, existing approaches tend to couple objects by leaking information from the neighboring contextual regions…

Cited by 2SourcePDFScholar
2024

Mitigating Object Dependencies: Improving Point Cloud Self-Supervised Learning through Object Exchange

CVPR 2024poster

In the realm of point cloud scene understanding particularly in indoor scenes objects are arranged following human habits resulting in objects of certain semantics being closely positioned and displaying notable inter-object correlations. This can create a tendency for neural networks to exploit the…

2024

Mitigating the Alignment Tax of RLHF

EMNLP 2024main

LLMs acquire a wide range of abilities during pre-training, but aligning LLMs under Reinforcement Learning with Human Feedback (RLHF) can lead to forgetting pretrained abilities, which is also known as the alignment tax. To investigate alignment tax, we conducted experiments with existing RLHF algor…

2024

Multi-Scale Prompt Memory-Augmented Model for Black-Box Scenarios

NAACL 2024long

Black-box few-shot text classification handles text classification in limited data without accessing the parameters and gradients of language models (LMs). Existing black-box optimization methods have demonstrated strong few-shot learning capabilities. However, they still require numerous LMs’ calls…

2024

On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization

EMNLP 2024finding

Reinforcement Learning from Human Feedback (RLHF) is an effective approach for aligning language models to human preferences. Central to RLHF is learning a reward function for scoring human preferences. Two main approaches for learning a reward model are 1) training an EXplicit Reward Model (EXRM) a…

2024

Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

NeurIPS 2024poster

We investigate Reinforcement Learning from Human Feedback (RLHF) in the context of a general preference oracle. In particular, we do not assume the existence of a reward function and an oracle preference signal drawn from the Bradley-Terry model as most of the prior works do. We consider a standard…

2024

PerceptionGPT: Effectively Fusing Visual Perception into LLM

CVPR 2024highlight

The integration of visual inputs with large language models (LLMs) has led to remarkable advancements in multi-modal capabilities giving rise to vision large language models (VLLMs). However effectively harnessing LLMs for intricate visual perception tasks such as detection and segmentation remains…

Cited by 28SourcePDFScholar
2024

Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning

ICML 2024spotlight

We study risk-sensitive reinforcement learning (RL), a crucial field due to its ability to enhance decision-making in scenarios where it is essential to manage uncertainty and minimize potential adverse outcomes. Particularly, our work focuses on applying the entropic risk measure to RL problems. Wh…

Cited by 2SourcePDFScholar
2024

Plum: Prompt Learning using Metaheuristics

ACL 2024findings

Since the emergence of large language models, prompt learning has become a popular method for optimizing and customizing these models. Special prompts, such as Chain-of-Thought, have even revealed previously unknown reasoning capabilities within these models. However, the progress of discovering eff…

Cited by 13SourcePDFScholar
2024

R-Tuning: Instructing Large Language Models to Say ‘I Don’t Know’

NAACL 2024long

Large language models (LLMs) have revolutionized numerous domains with their impressive performance but still face their challenges. A predominant issue is the propensity for these models to generate non-existent facts, a concern termed hallucination. Our research is motivated by the observation tha…

2024

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models

ACL 2024long

Retrieval-augmented generation (RAG) has become a main technique for alleviating hallucinations in large language models (LLMs). Despite the integration of RAG, LLMs may still present unsupported or contradictory claims to the retrieved contents. In order to develop effective hallucination preventio…

2024

RMSC-VIO: Robust Multi-Stereoscopic Visual-Inertial Odometry for Local Visually Challenging Scenarios

RA-L 2024

We present a Multi-Stereoscopic Visual-Inertial Odometry (VIO) system capable of integrating an arbitrary number of stereo cameras, exhibiting excellent robustness in the face of visually challenging scenarios. During system initialization, we introduce multi-view keyframes for simultaneous processi

Cited by 6SourceScholar
2024

Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

NeurIPS 2024poster

Reward models trained on human preference data have been proven to effectively align Large Language Models (LLMs) with human intent within the framework of reinforcement learning from human feedback (RLHF). However, current reward models have limited generalization capabilities to unseen prompts and…

2024

Reinforcement Learning with Foundation Priors: Let Embodied Agent Efficiently Learn on Its Own

CoRL 2024poster

Reinforcement learning (RL) is a promising approach for solving robotic manipulation tasks. However, it is challenging to apply the RL algorithms directly in the real world. For one thing, RL is data-intensive and typically requires millions of interactions with environments, which are impractical i…

Cited by 25SourceScholar
2024

Reverse Transition Kernel: A Flexible Framework to Accelerate Diffusion Inference

NeurIPS 2024spotlight

To generate data from trained diffusion models, most inference algorithms, such as DDPM, DDIM, and other variants, rely on discretizing the reverse SDEs or their equivalent ODEs. In this paper, we view such approaches as decomposing the entire denoising diffusion process into several segments, each…

Cited by 8SourcePDFScholar
2024

Spurious Feature Diversification Improves Out-of-distribution Generalization

ICLR 2024poster

Generalization to out-of-distribution (OOD) data is a critical challenge in machine learning. Ensemble-based methods, like weight space ensembles that interpolate model parameters, have been shown to achieve superior OOD performance. However, the underlying mechanism for their effectiveness remains…

Cited by 31SourcePDFScholar
2024

Strength Lies in Differences! Improving Strategy Planning for Non-collaborative Dialogues via Diversified User Simulation

EMNLP 2024main

We investigate non-collaborative dialogue agents, which are expected to engage in strategic conversations with diverse users, for securing a mutual agreement that leans favorably towards the system’s objectives. This poses two main challenges for existing dialogue agents: 1) The inability to integra…

Cited by 4SourcePDFScholar
2024

Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization

ECCV 2024oral

"Multimodal Large Language Models (MLLMs) excel in generating responses based on visual inputs. However, they often suffer from a bias towards generating responses similar to their pretraining corpus, overshadowing the importance of visual information. We treat this bias as a “preference” for pretra…

2024

Submodular-based In-context Example Selection for LLMs-based Machine Translation

COLING 2024main

Large Language Models (LLMs) have demonstrated impressive performances across various NLP tasks with just a few prompts via in-context learning. Previous studies have emphasized the pivotal role of well-chosen examples in in-context learning, as opposed to randomly selected instances that exhibits u…

2024

TagFog: Textual Anchor Guidance and Fake Outlier Generation for Visual Out-of-Distribution Detection

AAAI 2024technical

Out-of-distribution (OOD) detection is crucial in many real-world applications. However, intelligent models are often trained solely on in-distribution (ID) data, leading to overconfidence when misclassifying OOD data as ID classes. In this study, we propose a new learning framework which leverage…

2024

TensorOpera Router: A Multi-Model Router for Efficient LLM Inference

EMNLP 2024industry

With the rapid growth of Large Language Models (LLMs) across various domains, numerous new LLMs have emerged, each possessing domain-specific expertise. This proliferation has highlighted the need for quick, high-quality, and cost-effective LLM query response methods. Yet, no single LLM exists to ef…

Cited by 10SourcePDFScholar
2024

The Instinctive Bias: Spurious Images lead to Illusion in MLLMs

EMNLP 2024main

Large language models (LLMs) have recently experienced remarkable progress, where the advent of multi-modal large language models (MLLMs) has endowed LLMs with visual capabilities, leading to impressive performances in various multi-modal tasks. However, those powerful MLLMs such as GPT-4V still fai…

2024

The Non-linear $F$-Design and Applications to Interactive Learning

ICML 2024poster

We propose a generalization of the classical G-optimal design concept to non-linear function classes. The criterion, termed F -design, coincides with G-design in the linear case. We compute the value of the optimal design, termed the F-condition number, for several non-linear function classes. We fu…

Cited by 1SourcePDFScholar
2024

TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts

EMNLP 2024main

Proving mathematical theorems using computer-verifiable formal languages like Lean significantly impacts mathematical reasoning. One approach to formal theorem proving involves generating complete proofs using Large Language Models (LLMs) based on Natural Language (NL) proofs. However, due to the sc…

2024

Towards Better Generalization in Open-Domain Question Answering by Mitigating Context Memorization

NAACL 2024findings

Open-domain Question Answering (OpenQA) aims at answering factual questions with an external large-scale knowledge corpus. However, real-world knowledge is not static; it updates and evolves continually. Such a dynamic characteristic of knowledge poses a vital challenge for these models, as the trai…

2024

Towards Robust Model-Based Reinforcement Learning Against Adversarial Corruption

ICML 2024poster

This study tackles the challenges of adversarial corruption in model-based reinforcement learning (RL), where the transition dynamics can be corrupted by an adversary. Existing studies on corruption-robust RL mostly focus on the setting of model-free RL, where robust least-square regression is often…

Cited by 6SourcePDFScholar
2024

Towards Robust Offline Reinforcement Learning under Diverse Data Corruption

ICLR 2024spotlight

Offline reinforcement learning (RL) presents a promising approach for learning reinforced policies from offline datasets without the need for costly or unsafe interactions with the environment. However, datasets collected by humans in real-world environments are often noisy and may even be malicious…

2024

VFD-Net: Vocoder Fingerprints Detection for Fake Audio

ICASSP 2024accepted

With the rapid development of audio deepfake technology, the credibility and authenticity of public opinion is facing a formidable challenge. Since vocoder is the key component of audio deepfake and leaves distinctive fingerprint features, we propose VFD-Net (Vocoder Fingerprints Detection Net), a n…

Cited by 0SourceScholar
2024

VeraCT Scan: Retrieval-Augmented Fake News Detection with Justifiable Reasoning

ACL 2024system demonstrations

The proliferation of fake news poses a significant threat not only by disseminating misleading information but also by undermining the very foundations of democracy. The recent advance of generative artificial intelligence has further exacerbated the challenge of distinguishing genuine news from fab…

Cited by 2SourcePDFScholar
2023

A Theoretical Analysis of Optimistic Proximal Policy Optimization in Linear Markov Decision Processes

NeurIPS 2023poster

The proximal policy optimization (PPO) algorithm stands as one of the most prosperous methods in the field of reinforcement learning (RL). Despite its success, the theoretical understanding of PPO remains deficient. Specifically, it is unclear whether PPO or its optimistic variants can effectively s…

Cited by 37SourcePDFScholar
2023

A Universal Semantic-Geometric Representation for Robotic Manipulation

CoRL 2023poster

Robots rely heavily on sensors, especially RGB and depth cameras, to perceive and interact with the world. RGB cameras record 2D images with rich semantic information while missing precise spatial information. On the other side, depth cameras offer critical 3D geometry data but capture limited seman…

Cited by 22SourcecodeScholar
2023

ADAPT: Action-aware Driving Caption Transformer

ICRA 2023poster

End-to-end autonomous driving has great potential in the transportation industry. However, the lack of transparency and interpretability of the automatic decision-making process hinders its industrial adoption in practice. There have been some early attempts to use attention maps or cost volume for…

Cited by 90SourcecodeScholar
2023

Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data

EMNLP 2023long findings

Chain-of-thought (CoT) advances the reasoning abilities of large language models (LLMs) and achieves superior performance in complex reasoning tasks. However, most CoT studies rely on carefully designed human-annotated rational chains to prompt LLMs, posing challenges for real-world applications whe…

Cited by 0SourceScholar
2023

Beyond Uniform Lipschitz Condition in Differentially Private Optimization

ICML 2023poster

Most prior results on differentially private stochastic gradient descent (DP-SGD) are derived under the simplistic assumption of uniform Lipschitzness, i.e., the per-sample gradients are uniformly bounded. We generalize uniform Lipschitzness by assuming that the per-sample gradients have sample-depe…

Cited by 25SourcePDFScholar
2023

Catalyst Acceleration of Error Compensated Methods Leads to Better Communication Complexity

AISTATS 2023poster

Communication overhead is well known to be a key bottleneck in large scale distributed learning, and a particularly successful class of methods which help to overcome this bottleneck is based on the idea of communication compression. Some of the most practically effective gradient compressors, such…

Cited by 2SourcePDFScholar
2023

Corruption-Robust Algorithms with Uncertainty Weighting for Nonlinear Contextual Bandits and Markov Decision Processes

ICML 2023poster

Despite the significant interest and progress in reinforcement learning (RL) problems with adversarial corruption, current works are either confined to the linear setting or lead to an undesired $\tilde{\mathcal O}(\sqrt{T}\zeta)$ regret bound, where $T$ is the number of rounds and $\zeta$ is the to…

Cited by 29SourcePDFScholar
2023

Corruption-Robust Offline Reinforcement Learning with General Function Approximation

NeurIPS 2023poster

We investigate the problem of corruption robustness in offline reinforcement learning (RL) with general function approximation, where an adversary can corrupt each sample in the offline dataset, and the corruption level $\zeta\geq0$ quantifies the cumulative corruption amount over $n$ episodes and $…

2023

Covariate-Shift Generalization via Random Sample Weighting

AAAI 2023technical

Shifts in the marginal distribution of covariates from training to the test phase, named covariate-shifts, often lead to unstable prediction performance across agnostic testing data, especially under model misspecification. Recent literature on invariant learning attempts to learn an invariant predi…

Cited by 6SourcePDFScholar
2023

Deep Graph Structural Infomax

AAAI 2023technical

In the scene of self-supervised graph learning, Mutual Information (MI) was recently introduced for graph encoding to generate robust node embeddings. A successful representative is Deep Graph Infomax (DGI), which essentially operates on the space of node features but ignores topological structures,…

2023

Doolittle: Benchmarks and Corpora for Academic Writing Formalization

EMNLP 2023long main

Improving the quality of academic writing is a meaningful but challenging task. Conventional methods of language refinement focus on narrow, specific linguistic features within isolated sentences, such as grammatical errors and improper word use. We propose a more general task, Academic Writing Form…

Cited by 0SourceScholar
2023

Double Pessimism is Provably Efficient for Distributionally Robust Offline Reinforcement Learning: Generic Algorithm and Robust Partial Coverage

NeurIPS 2023poster

We study distributionally robust offline reinforcement learning (RL), which seeks to find an optimal robust policy purely from an offline dataset that can perform well in perturbed environments. We propose a generic algorithm framework Doubly Pessimistic Model-based Policy Optimization ($\texttt{P}^…

Cited by 41SourcePDFScholar
2023

Double Randomized Underdamped Langevin with Dimension-Independent Convergence Guarantee

NeurIPS 2023poster

This paper focuses on the high-dimensional sampling of log-concave distributions with composite structures: $p^*(\mathrm{d}x)\propto \exp(-g(x)-f(x))\mathrm{d}x$. We develop a double randomization technique, which leads to a fast underdamped Langevin algorithm with a dimension-independent convergenc…

Cited by 1SourcePDFScholar
2023

DyNCA: Real-Time Dynamic Texture Synthesis Using Neural Cellular Automata

CVPR 2023poster

Current Dynamic Texture Synthesis (DyTS) models can synthesize realistic videos. However, they require a slow iterative optimization process to synthesize a single fixed-size short video, and they do not offer any post-training control over the synthesis process. We propose Dynamic Neural Cellular A…

Cited by 18SourcePDFScholar
2023

ECHO: An Efficient Heuristic Viewpoint Determination Method on Frontier-Based Autonomous Exploration for Quadrotors

RA-L 2023

As a popular drone application, autonomous exploration suffers from low efficiency. To address the issue of repeated and unnecessary exploration, especially in a large-scale and cluttered environment, this letter proposes an efficient heuristic viewpoint determination method on frontier-based autono

Cited by 45SourceScholar
2023

Generalized Polyak Step Size for First Order Optimization with Momentum

ICML 2023poster

In machine learning applications, it is well known that carefully designed learning rate (step size) schedules can significantly improve the convergence of commonly used first-order optimization algorithms. Therefore how to set step size adaptively becomes an important research question. A popular a…

Cited by 27SourcePDFScholar
2023

Hierarchical Interactive Reconstruction Network for Video Compressive Sensing

ICASSP 2023accepted

Deep network-based image and video Compressive Sensing (CS) has attracted increasing attentions in recent years. However, in the existing deep network-based CS methods, a simple stacked convolutional network is usually adopted, which not only weakens the perception of rich contextual prior knowledge…

Cited by 0SourceScholar
2023

Learn and Sample Together: Collaborative Generation for Graphic Design Layout

IJCAI 2023poster

In the process of graphic layout generation, user specifications including element attributes and their relationships are commonly used to constrain the layouts (e.g.,"put the image above the button''). It is natural to encode spatial constraints between elements using a graph. This paper presents a…

2023

Learning in POMDPs is Sample-Efficient with Hindsight Observability

ICML 2023poster

POMDPs capture a broad class of decision making problems, but hardness results suggest that learning is intractable even in simple settings due to the inherent partial observability. However, in many realistic problems, more information is either revealed or can be computed during some point of the…

Cited by 27SourcePDFScholar
2023

Mixture-of-Domain-Adapters: Decoupling and Injecting Domain Knowledge to Pre-trained Language Models’ Memories

ACL 2023long

Pre-trained language models (PLMs) demonstrate excellent abilities to understand texts in the generic domain while struggling in a specific domain. Although continued pre-training on a large domain-specific corpus is effective, it is costly to tune all the parameters on the domain. In this paper, we…

2023

NEMTO: Neural Environment Matting for Novel View and Relighting Synthesis of Transparent Objects

ICCV 2023poster

We propose NEMTO, the first end-to-end neural rendering pipeline to model 3D transparent objects with complex geometry and unknown indices of refraction. Commonly used appearance modeling such as the Disney BSDF model cannot accurately address this challenging problem due to the complex light paths…

Cited by 12PDFcodeScholar
2023

Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov Game

ICLR 2023poster

Offline reinforcement learning (RL) aims at learning an optimal strategy using a pre-collected dataset without further interactions with the environment. While various algorithms have been proposed for offline RL in the previous literature, the minimax optimality has only been (nearly) established f…

Cited by 54SourcePDFScholar
2023

On the Convergence of Federated Averaging with Cyclic Client Participation

ICML 2023poster

Federated Averaging (FedAvg) and its variants are the most popular optimization algorithms in federated learning (FL). Previous convergence analyses of FedAvg either assume full client participation or partial client participation where the clients can be uniformly sampled. However, in practical cro…

Cited by 37SourcePDFScholar
2023

Particle-based Variational Inference with Preconditioned Functional Gradient Flow

ICLR 2023poster

Particle-based variational inference (VI) minimizes the KL divergence between model samples and the target posterior with gradient flow estimates. With the popularity of Stein variational gradient descent (SVGD), the focus of particle-based VI algorithms has been on the properties of functions in Re…

Cited by 22SourcePDFScholar
2023

Posterior Sampling for Competitive RL: Function Approximation and Partial Observation

NeurIPS 2023poster

This paper investigates posterior sampling algorithms for competitive reinforcement learning (RL) in the context of general function approximations. Focusing on zero-sum Markov games (MGs) under two critical settings, namely self-play and adversarial learning, we first propose the self-play and adve…

Cited by 4SourcePDFScholar
2023

Spatiotemporal Self-Supervised Learning for Point Clouds in the Wild

CVPR 2023poster

Self-supervised learning (SSL) has the potential to benefit many applications, particularly those where manually annotating data is cumbersome. One such situation is the semantic segmentation of point clouds. In this context, existing methods employ contrastive learning strategies and define positiv…

2023

TempSAL - Uncovering Temporal Information for Deep Saliency Prediction

CVPR 2023poster

Deep saliency prediction algorithms complement the object recognition features, they typically rely on additional information such as scene context, semantic relationships, gaze direction, and object dissimilarity. However, none of these models consider the temporal nature of gaze shifts during imag…

2023

VolRecon: Volume Rendering of Signed Ray Distance Functions for Generalizable Multi-View Reconstruction

CVPR 2023poster

The success of the Neural Radiance Fields (NeRF) in novel view synthesis has inspired researchers to propose neural implicit scene reconstruction. However, most existing neural implicit reconstruction methods optimize per-scene parameters and therefore lack generalizability to new scenes. We introdu…

2023

What is Essential for Unseen Goal Generalization of Offline Goal-conditioned RL?

ICML 2023poster

Offline goal-conditioned RL (GCRL) offers a way to train general-purpose agents from fully offline datasets. In addition to being conservative within the dataset, the generalization ability to achieve unseen goals is another fundamental challenge for offline GCRL. However, to the best of our knowled…

2022

A Self-Play Posterior Sampling Algorithm for Zero-Sum Markov Games

ICML 2022spotlight

Existing studies on provably efficient algorithms for Markov games (MGs) almost exclusively build on the “optimism in the face of uncertainty” (OFU) principle. This work focuses on a distinct approach of posterior sampling, which is celebrated in many bandits and reinforcement learning settings but…

Cited by 25SourcePDFScholar
2022

A Theoretical Analysis on Independence-driven Importance Weighting for Covariate-shift Generalization

ICML 2022spotlight

Covariate-shift generalization, a typical case in out-of-distribution (OOD) generalization, requires a good performance on the unknown test distribution, which varies from the accessible training distribution in the form of covariate shift. Recently, independence-driven importance weighting algorith…

2022

Benefits of Overparameterized Convolutional Residual Networks: Function Approximation under Smoothness Constraint

ICML 2022spotlight

Overparameterized neural networks enjoy great representation power on complex data, and more importantly yield sufficiently smooth output, which is crucial to their generalization and robustness. Most existing function approximation theories suggest that with sufficiently many parameters, neural net…

Cited by 18SourcePDFScholar
2022

Eigencurve: Optimal Learning Rate Schedule for SGD on Quadratic Objectives with Skewed Hessian Spectrums

ICLR 2022poster

Learning rate schedulers have been widely adopted in training deep neural networks. Despite their practical importance, there is a discrepancy between its practice and its theoretical analysis. For instance, it is not known what schedules of SGD achieve best convergence, even for simple problems suc…

2022

Exploiting Hybrid Semantics of Relation Paths for Multi-hop Question Answering over Knowledge Graphs

COLING 2022main

Answering natural language questions on knowledge graphs (KGQA) remains a great challenge in terms of understanding complex questions via multi-hop reasoning. Previous efforts usually exploit large-scale entity-related text corpus or knowledge graph (KG) embeddings as auxiliary information to facili…

Cited by 11SourcePDFScholar
2022

Frequency-Aware Contrastive Learning for Neural Machine Translation

AAAI 2022technical

Low-frequency word prediction remains a challenge in modern neural machine translation (NMT) systems. Recent adaptive training methods promote the output of infrequent words by emphasizing their weights in the overall training objectives. Despite the improved recall of low-frequency words, their pre…

2022

History-Aware Hierarchical Transformer for Multi-session Open-domain Dialogue System

EMNLP 2022finding

With the evolution of pre-trained language models, current open-domain dialogue systems have achieved great progress in conducting one-session conversations. In contrast, Multi-Session Conversation (MSC), which consists of multiple sessions over a long term with the same user, is under-investigated.…

Cited by 14SourcePDFScholar
2022

HyperDQN: A Randomized Exploration Method for Deep Reinforcement Learning

ICLR 2022poster

Randomized least-square value iteration (RLSVI) is a provably efficient exploration method. However, it is limited to the case where (1) a good feature is known in advance and (2) this feature is fixed during the training. If otherwise, RLSVI suffers an unbearable computational burden to obtain the…

2022

Increasing Visual Awareness in Multimodal Neural Machine Translation from an Information Theoretic Perspective

EMNLP 2022main

Multimodal machine translation (MMT) aims to improve translation quality by equipping the source sentence with its corresponding image. Despite the promising performance, MMT models still suffer the problem of input degradation: models focus more on textual information while visual information is ge…

Cited by 13SourcePDFScholar
2022

Leverage Your Local and Global Representations: A New Self-Supervised Learning Strategy

CVPR 2022poster

Self-supervised learning (SSL) methods aim to learn view-invariant representations by maximizing the similarity between the features extracted from different crops of the same image regardless of cropping size and content. In essence, this strategy ignores the fact that two crops may truly contain d…

Cited by 40PDFcodeScholar
2022

MICO: A Multi-alternative Contrastive Learning Framework for Commonsense Knowledge Representation

EMNLP 2022finding

Commonsense reasoning tasks such as commonsense knowledge graph completion and commonsense question answering require powerful representation learning. In this paper, we propose to learn commonsense knowledge representation by MICO, a Multi-alternative contrastIve learning framework on COmmonsense k…

2022

Model Agnostic Sample Reweighting for Out-of-Distribution Learning

ICML 2022spotlight

Distributionally robust optimization (DRO) and invariant risk minimization (IRM) are two popular methods proposed to improve out-of-distribution (OOD) generalization performance of machine learning models. While effective for small models, it has been observed that these methods can be vulnerable to…

2022

Model-based RL with Optimistic Posterior Sampling: Structural Conditions and Sample Complexity

NeurIPS 2022accept

We propose a general framework to design posterior sampling methods for model-based RL. We show that the proposed algorithms can be analyzed by reducing regret to Hellinger distance in conditional probability estimation. We further show that optimistic posterior sampling can control this Hellinger d…

Cited by 37SourcePDFScholar
2022

MulT: An End-to-End Multitask Learning Transformer

CVPR 2022poster

We propose an end-to-end Multitask Learning Transformer framework, named MulT, to simultaneously learn multiple high-level vision tasks, including depth estimation, semantic segmentation, reshading, surface normal estimation, 2D keypoint detection, and edge detection. Based on the Swin transformer m…

Cited by 105PDFScholar
2022

Multilingual Word Sense Disambiguation with Unified Sense Representation

COLING 2022main

As a key natural language processing (NLP) task, word sense disambiguation (WSD) evaluates how well NLP models can understand the fine-grained semantics of words under specific contexts. Benefited from the large-scale annotation, current WSD systems have achieved impressive performances in English b…

2022

Nearly Optimal Algorithms for Linear Contextual Bandits with Adversarial Corruptions

NeurIPS 2022accept

We study the linear contextual bandit problem in the presence of adversarial corruption, where the reward at each round is corrupted by an adversary, and the corruption level (i.e., the sum of corruption magnitudes over the horizon) is $C\geq 0$. The best-known algorithms in this setting are limited…

Cited by 61SourcePDFScholar
2022

Optimizing Latent Space Directions for Gan-Based Local Image Editing

ICASSP 2022accepted

Generative Adversarial Network (GAN) based localized image editing can suffer from ambiguity between semantic at-tributes. We thus present a novel objective function to evaluate the locality of an image edit. By introducing the super-vision from a pre-trained segmentation network and optimizing the…

Cited by 0SourceScholar
2022

Pessimistic Minimax Value Iteration: Provably Efficient Equilibrium Learning from Offline Datasets

ICML 2022spotlight

We study episodic two-player zero-sum Markov games (MGs) in the offline setting, where the goal is to find an approximate Nash equilibrium (NE) policy pair based on a dataset collected a priori. When the dataset does not have uniform coverage over all policy pairs, finding an approximate NE involves…

Cited by 52SourcePDFScholar
2022

Probabilistic Bilevel Coreset Selection

ICML 2022spotlight

The goal of coreset selection in supervised learning is to produce a weighted subset of data, so that training only on the subset achieves similar performance as training on the entire dataset. Existing methods achieved promising results in resource-constrained scenarios such as continual learning a…

Cited by 40SourcePDFScholar