← Search

Zheng Zhang

227 accepted papers

2026

AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Models

CVPR 2026

Pre-trained vision-language models (VLMs) exhibit strong zero-shot generalization but remain vulnerable to adversarial perturbations. Existing classification-guided adversarial fine-tuning methods often disrupt pre-trained cross-modal alignment, weakening visual-textual correspondence and degrading

Cited by 0SourcecodeScholar
2026

AutoPP: Towards Automated Product Poster Generation and Optimization

AAAI 2026technical

Product posters blend striking visuals with informative text to highlight the product and capture customer attention. However, crafting appealing posters and manually optimizing them based on online performance is laborious and resource-consuming. To address this, we introduce AutoPP, an automated p

Cited by 0SourcePDFScholar
2026

BEYOND VISUAL REALISM: TOWARD RELIABLE FINANCIAL TIME SERIES GENERATION

ICASSP 2026poster

Generative models for financial time series often create data that look realistic and even reproduce stylized facts such as fat tails or volatility clustering. However, these apparent successes break down under trading backtests: models like GANs or WGAN-GP frequently collapse, yielding extreme and…

Cited by 0SourcePDFScholar
2026

BTSP-CAM: A Brain-Inspired Geometric Memory for Class-Incremental Learning

ICML 2026poster

Gradient-based optimization in class-incremental learning (CIL) often faces the plasticity–stability dilemma, since continuous weight updates can distort decision boundaries learned from earlier tasks. We revisit this problem from the viewpoint of stochastic geometric memory allocation and propose B…

Cited by 0SourceScholar
2026

Bolster Hallucination Detection via Prompt-Guided Data Augmentation

AAAI 2026technical

Large language models (LLMs) have garnered significant interest in AI community. Despite their impressive generation capabilities, they have been found to produce misleading or fabricated information, a phenomenon known as hallucinations. Consequently, hallucination detection has become critical to

Cited by 0SourcePDFScholar
2026

CoFact: Conformal Factuality Guarantees for Language Models under Distribution Shift

ICLR 2026poster

Large Language Models (LLMs) excel in natural language processing (NLP) tasks but often generate false or misleading information, known as hallucinations, raising reliability concerns in high-stakes applications. To provide statistical guarantees on the factuality of LLM outputs, conformal predictio…

Cited by 0SourceScholar
2026

Contrastive Cross-Bag Augmentation for Multiple Instance Learning-based Whole Slide Image Classification

CVPR 2026

Recent pseudo-bag augmentation methods for Multiple Instance Learning (MIL)-based Whole Slide Image (WSI) classification sample instances from a limited number of bags, resulting in constrained diversity. To address this issue, we propose Contrastive Cross-Bag Augmentation (C2Aug) to sample instance

Cited by 0SourcecodeScholar
2026

Coverage-Constrained Human-AI Cooperation with Multiple Experts

AAAI 2026technical

Human-AI cooperative classification (HAI-CC) aims to develop hybrid intelligent systems that enhance decision-making in various high-stakes real-world scenarios by leveraging both human expertise and AI capabilities. Current HAI-CC methods primarily focus on learning-to-defer (L2D), where decisions

Cited by 0SourcePDFScholar
2026

Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition

AAAI 2026technical

Incomplete multi-modal emotion recognition (IMER) aims at understanding human intentions and sentiments by comprehensively exploring the partially-observed multi-source data. Although the multi-modal data is expected to provide more abundant information, the performance gap and modality under-optimi

Cited by 0SourcePDFScholar
2026

Decoupling Vision and Language: Codebook Anchored Visual Adaptation

CVPR 2026

Large Vision-Language Models (LVLMs) use their vision encoders to translate images into representations for downstream reasoning, but the encoders often underperform in domain-specific visual tasks such as medical image diagnosis or fine-grained classification, where representation errors can cascad

Cited by 0SourceScholar
2026

Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models

CVPR 2026

Generating realistic and user-preferred advertisements is a key challenge in e-commerce. Existing approaches utilize multiple independent models driven by click-through-rate (CTR) to controllably create attractive image or text advertisements. However, their pipelines lack cross-modal perception and

Cited by 0SourcecodeScholar
2026

Disturbance-Aware Adaptive Compensation in Hybrid Force-Position Locomotion Policy for Legged Robots

ICRA 2026poster

Reinforcement Learning (RL)-based methods have significantly improved the locomotion performance of legged robots. However, these motion policies face significant challenges when deployed in the real world. Robots operating in uncertain environments struggle to adapt to payload variations and extern…

2026

FlowPET: Physics-Informed Symplectic Flow Matching for Low-Count PET Reconstruction

ICML 2026poster

Low-count Positron Emission Tomography (PET) reconstruction is severely hindered by the dissipative nature of prevailing generative models, where the inherent phase-space contraction leads to the numerical extinction (``wash-out'') of weak but diagnostically critical lesion signals. To overcome this…

Cited by 0SourceScholar
2026

FourierPET: Deep Fourier-based Unrolled Network for Low-count PET Reconstruction

AAAI 2026technical

Low-count positron emission tomography (PET) reconstruction is a challenging inverse problem due to severe degradations arising from Poisson noise, photon scarcity, and attenuation correction errors. Existing deep learning methods typically address these in the spatial domain with an undifferentiate

Cited by 0SourcePDFScholar
2026

InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster Generation

CVPR 2026

E-commerce product poster generation aims to automatically synthesize a single image that effectively conveys product information by presenting a subject, text, and a designed style. Recent diffusion models with fine-grained and efficient controllability have advanced product poster synthesis, yet t

Cited by 0SourceScholar
2026

KANO: Kolmogorov-Arnold Neural Operator

ICLR 2026poster

We introduce Kolmogorov–Arnold Neural Operator (KANO), a dual‑domain neural operator jointly parameterized by both spectral and spatial bases with intrinsic symbolic interpretability. We theoretically demonstrate that KANO overcomes the pure-spectral bottleneck of Fourier Neural Operator (FNO): KANO…

Cited by 0SourcecodeScholar
2026

KD-CVG: A KNOWLEDGE-DRIVEN APPROACH FOR CREATIVE VIDEO GENERATION

ICASSP 2026poster

Creative Generation (CG) leverages generative models to automatically produce advertising content that highlights product features, and it has been a significant focus of recent research. However, while CG has advanced considerably, most efforts have concentrated on generating advertising text and i…

Cited by 0SourcePDFScholar
2026

Linking Process to Outcome: Conditional Reward Modeling for LLM Reasoning

ICLR 2026poster

Process Reward Models (PRMs) have emerged as a promising approach to enhance the reasoning capabilities of large language models (LLMs) by guiding their step-by-step reasoning toward a final answer. However, existing PRMs either treat each reasoning step in isolation, failing to capture inter-step d…

Cited by 0SourceScholar
2026

MAKP: Multi-Mode Accurate Kicking Policy for Humanoid Robots

ICRA 2026poster

Humanoid robot soccer players face fundamental challenges in achieving stable motion execution and ball trajectory control, particularly under balance constraints during single-leg support phases. In this paper, we introduce MAKP (Multi-mode Accurate Kicking Policy), a novel motion generation-based …

Cited by 0Scholar
2026

MV2UV: Generating High-quality UV Texture Maps with Multiview Prompts

CVPR 2026

Generating high-quality textures for 3D assets is a challenging task. Existing multiview texture generation methods suffer from the multiview inconsistency and missing textures on unseen parts, while UV inpainting texture methods do not generalize well due to insufficient UV data and cannot well uti

Cited by 0SourceScholar
2026

MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation

AAAI 2026technical

Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale inconsistency, where variations in subject size lead to unnatural g

Cited by 0SourcePDFScholar
2026

Norm$\times$Direction: Restoring the Missing Query Norm in Vision Linear Attention

ICML 2026poster

Linear attention mitigates the quadratic complexity of softmax attention but suffers from a critical loss of expressiveness. We identify two primary causes: (1) The normalization operation cancels the query norm, which breaks the correlation between a query's norm and the spikiness (entropy) of the …

Cited by 0SourceScholar
2026

Rethinking LLM Reasoning: From Explicit Trajectories to Latent Representations

ICLR 2026poster

Large Language Models (LLMs) have achieved impressive performance on complex tasks by generating human-like, step-by-step rationales, referred to as \textit{reasoning trajectory}, before arriving at final answers. However, the length of these reasoning trajectories often far exceeds that of the fina…

Cited by 0SourcecodeScholar
2026

SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs

ICML 2026poster

In large-scale LLM pretraining systems with $100\mathrm{k}+$ GPUs, failures become the norm rather than the exception, and restart costs can dominate wall-clock training time. However, existing fault-tolerance mechanisms are largely unprepared for this restart-dominant regime. To address this challe…

Cited by 0SourceScholar
2026

SinBasis Networks: Matrix-Equivalent Feature Extraction for Wave-Like Optical Spectrograms

AAAI 2026technical

Wave-like images—from attosecond streaking spectrograms to optical spectra, audio mel-spectrograms and periodic video frames—encode critical harmonic structures that elude conventional feature extractors. We propose a unified, matrix-equivalent framework that reinterprets convolution and attention a

Cited by 0SourcePDFScholar
2026

TALO: Pushing 3D Vision Foundation Models Towards Globally Consistent Online Reconstruction

CVPR 2026

3D vision foundation models have shown strong generalization in reconstructing key 3D attributes from uncalibrated images through a single feed-forward pass. However, when deployed in online settings such as driving scenarios, predictions are made over temporal windows, making it non-trivial to main

Cited by 0SourcecodeScholar
2026

TDSS: Task Dynamic-Synergistic Skill Adaptation for Boosting Efficient and Scalable Multi-Task Learning in Dense Visual Prediction

AAAI 2026technical

The transfer of knowledge from large-scale pre-trained models to diverse downstream tasks has achieved remarkable success. Beyond the traditional full fine-tuning paradigm, Parameter-Efficient Fine-Tuning (PEFT) has emerged as a more efficient model adaptation approach. However, applying existing PE

Cited by 0SourcePDFScholar
2026

Themis: Automated Constraint-Aware Test Synthesis Framework for Code Reinforcement Learning

AAAI 2026technical

Reinforcement learning (RL) has shown promise for enhancing code generation capabilities in large language models (LLMs), yet its effectiveness critically depends on high-quality test suites for reliable reward signals. Current approaches suffer from inadequate test case quantity and quality, leadin

Cited by 0SourcePDFScholar
2026

TreePO: Enhancing Policy Efficacy and Inference Efficiency with Tree Modeling

ICML 2026poster

Recent advancements in aligning large language models via reinforcement learning have achieved remarkable gains in solving complex reasoning problems, but at the cost of expensive on-policy rollouts and limited exploration of diverse reasoning paths. In this work, we introduce TreePO, involving a se…

Cited by 0SourceScholar
2026

UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Pass

CVPR 2026

We present UniSH, a unified, feed-forward framework for joint metric-scale 3D scene and human reconstruction. A key challenge in this domain is the scarcity of large-scale, annotated real-world data, forcing a reliance on synthetic datasets. This reliance introduces a significant sim-to-real domain

Cited by 0SourcecodeScholar
2026

Where Concept Erasure Should Occur: Concept–Layer Alignment in Text-to-Video Diffusion Models

ICML 2026poster

Text-to-video diffusion transformers encode semantic information unevenly across model depth, which constrains effective concept erasure. We identify a representational bottleneck, termed concept–layer topological alignment, under which target concepts exhibit higher separability at certain represen…

Cited by 0SourceScholar
2025

A Novel Wavy Soft Pneumatic Actuator Combining Variable Thickness and an Unconstrained Base

RA-L 2025

This letter proposes novel wavy soft pneumatic actuators (WSPAs) that integrates an unconstrained base with an elastic chamber of varying wall thicknesses. The base plate design eliminates bottom constraints, thereby enhancing the bending performance of WSPAs. Wavy elastomer cavities with varying wa

Cited by 4SourceScholar
2025

A Semantic Knowledge Complementarity based Decoupling Framework for Semi-supervised Class-imbalanced Medical Image Segmentation

CVPR 2025poster

The limited data annotations have made semi-supervised learning (SSL) increasingly popular in medical image analysis. However, the use of pseudo labels in SSL degrades the performance of decoders that heavily rely on high-accuracy annotations. This issue is particularly pronounced in class-imbalance…

2025

A Statistical Approach for Controlled Training Data Detection

ICLR 2025poster

Detecting training data for large language models (LLMs) is receiving growing attention, especially in applications requiring high reliability. While numerous efforts have been made to address this issue, they typically focus on accuracy without ensuring controllable results. To fill this gap, we pr…

Cited by 0SourcePDFScholar
2025

Agent4Edu: Generating Learner Response Data by Generative Agents for Intelligent Education Systems

AAAI 2025technical

Personalized learning represents a promising educational strategy within intelligent educational systems, aiming to enhance learners' practice efficiency. However, the scarcity of offline practice response data (e.g., answer correctness) and potential biases in human online practice create a signifi…

2025

Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach

ICLR 2025poster

As online video content rapidly grows, the task of text-video retrieval (TVR) becomes increasingly important. A key challenge in TVR is the information asymmetry between video and text: videos are inherently richer in information, while their textual descriptions often capture only fragments of this…

Cited by 0SourcePDFScholar
2025

CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank Activation

EMNLP 2025

The full-size MLPs and the projection layers in attention introduce tremendous model sizes of large language models (LLMs), consuming extensive computational resources in pre-training. We empirically observe that the activations of pre-trained LLMs exhibit low-rank property. Motivated by such observ

2025

CoderAgent: Simulating Student Behavior for Personalized Programming Learning with Large Language Models

IJCAI 2025

Personalized programming tutoring, such as exercise recommendation, can enhance learners' efficiency, motivation, and outcomes, which is increasingly important in modern digital education. However, the lack of sufficient and high-quality programming data, combined with the mismatch between offline e

2025

Communication-Efficient and Tensorized Federated Fine-Tuning of Large Language Models

ACL 2025finding

Parameter-efficient fine-tuning (PEFT) methods typically assume that Large Language Models (LLMs) are trained on data from a single device or client. However, real-world scenarios often require fine-tuning these models on private data distributed across multiple devices. Federated Learning (FL) offe…

Cited by 0SourcePDFScholar
2025

Connector-S: A Survey of Connectors in Multi-modal Large Language Models

IJCAI 2025

With the rapid advancements in multi-modal large language models (MLLMs), connectors play a pivotal role in bridging diverse modalities and enhancing model performance. However, the design and evolution of connectors have not been comprehensively analyzed, leaving gaps in understanding how these com

2025

Correlated Multiple IHC Virtual Staining for Breast Histopathological Images

ICASSP 2025accepted

Immunohistochemistry (IHC) examination is essential for determining breast cancer subtypes and provides critical prognostic factors to guide treatment decisions. However, the complex and expensive preparation of IHC staining limits its widespread use in clinical practice. Recent advancements in gene…

Cited by 0SourceScholar
2025

DAPO: An Open-Source LLM Reinforcement Learning System at Scale

NeurIPS 2025poster

Inference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the…

Cited by 0SourceScholar
2025

DVM: Towards Controllable LLM Agents in Social Deduction Games

ICASSP 2025accepted

Large Language Models (LLMs) have advanced the capability of game agents in social deduction games (SDGs). These games rely heavily on conversation-driven interactions and require agents to infer, make decisions, and express based on such information. While this progress leads to more sophisticated…

Cited by 0SourceScholar
2025

Distribution-Driven Dense Retrieval: Modeling Many-to-One Query-Document Relationship

AAAI 2025technical

Dense retrieval has emerged as the leading approach in information retrieval, aiming to find semantically relevant documents based on natural language queries. Given that a single document can be retrieved by multiple distinct queries, existing methods aim to represent a document with multiple vecto…

2025

GRAG: Graph Retrieval-Augmented Generation

NAACL 2025findings

Naive Retrieval-Augmented Generation (RAG) focuses on individual documents during retrieval and, as a result, falls short in handling networked documents which are very popular in many applications such as citation graphs, social media, and knowledge graphs. To overcome this limitation, we introduce…

2025

Generate E-commerce Product Background by Integrating Category Commonality and Personalized Style

ICASSP 2025accepted

The state-of-the-art methods for e-commerce product background generation suffer from the inefficiency of designing product-wise prompts when scaling up the production, as well as the ineffectiveness of describing fine-grained styles when customizing personalized backgrounds for some specific brands…

Cited by 0SourceScholar
2025

GraphNarrator: Generating Textual Explanations for Graph Neural Networks

ACL 2025long

Graph representation learning has garnered significant attention due to its broad applications in various domains, such as recommendation systems and social network analysis. Despite advancements in graph learning methods, challenges still remain in explainability when graphs are associated with sem…

Cited by 0SourcePDFScholar
2025

Intent Oriented Contrastive Learning for Sequential Recommendation

AAAI 2025technical

Sequential recommendation aims to predict the next item a user is likely to interact with based on their historical interaction sequence. Capturing user intent is crucial in this process, as each interaction is typically driven by specific intentions (e.g., buying skincare products for skin maintena…

Cited by 0SourcePDFScholar
2025

LEGEND: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets

AAAI 2025technical

The success of the reward model in distinguishing between responses with subtle safety differences depends critically on the high-quality preference dataset, which should capture the fine-grained nuances of harmful and harmless responses. This motivates the need to develop the datasets involving pre…

2025

LaX: Boosting Low-Rank Training of Foundation Models via Latent Crossing

NeurIPS 2025poster

Training foundation models such as ViTs and LLMs requires tremendous computing cost. Low-rank matrix or tensor factorization offers a parameter-efficient alternative, but often downgrades performance due to the restricted parameter space. In this work, we introduce ${\textbf{Latent Crossing (LaX)}}$…

Cited by 0SourceScholar
2025

Learning to Select In-Context Demonstration Preferred by Large Language Model

ACL 2025finding

In-context learning (ICL) enables large language models (LLMs) to adapt to new tasks during inference using only a few demonstrations. However, ICL performance is highly dependent on the selection of these demonstrations. Recent work explores retrieval-based methods for selecting query-specific demo…

Cited by 0SourcePDFScholar
2025

Leveraging Variation Theory in Counterfactual Data Augmentation for Optimized Active Learning

ACL 2025finding

Active Learning (AL) allows models to learn interactively from user feedback. However, only annotating existing samples may hardly benefit the model’s generalization. Moreover, AL commonly faces a cold start problem due to insufficient annotated data for effective sample selection. To address this,…

2025

MaZO: Masked Zeroth-Order Optimization for Multi-Task Fine-Tuning of Large Language Models

EMNLP 2025

Large language models have demonstrated exceptional capabilities across diverse tasks, but their fine-tuning demands significant memory, posing challenges for resource-constrained environments. Zeroth-order (ZO) optimization provides a memory-efficient alternative by eliminating the need for backpro

Cited by 0SourcePDFScholar
2025

Multi-Agent Generative Adversarial Interactive Self-Imitation Learning for AUV Formation Control and Obstacle Avoidance

RA-L 2025

Multiple autonomous underwater vehicles (multi-AUVs) can cooperatively accomplish tasks that a single AUV cannot complete. Recently, multi-agent reinforcement learning has been introduced to control of multi-AUV. However, designing efficient reward functions for various tasks of multi-AUV control is

Cited by 8SourceScholar
2025

Multi-Document Event Extraction Using Large and Small Language Models

EMNLP 2025

Multi-document event extraction aims to aggregate event information from diverse sources for a comprehensive understanding of complex events. Despite its practical significance, this task has received limited attention in existing research. The inherent challenges include handling complex reasoning

2025

NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens

ICLR 2025poster

Recent advancements in Large Language Models (LLMs) have pushed the boundaries of natural language processing, especially in long-context understanding. However, the evaluation of these models' long-context abilities remains a challenge due to the limitations of current benchmarks. To address this g…

2025

OT-StainNet: Optimal Transport Driven Semantic Matching for Weakly Paired H&E-to-IHC Stain Transfer

AAAI 2025technical

Immunohistochemistry (IHC) examination is essential for characterizing tumor subtypes, providing prognostic information, and developing personalized treatment plans. However, IHC staining preparation is more complex and expensive compared to Hematoxylin and Eosin (H&E) staining, limiting its widespr…

Cited by 0SourcePDFScholar
2025

Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing

CVPR 2025poster

Developing a face anti-spoofing model that meets the security requirements of clients worldwide is challenging due to the domain gap between training datasets and the diverse end-user test data. Moreover, for security and privacy reasons, it is undesirable for clients to share a large amount of thei…

Cited by 0SourcePDFScholar
2025

PCToolkit: A Unified Plug-and-Play Prompt Compression Toolkit of Large Language Models

IJCAI 2025

Prompt engineering enables Large Language Models (LLMs) to perform a variety of tasks. However, lengthy prompts significantly increase computational complexity and economic costs. To address this issue, prompt compression reduces prompt length while maintaining LLM response quality. To support rapid

2025

PQR: Improving Dense Retrieval via Potential Query Modeling

ACL 2025long

Dense retrieval has now become the mainstream paradigm in information retrieval. The core idea of dense retrieval is to align document embeddings with their corresponding query embeddings by maximizing their dot product. The current training data is quite sparse, with each document typically associa…

Cited by 0SourcePDFScholar
2025

PerSphere: A Comprehensive Framework for Multi-Faceted Perspective Retrieval and Summarization

ACL 2025long

As online platforms and recommendation algorithms evolve, people are increasingly trapped in echo chambers, leading to biased understandings of various issues. To combat this issue, we have introduced PerSphere, a benchmark designed to facilitate multi-faceted perspective retrieval and summarization…

2025

PolaFormer: Polarity-aware Linear Attention for Vision Transformers

ICLR 2025poster

Linear attention has emerged as a promising alternative to softmax-based attention, leveraging kernelized feature maps to reduce complexity from quadratic to linear in sequence length. However, the non-negative constraint on feature maps and the relaxed exponential function used in approximation lea…

Cited by 2SourcePDFScholar
2025

Portcullis: A Scalable and Verifiable Privacy Gateway for Third-Party LLM Inference

AAAI 2025technical

Businesses using third-party LLMs face privacy risks from exposed prompts. This paper presents Portcullis, a privacy-preserving gateway that safeguards sensitive data while supporting efficient and accurate LLM responses. Portcullis functions as a mediator, anonymizing sensitive data in prompts thro…

Cited by 0SourcePDFScholar
2025

QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language Models

EMNLP 2025

Large Language Models (LLMs) are often quantized to lower precision to reduce the memory cost and latency in inference. However, quantization often degrades model performance, thus fine-tuning is required for various downstream tasks. Traditional fine-tuning methods such as stochastic gradient desce

Cited by 0SourcePDFScholar
2025

Residual Reweighted Conformal Prediction for Graph Neural Networks

UAI 2025

Graph Neural Networks (GNNs) excel at modeling relational data but face significant challenges in high-stakes domains due to unquantified uncertainty. Conformal prediction (CP) offers statistical coverage guarantees, but existing methods often produce overly conservative prediction intervals that fa

Cited by 0SourcePDFScholar
2025

Saten: Sparse Augmented Tensor Networks for Post-Training Compression of Large Language Models

EMNLP 2025

The efficient implementation of large language models (LLMs) is crucial for deployment on resource-constrained devices. Low-rank tensor compression techniques, such as tensor-train (TT) networks, have been widely studied for over-parameterized neural networks. However, their applications to compress

2025

SharpZO: Hybrid Sharpness-Aware Vision Language Model Prompt Tuning via Forward-Only Passes

NeurIPS 2025poster

Fine-tuning vision language models (VLMs) has achieved remarkable performance across various downstream tasks; yet, it requires access to model gradients through backpropagation (BP), making them unsuitable for memory-constrained, inference-only edge devices. To address this limitation, previous wo…

Cited by 0SourcecodeScholar
2025

ShortListing Model: A Streamlined Simplex Diffusion for Discrete Variable Generation

NeurIPS 2025poster

Generative modeling of discrete variables is challenging yet crucial for applications in natural language processing and biological sequence design. We introduce the Shortlisting Model (SLM), a novel simplex-based diffusion model inspired by progressive candidate pruning. SLM operates on simplex cen…

Cited by 0SourcecodeScholar
2025

SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking

ICLR 2025poster

Recent advancements in large language models (LLMs) with billions of parameters have improved performance in various applications, but their inference processes demand significant energy and computational resources. In contrast, the human brain, with approximately 86 billion neurons, is much more en…

2025

StableDepth: Scene-Consistent and Scale-Invariant Monocular Depth

ICCV 2025poster

Recent advances in monocular depth estimation significantly improve robustness and accuracy. However, relative depth models exhibit flickering and 3D inconsistency in video data, limiting 3D reconstruction applications. We introduce StableDepth, a scene-consistent and scale-invariant depth estimatio…

Cited by 0SourcePDFScholar
2025

Transferable Adversarial Face Attack with Text Controlled Attribute

AAAI 2025technical

Traditional adversarial attacks typically produce adversarial examples under norm-constrained conditions, whereas unrestricted adversarial examples are free-form with semantically meaningful perturbations. Current unrestricted adversarial impersonation attacks exhibit limited control over adversaria…

2025

UniRAG: Unified Query Understanding Method for Retrieval Augmented Generation

ACL 2025long

Retrieval-Augmented Generation (RAG) technology effectively addresses the issues of knowledge update lag and hallucinations in large language models (LLMs) by integrating internal and external knowledge. Existing query augmentation methods improve RAG’s performance in handling complex queries but fa…

Cited by 0SourcePDFScholar
2025

Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models

COLING 2025main

Multimodal large language models (MLLMs) combine visual and textual data for tasks like image captioning and visual question answering. Proper uncertainty calibration is crucial but challenging for reliable use in areas like healthcare and autonomous driving. This paper investigates several MLLMs, f…

2025

VascularPilot3D: Toward a 3D Fully Autonomous Navigation for Endovascular Robotics

ICRA 2025

This research reports VascularPilot3D, the first 3D fully autonomous endovascular robot navigation system. As an exploration toward autonomous guidewire navigation, VascularPilot3D is developed as a complete navigation system based on intra-operative imaging systems (fluoroscopic X-ray in this study

Cited by 6SourceScholar
2025

Wanda++: Pruning Large Language Models via Regional Gradients

ACL 2025finding

Large Language Models (LLMs) pruning seeks to remove unimportant weights for inference speedup with minimal accuracy impact. However, existing methods often suffer from accuracy degradation without full-model sparsity-aware fine-tuning. This paper presents Wanda++, a novel pruning framework that out…

Cited by 0SourcePDFScholar
2024

4DBInfer: A 4D Benchmarking Toolbox for Graph-Centric Predictive Modeling on RDBs

NeurIPS 2024poster

Given a relational database (RDB), how can we predict missing column values in some target table of interest? Although RDBs store vast amounts of rich, informative data spread across interconnected tables, the progress of predictive machine learning models as applied to such tasks arguably falls we…

2024

AdaZeta: Adaptive Zeroth-Order Tensor-Train Adaption for Memory-Efficient Large Language Models Fine-Tuning

EMNLP 2024main

Fine-tuning large language models (LLMs) has achieved remarkable performance across various natural language processing tasks, yet it demands more and more memory as model sizes keep growing. To address this issue, the recently proposed Memory-efficient Zeroth-order (MeZO) methods attempt to fine-tu…

2024

Adaptive Slot Attention: Object Discovery with Dynamic Slot Number

CVPR 2024poster

Object-centric learning (OCL) extracts the representation of objects with slots offering an exceptional blend of flexibility and interpretability for abstracting low-level perceptual features. A widely adopted method within OCL is slot attention which utilizes attention mechanisms to iteratively ref…

2024

Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms

NeurIPS 2024poster

Modern vision models are trained on very large noisy datasets. While these models acquire strong capabilities, they may not follow the user's intent to output the desired results in certain aspects, e.g., visual aesthetic, preferred style, and responsibility. In this paper, we target the realm of vi…

Cited by 3SourcePDFScholar
2024

BiPFT: Binary Pre-trained Foundation Transformer with Low-Rank Estimation of Binarization Residual Polynomials

AAAI 2024technical

Pretrained foundation models offer substantial benefits for a wide range of downstream tasks, which can be one of the most potential techniques to access artificial general intelligence. However, scaling up foundation transformers for maximal task-agnostic knowledge has brought about computational c…

2024

CONSIDER: Commonalities and Specialties Driven Multilingual Code Retrieval Framework

AAAI 2024technical

Multilingual code retrieval aims to find code snippets relevant to a user's query from a multilingual codebase, which plays a crucial role in software development and expands their application scenarios compared to classical monolingual code retrieval. Despite the performance improvements achieved b…

2024

Can Language Models Learn to Skip Steps?

NeurIPS 2024poster

Trained on vast corpora of human language, language models demonstrate emergent human-like reasoning abilities. Yet they are still far from true intelligence, which opens up intriguing opportunities to explore the parallels of humans and model behaviors. In this work, we study the ability to skip st…

2024

CariesXrays: Enhancing Caries Detection in Hospital-Scale Panoramic Dental X-rays via Feature Pyramid Contrastive Learning

AAAI 2024technical

Dental caries has been widely recognized as one of the most prevalent chronic diseases in the field of public health. Despite advancements in automated diagnosis across various medical domains, it remains a substantial challenge for dental caries detection due to its inherent variability and intrica…

2024

CoMERA: Computing- and Memory-Efficient Training via Rank-Adaptive Tensor Optimization

NeurIPS 2024poster

Training large AI models such as LLMs and DLRMs costs massive GPUs and computing time. The high training cost has become only affordable to big tech companies, meanwhile also causing increasing concerns about the environmental impact. This paper presents CoMERA, a **Co**mputing- and **M**emory-**E**…

2024

Collaborative Cognitive Diagnosis with Disentangled Representation Learning for Learner Modeling

NeurIPS 2024poster

Learners sharing similar implicit cognitive states often display comparable observable problem-solving performances. Leveraging collaborative connections among such similar learners proves valuable in comprehending human learning. Motivated by the success of collaborative modeling in various domains…

2024

DACOOP-A: Decentralized Adaptive Cooperative Pursuit via Attention

RA-L 2024

Integrating rule-based policies into reinforcement learning promises to improve data efficiency and generalization in cooperative pursuit problems. However, most implementations do not properly distinguish the influence of neighboring robots in observation embedding or inter-robot interaction rules,

Cited by 13SourcecodeScholar
2024

Decoupled Self-Adaptive Distribution Regularization for Few-Shot Image Classification

ICASSP 2024accepted

The feature dispersion, arising from the inherent constraints of data scarcity, has emerged as a prominent challenge in the domain of few-shot learning. In this paper, we propose a novel Self-adaptive Distribution Regularization (SADR) approach, which can adaptively bridge the semantic gaps across d…

Cited by 0SourceScholar
2024

DeepZero: Scaling Up Zeroth-Order Optimization for Deep Model Training

ICLR 2024poster

Zeroth-order (ZO) optimization has become a popular technique for solving machine learning (ML) problems when first-order (FO) information is difficult or impossible to obtain. However, the scalability of ZO optimization remains an open problem: Its use has primarily been limited to relatively small…

2024

ECON: On the Detection and Resolution of Evidence Conflicts

EMNLP 2024main

The rise of large language models (LLMs) has significantly influenced the quality of information in decision-making systems, leading to the prevalence of AI-generated content and challenges in detecting misinformation and managing conflicting information, or “inter-evidence conflicts.” This study in…

2024

Enhancing Cross-Modal Retrieval via Visual-Textual Prompt Hashing

IJCAI 2024poster

Cross-modal hashing has garnered considerable research interest due to its rapid retrieval and low storage costs. However, the majority of existing methods suffer from the limitations of context loss and information redundancy, particularly in simulated textual environments enriched with manually an…

Cited by 3SourcePDFScholar
2024

EventGround: Narrative Reasoning by Grounding to Eventuality-centric Knowledge Graphs

COLING 2024main

Narrative reasoning relies on the understanding of eventualities in story contexts, which requires a wealth of background world knowledge. To help machines leverage such knowledge, existing solutions can be categorized into two groups. Some focus on implicitly modeling eventuality knowledge by pretr…

2024

Exploiting Descriptive Completeness Prior for Cross Modal Hashing with Incomplete Labels

NeurIPS 2024poster

In this paper, we tackle the challenge of generating high-quality hash codes for cross-modal retrieval in the presence of incomplete labels, which creates uncertainty in distinguishing between positive and negative pairs. Vision-language models such as CLIP offer a potential solution by providing ge…

2024

FlightBERT++: A Non-autoregressive Multi-Horizon Flight Trajectory Prediction Framework

AAAI 2024technical

Flight Trajectory Prediction (FTP) is an essential task in Air Traffic Control (ATC), which can assist air traffic controllers in managing airspace more safely and efficiently. Existing approaches generally perform multi-horizon FTP tasks in an autoregressive manner, thereby suffering from error acc…

Cited by 39SourcePDFScholar
2024

GroupCover: A Secure, Efficient and Scalable Inference Framework for On-device Model Protection based on TEEs

ICML 2024poster

Due to the high cost of training DNN models, how to protect the intellectual property of DNN models, especially when the models are deployed to users' devices, is becoming an important topic. One practical solution is to use Trusted Execution Environments (TEEs) and researchers have proposed various…

Cited by 2SourcePDFScholar
2024

InstructDiffusion: A Generalist Modeling Interface for Vision Tasks

CVPR 2024poster

We present InstructDiffusion a unified and generic framework for aligning computer vision tasks with human instructions. Unlike existing approaches that integrate prior knowledge and pre-define the output space (e.g. categories and coordinates) for each vision task we cast diverse vision tasks into…

Cited by 109SourcePDFScholar
2024

Investigating and Mitigating the Side Effects of Noisy Views for Self-Supervised Clustering Algorithms in Practical Multi-View Scenarios

CVPR 2024poster

Multi-view clustering (MVC) aims at exploring category structures among multi-view data in self-supervised manners. Multiple views provide more information than single views and thus existing MVC methods can achieve satisfactory performance. However their performance might seriously degenerate when…

2024

Iterative PnP and its application in 3D-2D vascular image registration for robot navigation

ICRA 2024poster

This paper reports on a new real-time robotcentered 3D-2D vascular image alignment algorithm, which is robust to outliers and can align nonrigid shapes. Few works have managed to achieve both real-time and accurate performance for vascular intervention robots. This work bridges high-accuracy 3D-2D r…

Cited by 6SourceScholar
2024

Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image Recognition

ICML 2024poster

Large-scale pre-trained vision-language models (e.g., CLIP) have shown powerful zero-shot transfer capabilities in image recognition tasks. Recent approaches typically employ supervised fine-tuning methods to adapt CLIP for zero-shot multi-label image recognition tasks. However, obtaining sufficient…

Cited by 3SourcePDFScholar
2024

Learning to Complement and to Defer to Multiple Users

ECCV 2024poster

"With the development of Human-AI Collaboration in Classification (HAI-CC), integrating users and AI predictions becomes challenging due to the complex decision-making process. This process has three options: 1) AI autonomously classifies, 2) learning to complement, where AI collaborates with users,…

2024

LoRETTA: Low-Rank Economic Tensor-Train Adaptation for Ultra-Low-Parameter Fine-Tuning of Large Language Models

NAACL 2024long

Various parameter-efficient fine-tuning (PEFT) techniques have been proposed to enable computationally efficient fine-tuning while maintaining model performance. However, existing PEFT methods are still limited by the growing number of trainable parameters with the rapid deployment of Large Language…

2024

Masked Structural Growth for 2x Faster Language Model Pre-training

ICLR 2024poster

Accelerating large language model pre-training is a critical issue in present research. In this paper, we focus on speeding up pre-training by progressively growing from a small Transformer structure to a large one. There are two main research problems associated with progressive growth: determining…

2024

Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging

EMNLP 2024main

Supervised fine-tuning (SFT) is crucial for adapting Large Language Models (LLMs) to specific tasks. In this work, we demonstrate that the order of training data can lead to significant training imbalances, potentially resulting in performance degradation. Consequently, we propose to mitigate this i…

Cited by 1SourcePDFScholar
2024

Norma: A Noise Robust Memory-Augmented Framework for Whole Slide Image Classification

ECCV 2024poster

"In recent years, the Whole Slide Image (WSI) classification task has achieved great advancement due to the success of Multiple Instance Learning (MIL). However, the MIL-based studies usually consider instances within each bag as unordered, potentially resulting in the missing of local and global co…

2024

One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

NeurIPS 2024poster

We introduce VideoLISA, a video-based multimodal large language model designed to tackle the problem of language-instructed reasoning segmentation in videos. Leveraging the reasoning capabilities and world knowledge of large language models, and augmented by the Segment Anything Model, VideoLISA gen…

2024

Optimizing Code Retrieval: High-Quality and Scalable Dataset Annotation through Large Language Models

EMNLP 2024main

Code retrieval aims to identify code from extensive codebases that semantically aligns with a given query code snippet. Collecting a broad and high-quality set of query and code pairs is crucial to the success of this task. However, existing data collection methods struggle to effectively balance sc…

Cited by 5SourcePDFScholar
2024

PSALM: Pixelwise Segmentation with Large Multi-modal Model

ECCV 2024poster

"PSALM is a powerful extension of the Large Multi-modal Model (LMM) to address the segmentation task challenges. To overcome the limitation of the LMM being limited to textual output, PSALM incorporates a mask decoder and a well-designed input schema to handle a variety of segmentation tasks. This s…

2024

Pixel-GS Density Control with Pixel-aware Gradient for 3D Gaussian Splatting

ECCV 2024poster

"3D Gaussian Splatting (3DGS) has demonstrated impressive novel view synthesis results and advancing real-time rendering performance. However, the effectiveness of 3DGS heavily relies on the quality of the initial point cloud, as poor initialization can result in blurring and needle-like artifacts.…

2024

RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation

NeurIPS 2024poster

Despite Retrieval-Augmented Generation (RAG) has shown promising capability in leveraging external knowledge, a comprehensive evaluation of RAG systems is still challenging due to the modular nature of RAG, evaluation of long-form responses and reliability of measurements. In this paper, we propose…

2024

Rethinking The Training And Evaluation of Rich-Context Layout-to-Image Generation

NeurIPS 2024poster

Recent advancements in generative models have significantly enhanced their capacity for image generation, enabling a wide range of applications such as image editing, completion and video editing. A specialized area within generative modeling is layout-to-image (L2I) generation, where predefined lay…

2024

Segment and Caption Anything

CVPR 2024poster

We propose a method to efficiently equip the Segment Anything Model (SAM) with the ability to generate regional captions. SAM presents strong generalizability to segment anything while is short for semantic understanding. By introducing a lightweight query-based feature mixer we align the region-spe…

2024

Self-Supervised Heterogeneous Graph Learning: a Homophily and Heterogeneity View

ICLR 2024poster

Self-supervised heterogeneous graph learning has achieved promising results in various real applications, but it still suffers from the following issues: (i) meta-paths can be employed to capture the homophily in the heterogeneous graph, but meta-paths are human-defined, requiring substantial exper…

Cited by 9SourcePDFScholar
2024

SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking Mechanisms

ICML 2024poster

Towards energy-efficient artificial intelligence similar to the human brain, the bio-inspired spiking neural networks (SNNs) have advantages of biological plausibility, event-driven sparsity, and binary activation. Recently, large-scale language models exhibit promising generalization capability, ma…

2024

Synergetic Event Understanding: A Collaborative Approach to Cross-Document Event Coreference Resolution with Large Language Models

ACL 2024long

Cross-document event coreference resolution (CDECR) involves clustering event mentions across multiple documents that refer to the same real-world events. Existing approaches utilize fine-tuning of small language models (SLMs) like BERT to address the compatibility among the contexts of event mentio…

2024

TEG-DB: A Comprehensive Dataset and Benchmark of Textual-Edge Graphs

NeurIPS 2024poster

Text-Attributed Graphs (TAGs) augment graph structures with natural language descriptions, facilitating detailed depictions of data and their interconnections across various real-world settings. However, existing TAG datasets predominantly feature textual information only at the nodes, with edges ty…

2024

Towards Accurate and Fair Cognitive Diagnosis via Monotonic Data Augmentation

NeurIPS 2024poster

Intelligent education stands as a prominent application of machine learning. Within this domain, cognitive diagnosis (CD) is a key research focus that aims to diagnose students' proficiency levels in specific knowledge concepts. As a crucial task within the field of education, cognitive diagnosis en…

Cited by 2SourcePDFScholar
2024

Towards Reliable Advertising Image Generation Using Human Feedback

ECCV 2024poster

"In the e-commerce realm, compelling advertising images are pivotal for attracting customer attention. While generative models automate image generation, they often produce substandard images that may mislead customers and require significant labor costs to inspect. This paper delves into increasing…

2024

Unified Lexical Representation for Interpretable Visual-Language Alignment

NeurIPS 2024poster

Visual-Language Alignment (VLA) has gained a lot of attention since CLIP's groundbreaking work. Although CLIP performs well, the typical direct latent feature alignment lacks clarity in its representation and similarity scores. On the other hand, lexical representation, a vector whose element repr…

2024

Zero-1-to-3: Domain-Level Zero-Shot Cognitive Diagnosis via One Batch of Early-Bird Students towards Three Diagnostic Objectives

AAAI 2024technical

Cognitive diagnosis seeks to estimate the cognitive states of students by exploring their logged practice quiz data. It plays a pivotal role in personalized learning guidance within intelligent education systems. In this paper, we focus on an important, practical, yet often underexplored task: domai…

2023

All in Tokens: Unifying Output Space of Visual Tasks via Soft Token

ICCV 2023oral

We introduce AiT, a unified output representation for various vision tasks, which is a crucial step towards general-purpose vision task solvers. Despite the challenges posed by the high-dimensional and task-specific outputs, we showcase the potential of using discrete representation (VQ-VAE) to mode…

Cited by 48PDFcodeScholar
2023

An AMR-based Link Prediction Approach for Document-level Event Argument Extraction

ACL 2023long

Recent works have introduced Abstract Meaning Representation (AMR) for Document-level Event Argument Extraction (Doc-level EAE), since AMR provides a useful interpretation of complex semantic structures and helps to capture long-distance dependency. However, in these works AMR is used only implicitl…

2023

AugESC: Dialogue Augmentation with Large Language Models for Emotional Support Conversation

ACL 2023findings

Crowdsourced dialogue corpora are usually limited in scale and topic coverage due to the expensive cost of data curation. This would hinder the generalization of downstream dialogue models to open-domain topics. In this work, we leverage large language models for dialogue augmentation in the task of…

2023

Bridging the Gap to Real-World Object-Centric Learning

ICLR 2023poster

Humans naturally decompose their environment into entities at the appropriate level of abstraction to act in the world. Allowing machine learning algorithms to derive this decomposition in an unsupervised way has become an important line of research. However, current methods are restricted to simula…

Cited by 144SourcePDFScholar
2023

Building Multi-domain Dialog State Trackers from Single-domain Dialogs

EMNLP 2023long main

Existing multi-domain dialog state tracking (DST) models are developed based on multi-domain dialogs, which require significant manual effort to define domain relations and collect data. This process can be challenging and expensive, particularly when numerous domains are involved. In this paper, we…

Cited by 0SourceScholar
2023

CASE: Aligning Coarse-to-Fine Cognition and Affection for Empathetic Response Generation

ACL 2023long

Empathetic conversation is psychologically supposed to be the result of conscious alignment and interaction between the cognition and affection of empathy. However, existing empathetic dialogue models usually consider only the affective aspect or treat cognition and affection in isolation, which lim…

2023

Click: Controllable Text Generation with Sequence Likelihood Contrastive Learning

ACL 2023findings

It has always been an important yet challenging problem to control language models to avoid generating texts with undesirable attributes, such as toxic language and unnatural repetition. We introduce Leo for controllable text generation, which needs no modification to the model architecture and faci…

2023

Coarse-to-Fine Amodal Segmentation with Shape Prior

ICCV 2023poster

Amodal object segmentation is a challenging task that involves segmenting both visible and occluded parts of an object. In this paper, we propose a novel approach, called Coarse-to-Fine Segmentation (C2F-Seg), that addresses this problem by progressively modeling the amodal segmentation. C2F-Seg…

Cited by 24PDFcodeScholar
2023

Curriculum Learning for Graph Neural Networks: Which Edges Should We Learn First

NeurIPS 2023poster

Graph Neural Networks (GNNs) have achieved great success in representing data with dependencies by recursively propagating and aggregating messages along the edges. However, edges in real-world graphs often have varying degrees of difficulty, and some edges may even be noisy to the downstream tasks.…

2023

DETR Does Not Need Multi-Scale or Locality Design

ICCV 2023poster

This paper presents an improved DETR detector that maintains a "plain" nature: using a single-scale feature map and global cross-attention calculations without specific locality constraints, in contrast to previous leading DETR-based detectors that reintroduce architectural inductive biases of multi…

Cited by 30PDFcodeScholar
2023

Distributed Marker Representation for Ambiguous Discourse Markers and Entangled Relations

ACL 2023long

Discourse analysis is an important task because it models intrinsic semantic structures between sentences in a document. Discourse markers are natural representations of discourse in our daily language. One challenge is that the markers as well as pre-defined and human-labeled discourse relations ca…

Cited by 2SourcePDFScholar
2023

Dual Cache for Long Document Neural Coreference Resolution

ACL 2023long

Recent works show the effectiveness of cache-based neural coreference resolution models on long documents. These models incrementally process a long document from left to right and extract relations between mentions and entities in a cache, resulting in much lower memory and computation cost compare…

2023

Dual Learning with Dynamic Knowledge Distillation for Partially Relevant Video Retrieval

ICCV 2023poster

Almost all previous text-to-video retrieval works assume that videos are pre-trimmed with short durations. However, in practice, videos are generally untrimmed containing much background content. In this work, we investigate the more practical but challenging Partially Relevant Video Retrieval (PRVR…

Cited by 19PDFcodeScholar
2023

Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus

EMNLP 2023long main

Large Language Models (LLMs) have gained significant popularity for their impressive performance across diverse fields. However, LLMs are prone to hallucinate untruthful or nonsensical outputs that fail to meet user expectations in many real-world applications. Existing works for detecting hallucina…

Cited by 0SourcecodeScholar
2023

Evaluating Open-QA Evaluation

NeurIPS 2023poster

This study focuses on the evaluation of the Open Question Answering (Open-QA) task, which can directly estimate the factuality of large language models (LLMs). Current automatic evaluation methods have shown limitations, indicating that human evaluation still remains the most reliable approach. We i…

2023

Exploiting Abstract Meaning Representation for Open-Domain Question Answering

ACL 2023findings

The Open-Domain Question Answering (ODQA) task involves retrieving and subsequently generating answers from fine-grained relevant passages within a database. Current systems leverage Pretrained Language Models (PLMs) to model the relationship between questions and passages. However, the diversity in…

2023

FairLISA: Fair User Modeling with Limited Sensitive Attributes Information

NeurIPS 2023poster

User modeling techniques profile users' latent characteristics (e.g., preference) from their observed behaviors, and play a crucial role in decision-making. Unfortunately, traditional user models may unconsciously capture biases related to sensitive attributes (e.g., gender) from behavior data, even…

2023

KECOR: Kernel Coding Rate Maximization for Active 3D Object Detection

ICCV 2023poster

Achieving a reliable LiDAR-based object detector in autonomous driving is paramount, but its success hinges on obtaining large amounts of precise 3D annotations. Active learning (AL) seeks to mitigate the annotation burden through algorithms that use fewer labels and can attain performance comparabl…

Cited by 19PDFScholar
2023

KPT: Keyword-Guided Pre-training for Grounded Dialog Generation

AAAI 2023technical

Incorporating external knowledge into the response generation process is essential to building more helpful and reliable dialog agents. However, collecting knowledge-grounded conversations is often costly, calling for a better pre-trained model for grounded dialog generation that generalizes well w.…

Cited by 3SourcePDFScholar
2023

On Data Scaling in Masked Image Modeling

CVPR 2023poster

Scaling properties have been one of the central issues in self-supervised pre-training, especially the data scalability, which has successfully motivated the large-scale self-supervised pre-trained language models and endowed them with significant modeling capabilities. However, scaling properties s…

2023

Optical Flow-Based Vascular Respiratory Motion Compensation

RA-L 2023

This letter develops a new vascular respiratory motion compensation algorithm, Motion-Related Compensation (MRC), to conduct vascular respiratory motion compensation by extrapolating the correlation between invisible vascular and visible non-vascular. Robot-assisted vascular intervention can signifi

Cited by 9SourceScholar
2023

Plan, Verify and Switch: Integrated Reasoning with Diverse X-of-Thoughts

EMNLP 2023long main

As large language models (LLMs) have shown effectiveness with different prompting methods, such as Chain of Thought, Program of Thought, we find that these methods have formed a great complementarity to each other on math reasoning tasks. In this work, we propose XoT, an integrated problem solving f…

Cited by 0SourcecodeScholar
2023

Reinforced Potential Field for Multi-Robot Motion Planning in Cluttered Environments

IROS 2023poster

Motion planning is challenging for multiple robots in cluttered environments without communication, especially in view of real-time efficiency, motion safety, distributed computation, and trajectory optimality, etc. In this paper, a reinforced potential field method is developed for distributed mult…

Cited by 6SourceScholar
2023

Rethinking Amodal Video Segmentation from Learning Supervised Signals with Object-centric Representation

ICCV 2023poster

Video amodal segmentation is a particularly challenging task in computer vision, which requires to deduce the full shape of an object from the visible parts of it. Recently, some studies have achieved promising performance by using motion flow to integrate information across frames under a self-supe…

Cited by 10PDFcodeScholar
2023

Revealing the Dark Secrets of Masked Image Modeling

CVPR 2023poster

Masked image modeling (MIM) as pre-training is shown to be effective for numerous vision downstream tasks, but how and where MIM works remain unclear. In this paper, we compare MIM with the long-dominant supervised pre-trained models from two perspectives, the visualizations and the experiments, to…

2023

Side Adapter Network for Open-Vocabulary Semantic Segmentation

CVPR 2023highlight

This paper presents a new framework for open-vocabulary semantic segmentation with the pre-trained vision-language model, named SAN. Our approach models the semantic segmentation task as a region recognition problem. A side network is attached to a frozen CLIP model with two branches: one for predic…

2023

StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical Understanding

EMNLP 2023long main

Analogy-making between narratives is crucial for human reasoning. In this paper, we evaluate the ability to identify and generate analogies by constructing a first-of-its-kind large-scale story-level analogy corpus, \textsc{StoryAnalogy}, which contains 24K story pairs from diverse domains with huma…

Cited by 0SourcecodeScholar
2023

TinyMIM: An Empirical Study of Distilling MIM Pre-Trained Models

CVPR 2023poster

Masked image modeling (MIM) performs strongly in pre-training large vision Transformers (ViTs). However, small models that are critical for real-world applications cannot or only marginally benefit from this pre-training approach. In this paper, we explore distillation techniques to transfer the suc…

2023

Unsupervised Open-Vocabulary Object Localization in Videos

ICCV 2023poster

In this paper, we show that recent advances in video representation learning and pre-trained vision-language models allow for substantial improvements in self-supervised video object localization. We propose a method that first localizes objects in videos via a slot attention approach and then assig…

Cited by 7PDFcodeScholar
2023

iCLIP: Bridging Image Classification and Contrastive Language-Image Pre-Training for Visual Recognition

CVPR 2023poster

This paper presents a method that effectively combines two prevalent visual recognition methods, i.e., image classification and contrastive language-image pre-training, dubbed iCLIP. Instead of naive multi-task learning that use two separate heads for each task, we fuse the two tasks in a deep fashi…

2022

"A Simple Approach and Benchmark for 21,000-Category Object Detection"

ECCV 2022poster

"Current object detection systems and benchmarks typically handle a limited number of categories, up to about a thousand categories. This paper scales the number of categories for object detection systems and benchmarks up to 21,000, by leveraging existing object detection and image classification d…

2022

A Simple Baseline for Open-Vocabulary Semantic Segmentation with Pre-trained Vision-Language Model

ECCV 2022poster

"Recently, open-vocabulary image classification by vision language pre-training has demonstrated incredible achievements, that the model can classify arbitrary categories without seeing additional annotated images of that category. However, it is still unclear how to make the open-vocabulary recogni…

2022

A Unified Dialogue User Simulator for Few-shot Data Augmentation

EMNLP 2022finding

Pre-trained language models have shown superior performance in task-oriented dialogues. However, existing datasets are on limited scales, which cannot support large-scale pre-training. Fortunately, various data augmentation methods have been developed to augment large-scale task-oriented dialogue co…

Cited by 27SourcePDFScholar
2022

Could Giant Pre-trained Image Models Extract Universal Representations?

NeurIPS 2022accept

Frozen pretrained models have become a viable alternative to the pretraining-then-finetuning paradigm for transfer learning. However, with frozen models there are relatively few parameters available for adapting to downstream tasks, which is problematic in computer vision where tasks vary significan…

Cited by 11SourcePDFScholar
2022

DORE: Document Ordered Relation Extraction based on Generative Framework

EMNLP 2022finding

In recent years, there is a surge of generation-based information extraction work, which allows a more direct use of pre-trained language models and efficiently captures output dependencies. However, previous generative methods using lexical representation do not naturally fit document-level relatio…

2022

Dialogue Meaning Representation for Task-Oriented Dialogue Systems

EMNLP 2022finding

Dialogue meaning representation formulates natural language utterance semantics in their conversational context in an explicit and machine-readable form. Previous work typically follows the intent-slot framework, which is easy for annotation yet limited in scalability for complex linguistic expressi…

2022

Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning

NeurIPS 2022accept

Vision transformers have recently achieved competitive results across various vision tasks but still suffer from heavy computation costs when processing a large number of tokens. Many advanced approaches have been developed to reduce the total number of tokens in the large-scale vision transformers,…

Cited by 32SourcePDFScholar
2022

Fantastic Questions and Where to Find Them: FairytaleQA – An Authentic Dataset for Narrative Comprehension

ACL 2022long

Question answering (QA) is a fundamental means to facilitate assessment and training of narrative comprehension skills for both machines and young children, yet there is scarcity of high-quality QA datasets carefully designed to serve this purpose. In particular, existing datasets rarely distinguish…

2022

Inductive Relation Prediction Using Analogy Subgraph Embeddings

ICLR 2022poster

Prevailing methods for relation prediction in heterogeneous graphs aim at learning latent representations (i.e., embeddings) of observed nodes and relations, and thus are limited to the transductive setting where the relation types must be known during training. Here, we propose ANalogy SubGraphE…

Cited by 7SourcePDFScholar
2022

Interactive Information Extraction by Semantic Information Graph

IJCAI 2022poster

Information extraction (IE) mainly focuses on three highly correlated subtasks, i.e., entity extraction, relation extraction and event extraction. Recently, there are studies using Abstract Meaning Representation (AMR) to utilize the intrinsic correlations among these three subtasks. AMR based model…

2022

It is AI’s Turn to Ask Humans a Question: Question-Answer Pair Generation for Children’s Story Books

ACL 2022long

Existing question answering (QA) techniques are created mainly to answer questions asked by humans. But in educational applications, teachers often need to decide what questions they should ask, in order to help students to improve their narrative understanding capabilities. We design an automated q…

2022

Learning Enhanced Representation for Tabular Data via Neighborhood Propagation

NeurIPS 2022accept

Prediction over tabular data is an essential and fundamental problem in many important downstream tasks. However, existing methods either take a data instance of the table independently as input or do not fully utilize the multi-row features and labels to directly change and enhance the target data…

2022

Multi-robot Cooperative Pursuit via Potential Field-Enhanced Reinforcement Learning

ICRA 2022poster

It is of great challenge, though promising, to coordinate collective robots for hunting an evader in a decentralized manner purely in light of local observations. In this paper, this challenge is addressed by a novel hybrid cooperative pursuit algorithm that combines reinforcement learning with the…

Cited by 57SourceScholar
2022

PSS: Progressive Sample Selection for Open-World Visual Representation Learning

ECCV 2022poster

"We propose a practical open-world representation learning setting where the objective is to learn the representations for unseen categories without prior knowledge or access to images associated with these novel categories during training. Existing open-world representation learning methods, howeve…

2022

RLET: A Reinforcement Learning Based Approach for Explainable QA with Entailment Trees

EMNLP 2022main

Interpreting the reasoning process from questions to answers poses a challenge in approaching explainable QA. A recently proposed structured reasoning format, entailment tree, manages to offer explicit logical deductions with entailment steps in a tree structure. To generate entailment trees, prior…

2022

SSEGCN: Syntactic and Semantic Enhanced Graph Convolutional Network for Aspect-based Sentiment Analysis

NAACL 2022long

Aspect-based Sentiment Analysis (ABSA) aims to predict the sentiment polarity towards a particular aspect in a sentence. Recently, graph neural networks based on dependency tree convey rich structural information which is proven to be utility for ABSA. However, how to effectively harness the semanti…

2022

Self-supervised Amodal Video Object Segmentation

NeurIPS 2022accept

Amodal perception requires inferring the full shape of an object that is partially occluded. This task is particularly challenging on two levels: (1) it requires more information than what is contained in the instant retina or imaging sensor, (2) it is difficult to obtain enough well-annotated amoda…

Cited by 11SourcePDFScholar
2022

SimMIM: A Simple Framework for Masked Image Modeling

CVPR 2022poster

This paper presents SimMIM, a simple framework for masked image modeling. We have simplified recently proposed relevant approaches, without the need for special designs, such as block-wise masking and tokenization via discrete VAE or clustering. To investigate what makes a masked image modeling task…

Cited by 1610PDFcodeScholar
2022

Swin Transformer V2: Scaling Up Capacity and Resolution

CVPR 2022poster

We present techniques for scaling Swin Transformer [??] up to 3 billion parameters and making it capable of training with images of up to 1,536x1,536 resolution. By scaling up capacity and resolution, Swin Transformer sets new records on four representative vision benchmarks: 84.0% top-1 accuracy on…

Cited by 2410PDFcodeScholar
2022

Why Propagate Alone? Parallel Use of Labels and Features on Graphs

ICLR 2022poster

One of the challenges of graph-based semi-supervised learning over ordinary supervised learning for classification tasks lies in label utilization. The direct use of ground-truth labels in graphs for training purposes can result in a parametric model learning trivial degenerate solutions (e.g., an…

Cited by 12SourcePDFScholar
2021

A Unified Generative Framework for Aspect-based Sentiment Analysis

ACL 2021long

Aspect-based Sentiment Analysis (ABSA) aims to identify the aspect terms, their corresponding sentiment polarities, and the opinion terms. There exist seven subtasks in ABSA. Most studies only focus on the subsets of these subtasks, which leads to various complicated ABSA models while hard to solve…

2021

A Unified Generative Framework for Various NER Subtasks

ACL 2021long

Named Entity Recognition (NER) is the task of identifying spans that represent entities in sentences. Whether the entity spans are nested or discontinuous, the NER task can be categorized into the flat NER, nested NER, and discontinuous NER subtasks. These subtasks have been mainly solved by the tok…

2021

Adapting Language Models for Zero-shot Learning by Meta-tuning on Dataset and Prompt Collections

EMNLP 2021finding

Large pre-trained language models (LMs) such as GPT-3 have acquired a surprising ability to perform zero-shot learning. For example, to classify sentiment without any training examples, we can “prompt” the LM with the review and the label description “Does the user like this movie?”, and ask whether…

2021

Bayesian Inference with Certifiable Adversarial Robustness

AISTATS 2021poster

We consider adversarial training of deep neural networks through the lens of Bayesian learning and present a principled framework for adversarial training of Bayesian Neural Networks (BNNs) with certifiable guarantees. We rely on techniques from constraint relaxation of non-convex optimisation probl…

2021

Bootstrap Your Object Detector via Mixed Training

NeurIPS 2021spotlight

We introduce MixTraining, a new training paradigm for object detection that can improve the performance of existing detectors for free. MixTraining enhances data augmentation by utilizing augmentations of different strengths while excluding the strong augmentations of certain training samples that m…

2021

End-to-End Semi-Supervised Object Detection With Soft Teacher

ICCV 2021poster

Previous pseudo-label approaches for semi-supervised object detection typically follow a multi-stage schema, with the first stage to train an initial detector on a few labeled data, followed by the pseudo labeling and re-training stage on unlabeled data. These multi-stage methods complicate the trai…

Cited by 620PDFcodeScholar
2021

Fork or Fail: Cycle-Consistent Training with Many-to-One Mappings

AISTATS 2021poster

Cycle-consistent training is widely used for jointly learning a forward and inverse mapping between two domains of interest without the cumbersome requirement of collecting matched pairs within each domain. In this regard, the implicit assumption is that there exists (at least approximately) a groun…

2021

GRIN: Generative Relation and Intention Network for Multi-agent Trajectory Prediction

NeurIPS 2021poster

Learning the distribution of future trajectories conditioned on the past is a crucial problem for understanding multi-agent systems. This is challenging because humans make decisions based on complex social relations and personal intents, resulting in highly complex uncertainties over trajectories.…

Cited by 50SourcePDFScholar
2021

Graph Neural Networks Inspired by Classical Iterative Algorithms

ICML 2021oral

Despite the recent success of graph neural networks (GNN), common architectures often exhibit significant limitations, including sensitivity to oversmoothing, long-range dependencies, and spurious edges, e.g., as can occur as a result of graph heterophily or adversarial attacks. To at least partiall…

2021

Hierarchical Network Based on the Fusion of Static and Dynamic Features for Speech Emotion Recognition

ICASSP 2021accepted

Many studies on automatic speech emotion recognition (SER) have been devoted to extracting meaningful emotional features for generating emotion-relevant representations. However, they generally ignore the complementary learning of static and dynamic features, leading to limited performances. In this…

Cited by 0SourceScholar
2021

Learning Hierarchical Graph Neural Networks for Image Clustering

ICCV 2021poster

We propose a hierarchical graph neural network (GNN) model that learns how to cluster a set of images into an unknown number of identities using a training set of images annotated with labels belonging to a disjoint set of identities. Our hierarchical GNN uses a novel approach to merge connected com…

Cited by 53PDFcodeScholar
2021

Partial-Label and Structure-constrained Deep Coupled Factorization Network

AAAI 2021technical

In this paper, we technically propose an enriched prior guided framework, called Dual-constrained Deep Semi-Supervised Coupled Factorization Network (DS2CF-Net), for discovering hierarchical coupled data representation. To extract hidden deep features, DS2CF-Net is formulated as a partial-label and…

Cited by 6SourcePDFScholar
2021

Propagate Yourself: Exploring Pixel-Level Consistency for Unsupervised Visual Representation Learning

CVPR 2021poster

Contrastive learning methods for unsupervised visual representation learning have reached remarkable levels of transfer performance. We argue that the power of contrastive learning has yet to be fully unleashed, as current methods are trained only on instance-level pretext tasks, leading to represen…

Cited by 498PDFcodeScholar
2021

Prototype-Supervised Adversarial Network for Targeted Attack of Deep Hashing

CVPR 2021poster

Due to its powerful capability of representation learning and high-efficiency computation, deep hashing has made significant progress in large-scale image retrieval. However, deep hashing networks are vulnerable to adversarial examples, which is a practical secure problem but seldom studied in hashi…

Cited by 62PDFcodeScholar
2021

Semantics Disentangling for Generalized Zero-Shot Learning

ICCV 2021poster

Generalized zero-shot learning (GZSL) aims to classify samples under the assumption that some classes are not observable during training. To bridge the gap between the seen and unseen classes, most GZSL methods attempt to associate the visual features of seen classes with attributes or to generate u…

Cited by 149PDFcodeScholar
2021

Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows

ICCV 2021poster

This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision. Challenges in adapting Transformer from language to vision arise from differences between the two domains, such as large variations in the scale of visual ent…

Cited by 29925PDFcodeScholar