← Search

Yi ZHANG

210 accepted papers

2026

AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis

CVPR 2026

Generating human grasping poses that accurately reflect both object geometry and user-specified interaction semantics is essential for natural hand-object interactions in AR/VR and embodied AI. However, existing semantic grasping approaches struggle with the large modality gap between 3D object repr

Cited by 0SourceScholar
2026

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning

ICLR 2026poster

Large language models (LLMs), despite possessing latent safety understanding from their vast pretraining data, remain vulnerable to generating harmful content and exhibit issues such as over-refusal and utility degradation after safety alignment. Current safety alignment methods often result in supe…

Cited by 0SourceScholar
2026

Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs

ICLR 2026poster

Fully open multimodal large language models (MLLMs) currently lag behind proprietary counterparts, primarily due to a significant gap in data quality for supervised fine-tuning (SFT). Existing open-source datasets are often plagued by widespread noise and a critical deficit in complex reasoning dat…

Cited by 0SourceScholar
2026

CAFU: Constrained Alignment and Filtered Uniformity for Denoising Recommendation

AAAI 2026technical

In recommender systems, recent advances highlight the critical role of alignment and uniformity (AU) in representation learning. Specifically, AU-based methods pull positive user-item pairs closer (alignment) and spread the overall representation distribution (uniformity), typically relying on obser

Cited by 0SourcePDFScholar
2026

Computer Agent Arena: Toward Human-Centric Evaluation and Analysis of Computer-Use Agents

ICLR 2026poster

As Computer-Use Agents (CUAs) proliferate and grow increasingly capable, evaluation has become more challenging: static, manually curated benchmarks are narrow in domain, contamination-prone, and environment-heavy, and they diverge substantially from user-driven, real-world evaluation. We present Co…

Cited by 0SourcecodeScholar
2026

Consistent Noisy Latent Rewards for Trajectory Preference Optimization in Diffusion Models

ICLR 2026poster

Recent advances in diffusion models for visual generation have sparked interest in human preference alignment, similar to developments in Large Language Models. While reward model (RM) based approaches enable trajectory-aware optimization by evaluating intermediate timesteps, they face two critical…

Cited by 0SourceScholar
2026

Distributed Virtual Model Control for Scalable Human-Robot Collaboration in Shared Workspace

ICRA 2026poster

We present a decentralized, agent agnostic, and safety-aware control framework for human–robot collaboration based on Virtual Model Control (VMC). In our approach, both humans and robots are embedded in the same virtual-component-shaped workspace, where motion is the result of the interaction with v…

2026

Do Large Language Models Think like the Brain? Sentence-Level Evidences from Layer-Wise Embeddings and fMRI

AAAI 2026technical

Understanding whether large language models (LLMs) and the human brain converge on similar computational principles remains a fundamental and important question in cognitive neuroscience and AI. Do the brain-like patterns observed in LLMs emerge simply from scaling, or do they reflect deeper alignme

Cited by 0SourcePDFScholar
2026

Evolutionary Generation of Multi-Agent Systems

ICML 2026poster

Large language model (LLM)–based multi-agent systems (MAS) show strong promise for complex reasoning, planning, and tool-augmented tasks, but designing effective MAS architectures remains labor-intensive, brittle, and hard to generalize. Existing automatic MAS generation methods either rely on code …

Cited by 0SourceScholar
2026

ExpAlign: Expectation-Guided Vision–Language Alignment for Open-Vocabulary Grounding

ICML 2026poster

Open-vocabulary grounding requires accurate vision-language alignment under weak supervision, yet existing methods either rely on global sentence embeddings that lack fine-grained expressiveness or introduce token-level alignment with explicit supervision or heavy cross-attention designs. We propose…

Cited by 0SourceScholar
2026

FACT: Fuzzy Alignment with Comorbidity Topology for Reliable Multi-Label Medical Image Diagnosis

ICML 2026poster

In clinical practice, patients often present with multiple co-occurring diseases, yet most existing Multi-Label-Diagnosis (MLD) methods treat diagnosis as a rigid discriminative partitioning task, implicitly assuming that overlapping pathologies are separable. This assumption is problematic in medic…

Cited by 0SourceScholar
2026

Fragile by Design: On the Limits of Adversarial Defenses in Personalized DreamBooth Generation

AAAI 2026technical

Personalized AI applications such as DreamBooth enable the generation of customized content from user images, but also raise significant privacy concerns, particularly the risk of facial identity leakage. Recent defense mechanisms like Anti-DreamBooth attempt to mitigate this risk by injecting adver

Cited by 0SourcePDFScholar
2026

From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlas

ICML 2026poster

Public health reasoning requires population-level inference grounded in scientific evidence, expert consensus, and safety constraints. However, it remains underexplored as a structured machine learning problem with limited supervised signals and benchmarks. We introduce GlobalHealthAtlas, a large-sc…

Cited by 0SourceScholar
2026

GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and Reasoning

CVPR 2026

Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable perceptual and reasoning abilities. However, they struggle to perceive fine-grained geometric structures, constraining their ability of geometric understanding and visual reasoning. To address this, we propose GeoTikzBrid

Cited by 0SourcecodeScholar
2026

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation

CVPR 2026

With the rapid advancement of diffusion models, talking face generation has made remarkable progress. However, existing diffusion-based methods still require task-specific fine-tuning and large-scale audiovisual datasets, resulting in high computational costs that hinder accessibility for resource-c

Cited by 0SourcecodeScholar
2026

Internalizing Safety Understanding in Large Reasoning Models via Verification

ICML 2026poster

While explicit Chain-of-Thought (CoT) empowers large reasoning models (LRMs), it enables the generation of riskier final answers. Current alignment paradigms primarily rely on externally enforced compliance, optimizing models to detect malicious prompts rather than evaluating the safety of their own…

Cited by 0SourceScholar
2026

LLM-Orchestrated Diagnose–Plan–Treat for Mixed-Degradation CT Reconstruction

IJCAI 2026

Clinical Computed Tomography (CT) reconstruction often faces mixed degradations, where quantum noise, streak artifacts, and geometric distortions co-occur with various compositions and severities. Recently, all-in-one frameworks have outperformed traditional single-task models through degradation-sp

Cited by 0Scholar
2026

LURE: Latent Space Unblocking for Multi-Concept Reawakening in Diffusion Models

IJCAI 2026

Concept erasure aims to suppress sensitive content in diffusion models, but recent studies show that erased concepts can still be reawakened, revealing vulnerabilities in erasure methods. Existing reawakening methods mainly rely on prompt-level optimization to manipulate sampling trajectories, negle

Cited by 0Scholar
2026

Language Does Matter for Cross-Domain Few-Shot Visual Feature Enhancement

CVPR 2026

Cross-domain few-shot image interpretation (CD-FSII) has been significantly advanced by fine-tuning pre-trained visual feature models using limited labeled samples in target domains. However, profound cross-domain distribution discrepancies, along with inherent conflicts between extensive object vis

Cited by 0SourcecodeScholar
2026

Learning to Self-Verify Makes Language Models Better Reasoners

ICML 2026poster

Recent large language models (LLMs) achieve strong performance in generating promising reasoning paths for complex tasks. However, despite powerful generation ability, LLMs remain weak at verifying their own answers, revealing a persistent capability asymmetry between generation and self-verificatio…

Cited by 0SourceScholar
2026

Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation

ICLR 2026poster

While humans naturally learn and adapt from past experiences, large language models (LLMs) and their agentic counterparts often fail to retain reasoning from previous tasks and apply it in future contexts. We introduce **L**og-**A**ugmented **G**eneration (LAG), a novel framework that *directly reus…

Cited by 0SourcecodeScholar
2026

Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models

CVPR 2026

Group Relative Policy Optimization (GRPO) has shown promise in aligning image and video generative models with human preferences. However, applying it to modern flow matching models is challenging because of its deterministic sampling paradigm. Current methods address this issue by converting Ordina

Cited by 0SourceScholar
2026

Non-Parametric Probabilistic Robustness: A Conservative Risk Estimator under Unknown Perturbation Distributions

ICML 2026poster

Deep learning (DL) models, despite their remarkable success, remain vulnerable to small input perturbations that can cause erroneous outputs, motivating probabilistic robustness (PR) as a complementary notion to adversarial robustness (AR) for stochastic reliability assessment. However, existing PR …

Cited by 0SourceScholar
2026

Poisoned Distillation: Injecting Backdoors into Distilled Datasets Without Raw Data Access

AAAI 2026technical

Dataset distillation (DD) condenses large datasets into smaller synthetic ones to enhance training efficiency and reducing bandwidth. DD enables models to achieve comparable performance to those trained on the raw full dataset, making it popular for data sharing. Existing work shows that injecting b

Cited by 0SourcePDFScholar
2026

SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image Distillation

ICLR 2026poster

The Distribution Matching Distillation (DMD) has been successfully applied to text-to-image diffusion models such as Stable Diffusion (SD) 1.5. However, vanilla DMD suffers from convergence difficulties on large-scale flow-based text-to-image models, such as SD 3.5 and FLUX. In this paper, we first…

Cited by 0SourcecodeScholar
2026

TRAINING-FREE TEST-TIME ADAPTATION WITH BROWNIAN DISTANCE COVARIANCE IN VISION-LANGUAGE MODELS

ICASSP 2026poster

Vision-language models suffer performance degradation under domain shift, limiting real-world applicability. Existing test-time adaptation methods are computationally intensive, rely on back-propagation, and often focus on single modalities. To address these issues, we propose Training-free Test-Tim…

Cited by 0SourcePDFScholar
2026

TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding

AAAI 2026technical

Spatio-Temporal Video Grounding (STVG) aims to localize a spatio-temporal tube that corresponds to a given language query in an untrimmed video. This is a challenging task since it involves complex vision-language understanding and spatiotemporal reasoning. Recent works have explored weakly-superv

Cited by 0SourcePDFScholar
2026

Unlocking Dynamic Inter-Client Spatial Dependencies: A Federated Spatio-temporal Graph Learning Method for Traffic Flow Forecasting

AAAI 2026technical

Spatio-temporal graphs are powerful tools for modeling complex dependencies in traffic time series. However, the distributed nature of real-world traffic data across multiple stakeholders poses significant challenges in modeling and reconstructing inter-client spatial dependencies while adhering to

Cited by 0SourcePDFScholar
2026

VILTA: A VLM-in-the-Loop Adversary for Enhancing Driving Policy Robustness

AAAI 2026technical

The safe deployment of autonomous driving (AD) systems is fundamentally hindered by the long-tail problem, where rare yet critical driving scenarios are severely underrepresented in real-world data. Existing solutions including safety-critical scenario generation and closed-loop learning often rely

Cited by 0SourcePDFScholar
2025

ADC: Enhancing Function Calling Via Adversarial Datasets and Code Line-Level Feedback

ICASSP 2025accepted

Large Language Models (LLMs) have made significant strides in Natural Language Processing and coding, yet they struggle with robustness and accuracy in complex function calls. To tackle these challenges, this paper introduces ADC, an innovative approach that enhances LLMs’ ability to follow function…

Cited by 0SourceScholar
2025

Adversarial Training for Probabilistic Robustness

ICCV 2025poster

Deep learning (DL) has shown transformative potential across industries, yet its sensitivity to adversarial examples (AEs) limits its reliability and broader deployment. Research on DL robustness has developed various techniques, with adversarial training (AT) established as a leading approach to co…

2025

Attributive Reasoning for Hallucination Diagnosis of Large Language Models

AAAI 2025technical

In recent years, large language models (LLMs) have demonstrated outstanding capabilities in various tasks. However, LLMs also have various drawbacks, especially hallucination. Hallucination refers to the generation of content that does not align with the user input, contradicts previously generated…

2025

Automated Composition of Agents: A Knapsack Approach for Agentic Component Selection

NeurIPS 2025poster

Designing effective agentic systems requires the seamless composition and integration of agents, tools, and models within dynamic and uncertain environments. Most existing methods rely on static, semantic retrieval approaches for tool or agent discovery. However, effective reuse and composition of e…

Cited by 0SourceScholar
2025

CAW-CL: Cascaded Adaptive Weighted Contrastive Learning for Unsupervised Ultrasound Plane-Wave Image Reconstruction

ICASSP 2025accepted

Ultrasound plane-wave (PW) imaging achieves an ultra-high frame rate at the expense of degraded image quality. Deep learning-based methods have emerged as a promising avenue for enhancing PW imaging quality. Existing research predominantly focuses on supervised learning that relies on high-quality p…

Cited by 0SourceScholar
2025

CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions

ACL 2025long

We introduce Conversational Function-Calling Evaluation Through Turn-Level Interactions (CONFETTI), a conversational benchmark designed to evaluate the function-calling capabilities and response quality of large language models (LLMs). Current benchmarks lack comprehensive assessment of LLMs in comp…

2025

Can we Retrieve Everything All at Once? ARM: An Alignment-Oriented LLM-based Retrieval Method

ACL 2025long

Real-world open-domain questions can be complex, especially when answering them requires integrating information from multiple sources. Effectively identifying the necessary information involves *aligning* it with the available data and its organization. However, existing RAG solutions address the a…

Cited by 0SourcePDFScholar
2025

Correcting Large Language Model Behavior via Influence Function

AAAI 2025technical

Recent advancements in AI alignment techniques have significantly improved the alignment of large language models (LLMs) with static human preferences. However, the dynamic nature of human preferences can render some prior training data outdated or even erroneous, ultimately causing LLMs to deviate…

Cited by 0SourcePDFScholar
2025

CriSPO: Multi-Aspect Critique-Suggestion-guided Automatic Prompt Optimization for Text Generation

AAAI 2025technical

Existing automatic prompt engineering methods are typically designed for discriminative tasks, where new task prompts are iteratively refined with limited feedback from a single metric reflecting a single aspect. However, these approaches are suboptimal for generative tasks, which require more nuanc…

2025

Cross-Modal Few-Shot Learning with Second-Order Neural Ordinary Differential Equations

AAAI 2025technical

We introduce SONO, a novel method leveraging Second-Order Neural Ordinary Differential Equations (Second-Order NODEs) to enhance cross-modal few-shot learning. By employing a simple yet effective architecture consisting of a Second-Order NODEs model paired with a cross-modal classifier, SONO address…

Cited by 1SourcePDFScholar
2025

DLEFT-MKC: Dynamic Late Fusion Multiple Kernel Clustering with Robust Tensor Learning via Min-Max Optimization

ICLR 2025spotlight

Recent advancements in multiple kernel clustering (MKC) have highlighted the effectiveness of late fusion strategies, particularly in enhancing computational efficiency to near-linear complexity while achieving promising clustering performance. However, existing methods encounter three significant l…

Cited by 0SourcePDFScholar
2025

DepMGNN: Matrixial Graph Neural Network for Video-based Automatic Depression Assessment

AAAI 2025technical

Depression can be reflected by long-term human spatio-temporal facial behaviours. While human face videos recorded in real-world usually have long and variable lengths, existing video-based depression assessment approaches frequently re-sample/down-sample such videos to short and equal-length videos…

2025

Enhanced then Progressive Fusion with View Graph for Multi-View Clustering

CVPR 2025poster

Multi-view clustering aims to improve clustering accuracy by effectively integrating complementary information from multiple perspectives. However, existing methods often encounter challenges such as feature conflicts between views and insufficient enhancement of individual view features, which hind…

Cited by 0SourcePDFScholar
2025

Equal Truth: Rumor Detection with Invariant Group Fairness

EMNLP 2025

Due to the widespread dissemination of rumors on social media platforms, detecting rumors has been a long-standing concern for various communities. However, existing rumor detection methods rarely consider the fairness issues inherent in the model, which can lead to biased predictions across differe

Cited by 0SourcePDFScholar
2025

Faster and Stronger: When ANN-SNN Conversion Meets Parallel Spiking Calculation

ICML 2025poster

Spiking Neural Network (SNN), as a brain-inspired and energy-efficient network, is currently facing the pivotal challenge of exploring a suitable and efficient learning framework. The predominant training methodologies, namely Spatial-Temporal Back-propagation (STBP) and ANN-SNN Conversion, are encu…

2025

From Spectrum-free towards Baseline-view-free: Double-track Proximity Driven Multi-view Clustering

ICML 2025poster

Current multi-view clustering (MVC) techniques generally focus only on the relationship between anchors and samples, while overlooking that between anchors. Moreover, due to the lack of data labels, the cluster order is inconsistent across views and accordingly anchors encounter misalignment, whi…

Cited by 0SourcePDFScholar
2025

GenIR: Generative Visual Feedback for Mental Image Retrieval

NeurIPS 2025poster

Vision-language models (VLMs) have shown strong performance on text-to-image retrieval benchmarks. However, bridging this success to real-world applications remains a challenge. In practice, human search behavior is rarely a one-shot action. Instead, it is often a multi-round process guided by clues…

Cited by 0SourceScholar
2025

HetGCoT: Heterogeneous Graph-Enhanced Chain-of-Thought LLM Reasoning for Academic Question Answering

EMNLP 2025

Academic question answering (QA) in heterogeneous scholarly networks presents unique challenges requiring both structural understanding and interpretable reasoning. While graph neural networks (GNNs) capture structured graph information and large language models (LLMs) demonstrate strong capabilitie

Cited by 0SourcePDFScholar
2025

Hierarchical Optimization via LLM-Guided Objective Evolution for Mobility-on-Demand Systems

NeurIPS 2025poster

Online ride-hailing platforms aim to deliver efficient mobility-on-demand services, often facing challenges in balancing dynamic and spatially heterogeneous supply and demand. Existing methods typically fall into two categories: reinforcement learning (RL) approaches, which suffer from data ineffici…

Cited by 0SourceScholar
2025

Hire Me or Not? Examining Language Model’s Behavior with Occupation Attributes

COLING 2025main

With the impressive performance in various downstream tasks, large language models (LLMs) have been widely integrated into production pipelines, such as recruitment and recommendation systems. A known issue of models trained on natural language data is the presence of human biases, which can impact…

2025

Improving Retrieval-Augmented Generation through Multi-Agent Reinforcement Learning

NeurIPS 2025poster

Retrieval-augmented generation (RAG) is widely utilized to incorporate external knowledge into large language models, thereby enhancing factuality and reducing hallucinations in question-answering (QA) tasks. A standard RAG pipeline consists of several components, such as query rewriting, document r…

Cited by 0SourcecodeScholar
2025

Incorporating Improved Sinusoidal Threshold-based Semi-supervised Method and Diffusion Models for Osteoporosis Diagnosis

ICASSP 2025accepted

Osteoporosis is a common skeletal disease that seriously affects patients’ quality of life. Traditional osteoporosis diagnosis methods are expensive and complex. The semi-supervised model based on diffusion model and class threshold sinusoidal decay proposed in this paper can automatically diagnose…

Cited by 0SourceScholar
2025

InstantSticker: Realistic Decal Blending via Disentangled Object Reconstruction

AAAI 2025technical

We present InstantSticker, a disentangled reconstruction pipeline based on Image-Based Lighting (IBL), which focuses on highly realistic decal blending, simulates stickers attached to the reconstructed surface, and allows for instant editing and real-time rendering. To achieve stereoscopic impressio…

2025

Kernel Regression in Structured Non-IID Settings: Theory and Implications for Denoising Score Learning

NeurIPS 2025poster

Kernel ridge regression (KRR) is a foundational tool in machine learning, with recent work emphasizing its connections to neural networks. However, existing theory primarily addresses the i.i.d. setting, while real-world data often exhibits structured dependencies - particularly in applications like…

Cited by 0SourceScholar
2025

Language Representations Can be What Recommenders Need: Findings and Potentials

ICLR 2025oral

Recent studies empirically indicate that language models (LMs) encode rich world knowledge beyond mere semantics, attracting significant attention across various fields. However, in the recommendation domain, it remains uncertain whether LMs implicitly encode user preference information. Contrary to…

2025

Large-scale Multi-view Tensor Clustering with Implicit Linear Kernels

CVPR 2025poster

Multi-view clustering is a long-standing hot topic in machine learning communities, due to its capability of integrating data information from multiple sources and modalities. By utilizing tensor Singular Value Decomposition (t-SVD) technique with the tensor rotation trick, recent advances have achi…

2025

M-ABSA: A Multilingual Dataset for Aspect-Based Sentiment Analysis

EMNLP 2025

Aspect-based sentiment analysis (ABSA) is a crucial task in information extraction and sentiment analysis, aiming to identify aspects with associated sentiment elements in text. However, existing ABSA datasets are predominantly English-centric, limiting the scope for multilingual evaluation and rese

2025

MemInsight: Autonomous Memory Augmentation for LLM Agents

EMNLP 2025

Large language model (LLM) agents have evolved to intelligently process information, make decisions, and interact with users or tools. A key capability is the integration of long-term memory capabilities, enabling these agents to draw upon historical interactions and knowledge. However, the growing

Cited by 0SourcePDFScholar
2025

Modality Modulation and Dual Consistency for Multi-Modality Semi-Supervised Medical Image Segmentation

ICASSP 2025accepted

Multi-modality (MM) semi-supervised learning (SSL) based medical image segmentation has recently gained increasing attention due to its ability to utilize MM data and low dependency on labeled images. However, current MM-SSL methods face two major challenges: (1) Complex network designs make it diff…

Cited by 0SourceScholar
2025

MultiPDENet: PDE-embedded Learning with Multi-time-stepping for Accelerated Flow Simulation

ICML 2025poster

Solving partial differential equations (PDEs) by numerical methods meet computational cost challenge for getting the accurate solution since fine grids and small time steps are required. Machine learning can accelerate this process, but struggle with weak generalizability, interpretability, and data…

Cited by 0SourcePDFScholar
2025

On Reasoning Strength Planning in Large Reasoning Models

NeurIPS 2025poster

Recent studies empirically reveal that large reasoning models (LRMs) can automatically allocate more reasoning strengths (\ie the number of reasoning tokens) for harder problems, exhibiting difficulty-awareness for better task performance. While this automatic reasoning strength allocation phenomeno…

Cited by 0SourcecodeScholar
2025

On Synthetic Data Strategies for Domain-Specific Generative Retrieval

ACL 2025long

This paper investigates synthetic data generation strategies in developing generative retrieval models for domain-specific corpora, thereby addressing the scalability challenges inherent in manually annotating in-domain queries. We study the data strategies for a two-stage training framework: in the…

2025

Open Domain Question Answering with Conflicting Contexts

NAACL 2025findings

Open domain question answering systems frequently rely on information retrieved from large collections of text (such as the Web) to answer questions. However, such collections of text often contain conflicting information, and indiscriminately depending on this information may result in untruthful a…

Cited by 3SourcePDFScholar
2025

Patient-Level Anatomy Meets Scanning-Level Physics: Personalized Federated Low-Dose CT Denoising Empowered by Large Language Model

CVPR 2025poster

Reducing radiation doses benefits patients, but the resultant low-dose computed tomography (LDCT) images often suffer from clinically unacceptable noise and artifacts. While deep learning (DL) has shown promise in LDCT reconstruction, it requires large-scale data collection from multiple clients, ra…

2025

PhyMPGN: Physics-encoded Message Passing Graph Network for spatiotemporal PDE systems

ICLR 2025spotlight

Solving partial differential equations (PDEs) serves as a cornerstone for modeling complex dynamical systems. Recent progresses have demonstrated grand benefits of data-driven neural-based models for predicting spatiotemporal dynamics (e.g., tremendous speedup gain compared with classical numerical…

Cited by 4SourcePDFScholar
2025

Plaintext-Free Deep Learning for Privacy-Preserving Medical Image Analysis through Frequency Information Embedding

ICASSP 2025accepted

In the fast-evolving field of medical image analysis, deep Learning (DL)-based methods have achieved tremendous success. However, these methods require plaintext data for training and inference stages, raising privacy concerns, especially in the sensitive area of medical data. To tackle these concer…

Cited by 0SourceScholar
2025

RAGDiffusion: Faithful Cloth Generation via External Knowledge Assimilation

ICCV 2025poster

Standard clothing asset generation involves restoring forward-facing flat-lay garment images displayed on a clear background by extracting clothing information from diverse real-world contexts, which presents significant challenges due to highly standardized structure sampling distributions and clot…

Cited by 0SourcePDFScholar
2025

RBench-V: A Primary Assessment for Visual Reasoning Models with Multimodal Outputs

NeurIPS 2025poster

The rapid advancement of native multi-modal models and omni-models, exemplified by GPT-4o, Gemini and o3 with their capability to process and generate content across modalities such as text and images, marks a significant milestone in the evolution of intelligence. Systematic evaluation of their mul…

Cited by 0SourcecodeScholar
2025

RBench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

ICML 2025poster

Reasoning stands as a cornerstone of intelligence, enabling the synthesis of existing knowledge to solve complex problems. Despite remarkable progress, existing reasoning benchmarks often fail to rigorously evaluate the nuanced reasoning capabilities required for complex, real-world problemsolving,…

2025

Robotic Hand Tool Use with Contact-Based Demonstration: The Case of Cucumber Peeling

IROS 2025

Robotic hand tool use has garnered significant attention from robotics researchers, because it enhances dexterity beyond the limitations imposed by manipulators with fixed tool configurations and human-involved manual tool changes. Despite extensive research, current methodologies predominantly focu

Cited by 0SourceScholar
2025

SAMULE: Self-Learning Agents Enhanced by Multi-level Reflection

EMNLP 2025

Despite the rapid advancements in LLM agents, they still face the challenge of generating meaningful reflections due to inadequate error analysis and a reliance on rare successful trajectories, especially in complex tasks. In this work, we propose SAMULE, a new framework for self-learning agents pow

Cited by 0SourcePDFScholar
2025

See Further When Clear: Curriculum Consistency Model

CVPR 2025poster

Significant advances have been made in the sampling efficiency of diffusion and flow matching models, driven by Consistency Distillation (CD), which trains a student model to mimic the output of a teacher model at a later timestep. However, we found that the knowledge discrepancy between student and…

Cited by 0SourcePDFScholar
2025

Simple yet Effective Incomplete Multi-view Clustering: Similarity-level Imputation and Intra-view Hybrid-group Prototype Construction

ICLR 2025spotlight

Most of incomplete multi-view clustering (IMVC) methods typically choose to ignore the missing samples and only utilize observed unpaired samples to construct bipartite similarity. Moreover, they employ a single quantity of prototypes to extract the information of $\textbf{all}$ views. To elimina…

Cited by 0SourcePDFScholar
2025

Structured List-Grounded Question Answering

COLING 2025main

Document-grounded dialogue systems aim to answer user queries by leveraging external information. Previous studies have mainly focused on handling free-form documents, often overlooking structured data such as lists, which can represent a range of nuanced semantic relations. Motivated by the observa…

Cited by 0SourcePDFScholar
2025

TReMu: Towards Neuro-Symbolic Temporal Reasoning for LLM-Agents with Memory in Multi-Session Dialogues

ACL 2025finding

Temporal reasoning in multi-session dialogues presents a significant challenge which has been under-studied in previous temporal reasoning benchmarks. To bridge this gap, we propose a new evaluation task for temporal reasoning in multi-session dialogues and introduce an approach to construct a new b…

Cited by 0SourcePDFScholar
2025

Testing Conditional Mean Independence Using Generative Neural Networks

ICML 2025poster

Conditional mean independence (CMI) testing is crucial for statistical tasks including model determination and variable importance evaluation. In this work, we introduce a novel population CMI measure and a bootstrap-based testing procedure that utilizes deep generative neural networks to estimate t…

Cited by 0SourcePDFScholar
2025

Transfer Faster, Price Smarter: Minimax Dynamic Pricing under Cross-Market Preference Shift

NeurIPS 2025spotlight

We study contextual dynamic pricing when a target market can leverage $K$ auxiliary markets—offline logs or concurrent streams—whose *mean utilities differ by a structured preference shift*. We propose *Cross-Market Transfer Dynamic Pricing (CM-TDP)*, the first algorithm that *provably* handles such…

Cited by 0SourceScholar
2025

Two-Stream Spiking Neural Network for Event-based Action Recognition

ICASSP 2025accepted

Spiking neural networks (SNNs) are increasingly applied to event-based data generated by event cameras due to their asynchronous and sparse properties. Event cameras can inherently respond to the changes in the scene, which is a quite desirable property for action recognition tasks. However, existin…

Cited by 0SourceScholar
2025

Wavelet Movement Primitives: A Unified Framework for Learning Discrete and Rhythmic Movements

RA-L 2025

Real-world tasks often require combinations of both discrete and rhythmic movements. However, most of current methods can only address one of them. This letter proposes a unified framework, Wavelet Movement Primitives (WMPs), which are built on Probabilistic Movement Primitives (ProMPs) integrated w

Cited by 1SourceScholar
2025

Worse than Zero-shot? A Fact-Checking Dataset for Evaluating the Robustness of RAG Against Misleading Retrievals

NeurIPS 2025poster

Retrieval-augmented generation (RAG) has shown impressive capabilities in mitigating hallucinations in large language models (LLMs). However, LLMs struggle to maintain consistent reasoning when exposed to misleading or conflicting evidence, especially in real-world domains such as politics, where in…

Cited by 0SourceScholar
2024

Bootstrapping LLM-based Task-Oriented Dialogue Agents via Self-Talk

ACL 2024findings

Large language models (LLMs) are powerful dialogue agents, but specializing them towards fulfilling a specific function can be challenging. Instructing tuning, i.e. tuning models on instruction and sample responses generated by humans (Ouyang et al., 2022), has proven as an effective method to do so…

2024

Concept-Guided Prompt Learning for Generalization in Vision-Language Models

AAAI 2024technical

Contrastive Language-Image Pretraining (CLIP) model has exhibited remarkable efficacy in establishing cross-modal connections between texts and images, yielding impressive performance across a broad spectrum of downstream applications through fine-tuning. However, for generalization tasks, the curre…

2024

DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D Data

CVPR 2024poster

We present DIRECT-3D a diffusion-based 3D generative model for creating high-quality 3D assets (represented by Neural Radiance Fields) from text prompts. Unlike recent 3D generative models that rely on clean and well-aligned 3D data limiting them to single or few-class generation our model is direct…

2024

Domain Separation Graph Neural Networks for Saliency Object Ranking

CVPR 2024poster

Saliency object ranking (SOR) has attracted significant attention recently. Previous methods usually failed to explicitly explore the saliency degree-related relationships between objects. In this paper we propose a novel Domain Separation Graph Neural Network (DSGNN) which starts with separately ex…

2024

EasyDrag: Efficient Point-based Manipulation on Diffusion Models

CVPR 2024poster

Generative models are gaining increasing popularity and the demand for precisely generating images is on the rise. However generating an image that perfectly aligns with users' expectations is extremely challenging. The shapes of objects the poses of animals the structures of landscapes and more may…

2024

Exploring Regional Clues in CLIP for Zero-Shot Semantic Segmentation

CVPR 2024poster

CLIP has demonstrated marked progress in visual recognition due to its powerful pre-training on large-scale image-text pairs. However it still remains a critical challenge: how to transfer image-level knowledge into pixel-level understanding tasks such as semantic segmentation. In this paper to solv…

2024

Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational Pathology

CVPR 2024poster

Multiple instance learning (MIL) is the most widely used framework in computational pathology encompassing sub-typing diagnosis prognosis and more. However the existing MIL paradigm typically requires an offline instance feature extractor such as a pre-trained ResNet or a foundation model. This appr…

2024

FlattenQuant: Breaking through the Inference Compute-bound for Large Language Models with Per-tensor Quantization

COLING 2024main

Large language models (LLMs) have demonstrated state-of-the-art accuracies across various tasks. However, the latency of inference and the large GPU memory consumption of LLMs restrict their deployment performance. Recently, there have been some efficient attempts to quantize LLMs, yet inference wit…

Cited by 3SourcePDFScholar
2024

FocalDreamer: Text-Driven 3D Editing via Focal-Fusion Assembly

AAAI 2024technical

While text-3D editing has made significant strides in leveraging score distillation sampling, emerging approaches still fall short in delivering separable, precise and consistent outcomes that are vital to content creation. In response, we introduce FocalDreamer, a framework that merges base shape w…

Cited by 56SourcePDFScholar
2024

Generating Images with 3D Annotations Using Diffusion Models

ICLR 2024spotlight

Diffusion models have emerged as a powerful generative method, capable of producing stunning photo-realistic images from natural language descriptions. However, these models lack explicit control over the 3D structure in the generated images. Consequently, this hinders our ability to obtain detailed…

Cited by 6SourcePDFScholar
2024

Hawkes-Enhanced Spatial-Temporal Hypergraph Contrastive Learning Based on Criminal Correlations

AAAI 2024technical

Crime prediction is a crucial yet challenging task within urban computing, which benefits public safety and resource optimization. Over the years, various models have been proposed, and spatial-temporal hypergraph learning models have recently shown outstanding performances. However, three correlati…

Cited by 7SourcePDFScholar
2024

Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval

ACL 2024long

Retrieving relevant tables containing the necessary information to accurately answer a given question over tables is critical to open-domain question-answering (QA) systems. Previous methods assume the answer to such a question can be found either in a single table or multiple tables identified thro…

Cited by 15SourcePDFScholar
2024

Leveraging Opposite Gender Interaction Ratio as a Path towards Fairness in Online Dating Recommendations Based on User Sexual Orientation

AAAI 2024technical

Online dating platforms have gained widespread popularity as a means for individuals to seek potential romantic relationships. While recommender systems have been designed to improve the user experience in dating platforms by providing personalized recommendations, increasing concerns about fairness…

Cited by 4SourcePDFScholar
2024

LiteSAM is Actually what you Need for segment Everything

ECCV 2024poster

"The Segment Anything model (SAM) has brought significant changes to the segmentation field with its superior performance, but its extensive computational resource requirements remain a limiting factor. Many works, such as MobileSAM, Edge-SAM, and MobileSAM-v2, have explored lightweight solutions. H…

Cited by 7SourcePDFScholar
2024

MACM: Utilizing a Multi-Agent System for Condition Mining in Solving Complex Mathematical Problems

NeurIPS 2024poster

Recent advancements in large language models, such as GPT-4, have demonstrated remarkable capabilities in processing standard queries. Despite these advancements, their performance substantially declines in advanced mathematical problems requiring complex, multi-step logical reasoning. To enhance th…

2024

MARCO: Multi-Agent Real-time Chat Orchestration

EMNLP 2024industry

Large language model advancements have enabled the development of multi-agent frameworks to tackle complex, real-world problems such as to automate workflows that require interactions with diverse tools, reasoning, and human collaboration. We present MARCO, a Multi-Agent Real-time Chat Orchestration…

Cited by 1SourcePDFScholar
2024

MDCR: A Dataset for Multi-Document Conditional Reasoning

EMNLP 2024finding

The same real-life questions posed to different individuals may lead to different answers based on their unique situations. For instance, whether a student is eligible for a scholarship depends on eligibility conditions, such as major or degree required. ConditionalQA was proposed to evaluate models…

2024

P$^2$C$^2$Net: PDE-Preserved Coarse Correction Network for efficient prediction of spatiotemporal dynamics

NeurIPS 2024poster

When solving partial differential equations (PDEs), classical numerical methods often require fine mesh grids and small time stepping to meet stability, consistency, and convergence conditions, leading to high computational cost. Recently, machine learning has been increasingly utilized to solve PDE…

Cited by 5SourcePDFScholar
2024

PairingNet: A Learning-based Pair-searching and -matching Network for Image Fragments

ECCV 2024poster

"In this paper, we propose a learning-based image fragment pair-searching and -matching approach to solve the challenging restoration problem. Existing works use rule-based methods to match similar contour shapes or textures, which are always difficult to tune hyperparameters for extensive data and…

2024

Point Segment and Count: A Generalized Framework for Object Counting

CVPR 2024poster

Class-agnostic object counting aims to count all objects in an image with respect to example boxes or class names a.k.a few-shot and zero-shot counting. In this paper we propose a generalized framework for both few-shot and zero-shot object counting based on detection. Our framework combines the sup…

2024

ProTIP: Probabilistic Robustness Verification on Text-to-Image Diffusion Models against Stochastic Perturbation

ECCV 2024poster

"Text-to-Image (T2I) Diffusion Models (DMs) excel at creating high-quality images from text descriptions but, like many deep learning models, suffer from robustness issues. While there are attempts to evaluate the robustness of T2I DMs as a binary or worst-case problem, they cannot answer how robust…

2024

Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding

EMNLP 2024main

Graphical User Interfaces (GUIs) are central to our interaction with digital devices and growing efforts have been made to build models for various GUI understanding tasks. However, these efforts largely overlook an important GUI-referring task: screen reading based on user-indicated points, which w…

2024

Reverse Transition Kernel: A Flexible Framework to Accelerate Diffusion Inference

NeurIPS 2024spotlight

To generate data from trained diffusion models, most inference algorithms, such as DDPM, DDIM, and other variants, rely on discretizing the reverse SDEs or their equivalent ODEs. In this paper, we view such approaches as decomposing the entire denoising diffusion process into several segments, each…

Cited by 8SourcePDFScholar
2024

Right this way: Can VLMs Guide Us to See More to Answer Questions?

NeurIPS 2024poster

In question-answering scenarios, humans can assess whether the available information is sufficient and seek additional information if necessary, rather than providing a forced answer. In contrast, Vision Language Models (VLMs) typically generate direct, one-shot responses without evaluating the suff…

2024

SM3-Text-to-Query: Synthetic Multi-Model Medical Text-to-Query Benchmark

NeurIPS 2024poster

Electronic health records (EHRs) are stored in various database systems with different database models on heterogeneous storage architectures, such as relational databases, document stores, or graph databases. These different database models have a big impact on query complexity and performance. Whi…

2024

Sample-Level Cross-View Similarity Learning for Incomplete Multi-View Clustering

AAAI 2024technical

Incomplete multi-view clustering has attracted much attention due to its ability to handle partial multi-view data. Recently, similarity-based methods have been developed to explore the complete relationship among incomplete multi-view data. Although widely applied to partial scenarios, most of the…

2024

StatBot.Swiss: Bilingual Open Data Exploration in Natural Language

ACL 2024findings

The potential for improvements brought by Large Language Models (LLMs) in Text-to-SQL systems is mostly assessed on monolingual English datasets. However, LLMs’ performance for other languages remains vastly unexplored. In this work, we release the StatBot.Swiss dataset, the first bilingual benchmar…

2024

TARP-VP: Towards Evaluation of Transferred Adversarial Robustness and Privacy on Label Mapping Visual Prompting Models

NeurIPS 2024poster

Adversarial robustness and privacy of deep learning (DL) models are two widely studied topics in AI security. Adversarial training (AT) is an effective approach to improve the robustness of DL models against adversarial attacks. However, while models with AT demonstrate enhanced robustness, they be…

Cited by 0SourcePDFScholar
2024

Tackling Vision Language Tasks through Learning Inner Monologues

AAAI 2024technical

Visual language tasks such as Visual Question Answering (VQA) or Visual Entailment (VE) require AI models to comprehend and reason with both visual and textual content. Driven by the power of Large Language Models (LLMs), two prominent methods have emerged: (1) the hybrid integration between LLMs an…

2024

TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization

NAACL 2024long

Single document news summarization has seen substantial progress on faithfulness in recent years, driven by research on the evaluation of factual consistency, or hallucinations. We ask whether these advances carry over to other text summarization domains. We propose a new evaluation benchmark on top…

2024

When Pedestrian Detection Meets Multi-Modal Learning: Generalist Model and Benchmark Dataset

ECCV 2024poster

"Recent years have witnessed increasing research attention towards pedestrian detection by taking the advantages of different sensor modalities (RGB, IR, Depth, LiDAR and Event). However, designing a unified generalist model that can effectively process diverse sensor modalities remains a challenge.…

2024

mABC: Multi-Agent Blockchain-inspired Collaboration for Root Cause Analysis in Micro-Services Architecture

EMNLP 2024finding

Root cause analysis (RCA) in Micro-services architecture (MSA) with escalating complexity encounters complex challenges in maintaining system stability and efficiency due to fault propagation and circular dependencies among nodes. Diverse root cause analysis faults require multi-agents with diverse…

2023

3D-Aware Neural Body Fitting for Occlusion Robust 3D Human Pose Estimation

ICCV 2023poster

Regression-based methods for 3D human pose estimation directly predict the 3D pose parameters from a 2D image using deep networks. While achieving state-of-the-art performance on standard benchmarks, their performance degrades under occlusion. In contrast, optimization-based methods fit a parametric…

Cited by 42PDFcodeScholar
2023

A Simple Baseline for Video Restoration With Grouped Spatial-Temporal Shift

CVPR 2023poster

Video restoration, which aims to restore clear frames from degraded videos, has numerous important applications. The key to video restoration depends on utilizing inter-frame information. However, existing deep learning methods often rely on complicated network architectures, such as optical flow es…

2023

A Unified Conditional Framework for Diffusion-based Image Restoration

NeurIPS 2023poster

Diffusion Probabilistic Models (DPMs) have recently shown remarkable performance in image generation tasks, which are capable of generating highly realistic images. When adopting DPMs for image restoration tasks, the crucial aspect lies in how to integrate the conditional information to guide the DP…

2023

Abstract then Play: A Skill-centric Reinforcement Learning Framework for Text-based Games

ACL 2023findings

Text-based games present an exciting test-bed for reinforcement learning algorithms in the natural language environment. In these adventure games, an agent must learn to interact with the environment through text in order to accomplish tasks, facing large and combinational action space as well as pa…

Cited by 1SourcePDFScholar
2023

Aerial Vision-and-Dialog Navigation

ACL 2023findings

The ability to converse with humans and follow natural language commands is crucial for intelligent unmanned aerial vehicles (a.k.a. drones). It can relieve people’s burden of holding a controller all the time, allow multitasking, and make drone control more accessible for people with disabilities o…

2023

Animal3D: A Comprehensive Dataset of 3D Animal Pose and Shape

ICCV 2023poster

Accurately estimating the 3D pose and shape is an essential step towards understanding animal behavior, and can potentially benefit many downstream applications, such as wildlife conservation. However, research in this area is held back by the lack of a comprehensive and diverse dataset with high-qu…

Cited by 24PDFScholar
2023

Boosting Point Clouds Rendering via Radiance Mapping

AAAI 2023technical

Recent years we have witnessed rapid development in NeRF-based image rendering due to its high quality. However, point clouds rendering is somehow less explored. Compared to NeRF-based rendering which suffers from dense spatial sampling, point clouds rendering is naturally less computation intensive…

2023

DFEE: Interactive DataFlow Execution and Evaluation Kit

AAAI 2023technical

DataFlow has been emerging as a new paradigm for building task-oriented chatbots due to its expressive semantic representations of the dialogue tasks. Despite the availability of a large dataset SMCalFlow and a simplified syntax, the development and evaluation of DataFlow-based chatbots remain chall…

2023

Frequency-Modulated Point Cloud Rendering With Easy Editing

CVPR 2023highlight

We develop an effective point cloud rendering pipeline for novel view synthesis, which enables high fidelity local detail reconstruction, real-time rendering and user-friendly editing. In the heart of our pipeline is an adaptive frequency modulation module called Adaptive Frequency Net (AFNet), whic…

2023

From Association to Generation: Text-only Captioning by Unsupervised Cross-modal Mapping

IJCAI 2023poster

With the development of Vision-Language Pre-training Models (VLPMs) represented by CLIP and ALIGN, significant breakthroughs have been achieved for association-based visual tasks such as image classification and image-text retrieval by the zero-shot capability of CLIP without fine-tuning. However, C…

2023

Improved Event-Based Dense Depth Estimation via Optical Flow Compensation

ICRA 2023poster

Event cameras have the potential to overcome the limitations of classical computer vision in real-world applications. Depth estimation is a crucial step for high-level robotics tasks and has attracted much attention from the community. In this paper, we propose an event-based dense depth estimation…

Cited by 7SourceScholar
2023

Learning Query Adaptive Anchor Representation for Inductive Relation Prediction

ACL 2023findings

Relation prediction on knowledge graphs (KGs) attempts to infer the missing links between entities. Most previous studies are limited to the transductive setting where all entities must be seen during the training, making them unable to perform reasoning on emerging entities. Recently, the inductive…

Cited by 2SourcePDFScholar
2023

Learning Shape Primitives via Implicit Convexity Regularization

ICCV 2023poster

Shape primitives decomposition has been an important and long-standing task in 3D shape analysis. Prior arts heavily rely on 3D point clouds or voxel data for shape primitives extraction, which are less practical in real-world scenarios. This paper proposes to learn shape primitives from multi-view…

Cited by 3PDFcodeScholar
2023

Learning threshold neurons via edge of stability

NeurIPS 2023poster

Existing analyses of neural network training often operate under the unrealistic assumption of an extremely small learning rate. This lies in stark contrast to practical wisdom and empirical studies, such as the work of J. Cohen et al. (ICLR 2021), which exhibit startling new phenomena (the "edge of…

Cited by 47SourcePDFScholar
2023

Measuring and Mitigating Constraint Violations of In-Context Learning for Utterance-to-API Semantic Parsing

EMNLP 2023long findings

In executable task-oriented semantic parsing, the system aims to translate users' utterances in natural language to machine-interpretable programs (API calls) that can be executed according to pre-defined API specifications. With the popularity of Large Language Models (LLMs), in-context learning of…

Cited by 0SourceScholar
2023

Multiple Instance Learning Framework with Masked Hard Instance Mining for Whole Slide Image Classification

ICCV 2023oral

The whole slide image (WSI) classification is often formulated as a multiple instance learning (MIL) problem. Since the positive tissue is only a small fraction of the gigapixel WSI, existing MIL methods intuitively focus on identifying salient instances via attention mechanisms. However, this leads…

Cited by 71PDFcodeScholar
2023

NatCS: Eliciting Natural Customer Support Dialogues

ACL 2023findings

Despite growing interest in applications based on natural customer support conversations,there exist remarkably few publicly available datasets that reflect the expected characteristics of conversations in these settings. Existing task-oriented dialogue datasets, which were collected to benchmark di…

2023

Optimized Design and Analysis of Active Propeller-driven Capsule Endoscopic Robot for Gastric Examination

ICRA 2023poster

Capsule endoscopic robot holds great promise for the early diagnosis of gastrointestinal diseases without causing discomfort to patients. However, currently available active capsule endoscopic robots suffer from issues such as complex structure, poor mobility, large size, and high cost, which have h…

Cited by 6SourceScholar
2023

Pre-training Intent-Aware Encoders for Zero- and Few-Shot Intent Classification

EMNLP 2023long main

Intent classification (IC) plays an important role in task-oriented dialogue systems. However, IC models often generalize poorly when training without sufficient annotated examples for each user intent. We propose a novel pre-training method for text encoders that uses contrastive learning with inte…

Cited by 0SourcecodeScholar
2023

Stay In The Middle: A Semi-Supervised Model for CT Metal Artifact Reduction

ICASSP 2023accepted

Metal artifacts degrade CT image’s quality. Recently, some deep learning-based metal artifact reduction (MAR) methods have been developed. Supervised MAR methods don’t perform well in clinical due to the domain gap between simulated and clinical data. Although this problem can be avoided in an unsup…

Cited by 0SourceScholar
2023

The Euclidean Space is Evil: Hyperbolic Attribute Editing for Few-shot Image Generation

ICCV 2023poster

Few-shot image generation is a challenging task since it aims to generate diverse new images for an unseen category with only a few images. Existing methods suffer from the trade-off between the quality and diversity of generated images. To tackle this problem, we propose Hyperbolic Attribute Editin…

Cited by 19PDFcodeScholar
2023

What Makes Convolutional Models Great on Long Sequence Modeling?

ICLR 2023poster

Convolutional models have been widely used in multiple domains. However, most existing models only use local convolution, making the model unable to handle long-range dependencies efficiently. Attention overcomes this problem by aggregating global information based on the pair-wise attention score b…

2022

An Efficient and Accurate Solution to Camera Pose Estimation Problem from Point and Line Correspondences Based on Null Space Analysis

IROS 2022poster

In this paper, we propose an accurate and simultaneously efficient solution to perspective-n-point-and-line (PnPL) problem by null space analysis. Although many PnPL-like methods have been proposed, it is hard to obtain the optimal solution considering both calculation efficiency and accuracy at the…

Cited by 4SourceScholar
2022

Design Challenges for a Multi-Perspective Search Engine

NAACL 2022findings

Many users turn to document retrieval systems (e.g. search engines) to seek answers to controversial or open-ended questions. However, classical document retrieval systems fall short at delivering users a set of direct and diverse responses in such cases, which requires identifying responses within…

2022

Dialogue Meaning Representation for Task-Oriented Dialogue Systems

EMNLP 2022finding

Dialogue meaning representation formulates natural language utterance semantics in their conversational context in an explicit and machine-readable form. Previous work typically follows the intent-slot framework, which is easy for annotation yet limited in scalability for complex linguistic expressi…

2022

Explicit Occlusion Reasoning for Multi-Person 3D Human Pose Estimation

ECCV 2022poster

"Occlusion poses a great threat to monocular multi-person 3D human pose estimation due to large variability in terms of the shape, appearance, and position of occluders. While existing methods try to handle occlusion with pose priors/constraints, data augmentation, or implicit reasoning, they still…

2022

Fusion Multiple Kernel K-means

AAAI 2022technical

Multiple kernel clustering aims to seek an appropriate combination of base kernels to mine inherent non-linear information for optimal clustering. Late fusion algorithms generate base partitions independently and integrate them in the following clustering procedure, improving the overall efficiency.…

2022

IDR: Self-Supervised Image Denoising via Iterative Data Refinement

CVPR 2022poster

The lack of large-scale noisy-clean image pairs restricts supervised denoising methods' deployment in actual applications. While existing unsupervised methods are able to learn image denoising without ground-truth clean images, they either show poor performance or work under impractical settings (e.…

Cited by 86PDFcodeScholar
2022

Injecting Domain Knowledge in Language Models for Task-oriented Dialogue Systems

EMNLP 2022main

Pre-trained language models (PLM) have advanced the state-of-the-art across NLP applications, but lack domain-specific knowledge that does not naturally occur in pre-training data. Previous studies augmented PLMs with symbolic knowledge for different downstream NLP tasks. However, knowledge bases (K…

2022

Label Semantic Aware Pre-training for Few-shot Text Classification

ACL 2022long

In text classification tasks, useful information is encoded in the label names. Label semantic aware systems have leveraged this information for improved text classification performance during fine-tuning and prediction. However, use of label-semantics during pre-training has not been extensively ex…

2022

Learning Degradation Representations for Image Deblurring

ECCV 2022poster

"In various learning-based image restoration tasks, such as image denoising and image super-resolution, the degradation representations were widely used to model the degradation process and handle complicated degradation patterns. However, they are less explored in learning-based image deblurring as…

2022

Measuring and Reducing Model Update Regression in Structured Prediction for NLP

NeurIPS 2022accept

Recent advance in deep learning has led to rapid adoption of machine learning based NLP models in a wide range of applications. Despite the continuous gain in accuracy, backward compatibility is also an important aspect for industrial applications, yet it received little research attention. Backward…

Cited by 10SourcePDFScholar
2022

Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System

ACL 2022long

Pre-trained language models have been recently shown to benefit task-oriented dialogue (TOD) systems. Despite their success, existing methods often formulate this task as a cascaded generation problem which can lead to error accumulation across different sub-tasks and greater data annotation overhea…

2022

Optimal Transport for Label-Efficient Visible-Infrared Person Re-identification

ECCV 2022poster

"Visible-infrared person re-identification (VI-ReID) has been a key enabler for night intelligent monitoring system. However, the extensive laboring efforts significantly limit its applications. In this paper, we raise a new label-efficient training pipeline for VI-ReID. Our observation is: RGB ReID…

2022

Rail Vehicle Localization and Mapping With LiDAR-Vision-Inertial-GNSS Fusion

RA-L 2022

In this letter, we present a global navigation satellite system (GNSS) aided LiDAR-visual-inertial scheme, RailLoMer-V, for accurate and robust rail vehicle localization and mapping. RailLoMer-V is formulated atop a factor graph and consists of two subsystems: an odometer assisted LiDAR-inertial sys

Cited by 25SourceScholar
2022

Self-supervised representations for multi-view reinforcement learning

UAI 2022poster

Learning policies from raw, pixel images are quite important for the real-world application of deep reinforcement learning (RL). Standard model-free RL algorithms focus on single-view settings and unify the representation learning and policy learning into an end-to-end training process. However, suc…

2022

Semi-supervised Learning with Multi-Head Co-Training

AAAI 2022technical

Co-training, extended from self-training, is one of the frameworks for semi-supervised learning. Without natural split of features, single-view co-training works at the cost of training extra classifiers, where the algorithm should be delicately designed to prevent individual classifiers from collap…

2022

Stability and Generalization of Kernel Clustering: from Single Kernel to Multiple Kernel

NeurIPS 2022accept

Multiple kernel clustering (MKC) is an important research topic that has been widely studied for decades. However, current methods still face two problems: inefficient when handling out-of-sample data points and lack of theoretical study of the stability and generalization of clustering. In this pap…

Cited by 5SourcePDFScholar
2022

Tailor Versatile Multi-Modal Learning for Multi-Label Emotion Recognition

AAAI 2022technical

Multi-modal Multi-label Emotion Recognition (MMER) aims to identify various human emotions from heterogeneous visual, audio and text modalities. Previous methods mainly focus on projecting multiple modalities into a common latent space and learning an identical representation for all labels, which n…

2022

Vision-based Uneven BEV Representation Learning with Polar Rasterization and Surface Estimation

CoRL 2022poster

In this work, we propose PolarBEV for vision-based uneven BEV representation learning. To adapt to the foreshortening effect of camera imaging, we rasterize the BEV space both angularly and radially, and introduce polar embedding decomposition to model the associations among polar grids. Polar gri…

Cited by 26SourcecodeScholar
2021

A Global Past-Future Early Exit Method for Accelerating Inference of Pre-trained Language Models

NAACL 2021long

Early exit mechanism aims to accelerate the inference speed of large-scale pre-trained language models. The essential idea is to exit early without passing through all the inference layers at the inference stage. To make accurate predictions for downstream tasks, the hierarchical linguistic informat…

2021

Context-aware Cross-level Fusion Network for Camouflaged Object Detection

IJCAI 2021poster

Camouflaged object detection (COD) is a challenging task due to the low boundary contrast between the object and its surroundings. In addition, the appearance of camouflaged objects varies significantly, e.g., object size and shape, aggravating the difficulties of accurate COD. In this paper, we pro…

2021

DASZL: Dynamic Action Signatures for Zero-shot Learning

AAAI 2021technical

There are many realistic applications of activity recognition where the set of potential activity descriptions is combinatorially large. This makes end-to-end supervised training of a recognition system impractical as no training set is practically able to encompass the entire label set. In this pap…

Cited by 32SourcePDFScholar
2021

Learning to Decompose and Organize Complex Tasks

NAACL 2021long

People rely on digital task management tools, such as email or to-do apps, to manage their tasks. Some of these tasks are large and complex, leading to action paralysis and feelings of being overwhelmed on the part of the user. The micro-productivity literature has shown that such tasks could benefi…

2021

Mask4D: 4D Convolution Network for Light Field Occlusion Removal

ICASSP 2021accepted

Current light field (LF) occlusion removal approaches usually select only a part of sub-aperture images (SAIs) or simply stack all SAIs to reconstruct the center view, which destroys the spatial layout of SAIs. In this paper, we present a simple yet effective LF occlusion removal method name Mask4D,…

Cited by 0SourceScholar
2021

Novelty Detection via Contrastive Learning with Negative Data Augmentation

IJCAI 2021poster

Novelty detection is the process of determining whether a query example differs from the learned training distribution. Previous generative adversarial networks based methods and self-supervised approaches suffer from instability training, mode dropping, and low discriminative ability. We overcome s…

Cited by 17SourcePDFScholar
2021

ODIST: Open World Classification via Distributionally Shifted Instances

EMNLP 2021finding

In this work, we address the open-world classification problem with a method called ODIST, open world classification via distributionally shifted instances. This novel and straightforward method can create out-of-domain instances from the in-domain training instances with the help of a pre-trained g…

2021

One Pass Late Fusion Multi-view Clustering

ICML 2021spotlight

Existing late fusion multi-view clustering (LFMVC) optimally integrates a group of pre-specified base partition matrices to learn a consensus one. It is then taken as the input of the widely used k-means to generate the cluster labels. As observed, the learning of the consensus partition matrix and…

Cited by 127SourcePDFScholar
2021

Pointer Networks for Arbitrary-Shaped Text Spotting

ICASSP 2021accepted

Current text spotting methods perform text detection and text recognition separately. However, in complex scenes where bounding boxes of texts with various shapes are often overlapped, text detection becomes error-prone. By contrast, character detection is more non-ambiguous and easier to learn. In…

Cited by 0SourceScholar
2021

RESA: Recurrent Feature-Shift Aggregator for Lane Detection

AAAI 2021technical

Lane detection is one of the most important tasks in self-driving. Due to various complex scenarios (e.g., severe occlusion, ambiguous lanes, etc.) and the sparse supervisory signals inherent in lane annotations, lane detection task is still challenging. Thus, it is difficult for the ordinary convol…

2021

RGB-D Salient Object Detection via 3D Convolutional Neural Networks

AAAI 2021technical

RGB-D salient object detection (SOD) recently has attracted increasing research interest and many deep learning methods based on encoder-decoder architectures have emerged. However, most existing RGB-D SOD models conduct feature fusion either in the single encoder or the decoder stage, which hardly…

2021

Regression Bugs Are In Your Model! Measuring, Reducing and Analyzing Regressions In NLP Model Updates

ACL 2021long

Behavior of deep neural networks can be inconsistent between different versions. Regressions during model update are a common cause of concern that often over-weigh the benefits in accuracy or efficiency gain. This work focuses on quantifying, reducing and analyzing regression errors in the NLP mode…

Cited by 15SourcePDFScholar
2021

RevMan: Revenue-aware Multi-task Online Insurance Recommendation

AAAI 2021technical

Online insurance is a new type of e-commerce with exponential growth. An effective recommendation model that maximizes the total revenue of insurance products listed in multiple customized sales scenarios is crucial for the success of online insurance business. Prior recommendation models are ineffe…

Cited by 17SourcePDFScholar
2021

Salient Object Ranking With Position-Preserved Attention

ICCV 2021poster

Instance segmentation can detect where the objects are in an image, but hard to understand the relationship between them. We pay attention to a typical relationship, relative saliency. A closely related task, salient object detection, predicts a binary map highlighting a visually salient region whil…

Cited by 31PDFcodeScholar
2021

Three-dimensional Positioning of the Micropipette for Intracytoplasmic Sperm Injection

ICRA 2021poster

ICSI (Intracytoplasmic sperm injection) is one of the most effective treatments for severe male infertility. During the implementation of the ICSI, it is necessary to perform the three-dimensional positioning of the tip of the glass injection micropipette. At present, the process is mainly controlle…

Cited by 7SourceScholar
2021

Translation as Cross-Domain Knowledge: Attention Augmentation for Unsupervised Cross-Domain Segmenting and Labeling Tasks

EMNLP 2021finding

The nature of no word delimiter or inflection that can indicate segment boundaries or word semantics increases the difficulty of Chinese text understanding, and also intensifies the demand for word-level semantic knowledge to accomplish the tagging goal in Chinese segmenting and labeling tasks. Howe…

2021

Using Optimal Transport as Alignment Objective for fine-tuning Multilingual Contextualized Embeddings

EMNLP 2021finding

Recent studies have proposed different methods to improve multilingual word representations in contextualized settings including techniques that align between source and target embedding spaces. For contextualized embeddings, alignment becomes more complex as we additionally take context into consid…

Cited by 21SourcePDFScholar
2020

A Sample Complexity Separation between Non-Convex and Convex Meta-Learning

ICML 2020poster

One popular trend in meta-learning is to learn from many training tasks a common initialization that a gradient-based method can use to solve a new task with few samples. The theory of meta-learning is still in its early stages, with several recent learning-theoretic analyses of methods such as Rept…

Cited by 24SourcePDFScholar
2020

Calibration, Entropy Rates, and Memory in Language Models

ICML 2020poster

Building accurate language models that capture meaningful long-term dependencies is a core challenge in natural language processing. Towards this end, we present a calibration-based approach to measure long-term discrepancies between a generative sequence model and the true distribution, and use the…

Cited by 46SourcePDFScholar
2020

DNN-based Mask Estimation Integrating Spectral and Spatial Features for Robust Beamforming

ICASSP 2020accepted

Spectral mask based beamforming has showed competitive performance on multi-channel speech enhancement in recent years. However, such methods apply mask estimation on each channel and ensemble the masks from multiple channels into one for speech and noise covariance estimation. Spectral-spatial mask…

Cited by 0SourceScholar
2020

ICNet: Intra-saliency Correlation Network for Co-Saliency Detection

NeurIPS 2020poster

Intra-saliency and inter-saliency cues have been extensively studied for co-saliency detection (Co-SOD). Model-based methods produce coarse Co-SOD results due to hand-crafted intra- and inter-saliency features. Current data-driven models exploit inter-saliency cues, but undervalue the potential powe…

2020

Over-parameterized Adversarial Training: An Analysis Overcoming the Curse of Dimensionality

NeurIPS 2020poster

Adversarial training is a popular method to give neural nets robustness against adversarial perturbations. In practice adversarial training leads to low robust training loss. However, a rigorous explanation for why this happens under natural conditions is still missing. Recently a convergence theory…

Cited by 59SourcePDFScholar
2020

Synthesize then Compare: Detecting Failures and Anomalies for Semantic Segmentation

ECCV 2020poster

The ability to detect failures and anomalies are fundamental requirements for building reliable systems for computer vision applications, especially safety-critical applications of semantic segmentation, such as autonomous driving and medical image analysis. In this paper, we systematically study fa…

2019

Deep Convolutional Robust PCA with Application to Ultrasound Imaging

ICASSP 2019accepted

Sparse and low-rank decomposition, also known as robust principle component analysis, has been applied successfully in numerous applications. Typically, this approach leads to a minimization problem which is solved using iterative algorithms. Drawing inspiration from recurrent networks, in recent ye…

Cited by 17SourceScholar
2019

Efficient Full-Matrix Adaptive Regularization

ICML 2019oral

Adaptive regularization methods pre-multiply a descent direction by a preconditioning matrix. Due to the large number of parameters of machine learning problems, full-matrix preconditioning methods are prohibitively expensive. We show how to modify full-matrix adaptive regularization in order to mak…

Cited by 70SourcePDFScholar
2019

Explaining Landscape Connectivity of Low-cost Solutions for Multilayer Nets

NeurIPS 2019poster

Mode connectivity is a surprising phenomenon in the loss landscape of deep nets. Optima---at least those discovered by gradient-based optimization---turn out to be connected by simple paths on which the loss function is almost constant. Often, these paths can be chosen to be piece-wise linear, with…

2018

A Multi-Position Joint Particle Filtering Method for Vehicle Localization in Urban Area

IROS 2018poster

Robust localization is a prerequisite for autonomous vehicles. Traditional visual localization methods like visual odometry suffer error accumulation on long range navigation. In this paper, a flexible road map based probabilistic filtering method is proposed to tackle this problem. To effectively m…

Cited by 5SourceScholar