← Search

Yan Zhang

169 accepted papers

2026

Beyond Counting: Evaluating Abstract and Emotional Reasoning in Vision-Language Models

AAAI 2026technical

Despite the rapid progress of Vision Language Models (VLMs), existing benchmarks still concentrate on coarse-grained object recognition or simple relational reasoning, leaving the fine-grained and higher-order reasoning abilities of these systems largely unexamined. To bridge this critical evaluati

Cited by 0SourcePDFScholar
2026

Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling

ICML 2026spotlight

The rapid evolution of generative models has unlocked new potentials in protein binder design, a pivotal task in structural biology, by facilitating end-to-end generation via joint sequence-structure modeling or hallucination. However, existing approaches are predominantly implemented under a single…

Cited by 0SourceScholar
2026

CofactGVR: Counterfactual Intervention for Grounded Visual Reasoning

ICML 2026poster

Despite rapid progress in Grounded Visual Reasoning (GVR) with MLLMs and RL-style fine-tuning, existing approaches often lack effective learning signals for intermediate grounding decisions and are prone to shortcut solutions. In this work, we explicitly decompose GVR into Evidence Generation follow…

Cited by 0SourceScholar
2026

Controlled Collaboration Geometry for Personalized Federated Learning

ICML 2026poster

In personalized federated learning (PFL), collaboration graphs specify model aggregation among clients. However, without constraints on the collaboration geometry, training can drift into two degenerate regimes: global consensus or spontaneous clustering. This paper provides a unified dynamical anal…

Cited by 0SourceScholar
2026

Cross-Domain AI-Generated Image Quality Assessment via Content-Distortion Awareness

IJCAI 2026

With the expanding use of artificial intelligence generated images (AGIs) in scenarios such as gaming, art, and film production, evaluating their quality is essential to ensure their practical utility. To guarantee effective quality measurement, both content and distortion must be considered. Howeve

Cited by 0Scholar
2026

Cross-Tactile Sensor Representation Learning

ICML 2026poster

Visuo-tactile sensors have been widely adopted in robotic manipulation. However, inherent heterogeneity in sensor designs hinders the learning of unified tactile representations in cross-sensor scenarios. Existing methods that focus on reconstruction or task-specific supervision often fail to captur…

Cited by 0SourceScholar
2026

GPR-GSLAM: Gaussian Process Regression-Enhanced Real-Time RGB-D SLAM Using Gaussian Splatting

RA-L 2026

3D Gaussian Splatting (3DGS) has recently revolutionized novel view synthesis and provided a new paradigm for photorealistic Simultaneous Localization and Mapping (SLAM). However, current 3DGS-based RGB-D SLAM systems still face three key challenges: incomplete depth observations due to sensor noise

Cited by 0SourceScholar
2026

Graph-Driven Domain Co-Adaptation for Cross-Domain Image Quality Assessment

AAAI 2026technical

As a typical information medium, images are widely utilized across various scenarios. Measuring image quality accurately is meaningful for the subsequent usability of images. However, significant variations exist in image types and distortion types in different scenarios. And, acquiring labeled imag

Cited by 0SourcePDFScholar
2026

LECDPR:LLM Enhancement and Concept-Document Interactive Modeling for Prerequisite Relation Prediction

IJCAI 2026

Accurate prediction of prerequisite relations among concepts is important for course planning and intelligent tutoring systems. Existing text-based methods are frequently contaminated by noise such as redundant phrasing, ambiguous sentences, and domain-specific colloquialisms. Moreover, previous met

Cited by 0Scholar
2026

Learning Transferable Temporal Primitives for Video Reasoning via Synthetic Videos

CVPR 2026

The transition from image to video understanding requires vision-language models (VLMs) to shift from recognizing static patterns to reasoning over temporal dynamics such as motion trajectories, speed changes, and state transitions. Yet current post-training methods fall short due to two critical li

Cited by 0SourcecodeScholar
2026

LookFlow: Training-Free and Efficient High-Resolution Image Synthesis via Dynamic Lookahead Guidance Flow

AAAI 2026technical

Rectification flow Transformers (RFTs) have shown promising performance in diffusion-based image synthesis but are typically confined to lower-resolution scenarios, limiting their ability to generate high-resolution images. Existing resolution extrapolation approaches often suffer from excessive co

Cited by 0SourcePDFScholar
2026

MTVCraft: Tokenizing 4D Motion for Arbitrary Character Animation

ICLR 2026poster

Character image animation has rapidly advanced with the rise of digital humans. However, existing methods rely largely on 2D-rendered pose images for motion guidance, which limits generalization and discards essential 4D information for open-world animation. To address this, we propose MTVCraft (Mot…

Cited by 0SourcecodeScholar
2026

Mem-T: Densifying Rewards for Long-Horizon Memory Agents

ICML 2026poster

Memory agents, which depart from predefined memory-processing pipelines by endogenously managing the processing, storage, and retrieval of memories, have garnered increasing attention for their autonomy and adaptability. However, existing training paradigms remain constrained: agents often traverse …

Cited by 0SourceScholar
2026

Mitigating Error Accumulation in Knowledge Editing for Multi-Hop Question Answering

AAAI 2026technical

Knowledge editing (KE) has emerged as an effective approach for updating factual information in large language models (LLMs) without the need for full retraining. Most of the existing methods for addressing the "ripple effect" in KE adopt a chain-structured reasoning process, making them vulnerable

Cited by 0SourcePDFScholar
2026

Regime-Adaptive Bayesian Optimization via Dirichlet Process Mixtures of Gaussian Processes

ICML 2026poster

Standard Bayesian Optimization (BO) assumes uniform smoothness across the search space—an assumption violated in multi-regime problems such as molecular conformation search through distinct energy basins or drug discovery across heterogeneous molecular scaffolds. A single GP either oversmooths sharp…

Cited by 0SourceScholar
2026

Robustness-Aware Tool Selection and Manipulation Planning with Learned Energy-Informed Guidance

ICRA 2026poster

Humans subconsciously choose robust ways of selecting and using tools, for example, choosing a ladle over a flat spatula to serve meatballs. However, robustness under external disturbances remains underexplored in robotic tool-use planning. This paper presents a robustness-aware method that jointly …

2026

SIAM: Towards Generalizable Articulated Object Modeling via Single Robot-Object Interaction

AAAI 2026technical

Articulated object modeling, which represents interconnected rigid bodies with their geometry, part segmentation, articulation tree, and physical properties, is crucial for robotic perception and manipulation. Recently existing methods like SAGCI leverage Interactive Perception (IP) to refine models

Cited by 0SourcePDFScholar
2026

ScaleADFG: Affordance-Based Dexterous Functional Grasping via Scalable Dataset

RA-L 2026

Dexterous functional tool-use grasping is essential for effective robotic manipulation of tools. However, existing approaches face significant challenges in efficiently constructing large-scale datasets and ensuring generalizability to everyday object scales. These issues primarily arise from size m

Cited by 1SourcecodeScholar
2026

SegMem-RAG: Adaptive Memory for Retrieval-Augmented Generation in Open-Ended Knowledge Environments

AAAI 2026technical

Retrieval-Augmented Generation (RAG) improves the factual accuracy of large language models by grounding responses in external content. However, most RAG systems assume access to static and well-organized corpora with fixed retrieval logic. In practice, real-world sources are heterogeneous and unlab

Cited by 0SourcePDFScholar
2026

SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing

CVPR 2026

Sketch editing requires jointly handling high-level semantic changes and precise local redrawing, a combination that is particularly challenging for sparse, style-sensitive line art. Unlike natural images, sketches rely on minimal visual cues, making it difficult for existing methods to reconcile gl

Cited by 0SourceScholar
2026

UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment

AAAI 2026technical

Image-based virtual try-on (VTON) aims to synthesize photorealistic images of a person wearing specified garments. Despite significant progress, building a universal VTON framework that can flexibly handle diverse and complex tasks remains a major challenge. Recent methods explore multi-task VTON fr

Cited by 0SourcePDFScholar
2026

UniVBench: Towards Unified Evaluation for Video Foundation Models

CVPR 2026

Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-generation multimodal systems. However, existing evaluation benchmarks remain fragmented and limited in scope, as they each

Cited by 0SourcecodeScholar
2026

Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding

CVPR 2026

Frame selectoin is crucial due to high frame redundancy and limited context windows when applying Large Vision-Language Models (LVLMs) to long videos. Current methods typically select frames with high relevance to a given query, resulting a disjointed set of frames that disregard the narrative struc

Cited by 0SourcecodeScholar
2026

When Simple Problems Wear Complex Costumes: Improving Efficiency in LRM’s Adaptive Reasoning

ICML 2026poster

Recent Large Reasoning Models (LRMs) have demonstrated powerful multi-step problem-solving capabilities but often suffer from inefficiency due to an ``overthinking phenomenon", where they apply complex reasoning to simple tasks, resulting in unnecessary computational cost and latency. While adaptive…

Cited by 0SourceScholar
2025

A Theory for Conditional Generative Modeling on Multiple Data Sources

ICML 2025poster

The success of large generative models has driven a paradigm shift, leveraging massive multi-source data to enhance model capabilities. However, the interaction among these sources remains theoretically underexplored. This paper provides a first attempt to fill this gap by rigorously analyzing multi…

2025

APKGC: Noise-enhanced Multi-Modal Knowledge Graph Completion with Attention Penalty

AAAI 2025technical

Multimodal knowledge graphs (MMKG) store structured world knowledge enriched with multimodal descriptive information. However, MMKG often faces the challenge of incompleteness. The primary objective of multimodal knowledge graph completion (MMKGC) is to predict missing entities within MMKG. Current…

2025

Adaptive Layered-Trust Robust Defense Mechanism for Personalized Federated Learning

ICASSP 2025accepted

Personalized Federated Learning (PFL) is confronted with escalating security threats, yet existing defense strategies primarily concentrate on traditional federated learning, lacking robust defense mechanisms tailored for PFL. To fortify the robustness of PFL against stealthy malicious attacks, we p…

Cited by 0SourceScholar
2025

BUFF: Bayesian Uncertainty Guided Diffusion Probabilistic Model for Single Image Super-Resolution

AAAI 2025technical

Super-resolution (SR) techniques are critical for enhancing image quality, particularly in scenarios where high-resolution imagery is essential yet limited by hardware constraints. Existing diffusion models for SR have relied predominantly on Gaussian models for noise generation, which often fall sh…

Cited by 0SourcePDFScholar
2025

CFBench: A Comprehensive Constraints-Following Benchmark for LLMs

ACL 2025long

The adeptness of Large Language Models (LLMs) in comprehending and following natural language instructions is critical for their deployment in sophisticated real-world applications. Existing evaluations mainly focus on fragmented constraints or narrow scenarios, but they overlook the comprehensivene…

2025

CoA: Towards Real Image Dehazing via Compression-and-Adaptation

CVPR 2025poster

Learning-based image dehazing algorithms have shown remarkable success in synthetic domains. However, real image dehazing is still in suspense due to computational resource constraints and the diversity of real-world scenes. Therefore, there is an urgent need for an algorithm that excels in both eff…

2025

DGCPL: Dual Graph Distillation for Concept Prerequisite Relation Learning

IJCAI 2025

Concept prerequisite relations determine the learning order of knowledge concepts in one domain, which has an important impact on teachers' course design and students' personalized learning. Current research usually predicts concept prerequisite relations from the perspective of knowledge, and rarel

2025

DecoupleSearch: Decouple Planning and Search via Hierarchical Reward Modeling

EMNLP 2025

Retrieval-Augmented Generation (RAG) systems have emerged as a pivotal methodology for enhancing Large Language Models (LLMs) through the dynamic integration of external knowledge. To further improve RAG’s flexibility, Agentic RAG introduces autonomous agents into the workflow. However, Agentic RAG

Cited by 0SourcePDFScholar
2025

Distilling Spatially-Heterogeneous Distortion Perception for Blind Image Quality Assessment

CVPR 2025poster

In the Blind Image Quality Assessment (BIQA) field, accurately assessing the quality of authentically distorted images presents a substantial challenge due to the diverse distortion types in natural settings. Existing state-of-the-art IQA methods mix a sequence of distortions into entire images to e…

Cited by 0SourcePDFScholar
2025

Do Less and Achieve More: Free Condition Video Outpainting with Diffusion Model

ICASSP 2025accepted

Video outpainting aims to extend the content of a video beyond its original spatial boundaries. Existing methods tend to condition the generation process on a single frame or caption, failing to address the challenge in long videos with multiple video clips. To address this, we extend the diffusion-…

Cited by 0SourceScholar
2025

ESCNet:Edge-Semantic Collaborative Network for Camouflaged Object Detection

ICCV 2025poster

Camouflaged object detection (COD) faces unique challenges where target boundaries are intrinsically ambiguous due to their textural similarity to backgrounds. Existing methods relying on single-modality features often produce fragmented predictions due to insufficient boundary constraints.To addres…

2025

ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledge

EMNLP 2025

We introduce ESGenius , a comprehensive benchmark for evaluating and enhancing the proficiency of Large Language Models (LLMs) in Environmental, Social, and Governance (ESG) and sustainability-focused question answering. ESGenius comprises two key components: (i) ESGenius-QA , a collection of 1,136

2025

Enhancing Retrieval-Augmented Generation via Evidence Tree Search

ACL 2025long

Retrieval-Augmented Generation (RAG) is widely used to enhance Large Language Models (LLMs) by grounding responses in external knowledge. However, in real-world applications, retrievers often return lengthy documents with redundant or irrelevant content, confusing downstream readers. While evidence…

Cited by 0SourcePDFScholar
2025

FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering

ACL 2025finding

Large language models (LLMs) are prone to capturing biases from training corpus, leading to potential negative social impacts. Existing prompt-based debiasing methods exhibit instability due to their sensitivity to prompt changes, while fine-tuning-based techniques incur substantial computational ov…

2025

Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering

ACL 2025short

Multimodal large language models (MLLMs) still struggle with complex reasoning tasks in Visual Question Answering (VQA). While current methods have advanced by incorporating visual prompts, our study uncovers critical limitations: these approaches indiscriminately annotate all detected objects for e…

Cited by 0SourcePDFScholar
2025

Feature Denoising Diffusion Model for Blind Image Quality Assessment

AAAI 2025technical

Blind Image Quality Assessment (BIQA) aims to evaluate image quality in line with human perception, without reference benchmarks. Currently, deep learning BIQA methods typically depend on using features from high-level tasks for transfer learning. However, the inherent differences between BIQA and t…

Cited by 1SourcePDFScholar
2025

Few-Shot Image Quality Assessment via Adaptation of Vision-Language Models

ICCV 2025poster

Image Quality Assessment (IQA) remains an unresolved challenge in computer vision due to complex distortions, diverse image content, and limited data availability. Existing Blind IQA (BIQA) methods largely rely on extensive human annotations, which are labor-intensive and costly due to the demanding…

2025

HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models

ACL 2025long

Medical Vision-Language Models (Med-VLMs) have achieved success across various tasks, yet most existing methods overlook the modality misalignment issue that can lead to untrustworthy responses in clinical settings. In this paper, we propose Hierarchical Self-Contrastive Rewarding (HSCR), a novel ap…

2025

Improving Equivariant Networks with Probabilistic Symmetry Breaking

ICLR 2025poster

Equivariance encodes known symmetries into neural networks, often enhancing generalization. However, equivariant networks cannot *break* symmetries: the output of an equivariant network must, by definition, have at least the same self-symmetries as its input. This poses an important problem, both (1…

Cited by 8SourcePDFScholar
2025

LTD-Bench: Evaluating Large Language Models by Letting Them Draw

NeurIPS 2025poster

Current evaluation paradigms for large language models (LLMs) represent a critical blind spot in AI research—relying on opaque numerical metrics that conceal fundamental limitations in spatial reasoning while providing no intuitive understanding of model capabilities. This deficiency creates a dange…

Cited by 0SourcecodeScholar
2025

Learning Concept Prerequisite Relation via Global Knowledge Relation Optimization

AAAI 2025technical

Learning concept prerequisite relations helps better master and build a logically coherent knowledge structure. Many studies use graph neural networks to create heterogeneous knowledge networks that enhance concept representations. However, different types of relations in these networks can influenc…

2025

Learning Problem Decomposition for Efficient Sequential Multi-Object Manipulation Planning

RA-L 2025

We present an efficient task and motion replanning approach for sequential multi-object manipulation in dynamic environments. Conventional Task And Motion Planning (TAMP) solvers experience an exponential increase in planning time as the planning horizon and number of objects grow, limiting their ap

Cited by 0SourceScholar
2025

Learning Stroke-Order Dynamics in Few-Shot Font Generation via Sequential Awareness

ICASSP 2025accepted

Few-shot font generation has garnered significant attention due to its wide range of applications. The mainstream methods are based on the idea of the style and content disentangled representation learning and can be mainly categorized into two kinds of methods according to the prior used, i.e., the…

Cited by 0SourceScholar
2025

M-MAD: Multidimensional Multi-Agent Debate for Advanced Machine Translation Evaluation

ACL 2025long

Recent advancements in large language models (LLMs) have given rise to the LLM-as-a-judge paradigm, showcasing their potential to deliver human-like judgments. However, in the field of machine translation (MT) evaluation, current LLM-as-a-judge methods fall short of learned automatic metrics. In thi…

2025

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces

IJCAI 2025

In the latest advancements in multimodal learning, effectively addressing the spatial and semantic losses of visual data after encoding remains a critical challenge. This is because the performance of large multimodal models is positively correlated with the coupling between visual encoders and larg

2025

MLEP: Multi-granularity Local Entropy Patterns for Generalized AI-generated Image Detection

NeurIPS 2025poster

Advances in image generation technologies have raised growing concerns about their potential misuse, particularly in producing misinformation and deepfakes. This creates an urgent demand for effective methods to detect AI-generated images (AIGIs). While progress has been made, achieving reliable per…

Cited by 0SourceScholar
2025

MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning

EMNLP 2025

Large-scale reinforcement learning (RL) methods have proven highly effective in enhancing the reasoning abilities of large language models (LLMs), particularly for tasks with verifiable solutions such as mathematics and coding. However, applying this idea to machine translation (MT), where outputs a

2025

Mitigating Posterior Salience Attenuation in Long-Context LLMs with Positional Contrastive Decoding

ACL 2025short

While Large Language Models (LLMs) support long contexts, they struggle with performance degradation within the context window. Current solutions incur prohibitive training costs, leaving statistical behaviors and cost-effective approaches underexplored. From the decoding perspective, we identify th…

Cited by 0SourcePDFScholar
2025

Modality-Fair Preference Optimization for Trustworthy MLLM Alignment

IJCAI 2025

Multimodal large language models (MLLMs) have achieved remarkable success across various tasks. However, separate training of visual and textual encoders often results in a misalignment of the modality. Such misalignment may lead models to generate content that is absent from the input image, a phen

Cited by 0SourcePDFScholar
2025

M³GQA: A Multi-Entity Multi-Hop Multi-Setting Graph Question Answering Benchmark

ACL 2025long

Recently, GraphRAG systems have achieved remarkable progress in enhancing the performance and reliability of large language models (LLMs). However, most previous benchmarks are template-based and primarily focus on few-entity queries, which are monotypic and simplistic, failing to offer comprehensiv…

2025

N3C: Towards Replay-based Novelty Continual Clustering with Class-Overlapping

ICASSP 2025accepted

Deep clustering has excelled in batch settings, but little work has addressed the more practical and challenging continual clustering (CC) with shifting data distributions. Additionally, class-overlapping, also a challenging issue, where classes recur across tasks, is common in real-world scenarios.…

Cited by 0SourceScholar
2025

PRIMAL: Physically Reactive and Interactive Motor Model for Avatar Learning

ICCV 2025poster

We formulate the motor system of an interactive avatar as a generative motion model that can drive the body to move through 3D space in a perpetual, realistic, controllable, and responsive manner. Although human motion generation has been extensively studied, many existing methods lack the responsiv…

Cited by 0SourcePDFScholar
2025

ProtPainter: Draw or Drag Protein via Topology-guided Diffusion

ICLR 2025poster

Recent advances in protein backbone generation have achieved promising results under structural, functional, or physical constraints. However, existing methods lack the flexibility for precise topology control, limiting navigation of the backbone space. We present $\textbf{ProtPainter}$, a diffusion…

Cited by 0SourcePDFScholar
2025

Retrieval Augmented Instruction Tuning for Open NER with Large Language Models

COLING 2025main

The strong capability of large language models (LLMs) has been applied to information extraction (IE) through either retrieval augmented prompting or instruction tuning (IT). However, the best way to incorporate information with LLMs for IE remains an open question. In this paper, we explore Retriev…

2025

SCOUT: Semi-supervised Camouflaged Object Detection by Utilizing Text and Adaptive Data Selection

IJCAI 2025

The difficulty of pixel-level annotation has significantly hindered the development of the Camouflaged Object Detection (COD) field. To save on annotation costs, previous works leverage the semi-supervised COD framework that relies on a small number of labeled data and a large volume of unlabeled da

2025

SerialGen: Personalized Image Generation by First Standardization Then Personalization

CVPR 2025poster

In this work, we are interested in achieving both high text controllability and whole-body appearance consistency in the generation of personalized human characters. We propose a novel framework, named SerialGen, which is a serial generation method consisting of two stages: first, a standardization…

Cited by 1SourcePDFScholar
2025

SysBench: Can LLMs Follow System Message?

ICLR 2025poster

Large Language Models (LLMs) have become instrumental across various applications, with the customization of these models to specific scenarios becoming increasingly critical. System message, a fundamental component of LLMs, is consist of carefully crafted instructions that guide the behavior of mod…

Cited by 0SourcePDFScholar
2025

TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement

NAACL 2025findings

Large Language Models (LLMs) have achieved impressive results in Machine Translation (MT). However, human evaluations reveal that LLM-generated translations still contain various errors. Notably, feeding the error information back into the LLMs can facilitate self-refinement, leading to enhanced tra…

2025

Test-Time Code-Switching for Cross-lingual Aspect Sentiment Triplet Extraction

NAACL 2025long

Aspect Sentiment Triplet Extraction (ASTE) is a thriving research area with impressive outcomes being achieved on high-resource languages. However, the application of cross-lingual transfer to the ASTE task has been relatively unexplored, and current code-switching methods still suffer from term bou…

Cited by 0SourcePDFScholar
2025

Towards Macro-AUC Oriented Imbalanced Multi-Label Continual Learning

AAAI 2025technical

In Continual Learning (CL), while existing work primarily focuses on the multi-class classification task, there has been limited research on Multi-Label Learning (MLL). In practice, MLL datasets are often class-imbalanced, making it inherently challenging, a problem that is even more acute in CL.…

2025

Towards Micro-Action Recognition with Limited Annotations: An Asynchronous Pseudo Labeling and Training Approach

IJCAI 2025

Micro-Action Recognition (MAR) aims to classify subtle human actions in video. However, annotating MAR datasets is particularly challenging due to the subtlety of actions. To this end, we introduce the setting of Semi-Supervised MAR (SSMAR), where only a part of samples are labeled. We first evaluat

2025

Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues

AAAI 2025technical

Video text-based visual question answering (TextVQA) is a practical task that aims to answer questions by jointly reasoning textual and visual information in a given video. Inspired by the development of TextVQA in image domain, existing Video TextVQA approaches leverage a language model (e.g. T5) t…

2025

Tuning Less, Prompting More: In-Context Preference Learning Pipeline for Natural Language Transformation

EMNLP 2025

Natural language transformation (NLT) tasks, such as machine translation (MT) and text style transfer (TST), require models to generate accurate and contextually appropriate outputs. However, existing approaches face significant challenges, including the computational costs of leveraging large pre-t

2025

UCOD-DPL: Unsupervised Camouflaged Object Detection via Dynamic Pseudo-label Learning

CVPR 2025highlight

Unsupervised Camoflaged Object Detection (UCOD) has gained attention since it doesn't need to rely on extensive pixel-level labels. Existing UCOD methods typically generate pseudo-labels using fixed strategies and train 1 x1 convolutional layers as a simple decoder, leading to low performance compar…

2025

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval

ACL 2025long

Document retrieval in real-world scenarios faces significant challenges due to diverse document formats and modalities. Traditional text-based approaches rely on tailored parsing techniques that disregard layout information and are prone to errors, while recent parsing-free visual methods often stru…

Cited by 0SourcePDFScholar
2025

When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding

NeurIPS 2025poster

Large Multimodal Models (LMMs) have achieved impressive progress in visual perception and reasoning. However, when confronted with visually ambiguous or non-semantic scene text, they often struggle to accurately spot and understand the content, frequently generating semantically plausible yet visual…

Cited by 0SourceScholar
2025

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs

NeurIPS 2025poster

Multi-modal Large Language Models (MLLMs) excel at single-image tasks but struggle with multi-image understanding due to cross-modal misalignment, leading to hallucinations (context omission, conflation, and misinterpretation). Existing methods using Direct Preference Optimization (DPO) constrain op…

Cited by 0SourcecodeScholar
2024

3DSAM: Segment Anything in NeRF

ICASSP 2024accepted

Object segmentation within Neural Radiance Fields (NeRF) plays a pivotal role, holding potential to enrich a myriad of downstream applications like NeRF editing. Most existing methods, heavily reliant on feature similarity of 3D space, make it non-trivial to manipulate. Instead of intricate 3D inter…

Cited by 0SourceScholar
2024

AdaSwitch: Adaptive Switching between Small and Large Agents for Effective Cloud-Local Collaborative Learning

EMNLP 2024main

Recent advancements in large language models (LLMs) have been remarkable. Users face a choice between using cloud-based LLMs for generation quality and deploying local-based LLMs for lower computational cost. The former option is typically costly and inefficient, while the latter usually fails to de…

Cited by 2SourcePDFScholar
2024

Adaptive Feature Selection for No-Reference Image Quality Assessment by Mitigating Semantic Noise Sensitivity

ICML 2024poster

The current state-of-the-art No-Reference Image Quality Assessment (NR-IQA) methods typically rely on feature extraction from upstream semantic backbone networks, assuming that all extracted features are relevant. However, we make a key observation that not all features are beneficial, and some may…

Cited by 4SourcePDFScholar
2024

Complementary Fusion Network Based on Frequency Hybrid Attention for Pansharpening

ICASSP 2024accepted

Pansharpening is a feasible way to obtain the high-resolution (HR) multispectral (MS) images by using panchromatic (PAN) images to sharpen low-resolution MS images. Despite its great advances, most existing pansharpening methods neglect the importance of integrating local and non-local characteristi…

Cited by 0SourceScholar
2024

CrossTune: Black-Box Few-Shot Classification with Label Enhancement

COLING 2024main

Training or finetuning large-scale language models (LLMs) requires substantial computation resources, motivating recent efforts to explore parameter-efficient adaptation to downstream tasks. One approach is to treat these models as black boxes and use forward passes (Inference APIs) to interact with…

Cited by 3SourcePDFScholar
2024

DDN-Net: Deep Residual Shrinkage Denoising Networks with Channel-Wise Adaptively Soft Thresholds for Automated Major Depressive Disorder Identification

ICASSP 2024accepted

Major Depressive Disorder (MDD) is a severe mental illness that poses significant challenges to society and families. Recently, using rs-fMRI, several graph-based methods have been proposed for MDD diagnosis. However, these methods encode the entire braingraph directly, without considering the subgr…

Cited by 0SourceScholar
2024

Degrees of Freedom Matter: Inferring Dynamics from Point Trajectories

CVPR 2024poster

Understanding the dynamics of generic 3D scenes is fundamentally challenging in computer vision essential in enhancing applications related to scene reconstruction motion tracking and avatar creation. In this work we address the task as the problem of inferring dense long-range motion of 3D points.…

2024

DiffAIL: Diffusion Adversarial Imitation Learning

AAAI 2024technical

Imitation learning aims to solve the problem of defining reward functions in real-world decision-making tasks. The current popular approach is the Adversarial Imitation Learning (AIL) framework, which matches expert state-action occupancy measures to obtain a surrogate reward for forward reinforceme…

2024

DynaThink: Fast or Slow? A Dynamic Decision-Making Framework for Large Language Models

EMNLP 2024main

Large language models (LLMs) have demonstrated emergent capabilities across diverse reasoning tasks via popular Chains-of-Thought (COT) prompting. However, such a simple and fast COT approach often encounters limitations in dealing with complicated problems, while a thorough method, which considers…

2024

EgoGen: An Egocentric Synthetic Data Generator

CVPR 2024poster

Understanding the world in first-person view is fundamental in Augmented Reality (AR). This immersive perspective brings dramatic visual changes and unique challenges compared to third-person views. Synthetic data has empowered third-person-view vision models but its application to embodied egocentr…

Cited by 17SourcePDFScholar
2024

Graph Neural Networks for Learning Equivariant Representations of Neural Networks

ICLR 2024oral

Neural networks that process the parameters of other neural networks find applications in domains as diverse as classifying implicit neural representations, generating neural network weights, and predicting generalization errors. However, existing approaches either overlook the inherent permutation…

2024

Improved Generalization of Weight Space Networks via Augmentations

ICML 2024poster

Learning in deep weight spaces (DWS), where neural networks process the weights of other neural networks, is an emerging research direction, with applications to 2D and 3D neural fields (INRs, NeRFs), as well as making inferences about other types of neural networks. Unfortunately, weight space mode…

2024

Improving Large Language Models in Event Relation Logical Prediction

ACL 2024long

Event relations are crucial for narrative understanding and reasoning. Governed by nuanced logic, event relation extraction (ERE) is a challenging task that demands thorough semantic understanding and rigorous logical reasoning. In this paper, we conduct an in-depth investigation to systematically e…

2024

Integrating Global Context Contrast and Local Sensitivity for Blind Image Quality Assessment

ICML 2024spotlight

Blind Image Quality Assessment (BIQA) mirrors subjective made by human observers. Generally, humans favor comparing relative qualities over predicting absolute qualities directly. However, current BIQA models focus on mining the "local" context, i.e., the relationship between information among indiv…

Cited by 7SourcePDFScholar
2024

Ladder: A Model-Agnostic Framework Boosting LLM-based Machine Translation to the Next Level

EMNLP 2024main

General-purpose Large Language Models (LLMs) like GPT-4 have achieved remarkable advancements in machine translation (MT) by leveraging extensive web content. On the other hand, translation-specific LLMs are built by pre-training on domain-specific monolingual corpora and fine-tuning with human-anno…

2024

LiDAR-Net: A Real-scanned 3D Point Cloud Dataset for Indoor Scenes

CVPR 2024poster

In this paper we present LiDAR-Net a new real-scanned indoor point cloud dataset containing nearly 3.6 billion precisely point-level annotated points covering an expansive area of 30000m^2. It encompasses three prevalent daily environments including learning scenes working scenes and living scenes.…

Cited by 8SourcePDFScholar
2024

Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives

CVPR 2024poster

We propose Lodge a network capable of generating extremely long dance sequences conditioned on given music. We design Lodge as a two-stage coarse to fine diffusion architecture and propose the characteristic dance primitives that possess significant expressiveness as intermediate representations bet…

2024

Logic Learning From Demonstrations for Multi-Step Manipulation Tasks in Dynamic Environments

RA-L 2024

Learning from Demonstration (LfD) stands as an efficient framework for imparting human-like skills to robots. Nevertheless, designing an LfD framework capable of seamlessly imitating, generalizing, and reacting to disturbances for long-horizon manipulation tasks in dynamic environments remains a cha

Cited by 5SourceScholar
2024

Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models

EMNLP 2024finding

Recent advancements in general-purpose or domain-specific multimodal large language models (LLMs) have witnessed remarkable progress for medical decision-making. However, they are designated for specific classification or generative tasks, and require model training or finetuning on large-scale data…

2024

Motif-Matching Based Sub-Braingraph Level Networks for Noisy Resting-State fMRI Analysis

ICASSP 2024accepted

Biomarkers extracted from rs-fMRI based brain functional connectivity (FC) can assist in diagnosing various brain disorders. Recently, several graph-based methods have been proposed for modeling the braingraph of brain disorders and brain disorders diagnosis. However, those methods overlook the subg…

Cited by 0SourceScholar
2024

Object centric architectures enable efficient causal representation learning

ICLR 2024poster

Causal representation learning has showed a variety of settings in which we can disentangle latent variables with identifiability guarantees (up to some reasonable equivalence class). Common to all of these approaches is the assumption that (1) the latent variables are represented as $d$-dimensional…

2024

Quantifying and Mitigating Unimodal Biases in Multimodal Large Language Models: A Causal Perspective

EMNLP 2024finding

Recent advancements in Large Language Models (LLMs) have facilitated the development of Multimodal LLMs (MLLMs). Despite their impressive capabilities, MLLMs often suffer from over-reliance on unimodal biases (e.g., language bias and vision bias), leading to incorrect answers in complex multimodal t…

2024

RELI11D: A Comprehensive Multimodal Human Motion Dataset and Method

CVPR 2024poster

Comprehensive capturing of human motions requires both accurate captures of complex poses and precise localization of the human within scenes. Most of the HPE datasets and methods primarily rely on RGB LiDAR or IMU data. However solely using these modalities or a combination of them may not be adequ…

Cited by 8SourcePDFScholar
2024

RLE: A Unified Perspective of Data Augmentation for Cross-Spectral Re-Identification

NeurIPS 2024poster

This paper makes a step towards modeling the modality discrepancy in the cross-spectral re-identification task. Based on the Lambertain model, we observe that the non-linear modality discrepancy mainly comes from diverse linear transformations acting on the surface of different materials. From this…

2024

Representing Robot Geometry as Distance Fields: Applications to Whole-body Manipulation

ICRA 2024poster

In this work, we propose a novel approach to represent robot geometry as distance fields (RDF) that extends the principle of signed distance fields (SDFs) to articulated kinematic chains. Our method employs a combination of Bernstein polynomials to encode the signed distance for each robot link with…

Cited by 18SourcecodeScholar
2024

Retrieved In-Context Principles from Previous Mistakes

EMNLP 2024main

In-context learning (ICL) has been instrumental in adapting large language models (LLMs) to downstream tasks using correct input-output examples. Recent advances have attempted to improve model performance through principles derived from mistakes, yet these approaches suffer from lack of customizati…

Cited by 5SourcePDFScholar
2024

Self-Improving for Zero-Shot Named Entity Recognition with Large Language Models

NAACL 2024short

Exploring the application of powerful large language models (LLMs) on the named entity recognition (NER) task has drawn much attention recently. This work pushes the performance boundary of zero-shot NER with LLMs by proposing a training-free self-improving framework, which utilizes an unlabeled cor…

2024

Semi-Supervised Blind Image Quality Assessment through Knowledge Distillation and Incremental Learning

AAAI 2024technical

Blind Image Quality Assessment (BIQA) aims to simulate human assessment of image quality. It has a great demand for labeled data, which is often insufficient in practice. Some researchers employ unsupervised methods to address this issue, which is challenging to emulate the human subjective system.…

Cited by 8SourcePDFScholar
2024

Structure-Aware in-Air Handwritten Text Recognition with Graph-Guided Cross-Modality Translator

ICASSP 2024accepted

In-air handwriting as a new human-computer interaction way plays an important role in many virtual/mixed-reality applications. Existing methods for in-air handwritten text recognition (IAHTR) typically directly process handwriting trajectories with deep neural networks. However, those methods all si…

Cited by 0SourceScholar
2024

Towards Verifiable Text Generation with Evolving Memory and Self-Reflection

EMNLP 2024main

Despite the remarkable ability of large language models (LLMs) in language comprehension and generation, they often suffer from producing factually incorrect information, also known as hallucination. A promising solution to this issue is verifiable text generation, which prompts LLMs to generate con…

Cited by 16SourcePDFScholar
2024

UNO-DST: Leveraging Unlabelled Data in Zero-Shot Dialogue State Tracking

NAACL 2024findings

Previous zero-shot dialogue state tracking (DST) methods only apply transfer learning, but ignore unlabelled data in the target domain.We transform zero-shot DST into few-shot DST by utilising such unlabelled data via joint and self-training methods. Our method incorporates auxiliary tasks that gene…

2024

Unsupervised Concept Discovery Mitigates Spurious Correlations

ICML 2024poster

Models prone to spurious correlations in training data often produce brittle predictions and introduce unintended biases. Addressing this challenge typically involves methods relying on prior knowledge and group annotation to remove spurious correlations, which may not be readily available in many a…

2023

Allies: Prompting Large Language Model with Beam Search

EMNLP 2023long findings

With the advance of large language models (LLMs), the research field of LLM applications becomes more and more popular and the idea of constructing pipelines to accomplish complex tasks by stacking LLM API calls come true. However, this kind of methods face two limitations: narrow information covera…

Cited by 0SourcecodeScholar
2023

Annotating Covert Hazardous Driving Scenarios Online: Utilizing Drivers' Electroencephalography (EEG) Signals

ICRA 2023poster

As autonomous driving systems prevail, it is becoming increasingly critical that the systems learn from databases containing fine-grained driving scenarios. Most databases currently available are human-annotated; they are expensive, time-consuming, and subject to behavioral biases. In this paper, we…

Cited by 3SourceScholar
2023

CHEER: Centrality-aware High-order Event Reasoning Network for Document-level Event Causality Identification

ACL 2023long

Document-level Event Causality Identification (DECI) aims to recognize causal relations between events within a document. Recent studies focus on building a document-level graph for cross-sentence reasoning, but ignore important causal structures — there are one or two “central” events that prevail…

Cited by 19SourcePDFScholar
2023

Cascading Bandits: Optimizing Recommendation Frequency in Delayed Feedback Environments

NeurIPS 2023poster

Delayed feedback is a critical problem in dynamic recommender systems. In practice, the feedback result often depends on the frequency of recommendation. Most existing online learning literature fails to consider optimization of the recommendation frequency, and regards the reward from each successf…

Cited by 2SourcePDFScholar
2023

CrossSplit: Mitigating Label Noise Memorization through Data Splitting

ICML 2023poster

We approach the problem of improving robustness of deep learning algorithms in the presence of label noise. Building upon existing label correction and co-teaching methods, we propose a novel training procedure to mitigate the memorization of noisy labels, called CrossSplit, which uses a pair of neu…

2023

Data-Efficient Image Quality Assessment with Attention-Panel Decoder

AAAI 2023technical

Blind Image Quality Assessment (BIQA) is a fundamental task in computer vision, which however remains unresolved due to the complex distortion conditions and diversified image contents. To confront this challenge, we in this paper propose a novel BIQA pipeline based on the Transformer architecture,…

2023

DecomFormer: Decompose Self-Attention Via Fourier Transform for VHR Aerial Image Scene Classification

ICASSP 2023accepted

Very high-resolution (VHR) aerial image scene classification is an essential task for aerial image understanding. Although transformer-based models have demonstrated strong ability in natural image classification, transformer-based methods on VHR aerial image tasks are still lack of concern because…

Cited by 0SourceScholar
2023

Empirical Study of Zero-Shot NER with ChatGPT

EMNLP 2023long main

Large language models (LLMs) exhibited powerful capability in various natural language processing tasks. This work focuses on exploring LLM performance on zero-shot information extraction, with a focus on the ChatGPT and named entity recognition (NER) task. Inspired by the remarkable reasoning capab…

Cited by 0SourcecodeScholar
2023

Equivariance with Learned Canonicalization Functions

ICML 2023poster

Symmetry-based neural networks often constrain the architecture in order to achieve invariance or equivariance to a group of transformations. In this paper, we propose an alternative that avoids this architectural constraint by learning to produce canonical representations of the data. These canonic…

Cited by 84SourcePDFScholar
2023

FCIR: Rethink Aerial Image Super Resolution with Fourier Analysis

ICASSP 2023accepted

Recent years, deep-learning-based methods achieve remarkable improvements on the super-resolution (SR) task. However, recovering high-quality (HQ) texture from the low-quality (LQ) aerial image is still challenging due to the limited contextual modeling ability of current deep-learning methods as we…

Cited by 0SourceScholar
2023

Hierarchical Hypergraph Recurrent Attention Network for Temporal Knowledge Graph Reasoning

ICASSP 2023accepted

Temporal knowledge graph (TKG) serves as an essential tool in modeling complex event relations among real-world entities. A temporal knowledge graph can be viewed as a collection of knowledge graph snapshots ordered by time. Reasoning over such graphs remains nontrivial as temporal causal dependenci…

Cited by 0SourceScholar
2023

History Semantic Graph Enhanced Conversational KBQA with Temporal Information Modeling

ACL 2023long

Context information modeling is an important task in conversational KBQA. However, existing methods usually assume the independence of utterances and model them in isolation. In this paper, we propose a History Semantic Graph Enhanced KBQA model (HSGE) that is able to effectively model long-range se…

Cited by 2SourcePDFScholar
2023

How Well Do Text Embedding Models Understand Syntax?

EMNLP 2023long findings

Text embedding models have significantly contributed to advancements in natural language processing by adeptly capturing semantic properties of textual data. However, the ability of these models to generalize across a wide range of syntactic contexts remains under-explored. In this paper, we first d…

Cited by 0SourcecodeScholar
2023

Point Clouds Outlier Removal Method Based on Improved Mahalanobis and Completion

RA-L 2023

Point clouds have been regarded as a representative format for 3D visualization of real-world objects or scenes. However, point clouds acquired from depth cameras or laser scanning devices commonly contain outliers. Outlier removal performance will directly affect the downstream applications. Existi

Cited by 11SourceScholar
2023

Probabilistic Human Mesh Recovery in 3D Scenes from Egocentric Views

ICCV 2023oral

Automatic perception of human behaviors during social interactions is crucial for AR/VR applications, and an essential component is estimation of plausible 3D human pose and shape of our social partners from the egocentric view. One of the biggest challenges of this task is severe body truncation du…

Cited by 30PDFcodeScholar
2023

Self-Supervised Boundary Point Prediction Task for Point Cloud Domain Adaptation

RA-L 2023

Unsupervised domain adaptation (UDA) could significantly improve the cross-domain performance of current supervised 3D deep learning methods and have a widespread application prospect. However, the domain gap between source domain and target domain renders the UDA problem highly challenging. In this

Cited by 10SourceScholar
2023

Synthesizing Diverse Human Motions in 3D Indoor Scenes

ICCV 2023poster

We present a novel method for populating 3D indoor scenes with virtual humans that can navigate in the environment and interact with objects in a realistic manner. Existing approaches rely on high-quality training sequences that contain captured human motions and the 3D scenes they interact with. Ho…

Cited by 66PDFcodeScholar
2023

T5-SR: A Unified Seq-to-Seq Decoding Strategy for Semantic Parsing

ICASSP 2023accepted

Translating natural language queries into SQLs in a seq2seq manner has attracted much attention recently. However, compared with abstract-syntactic-tree-based SQL generation, seq2seq semantic parsers face much more challenges, including poor quality on schematical information prediction and poor sem…

Cited by 0SourceScholar
2023

Unlocking Slot Attention by Changing Optimal Transport Costs

ICML 2023poster

Slot attention is a powerful method for object-centric modeling in images and videos. However, its set-equivariance limits its ability to handle videos with a dynamic number of objects because it cannot break ties. To overcome this limitation, we first establish a connection between slot attention a…

2022

Analyzing and Evaluating Faithfulness in Dialogue Summarization

EMNLP 2022main

Dialogue summarization is abstractive in nature, making it suffer from factual errors. The factual correctness of summaries has the highest priority before practical applications. Many efforts have been made to improve faithfulness in text summarization. However, there is a lack of systematic study…

2022

Boundary-Aware Bias Loss for Transformer-Based Aerial Image Segmentation Model

ICASSP 2022accepted

Inspired by the tremendous success of the transformer-based model in natural language processing (NLP), many efforts introduce the transformer-based model into the image processing tasks. However, naive transformer models have to down-sample the image resolution to satisfy computational restrictions…

Cited by 0SourceScholar
2022

Compositional Human-Scene Interaction Synthesis with Semantic Control

ECCV 2022poster

"Synthesizing natural interactions between virtual humans and their 3D environments is critical for numerous applications, such as computer games and AR/VR experiences. Recent methods mainly focus on modeling geometric relations between 3D environments and humans, where the high-level semantics of t…

2022

ERGO: Event Relational Graph Transformer for Document-level Event Causality Identification

COLING 2022main

Document-level Event Causality Identification (DECI) aims to identify event-event causal relations in a document. Existing works usually build an event graph for global reasoning across multiple sentences. However, the edges between events have to be carefully designed through heuristic rules or ext…

2022

EgoBody: Human Body Shape and Motion of Interacting People from Head-Mounted Devices

ECCV 2022poster

"Understanding social interactions from egocentric views is crucial for many applications, ranging from assistive robotics to AR/VR. Key to reasoning about interactions is to understand the body pose and motion of the interaction partner from the egocentric view. However, research in this area is se…

2022

Feature Space Message Passing Network for Medical Image Semantic Segmentation

ICASSP 2022accepted

Accurate semantic segmentation of medical images is of significant importance for subsequent processing and analysis. The encoder-decoder deep learning framework has been widely applied for numerous medical image segmentation tasks. However, most existing approaches are restricted by the limited rec…

Cited by 0SourceScholar
2022

Formal Verification of Stochastic Systems with ReLU Neural Network Controllers

ICRA 2022poster

In this work, we address the problem of formal safety verification for stochastic cyber-physical systems (CPS) equipped with ReLU neural network (NN) controllers. Our goal is to find the set of initial states from where, with a predetermined confidence, the system will not reach an unsafe configurat…

Cited by 7SourceScholar
2022

Generate, Discriminate and Contrast: A Semi-Supervised Sentence Representation Learning Framework

EMNLP 2022main

Most sentence embedding techniques heavily rely on expensive human-annotated sentence pairs as the supervised signals. Despite the use of large-scale unlabeled data, the performance of unsupervised methods typically lags far behind that of the supervised counterparts in most downstream tasks. In thi…

2022

IAM: A Comprehensive and Large-Scale Dataset for Integrated Argument Mining Tasks

ACL 2022long

Traditionally, a debate usually requires a manual preparation process, including reading plenty of articles, selecting the claims, identifying the stances of the claims, seeking the evidence for the claims, etc. As the AI debate attracts more attention these years, it is worth exploring the methods…

2022

Med-DANet: Dynamic Architecture Network for Efficient Medical Volumetric Segmentation

ECCV 2022poster

"For 3D medical image (e.g. CT and MRI) segmentation, the difficulty of segmenting each slice in a clinical case varies greatly. Previous research on volumetric medical image segmentation in a slice-by-slice manner conventionally use the identical 2D deep neural network to segment all the slices of…

2022

Multiset-Equivariant Set Prediction with Approximate Implicit Differentiation

ICLR 2022poster

Most set prediction models in deep learning use set-equivariant operations, but they actually operate on multisets. We show that set-equivariant functions cannot represent certain functions on multisets, so we introduce the more appropriate notion of multiset-equivariance. We identify that the exist…

2022

SAGA: Stochastic Whole-Body Grasping with Contact

ECCV 2022poster

"The synthesis of human grasping has numerous applications including AR/VR, video games and robotics. While methods have been proposed to generate realistic hand-object interaction for object grasping and manipulation, these typically only consider interacting hand alone. Our goal is to synthesize w…

2022

The Wanderings of Odysseus in 3D Scenes

CVPR 2022poster

Our goal is to populate digital environments, in which digital humans have diverse body shapes, move perpetually, and have plausible body-scene contact. The core challenge is to generate realistic, controllable, and infinitely long motions for diverse 3D bodies. To this end, we propose generative mo…

Cited by 53PDFScholar
2021

Attention Is Not Enough: Mitigating the Distribution Discrepancy in Asynchronous Multimodal Sequence Fusion

ICCV 2021poster

Videos flow as the mixture of language, acoustic, and vision modalities. A thorough video understanding needs to fuse time-series data of different modalities for prediction. Due to the variable receiving frequency for sequences from each modality, there usually exists inherent asynchrony across the…

Cited by 74PDFScholar
2021

Bootstrapped Unsupervised Sentence Representation Learning

ACL 2021long

As high-quality labeled data is scarce, unsupervised sentence representation learning has attracted much attention. In this paper, we propose a new framework with a two-branch Siamese Network which maximizes the similarity between two augmented views of each sentence. Specifically, given one augment…

2021

DynaEval: Unifying Turn and Dialogue Level Evaluation

ACL 2021long

A dialogue is essentially a multi-turn interaction among interlocutors. Effective evaluation metrics should reflect the dynamics of such interaction. Existing automatic metrics are focused very much on the turn-level quality, while ignoring such dynamics. To this end, we propose DynaEval, a unified…

2021

Keep the Structure: A Latent Shift-Reduce Parser for Semantic Parsing

IJCAI 2021poster

Traditional end-to-end semantic parsing models treat a natural language utterance as a holonomic structure. However, hierarchical structures exist in natural languages, which also align with the hierarchical structures of logical forms. In this paper, we propose a latent shift-reduce parser, called…

Cited by 5SourcePDFScholar
2021

Learning Motion Priors for 4D Human Body Capture in 3D Scenes

ICCV 2021poster

Recovering high-quality 3D human motion in complex scenes from monocular videos is important for many applications, ranging from AR/VR to robotics. However, capturing realistic human-scene interactions, while dealing with occlusions and partial views, is challenging; current approaches are still far…

Cited by 116PDFcodeScholar
2021

Learning from History: Modeling Temporal Knowledge Graphs with Sequential Copy-Generation Networks

AAAI 2021technical

Large knowledge graphs often grow to store temporal facts that model the dynamic relations or interactions of entities along the timeline. Since such temporal knowledge graphs often suffer from incompleteness, it is important to develop time-aware representation learning models that help to infer th…

2021

Partial-Label and Structure-constrained Deep Coupled Factorization Network

AAAI 2021technical

In this paper, we technically propose an enriched prior guided framework, called Dual-constrained Deep Semi-Supervised Coupled Factorization Network (DS2CF-Net), for discovering hierarchical coupled data representation. To extract hidden deep features, DS2CF-Net is formulated as a partial-label and…

Cited by 6SourcePDFScholar
2021

Revisiting Self-training for Few-shot Learning of Language Model

EMNLP 2021main

As unlabeled data carry rich task-relevant information, they are proven useful for few-shot learning of language model. The question is how to effectively make use of such data. In this work, we revisit the self-training technique for language model fine-tuning and present a state-of-the-art prompt-…

2020

Better Set Representations For Relational Reasoning

NeurIPS 2020poster

Incorporating relational reasoning into neural networks has greatly expanded their capabilities and scope. One defining trait of relational reasoning is that it operates on a set of entities, as opposed to standard vector representations. Existing end-to-end approaches for relational reasoning typic…

2020

FSPool: Learning Set Representations with Featurewise Sort Pooling

ICLR 2020poster

Traditional set prediction models can struggle with simple datasets due to an issue we call the responsibility problem. We introduce a pooling method for sets of feature vectors based on sorting features across elements of the set. This can be used to construct a permutation-equivariant auto-encoder…

Cited by 94SourcecodeScholar
2019

Local Temporal Bilinear Pooling for Fine-Grained Action Parsing

CVPR 2019poster

Fine-grained temporal action parsing is important in many applications, such as daily activity understanding, human motion analysis, surgical robotics and others requiring subtle and precise operations over a long-term period. In this paper we propose a novel bilinear pooling operation, which is use…

Cited by 32PDFScholar
2019

PartNet: A Recursive Part Decomposition Network for Fine-Grained and Hierarchical Shape Segmentation

CVPR 2019poster

Deep learning approaches to 3D shape segmentation are typically formulated as a multi-class labeling problem. These models are trained for a fixed set of labels, which greatly limits their flexibility and adaptivity. We opt for top-down recursive decomposition and develop the first deep learning mod…

Cited by 121PDFScholar
2019

Robust Unsupervised Flexible Auto-weighted Local-coordinate Concept Factorization for Image Clustering

ICASSP 2019accepted

We investigate the high-dimensional data clustering problem by proposing a novel and unsupervised representation learning model called Robust Flexible Auto-weighted Local-coordinate Concept Factorization (RFA-LCF). RFA-LCF integrates the robust flexible CF, robust sparse local-coordinate coding and…

Cited by 0SourceScholar
2018

Feature Quantization for Defending Against Distortion of Images

CVPR 2018poster

In this work, we address the problem of improving robustness of convolutional neural networks (CNNs) to image distortion. We argue that higher moment statistics of feature distributions can be shifted due to image distortion, and the shift leads to performance decrease and cannot be reduced by ordin…

Cited by 35SourcePDFScholar
2018

Learning to Count Objects in Natural Images for Visual Question Answering

ICLR 2018poster

Visual Question Answering (VQA) models have struggled with counting objects in natural images so far. We identify a fundamental problem due to soft attention in these models as a cause. To circumvent this problem, we propose a neural network component that allows robust counting from object proposal…