← Search

Kai CHEN

232 accepted papers

2026

Achieving Expert-Level Agent from Foundation Model via Complexity Curriculum Reinforcement Learning with Synthetic Data

ICLR 2026poster

Large language model (LLM) agents exhibit strong mathematical problem-solving abilities and can even solve International Mathematical Olympiad (IMO) level problems with the assistance of formal proof systems. However, due to weak heuristics for auxiliary constructions, AI for geometry problem solvin…

Cited by 0SourceScholar
2026

Beyond Mode Collapse: Distribution Matching for Diverse Reasoning

ICML 2026poster

On-policy reinforcement learning methods like GRPO suffer from \emph{mode collapse}: they exhibit reduced solution diversity, concentrating probability mass on a single solution once discovered and ceasing exploration of alternative strategies. We show this stems from reverse KL minimization's mode-…

Cited by 0SourceScholar
2026

Context-level Language Modeling by Learning Predictive Context Embeddings

ICML 2026poster

We propose ContextLM, a framework that implicitly learns multi-token prediction by augmenting standard pretraining with an intrinsic next-context prediction objective. ContextLM builds a language model on top of context embeddings that span multiple tokens, enabling better next-token prediction by p…

Cited by 0SourceScholar
2026

EPS3D: End-to-End Feed-Forward 3D Panoptic Segmentation

ICML 2026poster

This paper introduces EPS3D, a new end-to-end feed-forward framework for open-vocabulary 3D panoptic segmentation. Unlike existing methods relying on additional preprocessing, we design an end-to-end architecture, with a distillation-based training strategy on diverse 3D scenes to predict 3D-aware s…

Cited by 0SourceScholar
2026

Embracing Bulky Objects with Humanoid Robots: Whole-Body Manipulation with Reinforcement Learning

ICRA 2026poster

Whole-body manipulation (WBM) for humanoid robots presents a promising approach for executing embracing tasks involving bulky objects, where traditional grasping relying on end-effectors only remains limited in such scenarios due to inherent stability and payload constraints. This paper introduces a…

2026

Enhancing Logical Expressiveness in Graph Neural Networks via Path-Neighbor Aggregation

AAAI 2026technical

Graph neural networks (GNNs) can effectively model structural information of graphs, making them widely used in knowledge graph (KG) reasoning. However, existing studies on the expressive power of GNNs mainly focuses on simple single-relation graphs, and there is still insufficient discussion on the

Cited by 0SourcePDFScholar
2026

Estimating Worst-Case Frontier Risks of Open-Weight LLMs

ICLR 2026poster

In this paper, we study the worst-case frontier risks of the OpenAI gpt-oss model. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning gpt-oss to be as capable as possible in two domains: biology and cybersecurity. To maximize biological risk (bio…

Cited by 0SourcecodeScholar
2026

ExpVid: A Benchmark for Experiment Video Understanding & Reasoning

ICLR 2026poster

Multimodal Large Language Models (MLLMs) hold promise for accelerating scientific discovery by interpreting complex experimental procedures. However, their true capabilities are poorly understood, as existing benchmarks neglect the fine-grained and long-horizon nature of authentic laboratory work, e…

Cited by 0SourcecodeScholar
2026

Exploring Visual Pretraining for Learning Language Intelligence

CVPR 2026

While the most fundamental pretraining paradigm typically trains modality-specific models on their respective datasets, the Platonic Representation Hypothesis that representations eventually align across modalities as data and model scale suggests an intriguing possibility: large language models (LL

Cited by 0SourcecodeScholar
2026

Geometry-Guided Modeling of Foundation Features Enables Generalizable Object Shape Deformation Learning

ICML 2026poster

Monocular 3D shape recovery is fundamental to geometric understanding, yet achieving robust generalization across arbitrary viewpoints and unseen object categories remains a significant challenge. In this paper, we present a generalizable deformation learning framework that reconstructs 3D objects b…

Cited by 0SourceScholar
2026

Hearing Without Noticing? Attention-Aware Stealthy Black-box Adversarial Audio Attacks

ICML 2026poster

Automatic Speech Recognition (ASR) systems, such as those in intelligent assistants, are vulnerable to adversarial examples (AEs). Benign audio clips like music, when embedded with small perturbations, can trick ASR models into recognizing attacker-specified commands. Prior studies focus on minimizi…

Cited by 0SourceScholar
2026

How to Fine-Tune a Reasoning Model? A Teacher–Student Cooperation Framework to Synthesize Student-Consistent SFT Data

ICML 2026poster

A widely adopted strategy for model enhancement is to use synthetic data generated by a stronger model for supervised fine-tuning (SFT). However, for emerging reasoning models like Qwen3-8B, this approach often fails to improve reasoning capabilities and can even lead to a substantial drop in perfor…

Cited by 0SourceScholar
2026

InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation

ICML 2026poster

While Large Language Models (LLMs) hold promise for automating science and education, generating interactive scientific demonstrations demands a complex synthesis of deep domain knowledge and precise reactive coding. Current benchmarks fail to capture this synergy, largely bifurcating into static co…

Cited by 0SourceScholar
2026

InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models

ICLR 2026poster

General SVG modeling remains challenging due to fragmented datasets, limited transferability of methods across tasks, and the difficulty of handling structural complexity. In response, we leverage the strong transfer and generalization capabilities of multimodal large language models (MLLMs) to achi…

Cited by 0SourcecodeScholar
2026

JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence

ICLR 2026poster

The scope of neural code intelligence is rapidly expanding beyond text-based source code to encompass the rich visual outputs that programs generate. This visual dimension is critical for advanced applications like flexible content generation and precise, program-driven editing of visualizations. Ho…

Cited by 0SourcecodeScholar
2026

LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries

ICML 2026poster

Vision-Language-Action (VLA) models have shown promise in robot manipulation but often struggle to generalize to new instructions or complex multi-task scenarios. We identify a critical pathology in current training paradigms where goal-driven data collection creates a dataset bias. In such datasets…

Cited by 0SourceScholar
2026

MM-HELIX: Boosting Multimodal Long-Chain Reflective Reasoning with Holistic Platform and Adaptive Hybrid Policy Optimization

ICLR 2026poster

While current Multimodal Large Language Models (MLLMs) have demonstrated proficiency in reasoning tasks such as mathematics and logic, their capacity for long-chain reflective reasoning, a prerequisite for solving complex real-world problems, remains largely underexplored. In this work, we first co…

Cited by 0SourcecodeScholar
2026

MindCopilot: Towards Formalizing and Evaluating Granular Human-LLM Co-Writing

IJCAI 2026

Recent writing assistants are increasingly shifting from passive, prompt-driven interaction to proactive, suggestion-based completion, which integrates localized continuations into the writing flow and reduces coordination burden. However, existing evaluations simply focus on output quality, failing

Cited by 0Scholar
2026

On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

ICLR 2026poster

In Vision–Language–Action (VLA) models, robustness to real-world perturbations is critical for deployment. Existing methods target simple visual disturbances, overlooking the broader multi-modal perturbations that arise in actions, instructions, environments, and observations. Here, we first evaluat…

Cited by 0SourcecodeScholar
2026

Parameter-efficient Continual Learning for Enhancing Plasticity without Forgetting under Limited Model Capacity

CVPR 2026

Avoiding catastrophic forgetting for previous tasks and maintaining model plasticity to support new tasks are two critical objectives of continual learning. However, existing methods usually neglect one of the two aspects and fail to support long task sequences with satisfactory performance, especia

Cited by 0SourceScholar
2026

RIG: Synergizing Reasoning and Imagination in End-to-End Generalist Policy

ICLR 2026poster

Reasoning before action and imagining potential outcomes (i.e., world models) are essential for embodied agents operating in complex open-world environments. Yet, prior work either incorporates only one of these abilities in an end-to-end agent or integrates multiple specialized models into an agent…

Cited by 0SourceScholar
2026

Reasoning Compartmentalization: Bridging the Concretization Gap via Abstraction-based Routing

ICML 2026poster

While previous research has documented the sensitivity of Large Language Models (LLMs) to surface-level performance degradation, the underlying impact on internal representations and learning dynamics remains under-explored. In this work, we study this question using a controlled setup with paired r…

Cited by 0SourceScholar
2026

Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning

ICLR 2026poster

Recent breakthroughs in reasoning language models have significantly advanced text-based reasoning. On the other hand, Multi-modal Large Language Models (MLLMs) still lag behind, hindered by their outdated internal LLMs. Upgrading these is often prohibitively expensive, as it requires complete visio…

Cited by 0SourcecodeScholar
2026

Rethinking Flow and Diffusion Bridge Models for Speech Enhancement

AAAI 2026technical

Flow matching and diffusion bridge models have emerged as leading paradigms in generative speech enhancement, modeling stochastic processes between paired noisy and clean speech signals based on principles such as flow matching, score matching, and Schrödinger bridge. In this paper, we present a fra

Cited by 0SourcePDFScholar
2026

ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data

ICLR 2026oral

Vision-Language Models (VLMs) have enabled computer use agents (CUAs) that operate GUIs autonomously, showing great potential, yet progress is limited by the lack of large-scale, open-source computer use data and foundation models. In this work, we introduce ScaleCUA, a step toward scaling open-sour…

Cited by 0SourcecodeScholar
2026

SciTS: Scientific Time Series Understanding and Generation with LLMs

ICLR 2026poster

The scientific reasoning ability of large language models (LLMs) has recently attracted significant attention. Time series, as a fundamental modality in scientific data, presents unique challenges that are often overlooked in current multimodal LLMs, which either encode numerical sequences as text o…

Cited by 0SourceScholar
2026

Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Understanding

ICLR 2026poster

Recurrent large language models (Recurrent LLMs) offer linear computational complexity as efficient alternatives to quadratic self-attention-based LLMs (Self-Attention LLMs). However, Recurrent LLMs underperform on long-context tasks due to limited fixed-size memory. Previous research focused on arc…

Cited by 0SourceScholar
2026

Task-Aware Image Signal Processor for Advanced Visual Perception

CVPR 2026

In recent years, there has been a growing trend in computer vision towards exploiting RAW sensor data, which preserves richer information compared to conventional low-bit RGB images. Early studies mainly focused on enhancing visual quality, while more recent efforts aim to leverage the abundant info

Cited by 0SourcecodeScholar
2026

The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner

ICLR 2026poster

Length generalization, the ability to solve problems of longer sequences than those observed during training, poses a core challenge of Transformer-based large language models (LLMs). Although existing studies have predominantly focused on data-driven approaches for particular arithmetic operations…

Cited by 0SourceScholar
2026

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning

ICML 2026poster

Large Reasoning Models (LRMs) have achieved remarkable progress thanks to Reinforcement Learning with Verifiable Rewards (RLVR) on Chain-of-Thoughts (CoTs). However, since long CoTs naturally contain trial and errors and mainstream RLVR approaches choose outcome-correct CoT trajectories for memoriza…

Cited by 0SourceScholar
2026

Unlocking Full Efficiency of Token Filtering in Large Language Model Training

ICLR 2026poster

Token filtering has been proposed to enhance the utility of large language models (LLMs) by eliminating inconsequential tokens during training. While using fewer tokens is expected to reduce computational workloads, existing methods have not yet achieved a real-world efficiency boost. This is primar…

Cited by 0SourceScholar
2026

Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data Perspective

ICLR 2026poster

Post-training algorithms such as Supervised Fine-Tuning (SFT) and Reinforcement Fine-Tuning (RFT) are widely used to adapt multimodal large language models to downstream tasks. While effective at task adaptation, their impact on prior knowledge remains unclear. In this paper, we introduce jigsaw puz…

Cited by 0SourceScholar
2025

A Model Stealing Attack Against Multi-Exit Networks

ICASSP 2025accepted

Compared to traditional neural networks with a single output channel, a multi-exit network has multiple exits that allow for early outputs from the model's intermediate layers, thus significantly improving computational efficiency while maintaining similar main task accuracy. Existing model stealing…

Cited by 0SourceScholar
2025

A Self-Evolving Framework for Multi-Agent Medical Consultation Based on Large Language Models

ICASSP 2025accepted

We propose a multi-agent approach (SeM-Agents) based on large language models for medical consultations. This framework incorporates various doctor roles and auxiliary roles, with agents communicating through natural language. Using a residual structure, the system conducts multi-round medical consu…

Cited by 0SourceScholar
2025

Are Your LLMs Capable of Stable Reasoning?

ACL 2025finding

The rapid advancement of large language models (LLMs) has shown remarkable progress in complex reasoning tasks. However, a significant disparity exists between benchmark performances and real-world applications. We attribute this gap primarily to current evaluation protocols and metrics, which inade…

2025

Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language

CVPR 2025poster

Diffusion models can generate realistic and diverse images, potentially facilitating data availability for data-intensive perception tasks. However, leveraging these models to boost performance on downstream tasks with synthetic data poses several challenges, including aligning with real data distri…

Cited by 2SourcePDFScholar
2025

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law

ACL 2025long

Scaling law builds the relationship between training computation and validation loss, enabling researchers to effectively predict the loss trending of models across different levels of computation. However, a gap still remains between validation loss and the model’s downstream capabilities, making i…

2025

ColaDex: Contact-guided Optimization and VLM-assisted Selection for Task-oriented Dexterous Grasp Generation

IROS 2025

Task-oriented dexterous grasp generation aims to generate stable and functional grasps that enable a robotic hand to effectively interact with objects to accomplish specific tasks. However, generating high-dimensional hand configurations that seamlessly adapt to diverse task requirements and object

Cited by 0SourceScholar
2025

CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward

EMNLP 2025

Answer verification is crucial not only for evaluating large language models (LLMs) by matching their unstructured outputs against standard answers, but also serves as the reward model to guide LLM optimization. Most evaluation frameworks rely on regularized matching or employ general LLMs for answe

2025

Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement

ACL 2025long

The quality of Supervised Fine-Tuning (SFT) data plays a critical role in enhancing the conversational capabilities of Large Language Models (LLMs). However, the availability of high-quality human-annotated SFT data has become a significant bottleneck for LLMs, necessitating a greater reliance on sy…

2025

Contact Map Transfer with Conditional Diffusion Model for Generalizable Dexterous Grasp Generation

NeurIPS 2025poster

Dexterous grasp generation is a fundamental challenge in robotics, requiring both grasp stability and adaptability across diverse objects and tasks. Analytical methods ensure stable grasps but are inefficient and lack task adaptability, while generative approaches improve efficiency and task integra…

Cited by 0SourcecodeScholar
2025

Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction Tuning

EMNLP 2025

Visual Instruction Tuning (VIT) aims to enhance Multimodal Large Language Models (MLLMs), yet its effectiveness is often compromised by corrupted datasets with issues such as hallucinated content, incorrect responses, and poor OCR quality. Previous approaches to address these challenges have focused

Cited by 0SourcePDFScholar
2025

Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLMs

ICCV 2025poster

Creativity is a fundamental aspect of intelligence, involving the ability to generate novel and appropriate solutions across diverse contexts. While Large Language Models (LLMs) have been extensively evaluated for their creative capabilities, the assessment of Multimodal Large Language Models (MLLMs…

Cited by 0SourcePDFScholar
2025

CritiQ: Mining Data Quality Criteria from Human Preferences

ACL 2025long

Language model heavily depends on high-quality data for optimal performance. Existing approaches rely on manually designed heuristics, the perplexity of existing models, training classifiers, orcareful prompt engineering, which require significant expert experience and human annotation effort while…

2025

CryoGEN: Generative Energy-based Models for Cryogenic Electron Tomography Reconstruction

ICLR 2025poster

Cryogenic electron tomography (Cryo-ET) is a powerful technique for visualizing subcellular structures in their native states. Nonetheless, its effectiveness is compromised by anisotropic resolution artifacts caused by the missing-wedge effect. To address this, IsoNet, a deep learning-based method,…

Cited by 0SourcePDFScholar
2025

DistillW2N: A Lightweight One-Shot Whisper to Normal Voice Conversion Model Using Distillation of Self-Supervised Features

ICASSP 2025accepted

Whisper to Normal voice conversion (W2N) holds great promise for assistive communication and healthcare, making it an exciting area of research and development. Recent advancements in W2N are predominantly driven by self-supervised speech representation learning (SSL) techniques. While effective, SS…

Cited by 0SourceScholar
2025

DuMo: Dual Encoder Modulation Network for Precise Concept Erasure

AAAI 2025technical

The exceptional generative capability of text-to-image models has raised substantial safety concerns regarding the generation of Not-Safe-For-Work (NSFW) content and potential copyright infringement. To address these concerns, previous methods safeguard the models by eliminating inappropriate concep…

2025

Dual-energy CT metal artifact reduction by combined material decomposition and projection domain threshold segmentation

ICASSP 2025accepted

Dual-energy CT exploits the different attenuation characteristics of substances under different energy X-rays and collects high- and low-energy data from the same area to differentiate and quantify specific substances, which is now widely used in clinical diagnosis, disease monitoring, and other fie…

Cited by 0SourceScholar
2025

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

CVPR 2025poster

GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language Models to perceive and generate images, texts, and speeches end-to-end with publicly available data remains challenging…

Cited by 23SourcePDFScholar
2025

FRTree Planner: Robot Navigation in Cluttered and Unknown Environments With Tree of Free Regions

RA-L 2025

In this work, we present FRTree planner, a novel robot navigation framework that leverages a tree structure of free regions, specifically designed for navigation in cluttered and unknown environments with narrow passages. The framework continuously incorporates real-time perceptive information to id

Cited by 6SourceScholar
2025

FaceShot: Bring Any Character into Life

ICLR 2025poster

In this paper, we present ***FaceShot***, a novel training-free portrait animation framework designed to bring any character into life from any driven video without fine-tuning or retraining. We achieve this by offering precise and robust reposed landmark sequences from an appearance-guided landmark…

2025

FedMKT: Federated Mutual Knowledge Transfer for Large and Small Language Models

COLING 2025main

Recent research in federated large language models (LLMs) has primarily focused on enabling clients to fine-tune their locally deployed homogeneous LLMs collaboratively or on transferring knowledge from server-based LLMs to small language models (SLMs) at downstream clients. However, a significant g…

2025

G2-SDF: Geometry-Guided Neural Signed Distance Fields for Scalable and Detailed Reconstruction

RA-L 2025

Effcient reconstruction methods, particularly capable of providing detailed information on obstacle distances across diverse environments, are crucial for effective robot motion planning. In this context, neural Signed Distance Fields (SDFs) offer a powerful solution by learning implicit representat

Cited by 1SourceScholar
2025

Gaussian Splatting with Reflectance Regularization for Endoscopic Scene Reconstruction

IROS 2025

Endoscopic reconstruction plays a crucial role in surgical robotics. The dynamic lighting conditions and integrated camera-light source in endoscopic scenes create a distinct reconstruction challenge: shape ambiguity. To mitigate this, we propose a Gaussian Splatting (GS) based framework for endosco

Cited by 1SourcecodeScholar
2025

Hybrid Reciprocal Transformer with Triplet Feature Alignment for Scene Graph Generation

CVPR 2025poster

Scene graph generation is a pivotal task in computer vision, focusing on comprehensive identification of visual relation tuples embedded within images. The advancement of methods involving triplets has sought to enhance task performance by integrating triplets as contextual features for more precise…

2025

Information Density Principle for MLLM Benchmarks

ICCV 2025poster

With the emergence of Multimodal Large Language Models (MLLMs), hundreds of benchmarks have been developed to ensure the reliability of MLLMs in downstream tasks. However, the evaluation mechanism itself may not be reliable. For developers of MLLMs, questions remain about which benchmark to use and…

2025

Interactive Navigation for Legged Manipulators with Learned Arm-Pushing Controller

IROS 2025

Interactive navigation is crucial in scenarios where proactively interacting with objects can yield shorter paths, thus significantly improving traversal efficiency. Existing methods primarily focus on using the robot body to relocate obstacles during navigation. However, they prove ineffective in n

Cited by 5SourcecodeScholar
2025

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

ACL 2025finding

Despite the promising performance of Large Vision Language Models (LVLMs) in visual understanding, they occasionally generate incorrect outputs. While reward models (RMs) with reinforcement learning or test-time scaling offer the potential for improving generation quality, a critical gap remains: pu…

2025

LLM-DR: A Novel LLM-Aided Diffusion Model for Rule Generation on Temporal Knowledge Graphs

AAAI 2025technical

Among various temporal knowledge graph (TKG) extrapolation methods, rule-based approaches stand out for their explicit rules and transparent reasoning paths. However, the vast search space for rule extraction poses a challenge in identifying high-quality logic rules. To navigate this challenge, we e…

Cited by 0SourcePDFScholar
2025

MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space

ACL 2025finding

Data quality and diversity are key to the construction of effective instruction-tuning datasets. With the increasing availability of open-source instruction-tuning datasets, it is advantageous to automatically select high-quality and diverse subsets from a vast amount of data. Existing methods typic…

2025

MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control

ICCV 2025poster

The rapid advancement of diffusion models has greatly improved video synthesis, especially in controllable video generation, which is vital for applications like autonomous driving. Although DiT with 3D VAE has become a standard framework for video generation, it introduces challenges in controllabl…

2025

Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs

ICLR 2025poster

Large language models (LLMs) exhibit hallucinations (i.e., unfaithful or nonsensical information) when serving as AI assistants in various domains. Since hallucinations always come with truthful content in the LLM responses, previous factuality alignment methods that conduct response-level preferenc…

2025

Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models

NeurIPS 2025poster

Large Language Models (LLMs) have shown strong abilities in general language tasks, yet adapting them to specific domains remains a challenge. Current method like Domain Adaptive Pretraining (DAPT) requires costly full-parameter training and suffers from catastrophic forgetting. Meanwhile, Retrieval…

Cited by 0SourceScholar
2025

MindSearch: Mimicking Human Minds Elicits Deep AI Searcher

ICLR 2025poster

Information seeking and integration is a complex cognitive task that consumes enormous time and effort. Inspired by the remarkable progress of Large Language Models, recent works attempt to solve this task by combining LLMs and search engines. However, these methods still obtain unsatisfying perform…

2025

Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of Go

NeurIPS 2025poster

Large language models (LLMs) have demonstrated exceptional performance in reasoning tasks such as mathematics and coding, matching or surpassing human capabilities. However, these impressive reasoning abilities face significant challenges in specialized domains. Taking Go as an example, although Alp…

Cited by 0SourceScholar
2025

Mixture of insighTful Experts (MoTE): The Synergy of Reasoning Chains and Expert Mixtures in Self-Alignment

ACL 2025long

As the capabilities of large language models (LLMs) continue to expand, aligning these models with human values remains a significant challenge. Recent studies show that reasoning abilities contribute significantly to model safety, while integrating Mixture-of-Experts (MoE) architectures can further…

Cited by 0SourcePDFScholar
2025

MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation

ICCV 2025poster

Existing text-to-video methods struggle to transfer motion smoothly from a reference object to a target object with significant differences in appearance or structure between them. To address this challenge, we introduce MotionShot, a training-free framework capable of parsing reference-target corre…

2025

MusKGC: A Flexible Multi-source Knowledge Enhancement Framework for Open-World Knowledge Graph Completion

EMNLP 2025

Open-world knowledge graph completion (KGC) aims to infer novel facts by enriching existing graphs with external knowledge sources while maintaining semantic consistency under the open-world assumption (OWA). Generation-based KGC methods leverage the inherent strengths of large language models (LLMs

2025

OOD-Barrier: Build a Middle-Barrier for Open-Set Single-Image Test Time Adaptation via Vision Language Models

NeurIPS 2025poster

In real-world environments, a well-designed model must be capable of handling dynamically evolving distributions, where both in-distribution (ID) and out-of-distribution (OOD) samples appear unpredictably and individually, making real-time adaptation particularly challenging. While open-set test-tim…

Cited by 0SourceScholar
2025

OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference

ACL 2025long

Recent advancements in open-source multi-modal large language models (MLLMs) have primarily focused on enhancing foundational capabilities, leaving a significant gap in human preference alignment. This paper introduces OmniAlign-V, a comprehensive dataset of 200K high-quality training samples featur…

2025

PatchScaler: An Efficient Patch-Independent Diffusion Model for Image Super-Resolution

ICCV 2025poster

While diffusion models significantly improve the perceptual quality of super-resolved images, they usually require a large number of sampling steps, resulting in high computational costs and long inference times. Recent efforts have explored reasonable acceleration schemes by reducing the number of…

2025

Physically Robust and Imperceptible Adversarial Examples Generation Based on Frequency

ICASSP 2025accepted

Adversarial examples generated in digital space may fail to work in the physical world because the recapture process will ruin the adversarial property of the examples. Several approaches have been proposed to generate adversarial examples that can survive in the physical world, they however either…

Cited by 0SourceScholar
2025

Pre-Trained Policy Discriminators are General Reward Models

NeurIPS 2025poster

We offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guiding the training policy towards a target policy with desired behaviors. Based on this conceptual insight, we propose a sc…

Cited by 0SourceScholar
2025

Promptable Anomaly Segmentation with SAM Through Self-Perception Tuning

AAAI 2025technical

Segment Anything Model (SAM) has made great progress in anomaly segmentation tasks due to its impressive generalization ability. However, existing methods that directly apply SAM through prompting often overlook the domain shift issue, where SAM performs well on natural images but struggles in indus…

2025

Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study

COLING 2025main

Recent advancements in large language models (LLMs) have significantly enhanced their coding capabilities. However, existing benchmarks predominantly focused on simplified or isolated aspects of coding, such as single-file code generation or repository issue debugging, falling short of measuring the…

2025

RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything

ICLR 2025oral

Recent segmentation methods, which adopt large-scale data training and transformer architecture, aim to create one foundation model that can perform multiple tasks. However, most of these methods rely on heavy encoder and decoder frameworks, hindering their performance in real-time scenarios.…

2025

Redundancy Principles for MLLMs Benchmarks

ACL 2025long

With the rapid iteration of Multi-modality Large Language Models (MLLMs) and the evolving demands of the field, the number of benchmarks produced annually has surged into the hundreds. The rapid growth has inevitably led to significant redundancy among benchmarks. Therefore, it is crucial to take a…

Cited by 0SourcePDFScholar
2025

RepeatLeakage: Leak Prompts from Repeating as Large Language Model Is a Good Repeater

AAAI 2025technical

With the development of large language models (LLMs), numerous online applications based on these models have emerged. As system prompts significantly influence the performance of LLMs, many such applications conceal their system prompts and regard them as intellectual property. Consequently, numero…

2025

Rethinking Verification for LLM Code Generation: From Generation to Testing

NeurIPS 2025poster

Large language models (LLMs) have recently achieved notable success in code‑generation benchmarks such as HumanEval and LiveCodeBench. However, a detailed examination reveals that these evaluation suites often comprise only a limited number of homogeneous test cases, resulting in subtle faults going…

Cited by 0SourcecodeScholar
2025

Robot Navigation in Unknown and Cluttered Workspace with Dynamical System Modulation in Starshaped Roadmap

ICRA 2025

Compared to conventional decomposition methods that use ellipses or polygons to represent free space, starshaped representation can better capture the natural distribution of sensor data, thereby exploiting a larger portion of traversable space. This paper introduces a novel motion planning and cont

Cited by 3SourcecodeScholar
2025

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models

EMNLP 2025

Steerability, or the ability of large language models (LLMs) to adapt outputs to align with diverse community-specific norms, perspectives, and communication styles, is critical for real-world applications but remains under-evaluated. We introduce STEER-BENCH, a benchmark for assessing population-sp

Cited by 0SourcePDFScholar
2025

SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution

ACL 2025finding

Large Language Models (LLMs) have demonstrated remarkable proficiency across a variety of complex tasks. One significant application of LLMs is in tackling software engineering challenges, particularly in resolving real-world tasks on GitHub by fixing code based on the issues reported by the users.…

2025

Semantic-guided Masked Mutual Learning for Multi-modal Brain Tumor Segmentation with Arbitrary Missing Modalities

AAAI 2025technical

Malignant brain tumors have become an aggressive and dangerous disease that leads to death worldwide. Multi-modal MRI data is crucial for accurate brain tumor segmentation, but missing modalities common in clinical practice can severely degrade the segmentation performance. While incomplete multi-mo…

Cited by 0SourcePDFScholar
2025

Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning

NeurIPS 2025poster

Enhancing large vision-language models (LVLMs) with visual slow-thinking reasoning is crucial for solving complex multimodal tasks. However, since LVLMs are mainly trained with vision-language alignment, it is difficult to adopt on-policy reinforcement learning (RL) to develop the slow thinking abil…

Cited by 0SourceScholar
2025

Social Recommendation via Graph-Level Counterfactual Augmentation

AAAI 2025technical

Traditional recommendation system focus more on the correlations between users and items (user-item relationships), while research on user-user relationships has received significant attention these years, which is also known as social recommendation. Graph-based models have achieved a great success…

Cited by 0SourcePDFScholar
2025

SocialMOIF: Multi-Order Intention Fusion for Pedestrian Trajectory Prediction

CVPR 2025poster

The analysis and prediction of agent trajectories are crucial for decision-making processes in intelligent systems, with precise short-term trajectory forecasting being highly significant across a range of applications. Agents and their social interactions have been quantified and modeled by researc…

2025

Sparse Meets Dense: Unified Generative Recommendations with Cascaded Sparse-Dense Representations

NeurIPS 2025poster

Generative models have recently gained attention in recommendation systems by directly predicting item identifiers from user interaction sequences. However, existing methods suffer from significant information loss due to the separation of stages such as quantization and sequence modeling, hindering…

Cited by 0SourceScholar
2025

Spherical Scissor-Like Reconfigurable Palm Design in Robotic Hands: Insights from Human Hand Functionality

IROS 2025

The human palm demonstrates spatial reconfigurability during the gripping process and forms a spherical grasping envelope. Based on these observations, this study designs a reconfigurable spherical palm that incorporates a spatial scissor mechanism, which only requires a single actuator to reshape t

Cited by 0SourceScholar
2025

TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language Models

CVPR 2025poster

Large pre-trained Vision-Language Models (VLMs) such as CLIP have demonstrated excellent zero-shot generalizability across various downstream tasks. However, recent studies have shown that the inference performance of CLIP can be greatly degraded by small adversarial perturbations, especially its vi…

2025

Training Language Models to Critique With Multi-agent Feedback

EMNLP 2025

Critique ability, a meta-cognitive capability of humans, presents significant challenges for LLMs to improve. While utilizing human annotation can enhance critique ability effectively, most recent works primarily rely on supervised fine-tuning (SFT) using critiques generated by a single LLM like GPT

2025

UnitCoder: Scalable Code Synthesis from Pre-training Corpora

EMNLP 2025

Large Language Models (LLMs) have demonstrated remarkable capabilities in various tasks, yet code generation remains a major challenge. Despite the abundant sources of code data, constructing high-quality training datasets at scale poses a significant challenge. Pre-training code data typically suff

2025

What are the Essential Factors in Crafting Effective Long Context Multi-Hop Instruction Datasets? Insights and Best Practices

ACL 2025long

Recent advancements in large language models (LLMs) with extended context windows have significantly improved various tasks. To improve long-context capabilities, much work focuses on augmenting LLM’s capabilities with synthetic data. Existing methods often leverage the Self-Instruct framework to ge…

2024

"Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation"

ECCV 2024poster

"Multimodal large language models (MLLMs) have shown impressive reasoning abilities. However, they are also more vulnerable to jailbreak attacks than their LLM predecessors. Although still capable of detecting the unsafe responses, we observe that safety mechanisms of the pre-aligned LLMs in MLLMs c…

Cited by 48SourcePDFScholar
2024

4D Contrastive Superflows are Dense 3D Representation Learners

ECCV 2024poster

"In the realm of autonomous driving, accurate 3D perception is the foundation. However, developing such models relies on extensive human annotations – a process that is both costly and labor-intensive. To address this challenge from a data representation learning perspective, we introduce SuperFlow,…

2024

A Light-Weight State Detection Model for Kalman-Filter-Based Acoustic Feedback Cancellation with Rapid Recovery from Abrupt Path Changes

ICASSP 2024accepted

The partitioned block frequency domain Kalman Filter (PBFDKF) has been applied in acoustic feedback cancellation (AFC) due to its fast convergence and low steady-state misalignment. However, in cases where the feedback path experiences abrupt changes, the Kalman filter, once it reaches a steady stat…

Cited by 0SourceScholar
2024

A Unified Temporal Knowledge Graph Reasoning Model Towards Interpolation and Extrapolation

ACL 2024long

Temporal knowledge graph (TKG) reasoning has two settings: interpolation reasoning and extrapolation reasoning. Both of them draw plenty of research interest and have great significance. Methods of the former de-emphasize the temporal correlations among facts sequences, while methods of the latter r…

2024

ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models

NeurIPS 2024poster

Large language models (LLMs) exhibit hallucinations in long-form question-answering tasks across various domains and wide applications. Current hallucination detection and mitigation datasets are limited in domain and size, which struggle to scale due to prohibitive labor costs and insufficient reli…

2024

ANAH: Analytical Annotation of Hallucinations in Large Language Models

ACL 2024long

Reducing the ‘hallucination' problem of Large Language Models (LLMs) is crucial for their wide applications. A comprehensive and fine-grained measurement of the hallucination is the first key step for the governance of this issue but is under-explored in the community.Thus, we present ANAH, a biling…

2024

Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks

NAACL 2024long

Recently, the large language model (LLM) community has shown increasing interest in enhancing LLMs’ capability to handle extremely long documents. As various long-text techniques and model architectures emerge, the precise and detailed evaluation of models’ long-text capabilities has become increasi…

2024

Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models

ACL 2024findings

Open-sourced Large Language Models (LLMs) have achieved great success in various NLP tasks, however, they are still far inferior to API-based models when acting as agents. How to integrate agent ability into general LLMs becomes a crucial and urgent problem.This paper first delivers three key observ…

2024

AlchemistCoder: Harmonizing and Eliciting Code Capability by Hindsight Tuning on Multi-source Data

NeurIPS 2024poster

Open-source Large Language Models (LLMs) and their specialized variants, particularly Code LLMs, have recently delivered impressive performance. However, previous Code LLMs are typically fine-tuned on single-source data with limited quality and diversity, which may insufficiently elicit the potentia…

2024

Any-point Trajectory Modeling for Policy Learning

RSS 2024poster

Learning from demonstration is a powerful method for teaching robots new skills, and having more demonstration data often improves policy learning. However, the high cost of collecting demonstration data is a significant bottleneck. Videos, as a rich data source, contain knowledge of behaviors, phys…

Cited by 102SourcePDFScholar
2024

AnyControl: Create Your Artwork with Versatile Control on Text-to-Image Generation

ECCV 2024poster

"The field of text-to-image (T2I) generation has made significant progress in recent years, largely driven by advancements in diffusion models. Linguistic control enables effective content creation, but struggles with fine-grained control over image generation. This challenge has been explored, to a…

2024

BotChat: Evaluating LLMs’ Capabilities of Having Multi-Turn Dialogues

NAACL 2024findings

In the realm of modern Large Language Models (LLMs), facilitating high-quality, multi-turn dialogues with humans represents a cornerstone feature. However, human-based evaluation of such a capability involves substantial manual effort. This study offers a formative assessment of current LLMs’ profic…

2024

Can AI Assistants Know What They Don't Know?

ICML 2024poster

AI assistants powered by Large Language Models (LLMs) have demonstrated impressive performance in various tasks. However, LLMs still make factual errors in knowledge-intensive tasks such as open-domain question answering. These untruthful responses from AI assistants can pose significant risks in pr…

2024

Collision-Free Trajectory Optimization in Cluttered Environments Using Sums-of-Squares Programming

RA-L 2024

In this work, we propose a trajectory optimization approach for robot navigation in cluttered 3D environments. We represent the robot's geometry as a semialgebraic set defined by polynomial inequalities such that robots with general shapes can be suitably characterized. We exploit the collision-free

Cited by 12SourcecodeScholar
2024

CriticEval: Evaluating Large-scale Language Model as Critic

NeurIPS 2024poster

Critique ability, i.e., the capability of Large Language Models (LLMs) to identify and rectify flaws in responses, is crucial for their applications in self-improvement and scalable oversight. While numerous studies have been proposed to evaluate critique ability of LLMs, their comprehensiveness and…

Cited by 6SourcePDFScholar
2024

DataElixir: Purifying Poisoned Dataset to Mitigate Backdoor Attacks via Diffusion Models

AAAI 2024technical

Dataset sanitization is a widely adopted proactive defense against poisoning-based backdoor attacks, aimed at filtering out and removing poisoned samples from training datasets. However, existing methods have shown limited efficacy in countering the ever-evolving trigger functions, and often leading…

2024

DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and Perception

CVPR 2024poster

Current perceptive models heavily depend on resource-intensive datasets prompting the need for innovative solutions. Leveraging recent advances in diffusion models synthetic data by constructing image inputs from various annotations proves beneficial for downstream tasks. While prior methods have se…

Cited by 26SourcePDFScholar
2024

Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained Optimization

NeurIPS 2024poster

Recent research indicates that large language models (LLMs) are susceptible to jailbreaking attacks that can generate harmful content. This paper introduces a novel token-level attack method, Adaptive Dense-to-Sparse Constrained Optimization (ADC), which has been shown to successfully jailbreak mult…

2024

EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI

CVPR 2024poster

In the realm of computer vision and robotics embodied agents are expected to explore their environment and carry out human instructions. This necessitates the ability to fully understand 3D scenes given their first-person observations and contextualize them into language for interaction. However tra…

2024

EpiGEN: An Efficient Multi-Api Code GENeration Framework under Enterprise Scenario

COLING 2024main

In recent years, Large Language Models (LLMs) have demonstrated exceptional performance in code-generation tasks. However, under enterprise scenarios where private APIs are pre-built, general LLMs often fail to meet expectations. Existing approaches are confronted with drawbacks of high resource con…

Cited by 0SourcePDFScholar
2024

Everything2Motion: Synchronizing Diverse Inputs via a Unified Framework for Human Motion Synthesis

AAAI 2024technical

In the dynamic field of film and game development, the emergence of human motion synthesis methods has revolutionized avatar animation. Traditional methodologies, typically reliant on single modality inputs like text or audio, employ modality-specific model frameworks, posing challenges for unified…

Cited by 3SourcePDFScholar
2024

From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models

CVPR 2024poster

Scene graph generation (SGG) aims to parse a visual scene into an intermediate graph representation for downstream reasoning tasks. Despite recent advancements existing methods struggle to generate scene graphs with novel visual relation concepts. To address this challenge we introduce a new open-vo…

2024

GTA: A Benchmark for General Tool Agents

NeurIPS 2024poster

In developing general-purpose agents, significant focus has been placed on integrating large language models (LLMs) with various tools. This poses a challenge to the tool-use capabilities of LLMs. However, there are evident gaps between existing tool evaluations and real-world scenarios. Current eva…

2024

Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis

ICLR 2024poster

The rapid development of large language models (LLMs) has not only provided numerous opportunities but also presented significant challenges. This becomes particularly evident when LLMs inadvertently generate harmful or toxic content, either unintentionally or because of intentional inducement. Exis…

Cited by 36SourcePDFScholar
2024

GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation

ICLR 2024poster

Diffusion models have attracted significant attention due to the remarkable ability to create content and generate data for tasks like image classification. However, the usage of diffusion models to generate the high-quality object detection data remains an underexplored area, where not only image-l…

Cited by 24SourcePDFScholar
2024

Geometry-Aware Safety-Critical Local Reactive Controller for Robot Navigation in Unknown and Cluttered Environments

RA-L 2024

This work proposes a safety-critical local reactive controller that enables the robot to navigate in unknown and cluttered environments. In particular, the trajectory tracking task is formulated as a constrained polynomial optimization problem. Then, safety constraints are imposed on the control var

Cited by 13SourceScholar
2024

How Susceptible are Large Language Models to Ideological Manipulation?

EMNLP 2024main

Large Language Models (LLMs) possess the potential to exert substantial influence on public perceptions and interactions with information. This raises concerns about the societal impact that could arise if the ideologies within these models can be easily manipulated. In this work, we investigate how…

2024

HumanVid: Demystifying Training Data for Camera-controllable Human Image Animation

NeurIPS 2024poster

Human image animation involves generating videos from a character photo, allowing user control and unlocking the potential for video and movie production. While recent approaches yield impressive results using high-quality training data, the inaccessibility of these datasets hampers fair and transpa…

2024

Implicit Concept Removal of Diffusion Models

ECCV 2024poster

"Text-to-image (T2I) diffusion models often inadvertently generate unwanted concepts such as watermarks and unsafe images. These concepts, termed “implicit concepts”, can be unintentionally learned during training and then be generated uncontrollably during inference. Existing removal methods still…

2024

InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

NeurIPS 2024poster

The Large Vision-Language Model (LVLM) field has seen significant advancements, yet its progression has been hindered by challenges in comprehending fine-grained visual content due to limited resolution. Recent efforts have aimed to enhance the high-resolution understanding capabilities of LVLMs, ye…

2024

LLM Factoscope: Uncovering LLMs’ Factual Discernment through Measuring Inner States

ACL 2024findings

Large Language Models (LLMs) have revolutionized various domains with extensive knowledge and creative capabilities. However, a critical issue with LLMs is their tendency to produce outputs that diverge from factual reality. This phenomenon is particularly concerning in sensitive applications such a…

2024

LLM-REDIAL: A Large-Scale Dataset for Conversational Recommender Systems Created from User Behaviors with LLMs

ACL 2024findings

The large-scale conversational recommendation dataset is pivotal for the development of conversational recommender systems (CRS). Most existing CRS datasets suffers from the problems of data inextensibility and semantic inconsistency. To tackle these limitations and establish a benchmark in the conv…

Cited by 7SourcePDFScholar
2024

LLaST: Improved End-to-end Speech Translation System Leveraged by Large Language Models

ACL 2024findings

We introduces ***LLaST***, a framework for building high-performance Large Language model based Speech-to-text Translation systems. We address the limitations of end-to-end speech translation (E2E ST) models by exploring model architecture design and optimization techniques tailored for LLMs. Our ap…

2024

LawBench: Benchmarking Legal Knowledge of Large Language Models

EMNLP 2024main

We present LawBench, the first evaluation benchmark composed of 20 tasks aimed to assess the ability of Large Language Models (LLMs) to perform Chinese legal-related tasks. LawBench is meticulously crafted to enable precise assessment of LLMs’ legal capabilities from three cognitive levels that corr…

2024

Lean Workbook: A large-scale Lean problem set formalized from natural language math problems

NeurIPS 2024poster

Large language models have demonstrated impressive capabilities across various natural language processing tasks, especially in solving mathematical problems. However, large language models are not good at math theorem proving using formal languages like Lean. A significant challenge in this area is…

2024

MCGMapper: Light-Weight Incremental Structure from Motion and Visual Localization with Planar Markers and Camera Groups

IROS 2024poster

Structure from Motion (SfM) and visual localization in indoor texture-less scenes and industrial scenarios present prevalent yet challenging research topics. Existing SfM methods designed for natural scenes typically yield low accuracy or map-building failures due to insufficient robust feature extr…

Cited by 1SourcecodeScholar
2024

MMBENCH: Is Your Multi-Modal Model an All-around Player?

ECCV 2024oral

"Large vision-language models (VLMs) have recently achieved remarkable progress, exhibiting impressive multimodal perception and reasoning abilities. However, effectively evaluating these large VLMs remains a major challenge, hindering future development in this domain. Traditional benchmarks like V…

2024

MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

NeurIPS 2024poster

The advent of large vision-language models (LVLMs) has spurred research into their applications in multi-modal contexts, particularly in video understanding. Traditional VideoQA benchmarks, despite providing quantitative metrics, often fail to encompass the full spectrum of video content and inadequ…

2024

MagicDrive: Street View Generation with Diverse 3D Geometry Control

ICLR 2024poster

Recent advancements in diffusion models have significantly enhanced the data synthesis with 2D control. Yet, precise 3D control in street view generation, crucial for 3D perception tasks, remains elusive. Specifically, utilizing Bird's-Eye View (BEV) as the primary condition often leads to challenge…

2024

Make-It-Vivid: Dressing Your Animatable Biped Cartoon Characters from Text

CVPR 2024poster

Creating and animating 3D biped cartoon characters is crucial and valuable in various applications. Compared with geometry the diverse texture design plays an important role in making 3D biped cartoon characters vivid and charming. Therefore we focus on automatic texture design for cartoon character…

Cited by 6SourcePDFScholar
2024

MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

ACL 2024findings

Recent advancements in large language models (LLMs) have showcased significant improvements in mathematics. However, traditional math benchmarks like GSM8k offer a unidimensional perspective, which fall short in providing a holistic assessment of the LLMs’ math capabilities. To address this gap, we…

2024

MotionBooth: Motion-Aware Customized Text-to-Video Generation

NeurIPS 2024spotlight

In this work, we present MotionBooth, an innovative framework designed for animating customized subjects with precise control over both object and camera movements. By leveraging a few images of a specific object, we efficiently fine-tune a text-to-video model to capture the object's shape and attri…

Cited by 34SourcePDFScholar
2024

Multi-objective Cross-task Learning via Goal-conditioned GPT-based Decision Transformers for Surgical Robot Task Automation

ICRA 2024poster

Surgical robot task automation has been a promising research topic for improving surgical efficiency and quality. Learning-based methods have been recognized as an interesting paradigm and been increasingly investigated. However, existing approaches encounter difficulties in long-horizon goal-condit…

Cited by 4SourcecodeScholar
2024

OMG-Seg: Is One Model Good Enough For All Segmentation?

CVPR 2024poster

In this work we address various segmentation tasks each traditionally tackled by distinct or partially unified models. We propose OMG-Seg One Model that is Good enough to efficiently and effectively handle all the segmentation tasks including image semantic instance and panoptic segmentation as well…

2024

Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively

ECCV 2024poster

"The CLIP and Segment Anything Model (SAM) are remarkable vision foundation models (VFMs). SAM excels in segmentation tasks across diverse domains, whereas CLIP is renowned for its zero-shot recognition capabilities. This paper presents an in-depth exploration of integrating these two models into a…

2024

PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image Models

CVPR 2024poster

Recent advancements in personalized text-to-image (T2I) models have revolutionized content creation empowering non-experts to generate stunning images with unique styles. While promising animating these personalized images with realistic motions poses significant challenges in preserving distinct st…

2024

Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs

NeurIPS 2024poster

Vision Language Models (VLMs) demonstrate remarkable proficiency in addressing a wide array of visual questions, which requires strong perception and reasoning faculties. Assessing these two competencies independently is crucial for model refinement, despite the inherent difficulty due to the intert…

2024

ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs

EMNLP 2024finding

Large language models (LLMs) have demonstrated impressive capabilities across various tasks, but their performance is highly sensitive to the prompts utilized. This variability poses challenges for accurate assessment and user satisfaction. Current research frequently overlooks instance-level prompt…

2024

RGBManip: Monocular Image-based Robotic Manipulation through Active Object Pose Estimation

ICRA 2024poster

Robotic manipulation requires accurate perception of the environment, which poses a significant challenge due to its inherent complexity and constantly changing nature. In this context, RGB image and point-cloud observations are two commonly used modalities in visual-based robotic manipulation, but…

Cited by 17SourcecodeScholar
2024

RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation

CVPR 2024poster

Real-time multi-person pose estimation presents significant challenges in balancing speed and precision. While two-stage top-down methods slow down as the number of people in the image increases existing one-stage methods often fail to simultaneously deliver high accuracy and real-time performance.…

2024

Reliable and Efficient Concept Erasure of Text-to-Image Diffusion Models

ECCV 2024poster

"Text-to-image models encounter safety issues, including concerns related to copyright and Not-Safe-For-Work (NSFW) content. Despite several methods have been proposed for erasing inappropriate concepts from diffusion models, they often exhibit incomplete erasure, consume a lot of computing resource…

2024

STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering

AAAI 2024technical

Recently we have witnessed the rapid development of video question answering models. However, most models can only handle simple videos in terms of temporal reasoning, and their performance tends to drop when answering temporal-reasoning questions on long and informative videos. To tackle this prob…

2024

ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities

ECCV 2024poster

"Although great progress has been made in 3D visual grounding, current models still rely on explicit textual descriptions for grounding and lack the ability to reason human intentions from implicit instructions. We propose a new task called and introduce a new benchmark ScanReason which provides ove…

Cited by 11SourcePDFScholar
2024

Shopping MMLU: A Massive Multi-Task Online Shopping Benchmark for Large Language Models

NeurIPS 2024poster

Online shopping is a complex multi-task, few-shot learning problem with a wide and evolving range of entities, relations, and tasks. However, existing models and benchmarks are commonly tailored to specific tasks, falling short of capturing the full complexity of online shopping. Large Language Mode…

2024

T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

ACL 2024long

Large language models (LLMs) have achieved remarkable performance on various NLP tasks and are augmented by tools for broader applications. Yet, how to evaluate and analyze the tool utilization capability of LLMs is still under-explored. In contrast to previous works that evaluate models holisticall…

2024

Temporal Knowledge Graph Extrapolation via Causal Subhistory Identification

IJCAI 2024poster

Temporal knowledge graph extrapolation has become a prominent area of study interest in recent years. Numerous methods for extrapolation have been put forth, mining query-relevant information from history to generate forecasts. However, existing approaches normally do not discriminate between causal…

Cited by 0SourcePDFScholar
2024

Temporal Relational Context Learning for Extrapolation Reasoning on Temporal Knowledge Graphs

ICASSP 2024accepted

Extrapolation reasoning on Temporal Knowledge Graphs (TKGs) aims to predict future events from a set of historical Knowledge Graphs (KGs) in a chronological order. The temporally adjacent facts in TKGs naturally form event sequences, implying informative temporal event dependencies. Recently, many e…

Cited by 0SourceScholar
2024

Towards Language-Driven Video Inpainting via Multimodal Large Language Models

CVPR 2024poster

We introduce a new task -- language-driven video inpainting which uses natural language instructions to guide the inpainting process. This approach overcomes the limitations of traditional video inpainting methods that depend on manually labeled binary masks a process often tedious and labor-intensi…

2024

UVEB: A Large-scale Benchmark and Baseline Towards Real-World Underwater Video Enhancement

CVPR 2024poster

Learning-based underwater image enhancement (UIE) methods have made great progress. However the lack of large-scale and high-quality paired training samples has become the main bottleneck hindering the development of UIE. The inter-frame information in underwater videos can accelerate or optimize th…

2024

Vision Foundation Model Enables Generalizable Object Pose Estimation

NeurIPS 2024poster

Object pose estimation plays a crucial role in robotic manipulation, however, its practical applicability still suffers from limited generalizability. This paper addresses the challenge of generalizable object pose estimation, particularly focusing on category-level object pose estimation for unseen…

Cited by 0SourcePDFScholar
2024

YOLOv10: Real-Time End-to-End Object Detection

NeurIPS 2024poster

Over the past years, YOLOs have emerged as the predominant paradigm in the field of real-time object detection owing to their effective balance between computational cost and detection performance. Researchers have explored the architectural designs, optimization objectives, data augmentation strate…

2023

A Low-Latency Hybrid Multi-Channel Speech Enhancement System For Hearing Aids

ICASSP 2023accepted

This paper summarizes a hybrid multi-channel speech enhancement system for the ICASSP Signal Processing Grand Challenge: Clarity Challenge (Speech Enhancement for Hearing Aids) 2023. The system consists of a rule-based dereverberation module, a multi-channel enhancement module, and a post-processing…

Cited by 0SourceScholar
2023

Boosting Point Clouds Rendering via Radiance Mapping

AAAI 2023technical

Recent years we have witnessed rapid development in NeRF-based image rendering due to its high quality. However, point clouds rendering is somehow less explored. Compared to NeRF-based rendering which suffers from dense spatial sampling, point clouds rendering is naturally less computation intensive…

2023

Consistent-Teacher: Towards Reducing Inconsistent Pseudo-Targets in Semi-Supervised Object Detection

CVPR 2023highlight

In this study, we dive deep into the inconsistency of pseudo targets in semi-supervised object detection (SSOD). Our core observation is that the oscillating pseudo-targets undermine the training of an accurate detector. It injects noise into the student's training, leading to severe overfitting pro…

2023

Convolutional Recurrent MetriCGAN With Spectral Dimension Compression For Full-Band Speech Enhancement

ICASSP 2023accepted

MetricGAN and its variations have been proven to be an effective wide-band speech enhancement model. In this paper, we expand it to full-band enhancement by combining our recently proposed learnable spectral dimension compression mapping strategy. The encoder-decoder structure with a time-frequency…

Cited by 0SourceScholar
2023

Deep Fusion Transformer Network with Weighted Vector-Wise Keypoints Voting for Robust 6D Object Pose Estimation

ICCV 2023poster

One critical challenge in 6D object pose estimation from a single RGBD image is efficient integration of two different modalities, i.e., color and depth. In this work, we tackle this problem by a novel Deep Fusion Transformer (DFTr) block that can aggregate cross-modality features for improving pose…

Cited by 42PDFcodeScholar
2023

Demonstration-Guided Reinforcement Learning with Efficient Exploration for Task Automation of Surgical Robot

ICRA 2023poster

Task automation of surgical robot has the potentials to improve surgical efficiency. Recent reinforcement learning (RL) based approaches provide scalable solutions to surgical automation, but typically require extensive data collection to solve a task if no prior knowledge is given. This issue is kn…

Cited by 29SourcecodeScholar
2023

Dense Distinct Query for End-to-End Object Detection

CVPR 2023poster

One-to-one label assignment in object detection has successfully obviated the need of non-maximum suppression (NMS) as a postprocessing and makes the pipeline end-to-end. However, it triggers a new dilemma as the widely used sparse queries cannot guarantee a high recall, while dense queries inevitab…

2023

Globally Consistent Federated Graph Autoencoder for Non-IID Graphs

IJCAI 2023poster

Graph neural networks (GNNs) have been applied successfully in many machine learning tasks due to their advantages in utilizing neighboring information. Recently, with the global enactment of privacy protection regulations, federated GNNs have gained increasing attention in academia and industry. Ho…

2023

GlyphControl: Glyph Conditional Control for Visual Text Generation

NeurIPS 2023poster

Recently, there has been an increasing interest in developing diffusion-based text-to-image generative models capable of generating coherent and well-formed visual text. In this paper, we propose a novel and efficient approach called GlyphControl to address this task. Unlike existing methods that re…

2023

Improving Pixel-based MIM by Reducing Wasted Modeling Capability

ICCV 2023poster

There has been significant progress in Masked Image Modeling (MIM). Existing MIM methods can be broadly categorized into two groups based on the reconstruction target: pixel-based and tokenizer-based approaches. The former offers a simpler pipeline and lower computational cost, but it is known to be…

Cited by 48PDFcodeScholar
2023

Learning Shape Primitives via Implicit Convexity Regularization

ICCV 2023poster

Shape primitives decomposition has been an important and long-standing task in 3D shape analysis. Prior arts heavily rely on 3D point clouds or voxel data for shape primitives extraction, which are less practical in real-world scenarios. This paper proposes to learn shape primitives from multi-view…

Cited by 3PDFcodeScholar
2023

Mixed Autoencoder for Self-Supervised Visual Representation Learning

CVPR 2023poster

Masked Autoencoder (MAE) has demonstrated superior performance on various vision tasks via randomly masking image patches and reconstruction. However, effective data augmentation strategies for MAE still remain open questions, different from those in contrastive learning that serve as the most impor…

Cited by 51SourcePDFScholar
2023

RIFormer: Keep Your Vision Backbone Effective but Removing Token Mixer

CVPR 2023poster

This paper studies how to keep a vision backbone effective while removing token mixers in its basic building blocks. Token mixers, as self-attention for vision transformers (ViTs), are intended to perform information communication between different spatial tokens but suffer from considerable computa…

Cited by 37SourcePDFScholar
2023

RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank

ACL 2023long

Unsupervised sentence representation learning is one of the fundamental problems in natural language processing with various downstream applications. Recently, contrastive learning has been widely adopted which derives high-quality sentence representations by pulling similar semantics closer and pus…

2023

Robo3D: Towards Robust and Reliable 3D Perception against Corruptions

ICCV 2023poster

The robustness of 3D perception systems under natural corruptions from environments and sensors is pivotal for safety-critical applications. Existing large-scale 3D perception datasets often contain data that are meticulously cleaned. Such configurations, however, cannot reflect the reliability of p…

Cited by 123PDFcodeScholar
2023

Segment Any Point Cloud Sequences by Distilling Vision Foundation Models

NeurIPS 2023spotlight

Recent advancements in vision foundation models (VFMs) have opened up new possibilities for versatile and efficient visual perception. In this work, we introduce Seal, a novel framework that harnesses VFMs for segmenting diverse automotive point cloud sequences. Seal exhibits three appealing propert…

Cited by 66SourcePDFScholar
2023

Semantics-Aware Dynamic Localization and Refinement for Referring Image Segmentation

AAAI 2023technical

Referring image segmentation segments an image from a language expression. With the aim of producing high-quality masks, existing methods often adopt iterative learning approaches that rely on RNNs or stacked attention layers to refine vision-language features. Despite their complexity, RNN-based me…

Cited by 28SourcePDFScholar
2023

StereoPose: Category-Level 6D Transparent Object Pose Estimation from Stereo Images via Back-View NOCS

ICRA 2023poster

Most existing methods for category-level pose estimation rely on object point clouds. However, when considering transparent objects, depth cameras are usually not able to capture high-quality data, resulting in point clouds with severe artifacts. Without a complete point cloud, existing methods are…

Cited by 23SourceScholar
2023

TG-VQA: Ternary Game of Video Question Answering

IJCAI 2023poster

Video question answering aims at answering a question about the video content by reasoning the alignment semantics within them. However, since relying heavily on human instructions, i.e., annotations or priors, current contrastive learning-based VideoQA methods remains challenging to perform fine-gr…

Cited by 17SourcePDFScholar
2023

Task-customized Masked Autoencoder via Mixture of Cluster-conditional Experts

ICLR 2023top-25%

Masked Autoencoder (MAE) is a prevailing self-supervised learning method that achieves promising results in model pre-training. However, when the various downstream tasks have data distributions different from the pre-training data, the semantically irrelevant pre-training information might result i…

Cited by 21SourcePDFScholar
2023

UMC: A Unified Bandwidth-efficient and Multi-resolution based Collaborative Perception Framework

ICCV 2023poster

Multi-agent collaborative perception (MCP) has recently attracted much attention. It includes three key processes: communication for sharing, collaboration for integration, and reconstruction for different downstream tasks. Existing methods pursue designing the collaboration process alone, ignoring…

Cited by 40PDFcodeScholar
2023

Value-Informed Skill Chaining for Policy Learning of Long-Horizon Tasks with Surgical Robot

IROS 2023poster

Reinforcement learning is still struggling with solving long-horizon surgical robot tasks which involve multiple steps over an extended duration of time due to the policy exploration challenge. Recent methods try to tackle this problem by skill chaining, in which the long-horizon task is decomposed…

Cited by 7SourcecodeScholar
2023

Visual-Kinematics Graph Learning for Procedure-Agnostic Instrument Tip Segmentation in Robotic Surgeries

IROS 2023poster

Accurate segmentation of surgical instrument tip is an important task for enabling downstream applications in robotic surgery, such as surgical skill assessment, tool-tissue interaction and deformation modeling, as well as surgical autonomy. However, this task is very challenging due to the small si…

Cited by 2SourceScholar
2022

A Priori SNR Estimation for Speech Enhancement Based on PESQ-Induced Reinforcement Learning

ICASSP 2022accepted

Perceptual evaluation of speech quality (PESQ) is widely accepted as an effective objective metric closely related to the speech quality sensed by human listening perception. Due to its evaluation complexity and non-differentiability, PESQ is difficult to include in the cost function for deep learni…

Cited by 0SourceScholar
2022

A Sim-to-Real Object Recognition and Localization Framework for Industrial Robotic Bin Picking

RA-L 2022

We present a generic and robust sim-to-real deep-learning-based framework, namely S2R-Pick, for fast and accurate object recognition and localization in industrial robotic bin picking. Unlike existing works designed for general everyday environments, objects for industrial bin picking are often text

Cited by 59SourceScholar
2022

Attacking Video Recognition Models with Bullet-Screen Comments

AAAI 2022technical

Recent research has demonstrated that Deep Neural Networks (DNNs) are vulnerable to adversarial patches which introduce perceptible but localized changes to the input. Nevertheless, existing approaches have focused on generating adversarial patches on images, their counterparts in videos have been l…

2022

CODA: A Real-World Road Corner Case Dataset for Object Detection in Autonomous Driving

ECCV 2022poster

"Contemporary deep-learning object detection methods for autonomous driving usually assume prefixed categories of common traffic participants, such as pedestrians and cars. Most existing detectors are unable to detect uncommon objects and corner cases (e.g., a dog crossing a street), which may lead…

2022

Deliberated Domain Bridging for Domain Adaptive Semantic Segmentation

NeurIPS 2022accept

In unsupervised domain adaptation (UDA), directly adapting from the source to the target domain usually suffers significant discrepancies and leads to insufficient alignment. Thus, many UDA works attempt to vanish the domain gap gradually and softly via various intermediate spaces, dubbed domain bri…

2022

Dense Siamese Network for Dense Unsupervised Learning

ECCV 2022poster

"This paper presents Dense Siamese Network (DenseSiam), a simple unsupervised learning framework for dense prediction tasks. It learns visual representations by maximizing the similarity between two views of one image with two types of consistency, i.e., pixel consistency and region consistency. Con…

2022

End-to-End ASR-Enhanced Neural Network for Alzheimer's Disease Diagnosis

ICASSP 2022accepted

This paper presents an approach to Alzheimer’s disease (AD) diagnosis from spontaneous speech using an end-to-end ASR-enhanced neural network. Under the condition that only audio data are provided and accurate transcripts are unavailable, this paper proposes a system that can analyze utterances to d…

Cited by 0SourceScholar
2022

GCFSR: A Generative and Controllable Face Super Resolution Method Without Facial and GAN Priors

CVPR 2022poster

Face image super resolution (face hallucination) usually relies on facial priors to restore realistic details and preserve identity information. Recent advances can achieve impressive results with the help of GAN prior. They either design complicated modules to modify the fixed GAN prior or adopt co…

Cited by 88PDFcodeScholar
2022

Group R-CNN for Weakly Semi-Supervised Object Detection With Points

CVPR 2022poster

We study the problem of weakly semi-supervised object detection with points (WSSOD-P), where the training data is combined by a small set of fully annotated images with bounding boxes and a large set of weakly-labeled images with only a single point annotated for each instance. The core of this task…

Cited by 55PDFcodeScholar
2022

LAVT: Language-Aware Vision Transformer for Referring Image Segmentation

CVPR 2022poster

Referring image segmentation is a fundamental vision-language task that aims to segment out an object referred to by a natural language expression from an image. One of the key challenges behind this task is leveraging the referring expression for highlighting relevant positions in the image. A para…

Cited by 385PDFcodeScholar
2022

Multitask Gaussian Process With Hierarchical Latent Interactions

ICASSP 2022accepted

Multitask Gaussian process (MTGP) is powerful for joint learning of multiple tasks with complicated correlation patterns. However, due to the assembling of additive independent latent functions (LFs), all current MTGPs including the salient linear model of coregionalization (LMC) and convolution fra…

Cited by 0SourceScholar
2022

OCSampler: Compressing Videos to One Clip With Single-Step Sampling

CVPR 2022poster

Videos incorporate rich semantics as well as redundant information. Seeking a compact yet effective video representation, e.g., sample informative frames from the entire video, is critical to efficient video recognition. There have been works that formulate frame sampling as a sequential decision ta…

Cited by 33PDFcodeScholar
2022

RotateQVS: Representing Temporal Information as Rotations in Quaternion Vector Space for Temporal Knowledge Graph Completion

ACL 2022long

Temporal factors are tied to the growth of facts in realistic applications, such as the progress of diseases and the development of political situation, therefore, research on Temporal Knowledge Graph (TKG) attracks much attention. In TKG, relation patterns inherent with temporality are required to…

2022

SESR: Self-Ensembling Sim-to-Real Instance Segmentation for Auto-Store Bin Picking

IROS 2022poster

Instance segmentation is an important task for supporting robotic grasping in auto-store scenarios. Accurate segmentation usually relies on the quantity and quality of available annotated training data. However, it requires tremendous cost to obtain these labels. In this work, without requiring any…

Cited by 2SourceScholar
2022

SMASH: Improving SMAll Language Models’ Few-SHot Ability with Prompt-Based Distillation

EMNLP 2022finding

Large-scale language models coupled with prompts have shown remarkable performance on few-shot learning. However, through systematic experiments, we find that the few-shot performance of small language models is poor, and using prompts on them brings fewer improvements than on larger ones. In this p…

2022

Sim-to-Real 6D Object Pose Estimation via Iterative Self-Training for Robotic Bin Picking

ECCV 2022poster

"6D object pose estimation is important for robotic bin-picking, and serves as a prerequisite for many downstream industrial applications. However, it is burdensome to annotate a customized dataset associated with each specific bin-picking scenario for training pose estimation models. In this paper,…

Cited by 31SourcePDFScholar
2022

Task-Customized Self-Supervised Pre-training with Scalable Dynamic Routing

AAAI 2022technical

Self-supervised learning (SSL), especially contrastive methods, has raised attraction recently as it learns effective transferable representations without semantic annotations. A common practice for self-supervised pre-training is to use as much data as possible. For a specific downstream task, howe…

Cited by 23SourcePDFScholar